#context
3 notes tagged “context”. All notes →
-
My agents adapt to where they work in the codebase. Here’s how.
agents
The same agent needs different instructions for a checkout form and a database migration. Here’s how I use AGENTS.md files to make that context follow the code.
-
Anchor on a fact the model can't see
agents
An agent's confidence is uncorrelated with whether it's right. The cheapest correction you can make is to hold out one ground-truth fact it never had in context — usually a number — and check its output against that. Works the same whether you're watching one session or auditing a fleet's reports.
-
Benchmarks measure a model you are not running
agents
Terminal-Bench 4.0 tasks have a median of 394 tokens. SWE-bench Pro: 707. HumanEval: 117. Every major coding benchmark evaluates a model operating with essentially an empty context window — which is almost never the condition you run in. Unless you are running cattle.