Our agents' context file claimed a thousand nodes. The real number was 77.
It gets injected into every agent session as fact, and nobody re-reads it, because a file written to sound authoritative reads as authoritative. Ours described one of our own diagrams as roughly a thousand nodes, for weeks. The largest version we have ever rendered is 77.
We make Backthread, which keeps the reasoning behind each change agents make to a system and shows a team where that understanding actually sits. A false premise in our own context file is the category of thing we are supposed to catch. We did not catch it.
What the file said, and what was true
The claim was about the map we generate from a repository: a graph of the system, rebuilt at every version. The context file that our agents load on every session described it in passing as being on the order of a thousand nodes. That single phrase then sat in the premise of every task anyone ran against that part of the code.
Measured against production, on the same day the claim was still in the file:
| Claimed | Measured | |
|---|---|---|
| Nodes in the latest version of the map | roughly 1,000 | 75 |
| Edges in the latest version | not stated | 222 |
| Largest version ever rendered, across all 348 | roughly 1,000 | 77 |
Not a rounding error. More than an order of magnitude, sustained across every version we have ever produced. The correction went into the file on 3 August 2026.
Why an order of magnitude is not a cosmetic error
A scale claim is an architecture decision wearing a description's clothes. At a thousand nodes you virtualise the canvas, cull off-screen geometry, sample rather than render, and think hard about whether layout can run in the request at all. At 75 you do none of that, and doing it anyway costs you a week and a pile of code that is harder to reason about than the problem it addressed.
Two consequences are in our own record.
The false premise mis-scoped a real validation before anyone checked it against production. The work was planned for a graph size that did not exist, which means the thing that got validated was not the thing that ships.
It also flipped which rendering path the code was actually taking. A branch chosen for large graphs was being exercised on graphs that were not large, so the path we believed was live was not the path running — a failure that reports nothing, because both branches produce a picture.
Neither of those is dramatic. That is the point. A wrong number in a context file does not throw; it quietly selects the wrong option, repeatedly, in work that otherwise looks fine.
Why this file rots differently from a stale comment
Everyone accepts that comments drift. This is a worse version of the same disease, for four reasons.
- Reach. A stale comment misleads whoever opens that file. A stale context file is prepended to every agent session in the repository, on every task, including tasks nowhere near the code it describes.
- Authority. It is written as instruction, not as commentary, and an agent treats it as a constraint rather than as a claim to verify. It will design to your number instead of measuring it.
- No test. Prose has no test. Nothing in your pipeline fails when the file stops being true, and nothing prompts a re-read, because the file is only edited when someone wants to add a rule.
- Inheritance. It is written once, usually early, by someone with the system fresh in mind, and is then read for years by people and agents that were not there. The estimate that was reasonable on the day it was typed becomes a fact nobody sourced.
The mechanism is the same one that makes an architecture diagram misleading rather than merely out of date — and as with a diagram that reshuffles on every commit, the fix is structural rather than a matter of trying harder.
The audit, in four steps
This took us under an hour once we accepted it needed doing. It is worth running on your own file today.
- List every quantitative claim in the file. Counts, sizes, latencies, version numbers, "we have about N services", "requests take roughly N milliseconds". Most files have between five and fifteen. Ours had fewer than we expected and still contained a wrong one.
- For each, write the query that produces the number now. A SQL statement, a script, a dashboard link. If you cannot write it in a couple of minutes, that is itself the finding: the claim was never measured, it was remembered.
- Delete every claim you cannot produce a number for. A missing specific costs an agent nothing. A confident wrong one costs you a week. This is the step people resist and it is the one that pays.
- Date the survivors, and re-run the queries on a schedule. A line reading "75 nodes, 222 edges (measured 2026-08-03)" tells the next reader both the number and how much to trust it. An undated number is an assertion.
The deeper fix is to stop typing numbers into prose at all: generate the paragraph from the query, or link to the dashboard and let the reader fetch it. Anything a human types by hand is a number that will be wrong later, and the only question is whether anyone will notice.
What this does not prove
Four honest limits.
- One file, one claim, one company. We found one wrong number in our own context file. Nothing here says most such files are wrong, only that ours was and that nothing in our process would have surfaced it.
- We cannot count the damage retroactively. Sessions are not tagged with which premises they acted on. We know of one mis-scoped validation and one wrong rendering path because we traced them; there may have been more and there may have been none.
- 348 versions is our render history, not a fact about codebases. A repository whose graph really does run to thousands of nodes exists. The error was claiming that ours was one.
- Correcting a line is not correcting a file. We fixed the number. We did not, that day, verify the rest of the document, which is why the audit above exists and why we now date the claims.
A newcomer is the last person who will ever read that file with fresh eyes, which is one reason it is worth handing them on their first morning as something to challenge rather than as orientation material. The broader version of the problem is the one we work on: what a team believes about a system, and whether any of it is still true after several thousand agent-written changes. That is what the record we keep is for, and our own file is in it, wrong number and correction date included.
Connect one repo and the reasoning behind each merged change starts accumulating alongside the areas that have none; the trial runs fourteen days.
In short
- A hand-typed number in a context file is injected into every session as fact
- A stale comment misleads whoever opens that file. A stale CLAUDE.md or AGENTS.md is prepended to every agent session in the repository, is written as instruction rather than commentary, and has no test that fails when it stops being true.
- Ours was wrong by more than an order of magnitude for weeks
- The file described one of our diagrams as roughly a thousand nodes. The latest version measured 75 nodes and 222 edges, and the largest across all 348 versions we have rendered is 77. The correction went in on 3 August 2026.
- A wrong scale claim silently picks an architecture
- It mis-scoped a validation against a graph size that does not exist, and it flipped which rendering path the code was taking. Neither failure throws an error, because both branches produce a plausible-looking result.
- Delete any claim you cannot produce a number for
- List the quantitative claims, write the query that returns each one today, delete the ones you cannot, and date the survivors. A missing specific costs an agent nothing; a confident wrong one costs a week.
Sources
Backthread shows how much of what your agents built your team really understands. See how it works