The harness is not enough, and neither is going back to reading every line

Because the models were never trained to keep a codebase maintainable, and no orchestration on top of them changes that. That is the claim in Dex Horthy's AI Engineer talk, and it is right. Where it stops short is the remedy: put the humans back to reading every line. For one engineer that works. For a team shipping with agents, it only works if the reasoning behind each merged change survives the merge, which is what Backthread is for.

The training fact under the slogan

The talk's argument runs on how coding models are rewarded. A task from a benchmark like SWE-bench is fifteen minutes long and scored one or zero: did the hidden test pass without breaking the others. Nothing in that loop can penalise a try-catch that should not be there, a cast that exists only to make the test go green, or a change that makes the next change harder. Horthy puts the reason in one sentence:

Verifying code quality and maintainability is orders of magnitude harder than the code runs and the test pass, because the cost function of bad architecture is measured in months and years.
Dex Horthy — Harness Engineering is not Enough: Why Software Factories Fail · AI Engineer, July 2026
The clip starts at 13:18, where the training argument is made.

A reward that arrives months later cannot be propagated back into a fifteen-minute episode. So the models got much better at solving the problem in front of them and, by his account of HumanLayer's own lights-off experiment in July 2025, no better at leaving the system in a state anyone can work in. The industry numbers point the same way. Faros AI's 2026 report, from two years of telemetry across 22,000 developers, has median time in PR review up 441.5 percent under high AI adoption and PRs skipping review altogether up 31.3 percent.

That is a stronger version of an argument we make about why a merged pull request no longer means anyone understood it. His version is better because it names the mechanism: the thing the agent cannot hold is the long-horizon shape of the system, and that is precisely the thing a team has to hold instead.

Where the remedy breaks

Horthy's answer is to turn the lights back on. Do a product review, an architecture doc, a program design with real types and call graphs, then vertical slices, so the agent's pull request arrives already aligned and a human can read every line quickly. All of that is good practice and none of it is new.

But look at what those artefacts are. Each one is reasoning written down before the code exists, by whoever was in the room. Three months later the code has moved, the doc has not, and the person who wrote it may be on another team. The plan gave you one good review. It did not give the rest of the team a model of the system that lasts.

And "read every line" rebuilds understanding in the one head doing the reading, the pattern behind who on your team understands each part of the codebase. Thirty reviewers reading every line end up with thirty private models and nothing shared.

What has to be held, and by whom

The training fact tells you what the human contribution is now. It is not correctness, which the diff and the tests carry. It is the long-horizon reasoning: what this change assumed about the rest of the system, what it rejected, what it knowingly left. Measured on our own repository, 62 percent of decisions captured from agent sessions record no alternative and no trade-off at all, because for cheap, reversible changes nobody weighed one. The other 38 percent are the system's architecture, written by the only party that was present when it was decided.

So the work is not to read more, it is to keep that minority and make sure more than one person holds it. Capture the reasoning at the moment of the change, while the session that made it still exists. Show, per area of the system, where that reasoning is on record and where the only artefact is a diff. Route the reading to the areas where the record is blank. That is the shape of what Backthread does, and the map of who understands what is where it shows up; Horthy's planning documents would make excellent input to it, and are wasted if they only ever produce one review.

The harness is not enough. Neither is a reviewer with more discipline. What is enough is a team that holds the reasoning the model structurally cannot, and can see where it does not.

Connect one repo and you get the per-area picture of where the reasoning was captured and where it is missing the same day; the trial runs fourteen days.

In short

Software factories fail on a training fact, not a skill issue
Coding models are rewarded for passing a hidden test in a fifteen-minute task. The cost of a bad design shows up months later and cannot reach that reward, so no harness on top of the model can buy maintainability the model was never taught.
Reading every line rebuilds understanding in one head
Horthy's remedy, aligned planning plus full human review, makes each review fast. It does not give the rest of a team a shared, lasting model of the system, and the planning documents rot as the code moves.
The human contribution is the long-horizon reasoning
What agents cannot hold is what a change assumed, rejected and left. On our repository 62 percent of agent decisions record none of that. Keeping the rest, and making sure more than one person holds it, is the job.

Sources

  1. Harness Engineering is not Enough: Why Software Factories Fail — Dex Horthy, HumanLayer, AI Engineer, 23 July 2026 (quote at ~13:18)
  2. AI Engineering Report 2026: The Acceleration Whiplash — Faros AI, 21 May 2026
  3. Comprehension Debt: The Hidden Cost of AI-Generated Code — Addy Osmani, O'Reilly Radar, April 2026

Backthread shows how much of what your agents built your team really understands. See how it works