We deleted our LLM judge instead of tuning it
Measure the judge's lift over the gate it already sits on before you touch the prompt. Ours scored at or below that gate, between 0.71 and 0.94 times, which means no amount of rewriting would have saved it. We deleted the stage.
The judge sat in the pipeline behind Backthread, which turns agent sessions into a record of the reasoning behind each change to a system, and decides which of those questions about the codebase are worth putting in front of a person. Here is the number that decided it.
What the stage was doing
The ingest takes a session, pulls out the changes and the deliberation around them, and produces candidate items: the questions about this part of the system that a person should actually be asked. Far more candidates exist than anyone wants, so there is a salience gate that scores and thresholds them.
On top of that gate we put a model call. Its job was to catch the items that scored well but were not worth anyone's attention: the restatements, the trivia, the questions with an answer visible in the diff. A quality stage above a numeric one. This is close to the default architecture people reach for, and it is why the result is worth writing down.
The measurement that ended it
The obvious thing to measure is what a filter keeps. We measured what it threw away.
We took a sample of rows the judge had culled after the salience gate had already passed them — the cull-override subset — and had a human rate each one against the same standard the stage claimed to apply.
| What we measured | Result |
|---|---|
| Sampled rows the judge culled after the gate passed them | 90 |
| Of those, wrongly hidden | 96.7% (confidence interval 90.7–98.9) |
| Judge's measured lift over the salience gate beneath it | 0.71–0.94× |
| Rows a simulated retune would have recovered | 5 of 90 |
| Model calls saved per ingest after deletion | at least 3 |
The second row is uncomfortable and the third row is the one that decided it. Lift below 1.0 means the stage was performing at or beneath the gate it was supposed to improve on. It was not a weak filter. On this evidence it was, within the resolution we could measure, an expensive coin.
That reframes the problem completely. A judge that is directionally right and badly calibrated is a tuning problem. A judge with no measurable lift over the thing underneath it is a design problem, and prompts do not fix design problems.
Why the retune would not have worked
Our first plan was the normal one: rewrite the rubric, add examples of the items it was wrongly killing, re-run. Before moving any code we simulated the retune against the same 90 rows.
It would have recovered five of them.
Five rows out of ninety, in exchange for a prompt that is now longer, an extra round of eval work, and a stage that still has to run on every ingest. Deleting the stage recovered all ninety and made the ingest at least three model calls cheaper. The deletion took less time than the simulation did.
The general lesson is not "LLM judges are bad". It is that the unit of evaluation is the stage, not the prompt. A prompt can only be as good as the stage's position in the pipeline allows, and when the stage sits on top of a gate that already does the work, the ceiling is low regardless of what the prompt says.
The fixtures had said it was fine
The stage had been certified against fixtures before it shipped, which is how it got there. Two of them, three rows and seven rows. Both passed. Both were too small and too easy to expose anything.
Worse, the evaluation setup itself changed the verdict. Judging candidates in a batch and judging them one at a time produced an identical kill rate — 25 of 25 either way — so by the metric we were watching, the setup did not matter. But judged alone, the same code false-rejected two of three items a human had already accepted. The kill rate was stable while the behaviour underneath it was not, and a stable number is exactly the thing that stops people looking further.
If you take one practice from this, take that one: batch composition is part of your judge's input, whether you intended it to be or not. An item that survives among nine others may not survive alone.
What the numbers do not show
Four honest limits.
- One pipeline, one domain, one period. Ninety rows from our own ingest. Nothing here says your judge has no lift; it says you probably have not measured whether it does.
- 96.7 percent is a statement about the cull, not about accuracy. It describes the rows the stage hid after the gate passed them. It does not describe the stage's behaviour across everything it saw, and reading it as an overall error rate would be wrong.
- The lift range is a range for a reason. 0.71–0.94× is bounded below 1.0 across the interval, which is what made the decision straightforward, but this was one comparison against one gate and not a general result about judges.
- We did not prove that no prompt could have worked. We proved that the retune we had planned would have recovered five rows, and that the stage had no measured lift to build on. Those are the grounds we acted on. A different architecture — the model call replacing the gate rather than sitting above it — was never tested.
What to measure before you add a judge
- The baseline you are improving on. If there is already a threshold, a score or a heuristic, its performance is the number your judge has to beat. Without it you will measure the judge against nothing and conclude it works.
- Lift, not accuracy. Accuracy on kept items is comfortable and nearly uninformative. Lift over the stage beneath it is the number that decides whether the stage should exist.
- What it discarded. Sample the rejects and have a person rate them. The failure mode is invisible from the keep side, which is why almost nobody catches it.
- The same items in isolation and in a batch. If the verdicts differ, your evaluation harness is part of the system and has to be described in the results.
- A fixture large enough to fail. Three rows cannot disagree with you. Ours did not, for weeks.
Getting this wrong is cheap to fix in a pipeline and expensive to fix in a codebase, where the same reflex — add a model call, tune it when it underperforms — produces stages nobody has measured and changes nobody can account for. We report our own deletions for that reason: the record of what we removed and why is part of what Backthread keeps about a system, and this stage is in ours with the numbers attached.
Connect one repo and the same kind of record starts accumulating for your own changes, including the ones you reverse; the trial runs fourteen days.
In short
- Measure a judge's lift over the gate it sits on, not its accuracy
- Accuracy on the items a filter kept says almost nothing about whether the filter earns its place. Ours measured 0.71–0.94 times the salience gate beneath it, which put the stage at or below what it was meant to improve.
- A simulated retune recovered five rows out of ninety
- Rewriting the rubric was the obvious next move, so we simulated it against the same sample before changing any code. Deleting the stage recovered all ninety rows and made each ingest at least three model calls cheaper.
- Sample what the filter threw away, not what it kept
- Of 90 rows the judge culled after the gate had already passed them, 96.7 percent were wrongly hidden, with a confidence interval of 90.7 to 98.9. That failure mode is invisible from the keep side.
- Batch composition is part of the judge's input
- Judging items in a batch and one at a time gave an identical kill rate of 25 of 25, while judged alone the same code false-rejected two of three items a human had accepted. A stable headline number hid completely different behaviour.
Sources
- Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias — arXiv, June 2026
- Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge — arXiv
- Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering — arXiv, April 2026
Backthread shows how much of what your agents built your team really understands. See how it works