The number in every slide deck

“LLM judges agree with humans over 80% of the time, same as humans agree with each other.” You have seen this sentence, probably more than once.

The figure is real. It’s also (1) computed with ties excluded, (2) not corrected for chance, (3) measured on a dataset where the human-human agreement is about 63% once ties are kept, and (4) conditional on a large quality gap between the two outputs being compared. Four asterisks on one sentence.

So here’s what happens to it when you do the correction properly, and what I think a team should actually do with a judge after that.


Raw agreement vs chance-corrected agreement

Two raters picking between A, B and tie will agree a lot just by luck, especially when the label distribution is lopsided. Cohen’s kappa removes that:

$$\kappa = \frac{p_o - p_e}{1 - p_e}$$

where $p_o$ is observed agreement and $p_e$ is the agreement you would expect from the label distributions alone.

A toy case (my numbers, not from any paper): binary labels, 50/50 split, 85% raw agreement. Then $p_e = 0.5$ and $\kappa = 0.70$. Fine, a bit of deflation. Now make the label set three-way and unbalanced and the same raw number collapses much harder.

That’s exactly what a large replication preprint reports: 21 judges, 9 providers, 3 benchmarks, 118 runs, around 541,000 judgments. One judge at 0.849 raw agreement sits at $\kappa = 0.511$. That is 33.8 percentage points gone. Other judges in the same study: 0.848 raw to 0.489, 0.836 to 0.457, 0.788 to 0.376. Deflation is 33.8 to 41.2 points across the board, and it tracks the label distribution: 38.6 points on a balanced A/B/tie set, only 10.4 on a binary set.

If you back out the chance term from the first judge, $p_e = (p_o - \kappa)/(1 - \kappa) \approx 0.69$. Two thirds of that “85% agreement” was available for free.

The same judge, before and after removing chance agreement

The same judge, before and after removing chance agreement

And this isn’t one preprint being grumpy. A peer-reviewed result (ACL 2025) across 20 NLP datasets with human annotations finds the best of 11 judges averages $\kappa = 0.28 \pm 0.32$, and kappa goes negative on safety tasks. So the honest range is about 0.28 to 0.51. On the usual reading scale that is “fair” to “moderate”. Not human-equivalent.

My take: any report of judge quality that doesn’t show a chance-corrected number should be treated as missing data.


When the judge does fine

There is a counter-result, and it matters. A model applying a fixed rubric agreed with experts 79% of the time while two experts agreed with each other 80%. Caveat, and it’s a big one: that was 12 questions, 2 systems, 2 raters per answer. Tiny. I’d call it suggestive.

But it fits the bigger picture. Free-form “which is better?” makes the judge supply the standard, and it doesn’t have one. Item-wise “does the answer contain X?” gives it the standard, it only applies it. The rubric carries the agreement and the judge inherits it. Which means the rubric’s own inter-rater agreement is the ceiling on what any judge can do, and the first number you should measure. It’s easy to skip straight to measuring the judge instead.


Test-retest is not reliability

The other number people quote is consistency: run the judge twice, see if it says the same thing. High is good, right?

A judge that always picks whatever is in position A is perfectly consistent and perfectly useless. The replication preprint measured exactly this paradox: test-retest above 0.95 coexisting with position bias above 0.10 in two production judges. For one of them: test-retest 0.992, position bias 0.192, benchmark kappa 0.289. And injecting bias raised the consistency score by 20 to 30 points. So a high consistency score can be evidence of a problem.

There’s also the other side. Another preprint ran 29 tasks, 50 repeated trials, 2 judges, and looked at how often a pairwise preference flips between identical runs. Average 13.6%. 28% of questions flipped more than 20% of the time. The worst question flipped 56%, basically a coin. Cross-judge agreement 76% ($\kappa = 0.51$, same number again, funny). Deterministic decoding “reduces but does not eliminate” it.

They suggest roughly 11 trials for a stable single verdict. A one-shot judge score is not a measurement. It’s a sample of size one.


Which biases still matter

Some of the classic ones have been tamed by better protocols. What I take from the source:

  • Verbosity: controlled under standard pairwise rubrics (below 0.011), but style attacks still succeed more than 65% of the time. The bias is handled, the attack isn’t.
  • Self-preference: largely gone once provenance labels are stripped. That’s a context hygiene fix, not a model fix.
  • Position: still material. Swap positions and re-run. There’s no other fix.
  • Prior-score anchoring: $d = 0.71$, a large effect. Showing the judge a previous score for the item moves its rating that much, on its own.

The practical residue is boring and cheap: strip tenant identifiers, model names, config names, and any prior score from the judge’s context. If you compare two configurations, the judge must not know which one produced which output. Shuffle, unlabel, no priors.


Panels, debate, and what to do with disagreement

This section surprised me most.

Debate between judges amplifies bias instead of correcting it. A panel of diverse-family small judges did better than one big judge on a QA task, and 7 to 8 times cheaper: kappa 0.763 vs 0.627 in one measurement. (Short QA only. I wouldn’t assume it transfers to long-form rubric grading, and neither does the source.)

Then majority vote. You’d think a panel that disagrees can at least be settled by counting. A preprint measured it:

ItemsAccuracyKappa
Panel agreed unanimously96.6%0.93
Panel disagreed, majority-voted69.5%0.39
Best single judge, same disputed items79.9%—
Staged strategy85.4%—

Majority vote on disputed items lands below the best single judge. The cases a panel disagrees on are the cases it is worst at, so the vote is weighted by the wrong thing.

A separate preprint finds the same shape when voting across rollouts of one model on long-horizon tasks: single rollout 47.2, majority vote 47.5, so +0.3 on a five-benchmark average and negative on 2 of 5. And one more uncomfortable line from that work: 34% of the values that all rollouts agreed on were wrong. Consensus is where errors hide.

So what do you do with disagreement? Not average it. Treat it as a map of which claims to go check against external evidence, or escalate to a human. The source reports that checking the contested claim against the environment gave roughly twice the gain of any voting scheme, and leaves a trail a reviewer can re-walk. I’d read that as directional. It’s one line of work, and the same source admits nobody has published the false-positive rate of such a verifier.

One more thing, a formal bound the source cites: if the judge is no more accurate than the system under test, no unbiased debiasing method cuts the human-label requirement by more than 2x. Prediction-powered inference gives tighter confidence intervals, it doesn’t get you out of annotation. Human labelling throughput stays the binding constraint.


Where a judge is legitimate

My short list, from the source:

Offline instrument vs closed-loop component

Offline instrument vs closed-loop component

The test is placement. Offline, output goes to a human, a wrong verdict costs an hour of reviewer attention. That is fine. A judge wrong 30% of the time still orders a queue usefully. Proposing a rubric change from failure patterns is fine because it is proposing, not deciding.

Inside the loop, gating an action, a wrong verdict reaches the user. Not fine. Also not fine: the judge producing the ground-truth label, or adjudicating a disagreement between reviewers (especially using the agent’s own output as the basis). Judging the lowest-agreement rubric dimension is “with caution” since judge noise compounds label noise.

One distinction I want to keep. “A second model call that declares the output sufficient” is the model judging itself with a different label. “A verifier that goes and looks at files, data, queries and emits per-claim findings with the evidence attached, no score” is a different animal. Roughly half of one 4.4-point selection gain came from just letting the verifier look instead of judging from memory.

Two loops that close by accident

Reviewer sees the judge’s prediction, then labels. Measured agreement goes up, quality doesn’t. The number improves, the instrument doesn’t. And if the prediction comes with cited evidence and nice formatting, reviewers defer more. Second loop: you tune the judge on labels that were produced after the labeller saw its predictions. Same drift, invisible. The cuts are simple: reviewers grade before seeing any prediction, and you never tune a judge on contaminated labels.


Prompt sensitivity

Same judge, same items, different prompt wording: kappa 0.518 vs 0.725. That’s +0.207 from phrasing alone. A single added instruction gave +0.10 on a strong model, and the same edit on a weaker model was severely negative.

Two consequences. The judge prompt is a versioned artifact, kept next to the rubric version. And tuning doesn’t transfer: change the judge model and the tuning is invalid, re-validate.


What I’d actually do

Measure rubric inter-rater agreement first. Report kappa next to raw agreement, always. Run enough trials per item to know the flip rate before trusting a verdict. Strip provenance and priors. Send disagreement to a human or to a check against evidence, not to a vote. And keep the judge offline.

None of this makes the judge useless. It makes it an instrument with a known, moderate accuracy, which is a perfectly good thing to have, as long as you stop quoting the 85%.


References

  • Large-scale judge reliability replication (21 judges, 118 runs), arXiv 2606.19544
  • Judging the judges across 20 NLP datasets, ACL 2025, arXiv 2406.18403
  • Pairwise preference instability over repeated trials, arXiv 2606.13685
  • Panel of diverse small judges and prompt-sensitivity results, arXiv 2404.18796
  • Majority voting on disputed judge items, arXiv 2603.25133
  • Majority vote across rollouts, arXiv 2610.00972