Two Errors, Opposite Directions

Everyone who builds an agent that investigates and diagnoses things eventually adds a way for it to say “I can’t determine this.” And then everyone measures accuracy and calls it a day. I think that is wrong, and the reason is simple.

There are two ways to get abstention wrong, and they point in opposite directions:

ConcludedAbstained
Evidence was sufficientcorrectover-abstention: the human still does all the work
Evidence was insufficientunder-abstention: a confident wrong answer, trust lostcorrect

Under-abstention is the one everybody fears. Over-abstention is the one nobody measures, and it’s the one that quietly makes the product useless. A system that abstains on everything has zero under-abstention errors and is worth nothing.

So you always report two numbers together: participation (how often it answers at all) and conditional correctness (accuracy on the answered subset). Either one alone is meaningless. A single accuracy figure is basically banned in my book.

There’s a very common, very avoidable source of over-abstention too. The agent sees a question it can’t fully answer, like “why did the connection pool end up configured that way, who set it, and when”, and it abstains. But the investigation goal was “why did latency rise”, and that was answered. Sufficiency has to be judged against the goal, not against complete knowledge of the trigger. Otherwise you get an agent that is technically honest and practically useless.


Abstention Is Risk Control, Not a Threshold

Here’s the reframing that made the rest of my notes click. Don’t think “pick a cut-off on a confidence score.” Think risk control: you want to bound the expected loss on the answers you do give, where the loss is monotone in how permissive you are. That’s the conformal risk control setup (arXiv 2208.02814; for several knobs at once, arXiv 2110.01052).

The consequence is something you should tell a stakeholder before they ask. There is an arithmetic floor on abstention.

Say the base error rate (the agent answers everything) is $\mu$, and you’re asked to hold the error among answered cases below $\alpha$. With a binary loss, even a perfect ranker that abstains only on its own errors has to decline at least

\begin{equation} a ;\ge; \frac{\mu - \alpha}{1 - \alpha} \end{equation}

of the cases. (The general form replaces $1$ with the maximum loss $M$.) Worked example: base error 25%, requested ceiling 5%:

$$\frac{0.25 - 0.05}{1 - 0.05} \approx 0.21$$

At least 21% of cases must be declined. No amount of engineering removes that. It’s arithmetic, not a product choice. And note the direction: the tighter the ceiling you promise, the lower participation is capped, so “how much under-abstention can the business accept?” is a business question that sets a minimum abstention rate. Better to say this up front than discover it during a pilot.


Every Convenient Score Is Unsafe

So you need a score to threshold on. The ones lying around are all bad:

  • Token log-probabilities. A provider’s own words, as quoted in my notes: post-training hurts calibration significantly. Not a score you can build a guarantee on.
  • Model-verbalised confidence. Overconfident, and near-discrete. You get 0.7, 0.9, 0.95 and nothing in between. There’s no smooth curve to cut.
  • Internal probes. They don’t transfer. Out-of-distribution performance approaches random.
  • Semantic entropy over samples. This one is subtle and I like it least for diagnosis. It catches arbitrary errors, the ones where samples disagree. It does not catch consistently wrong ones. Four samples all say “cause = pool exhaustion”, entropy is low, and all four are wrong. A confidently repeated wrong causal story is exactly the failure you care about, so this whole family is disqualified.
  • Reasoning fine-tuning as a fix. One benchmark measured it degrading abstention by about 24%, and scale did not help.

What’s left? A score the runtime computes, not the model. In my notes there are seven factors that confidence should reflect, and all seven can be read from metadata about the evidence without asking the model anything:

FactorComputed as
Evidence coveragerequired evidence present / required
Source reliabilityper the source qualification policy
Consistencyunresolved contradiction count
Contradicting evidenceper the counter-evidence rule
Freshnesssource timestamp vs decision time
Baselineis a pre-incident baseline available
Missing evidencecount of declared gaps

Zero model calls. Boring. Auditable. I’ll take boring.


Marginal Coverage Is Not a Safety Guarantee

This is the trap that gets sold to risk and compliance people as a guarantee.

Say your method is calibrated for 90% coverage. Now measure per group (customer, tenant, whatever your stratum is). Illustrative shape: four groups sit around 0.88 to 0.95, and one group sits near zero. The weighted average still looks fine. The guarantee is fully satisfied while one group receives no protection at all.

That’s not an implementation bug, it’s a formal result. Per-instance conditional validity is unachievable distribution-free (Machine Learning 92:349-376, 2012). What is achievable is per-group validity: calibrate separately per stratum (arXiv 2107.07511). One preprint measured on imbalanced data reports per-group calibration restoring minority-class coverage, an average improvement of 61.7 percentage points (arXiv 2607.27143).

And the price is sample size. The notes put the need at roughly 1,000 to 2,500 labelled cases per stratum for a per-group guarantee. If you have a couple dozen labelled cases in total, a statistical guarantee is simply not available yet, and promising one is the worst thing you could do. The right posture is: ship the rule gate now, accumulate measurements, and claim a guarantee only for strata that have the samples.

One more warning, because people forget it. Abstention can magnify group disparities. A uniform threshold over groups with different base rates declines more often for the harder group. That’s a fairness property, not only a quality one.


The Three-Stage Design

Rule gate always on; the statistical layers only cover the grey zone

Rule gate always on; the statistical layers only cover the grey zone

Stage 1, the rule gate. A handful of mandatory abstention conditions, deterministic, no parameters, no calibration data. The ones in my notes: the evidence needed for the claim isn’t qualified; the conclusion would rest on temporal coincidence with no mechanism; a contradiction between relevant items remains after collection has ended (name both sides); the evidence needed can’t be obtained; the hypotheses can’t be told apart. The first three are machine-checkable. The last two are judgement, and that’s your grey zone.

Pay attention to the timing in the contradiction rule. If a contradiction shows up while you can still collect more, you request the separating evidence. Abstaining there is over-abstention. Same contradiction after collection ends, abstain. One contradiction, two correct behaviours, depending purely on timing.

Strengths: explainable to an auditor, needs no calibration data, no parameter to drift. Weakness, stated plainly: it doesn’t cover the grey zone. A case that triggers nothing but rests on thin evidence will get proposed, and the implicit threshold is whatever the rule author picked, with no mechanism to detect that it was picked wrong.

Stage 2, record two signals, decide nothing. The seven-factor score, and the margin between the leading hypothesis and the best alternative. What stage 2 shows you decides stage 3: if the score separates right from wrong with enough samples, threshold it with conformal risk control. If it separates but your hand-written formula is weak, learn the weights. If the score doesn’t separate but the margin does, the information lives in hypothesis separation. If neither separates, you’re recording the wrong quantities, go back.

Margin deserves a caveat I found genuinely useful: read both margin and absolute score. Two hypotheses that tie high have a small margin, so abstain, they can’t be told apart. But if everything scores low and one is merely less bad, the margin is large and the answer is still not trustworthy. A system reading only the margin passes the first case and fails the second.

Also, and this is easy to miss: you can’t compute a margin if you only store the winning hypothesis. If your system keeps one hypothesis and a prose assessment, there’s nowhere to record alternatives considered, evidence per alternative, or why each was eliminated. That blocks the margin signal, blocks re-auditing the “can’t tell apart” condition after the run, and blocks telling “no evidence found in the checked scope” from “ruled out”. I think that storage change is the highest-leverage fix here, and its value doesn’t depend on any model branch.

Stage 3, thresholds, per stratum only. Enable them where there are enough samples, and report the guarantee per stratum, never pooled.

The layering is permanent. A statistical method can cover the grey zone. It can’t be trusted to re-derive a hard rule. The rule gate never retires.


The Oracle Is an Artifact

One result I keep thinking about. A preprint (arXiv 2609.25938) reports that for the same system on the same data, swapping the correctness oracle moved the certified risk by 2.73 to 10.23 points. And certificates issued at a nominal 10% risk carried 17 to 20 points of real risk under expert labels.

So your grading rubric version isn’t documentation. It’s part of the guarantee. A certificate issued under one rubric version doesn’t hold under another.


What to Report, and the Missing Test Cases

The risk-coverage curve is the report, not one point on it. Plot conditional correctness against participation, and always plot the full-coverage baseline (what accuracy would be if the system never abstained). Without the baseline, any curve looks like progress. Alongside it:

  • worst-stratum selective risk, because the marginal number hides the group problem
  • abstention precision and recall
  • over-abstention rate, the error nobody measures
  • useful-abstention rate: did it name the gap and a feasible next step

A normalised area-under-risk-coverage exists for comparing across case sets, subtracting the in-hindsight-optimal ranking so the result is unitless in $[0,1]$. Raw curves aren’t comparable across case sets, so that’s worth having.

Last point. In the public benchmarks surveyed in my notes, “knowing when to stop” is basically not rewarded: cases are constructed so an answer always exists, and an agent saying “cannot determine” scores zero. So build the case class yourself. Three cheap constructions, all with the correct answer “abstain”: truncate the observation window so the decisive signal falls outside it; remove the one metric that separates the two hypotheses; inject the fault outside the monitored perimeter so only the symptom is visible.

And keep a test that punishes over-abstention: evidence sufficient for the goal, one out-of-scope question unanswered, correct behaviour is propose. Without it, an abstain-everything system passes every other case.

I’m honestly not sure how far the per-stratum guarantee story can go with the labelled data most teams have. My guess is most of us will live in stages 1 and 2 for a long time, and that’s fine, as long as we say so.


References

  • Conformal risk control: arXiv 2208.02814; multiple-knob extension: arXiv 2110.01052
  • Distribution-free conditional validity limits: Machine Learning 92:349-376 (2012), arXiv 1209.2673
  • Per-group calibration: arXiv 2107.07511; minority-class coverage measurement: arXiv 2607.27143
  • Abstention floor: arXiv 2606.29054
  • Oracle sensitivity of certified risk: arXiv 2609.25938