The comparison table I stopped trusting

Before the actual argument, one caution, because it changes how I read every “pattern X beats pattern Y” claim in the agent world.

Configuration noise alone can move a benchmark score by up to 6 percentage points (p<0.01). Differences below 3 points deserve skepticism. Most harness comparisons I see in circulation sit inside that band. So before you adopt a pattern because some comparison favoured it, check whether the gap is bigger than what the infrastructure setup produces by itself. Usually it isn’t.

Once you do that, most of the framework debate goes quiet. What’s left is a much smaller set of things that actually discriminate. I think there’s really one axis, and one ceiling.


The axis: who decides the next step?

“Harness engineering” became a named thing in early 2026, with the working split agent = model + harness. People also separate an inner harness (what wraps one model call) from an outer harness (what wraps a whole run).

But the question that matters is simpler. Who drives the loop?

  • Code-driven: external code decides the next step. Graph and workflow frameworks, durable-execution platforms.
  • Model-driven: the model decides the next step. Every managed agent harness I found that shipped in 2026 is here.

In a code-driven loop the engine chooses the step, the agent produces one model call’s worth of output and may suggest a tool intent, and then the engine dispatches. The agent never dispatches anything itself.

People read this as conservative. I read it as a security property. The argument from the design-patterns literature goes like this: once an LLM agent has ingested untrusted input, it must be constrained so that the input cannot trigger consequential actions. A free-running observe-decide-act loop fails that, because every page the agent reads influences its next action. That’s your injection path.

Free-running loop vs action-selector. In the second, the model picks among predefined actions and nothing flows back into action selection.

Free-running loop vs action-selector. In the second, the model picks among predefined actions and nothing flows back into action selection.

The alternative is the action-selector pattern: the model picks among predefined actions, and there’s no feedback from those actions into the selection of the next one. Think of it as a program that accepts values without accepting actions. What I find funny is that the planning literature reached the same design independently, from the other direction.

So if someone later asks you to “make the agent more flexible”, write this down somewhere they will read it. More flexibility in who picks the step is exactly what you’d be trading away.


The ceiling: verifier fidelity

Now the part I think matters most. How much autonomy can you defend? Not “how smart is the model”. It’s how good is the thing that checks the model.

Three separate lines of evidence land in the same place:

  • The parallel-agent success everybody cites (sixteen agents building a very large codebase at a high test pass rate, modest budget) became possible once a known-good external oracle existed. The oracle was the enabler, not the parallelism. With it, all sixteen agents could be checked.
  • Peer-reviewed work (ICLR 2024) says intrinsic self-correction fails without external feedback.
  • Course lecture material I’ve been reading reports that more “overthinking” predicted less issue resolution across model types, with named patterns like analysis paralysis, rogue action chains and premature disengagement.

Add to that: reinforcement learning needs a checkable reward, and in domains like diagnosing incidents there often isn’t one. Public benchmarks in that area also don’t reward an agent for declining to answer. Different directions, same conclusion.

My take: verifier availability, not model capability, bounds defensible autonomy. And that’s a better argument for a constrained loop than any appeal to caution. “We’re scared” loses meetings. “We have no oracle, so a freer loop produces unchecked output faster” wins them.

If your domain has no automated oracle, that’s the number you should be working on, before you touch the loop.


One early error poisons everything after it

This is the failure mode I’d worry about most in any multi-step investigation, and it’s been measured.

A long-horizon business-simulation benchmark runs up to 2,000 messages, around 25M tokens, up to 222 simulated days. In it, one premature inference cascaded:

One premature inference, then a correct environment response, then a wrong conclusion that persists.

One premature inference, then a correct environment response, then a wrong conclusion that persists.

Look at step 2 again. The environment responded correctly. Nothing in the world lied to the agent. It just never waited for the fulfilment confirmation and never re-checked, so every later decision was conditioned on the earlier mistake.

The scaling result cuts both ways, and this is the bit that annoys people:

  • Higher per-step accuracy extends how long a horizon you can run. Helps.
  • Conditioning on your own earlier output makes later steps less accurate. Hurts.
  • Larger models self-condition more. Hurts more.

So a bigger model is not a fix for this failure mode. And it wasn’t a context-length failure either, so a longer context doesn’t fix it. A cross-step consistency check is worth more than extra window.

What I’d actually do with this:

ImplicationWhy
Keep an unverified inference labelled as an inferenceRestating it must never promote it to an observation, including on the agent’s own prior turns
Preserve “no signal found” vs “could not read”They look the same in a summary and mean opposite things later
Re-read the authoritative source instead of carrying a conclusion forwardSame instinct as reading config from the source, not from memory

Global constraints resist decomposition

Decomposition is the reflex answer to long tasks. It works until the constraint you care about isn’t local.

An ICML 2024 planning result, which I met via the same lecture material, puts it simply: flights, hotels, food and attractions decompose cleanly. Budget, transport and diet are global, and they get violated across the pieces.

Map that to an investigation agent. Checking metrics, checking topology, reading logs, scoring each rubric dimension: those decompose fine. What doesn’t:

  • the token and tool budget
  • the scope of the question (service, environment, time window)
  • “sufficient against the goal, not against complete knowledge”

Scope is the dangerous one. A sub-step reads outside the declared time window and produces an evidence item that is individually valid and collectively wrong. The sub-step cannot catch this, because from where it sits nothing is wrong. Only a check at the level above, a freshness rule or a per-step scope re-check, will catch it.

I don’t think there’s a clever trick here. You re-check scope per sub-step, not only at the end.


Do extra agents help at all?

The public discourse looks contradictory. One framework vendor said “don’t build multi-agents” in June 2025 and then published “what’s actually working” in April 2026, converging with another vendor on the same model. I think the underlying position is consistent:

Extra agents contribute intelligence, not actions. One writer, many readers.

Sub-agents gather, read and summarise. A single engine writes state. If you ever introduce sub-agents, enforce that structurally (a sub-agent has no write capability at all), not by convention.

Now the evidence on whether it helps:

  • A structured failure taxonomy over 1,600+ traces across 7 frameworks (NeurIPS 2025, inter-rater agreement κ=0.88, 14 failure modes) concludes multi-agent benchmark gains are “often minimal”.
  • Fixed-role decompositions have had limited success in coding, per the lecture material.
  • Where it does help, a stronger manager improves coordination and a weaker worker is lifted more than a strong one. Both point at manager quality as the lever.

The cost finding is the one I’d frame and hang on a wall. One vendor reports a +90.2% quality gain for multi-agent, at roughly 15x the tokens of a single chat. The same report says token usage alone explains 80% of the variance in quality. That’s a vendor measuring its own product, so directional only. But if it’s even roughly right, many “multi-agent wins” are compute wins. A result at 15x tokens isn’t evidence about architecture. Demand a comparison at matched budget.

Why fixed roles underperform is structural, I think:

  1. Roles are fixed before the task arrives. A “verifier” role has to localise the fault to check an answer, but localising is the gatherer’s job, and the verifier can’t do it.
  2. Handoffs are summaries. Agent B gets what survived summarisation. The dropped part is often exactly what B needed, and B has no way to know it’s missing.

So if a handoff carries only a summary, give it references as well, so the receiver can re-read what got dropped. And no sub-agent’s output counts as evidence. A sub-agent is a model, and model output alone isn’t evidence.


The question that comes first

Back to the sixteen agents. What people take from it is “parallelism works”. What actually enabled it was an oracle. Without a verifier, N agents produce N unchecked answers, and the human reviewer becomes the bottleneck N times over.

So the honest sequencing: the verifier question comes before the architecture question. A multi-agent proposal made before an instrument exists is a proposal to multiply unmeasured output.

Where I’m less sure: whether trained reasoning will absorb some of these scaffolds, so that parts of this become model capability rather than harness code. If it does, some of what I’ve written is a temporary patch. I don’t know. There’s also no general validator for a language plan or a sub-goal graph, so plan quality is only observable through outcomes. That’s an argument for measuring whether the evidence actually closes, rather than inspecting plans.

Short version: pick who drives the loop on purpose, put the verifier first, re-check global constraints per step, and don’t buy more agents until you’ve compared at matched budget.