The Wrong Question

Every long-running agent hits the same wall. The history doesn’t fit anymore, so something has to give. And the default instinct is “summarise the old stuff”, read the summary, think “yeah that looks fine”, move on.

I think that framing is the actual bug. Compaction is not shortening a transcript. It’s producing a bounded working state from which the same future decisions follow. Write it as ŝ = C(H; B), with |ŝ| ≤ B, and the thing you want is that the distribution over future actions given ŝ is about the same as given the full history H.

(I picked this framing up from a CMU lecture on the topic, and honestly it changed how I look at the whole problem.)

So the test is not “does ŝ read well”. It’s “do the same decisions follow from it”. A beautiful summary that makes the agent forget a constraint is a failed compaction. An ugly one that keeps the constraint is a good one.

Three properties fall out of that:

PropertyWhat it requires
Behavioural fidelitySame next decisions follow from ŝ as from H
Bounded representationŝ fits a stated budget B
RecoverabilityExact anchors copied verbatim; pointers to source evidence kept

The third one is the one people skip, and it matters most when your agent produces claims someone has to check later. “The log confirmed the port collision” is a claim with no provenance, nothing to check it against. “Cause: port collision, evidence: /tmp/ci.log” is a claim with a pointer. Same information, completely different usefulness.


What Survives, and in What Form

Not everything gets the same treatment. Roughly five buckets:

  • Keep exact: anchors. The goal and the constraints. These are what later decisions get checked against, and paraphrase loses them.
  • Encode: the checkpoint. Decisions taken, progress, identifiers. These can be restated compactly without losing meaning.
  • Keep exact: the recent tail. The active attempt, the fresh tool result. The model is mid-action, and the tail is the action.
  • Externalise: bulk evidence. A 24,000-line log becomes 3 failure lines, the command that produced them, and the artifact path. The content is still reachable, it’s just not in the window.
  • Discard: duplicates and superseded attempts. No real information loss.

Almost all of the saving is in the externalise row. And it’s the same row that dominates the budget in the first place, the source puts tool results at 37% of a typical session. So if you’re going to be clever about one thing, be clever about tool output, not about conversation.


Repeated Compaction Drifts, Silently

This is the part that makes compaction dangerous rather than just lossy.

Each pass inherits the previous checkpoint, not the original history.

Each pass inherits the previous checkpoint, not the original history.

Each compaction inherits the previous checkpoint, not the original history. So a constraint degrades one notch at a time. “Use CUDA 12.4” becomes “Use CUDA 12.x” becomes “Use recent CUDA” becomes nothing. And no single step looks wrong. If you diff summary 1 against summary 2, you’d shrug. And nothing in the transcript records that a constraint was lost.

That last bit is what bugs me. Lossy is fine, lossy you can budget for. Lossy-and-no-tombstone means the agent confidently continues on a state that is missing something, and neither it nor you knows.

There are three mitigations and, as far as I can tell, you need all three:

  1. Copy anchors exactly instead of re-summarising them. Restated once, they’re at risk. Restated three times, gone.
  2. Keep the recent tail verbatim, so at least the active work isn’t a summary of a summary.
  3. Retrieve source evidence rather than carrying a description of it, so the authoritative version is always one hop away.

For an agent that investigates and diagnoses things, the anchors are easy to name: the question being investigated, the scope (service, environment, time window), the pinned release. They get copied, never re-expressed.


The Policy Has Four Stages

When you look at how different coding harnesses publish their compaction behaviour, they decompose roughly the same way.

Trigger, select, replace, recover. The last stage is the one usually left unspecified.

Trigger, select, replace, recover. The last stage is the one usually left unspecified.

Trigger: token threshold, provider overflow, manual request, and reserve room for the next output. Select: protect the beginning, compact the older region, keep the recent tail. Replace: readable summary, structured checkpoint, opaque provider item, or a fresh context. Recover: what happens when it goes wrong.

Two details bite in practice.

Cut only at a valid boundary. If you cut between an assistant message that calls a tool and the tool’s result, you end up with a tool request with no result. Some providers reject that outright, and most models handle it badly. Cut after the result, not inside the pair.

Stage 4 is usually blank. What happens when the compacting call itself errors, or returns garbage? “Continue with the uncompacted history and hope” is not a behaviour, the provider will refuse the request because it’s over budget. You want something defined: retry, reset, or fail the run with a typed error. Any of those. Just pick one.

On stage 3, I lean hard toward a structured checkpoint over free text, and not for quality reasons. It’s because the record has to be reproducible and checkable. “We investigated the payment service and found the pool was likely saturated” can’t be checked against the original. A checkpoint with an investigation question, a scope object, retained evidence ids, open hypotheses and known gaps can, because every field maps to something you’d have to produce anyway.


Evaluating a Policy

You can’t evaluate a compaction policy by reading its output. I keep coming back to this. You evaluate it by what happens after:

plant state (constraints, decisions, artifacts), run the production policy unmodified, let the model continue, then measure the decisions it makes. Not the summary. The decisions.

Four readings come out of that:

ReadingWhat it measures
State recallDo the constraints and identifiers survive?
Task successIs the continuation correct?
EfficiencyTokens, latency, cost
StabilityWhat happens under repeated compaction?

Stability only shows up if the test compacts more than once. A single-compaction test happily passes a policy that loses a constraint on the third pass. A three-compaction test catches it. The source says this is the reading most often omitted, and from what I’ve seen that sounds right. Most eval setups I’ve looked at compact once, because that’s the cheap thing to build.

So the pair I’d insist on: one constraint stated once, 200 messages back, survives one compaction (most policies pass that), then the same run compacted three times (many fail). A policy tested with one compaction will pass the first and fail the second, and you won’t know.


Compaction vs Evidence Integrity

This is a genuine tension, not a solved problem.

Compaction wants fewer tokens, a paraphrase that reads well, and old attempts discarded. Evidence integrity wants every cited item present and verifiable, exact values, units, populations and windows, and a record of what was tried and didn’t work. Those are opposite pulls.

The resolution I find convincing has three parts:

Compaction never touches the qualified evidence bundle. The compactor works on the conversation: prior reasoning, tool call/result pairs, superseded attempts. The evidence enters by reference with provenance intact, and the retrieval path selects and bounds it, not the compactor.

A discarded attempt becomes a known gap, not a deletion. “Checked the adapter log, 09:00 to 09:30, no permission” is three lines and load-bearing. It’s the difference between “no signal found” and “could not read”, which are entirely different states and shouldn’t be collapsed.

The measurement unit survives, or the measurement is dropped. A metric reading needs its component, metric, unit, aggregation, population and window. All six kept, the reading can be used. Any one lost, drop the whole reading, because a number without its population and window can’t legitimately appear in a conclusion. Same rule as a budget rule: drop a whole unit, never a partial one.


The Replay Problem

One open question I think is the sharpest, and I don’t have a clean answer.

A compactor that calls a model is non-deterministic by default. Same history, compact twice, you get two different states. So the compacted state is not a function of the history alone.

Then you have two options. Store the compacted state as part of the run record and replay from storage: replay is exact, at a storage cost. Or don’t store it, and replay is not exact. If your evaluation setup depends on exact replay as its comparison basis, the second option quietly breaks everything.

My take: store it. Storage is cheap, and an unreplayable run is a run you can’t debug. The source frames this as one of the two having to give and says nobody has decided which, so treat my preference as mine.

Whether compaction should happen inside a single capability invocation or only between them is related and also open. Between is far easier to reason about and audit. I’d start there.