The method is not the problem

I’ve been reading a pile of papers on agents that improve from their own operation. Experience memory, evolving playbooks, prompt optimisers, selection policies, skill generation, weight updates. Different names, same shape: run, collect feedback, update something, run again.

And my honest take after all of it: the method matters much less than people think. Every one of these is an optimiser. An optimiser follows its signal, including into the wrong place. The ACE authors say it in their own limitations section: without ground-truth labels or reliable execution signals, the context can be “polluted by spurious or misleading signals”.

Here’s the part that makes it dangerous instead of merely useless. A lesson learned from noise gets stored in exactly the same shape as a lesson learned from evidence. Same schema, same id, same version, same confident wording. Nothing downstream can tell them apart. So the separation has to happen at the signal, before the write.

Most of this post is that idea unpacked: what goes wrong, what signals can be trusted, and in what order to put the gates.


Six ways a loop goes wrong

RiskWhat happens
MisevolutionSelf-improvement drifts in unintended directions. Safety alignment can degrade after memory accumulates, even on top-tier models
Context collapseRepeated rewriting of an accumulated context erodes its detail
Brevity biasSummarisation drops domain insight for short, generic rules
Learning the noiseA weak reflector or noisy signal turns spurious patterns into stored experience
Over-retrievalMore retrieved experience makes results worse past a peak
LeakageA case used to author an asset is later used to measure it

A few comments, because the table hides the interesting bits.

Misevolution comes from a paper called “Your Agent May Misevolve” (ICLR 2026). It describes four pathways for drift: model, memory, tool, workflow. Three of the four need no training run at all. So “we’re not training anything” is not a safety argument. It removes one pathway out of four, and not the one that was actually measured. I hear that sentence a lot and it bugs me every time.

Context collapse and brevity bias are the quiet ones. No single rewrite looks wrong. The loss only shows across versions, which is why the fix is structural: apply identified deltas (add bullet 41, remove bullet 12) instead of regenerating the whole context. Each change then has an id, an author, a reason, and can be revoked alone. A reviewer’s judgement is not going to catch gradual erosion.

Over-retrieval was observed in two independent studies: results go up, peak, then decline as you retrieve more. So a retrieval cap isn’t a cost control that happens to help quality. It is a quality control. And a library that only grows walks rightward along that curve by itself, without anyone deciding to.

Leakage is the boring classic. The lesson “works” on the case it was written from. Exclude by recorded source reference, not by anyone’s memory of which case produced which lesson.

One more sits under all six: reviewer throughput. A loop that generates candidates faster than people can review them does not get reviewed. It gets approved.


Which signals can teach, and which cannot

The sources I read support five kinds of signal.

KindExampleWeakness
CheckableExecution success, a known injected cause, a typed refusal codeOnly exists where an oracle exists, and for open-ended diagnosis it usually doesn’t
PairedSame case run with and without an asset, or by two agentsSays one branch differed, not which was right
IndirectA reviewer accepting or rejecting a conclusionAcceptance is not correctness
ProxyThe agent declaring which item it usedSelf-reported
Self-judgedThe model grading its own trajectoryA judge no more accurate than what it grades

Notice what’s strong and what’s free. The two strongest kinds, a recorded confirmation at case closure and a paired re-run, both have to be built. Everything that arrives for free is weak.

What each kind may change

This is the table I’d put on a wall.

KindMay raise an itemMay lower an itemMay label an eval case
Checkableyesyesyes
Pairedcandidate onlycandidate onlyno, unless one side is checkable
Indirectnoyesno
Proxynono, report onlyno
Self-judgedcandidate onlycandidate onlyno, never arbitrates its own output

Two rules carry most of the weight.

An indirect signal only moves things downward. A reviewer accepting a conclusion is a reason to look again, never a reason to trust more. And the failure mode isn’t a careless reviewer. It’s a reviewer past capacity, whose approval rate stays high precisely because they stopped reading. Raising on that signal rewards the overload. (Same reason a rising approval rate is not evidence the library improved. It’s equally consistent with overload.)

The agent’s own output never arbitrates the instrument that grades it. Self-judged loops are tempting because they need no labels, but they can’t be allowed to label eval cases.

Admission is a bet settled later

Saving a lesson has no reward at the moment it’s saved. Whether it was worth saving only shows on a later task. The best admission signal in the literature I read is successful later reuse, and you can record it with no training. So the store has to carry the bet, a per-item helpful/harmful counter (ACE does this per bullet), instead of a one-off decision at write time that is never revisited.

No confirmation field, no growth

This one is the practical takeaway. A cheap durable oracle is a retrospective confirmation field on closed cases. Without it, operation only produces indirect signals, and indirect signals may only lower. So a loop with no recorded confirmation can only shrink. The library can be pruned but not grown on evidence. That’s not a broken design. It’s a correct one running without its input.


What each kind of update is allowed to touch

The families of method differ by what they write to, and that decides cost and how long a fault takes to undo.

  • Experience memory: a store of lessons, weights frozen. Undone by revoking one item.
  • Evolving playbook: itemised bullets updated by delta. Undone by deleting a bullet.
  • Instruction optimisation: system prompt, tool descriptions. Undone by rolling back a prompt version.
  • Selection policy: what enters the context, not its content. Undone by rolling back and re-measuring.
  • Skill or procedure generation: undone by retiring the skill.
  • Weights: undone by a new model version, meaning a training run plus a re-release.

Five of the six sit below the training line. They change what the model sees. Only weights change the model. That single difference is basically the whole cost and risk difference between them, and it’s the argument for exhausting prompts and memory before reaching for training, per task. Prompt-level reflective optimisation has been reported as far cheaper than RL (GEPA, a preprint).

Evidence caveat, since I’m pointing at numbers: the gains reported for these methods are self-reported by authors on their own benchmarks, and the sources I read do not record the sample size behind them. Several are single digits. Treat them as directional. One example of how a figure gets stretched: a Dynamic Cheatsheet result of about 10% to 99% on Game of 24 is one task, and the mechanism was discovering a reusable Python solution, not better reasoning. Quote it with the task named or not at all.


Ordering the gates

Each gate stops something specific:

GateStops
Candidate, not writeUnreviewed experience reaching the context
Approver is not the authorAn author confirming their own error
Shape gate before content reviewCorrect content with loose conditions, the dangerous case
One change at a time, frozen baselineAttributing an effect to the wrong change
Shadow before serveAn item matching far more cases than it should
Pre-registered target and thresholdsMoving the goalposts after seeing results
Negative suite on every batchGains that hide a regression in safety
Delta updates with identityContext collapse, untraceable changes
Retrieval cap, coverage selection, retirementOver-retrieval, monotonic growth
Candidate cap per runReviewer overload, at the cost of some lessons lost

Mapped against the six risks, no single gate covers two columns well. Each risk is held by a different gate, so dropping one opens a specific hole rather than a general one.

The order matters as much as the set.

The gated loop: a typed signal produces a candidate, which passes every gate before any update is released.

The gated loop: a typed signal produces a candidate, which passes every gate before any update is released.

The shape gate is mechanical and cheap, so it goes before the expensive human review. Review comes before shadow, shadow before measurement, and the negative suite runs on every batch before release. Release is pinned per run and revocable.

And gate zero is the typed signal. A loop that skips it can’t be made safe by everything after, because every later gate is checking an update whose evidence nobody can weigh.


Where the evidence is weak

Honestly, a lot.

The one study that tested mitigations directly found them partial, and didn’t break down which failures survived. So I can’t tell you the shape gate catches misevolution rather than just sloppiness.

Nobody publishes a rejection rate. Every paper in this area reports the gain from updates that were kept; none reports how many candidates died at each stage. So the cost of running these gates can’t be estimated from the literature. It has to be measured in-house from the first batch, or the gates get dropped later because nobody can say what they’re worth. If you take one operational thing from this post, record the rejection count per batch.

No source measures a loop running with no gate at all, so “never unattended” is a position taken on principle, not on data. I hold it anyway. Given that one of the risks is safety degrading quietly as memory accumulates, I don’t want to be the one running the experiment.

Also: most of the measured work is on tasks with a checkable end state. Open-ended diagnosis, where a checkable signal usually doesn’t exist, is barely covered. That’s exactly where the “indirect signals only lower” rule bites hardest.


References

  • “Your Agent May Misevolve” (ICLR 2026)
  • ACE (Generator, Reflector, Curator), arXiv preprint
  • GEPA, arXiv preprint (cite the version)
  • Dynamic Cheatsheet (EACL 2026)
  • Memento, ReasoningBank, ExpeL, Reflexion, Agent Workflow Memory