The method is not the problem
I’ve been reading a pile of papers on agents that improve from their own operation. Experience memory, evolving playbooks, prompt optimisers, selection policies, skill generation, weight updates. Different names, same shape: run, collect feedback, update something, run again.
And my honest take after all of it: the method matters much less than people think. Every one of these is an optimiser. An optimiser follows its signal, including into the wrong place. The ACE authors say it in their own limitations section: without ground-truth labels or reliable execution signals, the context can be “polluted by spurious or misleading signals”.
Here’s the part that makes it dangerous instead of merely useless. A lesson learned from noise gets stored in exactly the same shape as a lesson learned from evidence. Same schema, same id, same version, same confident wording. Nothing downstream can tell them apart. So the separation has to happen at the signal, before the write.
Most of this post is that idea unpacked: what goes wrong, what signals can be trusted, and in what order to put the gates.
Six ways a loop goes wrong
| Risk | What happens |
|---|---|
| Misevolution | Self-improvement drifts in unintended directions. Safety alignment can degrade after memory accumulates, even on top-tier models |
| Context collapse | Repeated rewriting of an accumulated context erodes its detail |
| Brevity bias | Summarisation drops domain insight for short, generic rules |
| Learning the noise | A weak reflector or noisy signal turns spurious patterns into stored experience |
| Over-retrieval | More retrieved experience makes results worse past a peak |
| Leakage | A case used to author an asset is later used to measure it |
A few comments, because the table hides the interesting bits.
Misevolution comes from a paper called “Your Agent May Misevolve” (ICLR 2026). It describes four pathways for drift: model, memory, tool, workflow. Three of the four need no training run at all. So “we’re not training anything” is not a safety argument. It removes one pathway out of four, and not the one that was actually measured. I hear that sentence a lot and it bugs me every time.
Context collapse and brevity bias are the quiet ones. No single rewrite looks wrong. The loss only shows across versions, which is why the fix is structural: apply identified deltas (add bullet 41, remove bullet 12) instead of regenerating the whole context. Each change then has an id, an author, a reason, and can be revoked alone. A reviewer’s judgement is not going to catch gradual erosion.
Over-retrieval was observed in two independent studies: results go up, peak, then decline as you retrieve more. So a retrieval cap isn’t a cost control that happens to help quality. It is a quality control. And a library that only grows walks rightward along that curve by itself, without anyone deciding to.
Leakage is the boring classic. The lesson “works” on the case it was written from. Exclude by recorded source reference, not by anyone’s memory of which case produced which lesson.
One more sits under all six: reviewer throughput. A loop that generates candidates faster than people can review them does not get reviewed. It gets approved.
Which signals can teach, and which cannot
The sources I read support five kinds of signal.
| Kind | Example | Weakness |
|---|---|---|
| Checkable | Execution success, a known injected cause, a typed refusal code | Only exists where an oracle exists, and for open-ended diagnosis it usually doesn’t |
| Paired | Same case run with and without an asset, or by two agents | Says one branch differed, not which was right |
| Indirect | A reviewer accepting or rejecting a conclusion | Acceptance is not correctness |
| Proxy | The agent declaring which item it used | Self-reported |
| Self-judged | The model grading its own trajectory | A judge no more accurate than what it grades |
Notice what’s strong and what’s free. The two strongest kinds, a recorded confirmation at case closure and a paired re-run, both have to be built. Everything that arrives for free is weak.
What each kind may change
This is the table I’d put on a wall.
| Kind | May raise an item | May lower an item | May label an eval case |
|---|---|---|---|
| Checkable | yes | yes | yes |
| Paired | candidate only | candidate only | no, unless one side is checkable |
| Indirect | no | yes | no |
| Proxy | no | no, report only | no |
| Self-judged | candidate only | candidate only | no, never arbitrates its own output |
Two rules carry most of the weight.
An indirect signal only moves things downward. A reviewer accepting a conclusion is a reason to look again, never a reason to trust more. And the failure mode isn’t a careless reviewer. It’s a reviewer past capacity, whose approval rate stays high precisely because they stopped reading. Raising on that signal rewards the overload. (Same reason a rising approval rate is not evidence the library improved. It’s equally consistent with overload.)
The agent’s own output never arbitrates the instrument that grades it. Self-judged loops are tempting because they need no labels, but they can’t be allowed to label eval cases.
Admission is a bet settled later
Saving a lesson has no reward at the moment it’s saved. Whether it was worth saving only shows on a later task. The best admission signal in the literature I read is successful later reuse, and you can record it with no training. So the store has to carry the bet, a per-item helpful/harmful counter (ACE does this per bullet), instead of a one-off decision at write time that is never revisited.
No confirmation field, no growth
This one is the practical takeaway. A cheap durable oracle is a retrospective confirmation field on closed cases. Without it, operation only produces indirect signals, and indirect signals may only lower. So a loop with no recorded confirmation can only shrink. The library can be pruned but not grown on evidence. That’s not a broken design. It’s a correct one running without its input.
What each kind of update is allowed to touch
The families of method differ by what they write to, and that decides cost and how long a fault takes to undo.
- Experience memory: a store of lessons, weights frozen. Undone by revoking one item.
- Evolving playbook: itemised bullets updated by delta. Undone by deleting a bullet.
- Instruction optimisation: system prompt, tool descriptions. Undone by rolling back a prompt version.
- Selection policy: what enters the context, not its content. Undone by rolling back and re-measuring.
- Skill or procedure generation: undone by retiring the skill.
- Weights: undone by a new model version, meaning a training run plus a re-release.
Five of the six sit below the training line. They change what the model sees. Only weights change the model. That single difference is basically the whole cost and risk difference between them, and it’s the argument for exhausting prompts and memory before reaching for training, per task. Prompt-level reflective optimisation has been reported as far cheaper than RL (GEPA, a preprint).
Evidence caveat, since I’m pointing at numbers: the gains reported for these methods are self-reported by authors on their own benchmarks, and the sources I read do not record the sample size behind them. Several are single digits. Treat them as directional. One example of how a figure gets stretched: a Dynamic Cheatsheet result of about 10% to 99% on Game of 24 is one task, and the mechanism was discovering a reusable Python solution, not better reasoning. Quote it with the task named or not at all.
Ordering the gates
Each gate stops something specific:
| Gate | Stops |
|---|---|
| Candidate, not write | Unreviewed experience reaching the context |
| Approver is not the author | An author confirming their own error |
| Shape gate before content review | Correct content with loose conditions, the dangerous case |
| One change at a time, frozen baseline | Attributing an effect to the wrong change |
| Shadow before serve | An item matching far more cases than it should |
| Pre-registered target and thresholds | Moving the goalposts after seeing results |
| Negative suite on every batch | Gains that hide a regression in safety |
| Delta updates with identity | Context collapse, untraceable changes |
| Retrieval cap, coverage selection, retirement | Over-retrieval, monotonic growth |
| Candidate cap per run | Reviewer overload, at the cost of some lessons lost |
Mapped against the six risks, no single gate covers two columns well. Each risk is held by a different gate, so dropping one opens a specific hole rather than a general one.
The order matters as much as the set.

The gated loop: a typed signal produces a candidate, which passes every gate before any update is released.
The shape gate is mechanical and cheap, so it goes before the expensive human review. Review comes before shadow, shadow before measurement, and the negative suite runs on every batch before release. Release is pinned per run and revocable.
And gate zero is the typed signal. A loop that skips it can’t be made safe by everything after, because every later gate is checking an update whose evidence nobody can weigh.
Where the evidence is weak
Honestly, a lot.
The one study that tested mitigations directly found them partial, and didn’t break down which failures survived. So I can’t tell you the shape gate catches misevolution rather than just sloppiness.
Nobody publishes a rejection rate. Every paper in this area reports the gain from updates that were kept; none reports how many candidates died at each stage. So the cost of running these gates can’t be estimated from the literature. It has to be measured in-house from the first batch, or the gates get dropped later because nobody can say what they’re worth. If you take one operational thing from this post, record the rejection count per batch.
No source measures a loop running with no gate at all, so “never unattended” is a position taken on principle, not on data. I hold it anyway. Given that one of the risks is safety degrading quietly as memory accumulates, I don’t want to be the one running the experiment.
Also: most of the measured work is on tasks with a checkable end state. Open-ended diagnosis, where a checkable signal usually doesn’t exist, is barely covered. That’s exactly where the “indirect signals only lower” rule bites hardest.
References
- “Your Agent May Misevolve” (ICLR 2026)
- ACE (Generator, Reflector, Curator), arXiv preprint
- GEPA, arXiv preprint (cite the version)
- Dynamic Cheatsheet (EACL 2026)
- Memento, ReasoningBank, ExpeL, Reflexion, Agent Workflow Memory