Memory is a filing problem first

I have been reading a lot about agent memory lately, and almost everything starts at retrieval: embeddings, rerankers, hybrid search. Fine. But when I try to imagine how a memory system hurts you in production, retrieval is rarely the first thing. It’s “the agent wrote something down that it should not have, and now it’s reading it back as truth.”

So these notes are about the write side and the lifecycle, and about one idea that I think carries most of the weight: the item has to carry the conditions under which it applies, in a form a machine can check.


Three kinds of things, and the mis-filing problem

An agent run produces three kinds of content you might keep:

KindExampleWorth keeping when
EpisodeThe full trace, every observation and actionThe exact sequence, values or entities will matter again
Fact“This source only exposes latency at the caller”It recurs across tasks and can be updated
SkillA reusable procedureLater tasks repeat the same pattern

Sounds trivial. But mis-filing is the hardest error to fix afterwards, because the bad item looks exactly like a good one. An unverified hypothesis (“possibly the connection pool”) gets filed as a fact. A few weeks later it is retrieved as “this service’s pool saturates under load”, a known property of the environment, with no trace that it was ever a guess. The rule that says “don’t turn an inference into an observation” usually gets applied at conclusion time. It has to apply at write time too.

Or a diagnostic procedure that worked gets stored as a memory item. A procedure needs a version, a test and a release, and a memory item has none of the three. That’s a skill, and it goes through the skill release path.

Some content is not memory at all: current transaction state, deployed configuration, current access permissions, current metric values. Memory goes stale, the source doesn’t. “The service has pool maximum 20” is configuration, go read it. “At pool maximum 20, symptom X accompanied cause Y” is a conditional observation, and it can’t go stale the same way, because its condition carries the dependency.


The write gate: four checks, all on shape

Here’s the part that I find the most interesting. You don’t need a human to approve every write. You need a machine to check four things:

The write gate. All four checks are about the shape of the item, not whether its content is true.

The write gate. All four checks are about the shape of the item, not whether its content is true.

  • G1: the subject resolves to a canonical entity.
  • G2: the applicability conditions are predicates, each with an attribute, an operator and a value. Prose is rejected.
  • G3: evidence references are non-empty and every one resolves.
  • G4: the object is an observed value, not a rule.

Why is checking shape enough? Because an item only does harm when it’s used, and at read time the conditions get checked against present observations. Compare two bad items. One has wrong content but correct conditions: it matches only the situation it describes, and there it gets contradicted by present evidence. Self-limiting. The other has correct content and no conditions: it matches everywhere. That one is the dangerous one, and a machine can detect it without understanding a word of the content.

The argument holds only if G2 and G4 are enforced strictly. A lazy G2 and the whole thing is theatre.

G1 has a dependency people skip. The same component shows up as payment-api in telemetry, payment_api in topology, paymentapi-prod in one deployment. Without an alias-to-canonical mapping, two items about the same thing sit under different subjects and can never be compared. Where that mapping comes from is an open question.


Applicability conditions that a machine can check

Here’s the contrast I keep coming back to:

subject: payment-api
attribute: fix_method
object: "restart"            # rejected: this is a rule, not an observation

versus

subject: payment-api
attribute: restart_restored_service_at
object: "E-1041"             # an evidence reference
required_conditions:
  - config.connection_pool.max   eq   20
  - metrics.pool_utilisation     gte  0.8

The first form has no condition that can match or fail, so it is retrieved in every case with a similar symptom, including the ones where the cause is entirely different. The second excludes itself the moment pool.max is 50.

Two more details. required_conditions says when the item applies, excluded_conditions says when it does not, and they are not negations of each other. Both can hold at once: utilisation is high (required condition holds), but the environment has a different pool architecture (excluded condition holds too). Item doesn’t apply. Omitting exclusions is, I think, the common error, and it produces exactly the “applied everywhere that looks similar” failure.

And the open-world problem. Preconditions presuppose a typed world model, while incident telemetry is open-world and drifts. If nothing can answer “what’s config.connection_pool.max?” with a type, the predicate is UNKNOWN, and unknown is never granted. The sharp practical risk: you can build every table, role and state transition and never admit a single item, because the component that produces typed facts doesn’t exist. So build that fact generator first, for the two or three facts your first example items need. Everything else without it only demonstrates the ability to refuse.

How strong is the literature here? Weak. The one serious proposal I found is a preprint on “Atomic Knowledge Units”. Its premise is right, but its evaluation is a survey of 67 engineers (time saved, a Net Promoter Score), no task benchmark, single author, no venue. There is no established schema standard for a knowledge unit with machine-checkable preconditions. You’d be building ahead of the literature, which cuts both ways: nothing to copy, and no published reason it’s wrong.


Lifecycle: supersession, not decay

Frameworks with a time-to-live apply decay uniformly. A preprint argues this is a category error for factual claims, and I agree. “Slightly less true than yesterday” is not a state a claim can be in. A claim still holds, has been superseded, or has been refuted. So versions get marked superseded and stay readable but inactive.

Then freshness should be bound to what the item refers to, not a calendar. An item about a source whose contract changed yesterday is wrong immediately. An item about a general network failure mechanism may hold for months. One TTL can’t serve both. Instead, each content type gets a re-verification trigger:

Content typeRe-verification triggerSecondary TTL
Configuration, topologyA deploy or config-change event matching the subjectnone
General failure mechanismAn environment version changenone
Tool call refused for lack of permissionAn authorisation event for that principalvery short

The last row is the one that gets missed. “This call was refused” becomes wrong the moment the permission is granted, and nobody thinks of that as something living in memory.

The state machine, roughly:

Item lifecycle: case scope, shadow period, tenant scope, and the exits.

Item lifecycle: case scope, shadow period, tenant scope, and the exits.

A few pieces I like.

Shadow period. Between case scope and tenant scope, the item is retrieved and logged (did its conditions match, did that case conclude correctly) but never enters the context. It catches one specific error: content correct, conditions too loose. A real failure mechanism with a threshold at 0.5 instead of 0.9 isn’t wrong, it just injects an irrelevant precedent far too often. A human reviewer can’t see that either, because it only shows up against real data. The thresholds for graduating or dropping out of shadow need a real distribution before you set them.

Automatic demotion. Preventing a wrong item needs a human; detecting one doesn’t. When a case closes, every item that was in its context gets a support or counter signal depending on whether the conclusion was confirmed or rejected. Counter beats support, and the item goes to needs-re-verification. But the signal is indirect: a wrong conclusion doesn’t prove which item was responsible, maybe none. So it’s a reason to look, not evidence. Never auto-revoke.

Asymmetric authority on expiry. Match is deterministic (is this item’s subject among the event’s entities?), judge is semantic (does this event actually make it wrong?), act is deterministic. The semantic step may narrow the set and raise a flag, never clear one. “Not affected” does not return the item to active, it goes to verify-on-next-use, which is checked at the moment it would enter the context. So a wrong judgement can’t turn a stale item into a conclusion. Neatest detail in the design, I think.

Revocation is three enforcement points, not a status field. Refuse new items with the revoked one in their lineage (write). Filter revoked items before ranking (retrieval). Re-read each context item’s state before every model call. The third costs a re-read per call, deliberately: revoke at t=30s in a five-minute run and only that point catches it at the next model call. Items derived from a revoked one go to needs-re-verification, not revoked, since a lesson inferred from a wrong item may be right for another reason. Revocation is not deletion either; the record is kept.


The read path: predicates before ranking

At read time there are six predicates, all applied before any relevance ranking: tenant scope, capability permission, activated version, purpose of use, applicability conditions, and validity window.

The scenario that makes the fifth one load-bearing: an incident, tail latency spikes, cause is pool exhausted at max 20, restart recovers, verified by an SME. Twelve days later a deployment changes the max from 20 to 100. Nine days after that, the same tail-latency spike, symptoms identical, but measured pool utilisation is 0.12.

The old lesson is verified, correctly scoped, correctly versioned, and semantically near-identical. On relevance ranking alone it ranks first. Five of six predicates pass. Only applicability fails: max is 100, not 20; utilisation is 0.12, not at least 0.8. Freshness doesn’t catch it either, since the item isn’t old in time, it’s wrong in its conditions. Without that predicate the best case is a wasted step and the worst is trusting the precedent over the measurement.

This also flows backwards: an item without machine-checkable conditions isn’t eligible for reuse beyond its own case. That’s why G2 exists.


Where the evidence is thin

  • The approval-before-write idea: the survey I’m drawing on found it in zero of fifteen open-source memory frameworks and zero of three vendor primitives. First-party vendor memory tooling is client-side, so the model only requests operations and your application executes them. Approval, versioning, revocation and audit are entirely yours, and entirely within your power.
  • The one preprint proposing authorisation continuity and policy re-evaluation at retrieval reports no empirical evaluation, no numbers, no code. Treat it as a problem statement.
  • The published agent-memory leaderboard doesn’t measure the write path, revocation, multi-tenant isolation or abstention. Not usable as evidence for any of this.
  • Claim-evidence separation: peer-reviewed work (ALCE, EMNLP 2023) found even the best models lacking complete citation support about half the time on one long-form set. Nobody measured downstream accuracy with and without separation, so adopt it for auditability and don’t claim an accuracy gain.
  • Retraction: when a claim is retracted, the claim is marked refuted, but summaries, embeddings and graph structure derived from it have no answer. Two preprints name the problem and nobody measured it. Each of those is a path by which a retracted claim keeps influencing results.
  • Automatic scope expansion versus human approval is open. Approval load scales with the number of writes, which rises exactly when the system succeeds, but it gives you an audit record. Everything else here is identical either way.

If I had to compress it: get the kind right, gate the shape on write, make conditions checkable, expire by event instead of calendar, and run the predicates before ranking. The part I’d test first is the case where five of six predicates pass. A system that passes everything else and fails that one breaks exactly when it is trusted most.