Start with what the field already says

I’ve been reading a lot about prompt injection lately, and the thing that surprised me is how little disagreement there actually is. The current application-security guidance for LLM applications says it plainly: prompt injection is intrinsic to current generative AI, and no reliable prevention mechanism exists today. It also puts jailbreaking inside the same class, as the subset of prompt injection where the attacker wants the model to break its own safety rules.

So that is the starting position. Not “we need better detection”. The strongest published result in the area says detection doesn’t close the problem. Anything that promises to filter it away is contradicted by that.

“We added a guard model in front” is not a mitigation, and here’s why.


The result that kills filter-based defences

There’s a preprint called “The Attacker Moves Second” (arXiv 2510.09023, authors include people from Google DeepMind and ETH Zurich). The setup is simple. Take twelve recent defences, built on a diverse set of techniques. Look at what their own authors reported: attack success near zero. Now let the attacker adapt to the defence. Attack success goes above 90% for most of them.

Same defences. The only difference is who moved second.

A static benchmark tests a defence against attacks written before it existed. A real attacker reads your defence and writes the attack for it.

And the mechanism generalises, which is the part I find convincing. A classifier-based guard is itself a model. So it’s vulnerable to the same kind of manipulation it’s supposed to catch. The design-patterns literature says input and output filters remain fundamentally heuristic and cannot guarantee prevention of all attacks.

Does that make classifiers useless? No. A classifier that reduces nuisance volume is useful. But there’s a line: it’s defence in depth, never the control. Putting a classifier in front of a risk function and calling it the mitigation is a misrepresentation. That’s my reading, anyway, and I’d defend it in a review.


The architectural answer

The same design-patterns paper (arXiv 2506.08837, Beurer-Kellner, Tramèr and a long list of co-authors) states the guiding principle in one sentence: once an LLM agent has ingested untrusted input, it must be constrained so that it is impossible for that input to trigger any consequential actions.

Impossible. Not unlikely, not detected. Impossible by construction.

It names six patterns, and each one cuts something different:

  • Action-selector. The model is a switch choosing among predefined tool calls, with no feedback from those actions back into the agent. Retrieved content may inform a choice among fixed options, but it can’t add an option.
  • Plan-then-execute. The plan is fixed before untrusted content is read. Content can change values, not the sequence of actions.
  • Map-reduce over untrusted items. Each item is processed in isolation and only a constrained summary leaves.
  • Dual-model. A privileged model never sees untrusted content. A quarantined model sees it and cannot act.
  • Code-then-execute. The control flow is code, not a model decision.
  • Context minimisation. Untrusted content is removed from the context before the consequential step.

What I like is that they all say the same thing in different dialects. The model may supply values. It may not choose actions.

Compare that with the free-running observe-decide-act loop, where every page you read influences the next action. That loop is an injection path by design. A program that accepts values without accepting actions closes it: the page supplies a number, the program decides what to do with it.

If you’re building an agent where an engine picks the step and dispatches the tool, and the model only suggests an intent, you already have the action-selector pattern. The point I’d make is: write that down as a security property. Otherwise someone will trade it away later “for agent flexibility” and nobody will remember why it was there.


The strongest published defence, and what it costs

The strongest thing I’ve seen is CaMeL, “Defeating Prompt Injections by Design” (arXiv 2503.18813, Debenedetti, Shumailov, Carlini, Tramèr and others). Two things do the work.

First, control flow comes from the trusted query only. The system extracts control flow and data flow from the trusted query, so untrusted data retrieved by the LLM can never impact the program flow.

Second, every value carries a capability: its provenance (user, a transformation, a specific tool), its allowed readers, its origin within a tool. So a policy can be enforced at the moment a tool is called.

Control flow is extracted from the trusted query. Untrusted content only ever travels the data path, and policy is checked where the tool is called.

Control flow is extracted from the trusted query. Untrusted content only ever travels the data path, and policy is checked where the tool is called.

Now the number that matters. The paper reports tasks solved: 84% for the undefended system, 77% for CaMeL. That’s a seven-point utility cost for a provable property.

I want to be honest here, because people skip this. Architectural defence is not free. You lose capability. But it’s a named cost, and a named cost is a defensible one. Compare it with the alternative: other defence methods in the same comparison, with no explicit policies configured, still had dozens of successful attacks, while CaMeL was near zero. In one model’s setup it went from 300 successful attacks to 0.

What I’d steal without building the whole thing

Two ideas are portable and cheap.

  1. Provenance as metadata, not prose. Most systems tell the model “this part is from a user, this part is from a tool” in natural language. CaMeL’s contribution is making it per-value and machine-checkable, so a policy gets evaluated at the point the tool is called instead of being argued about in code review.
  2. Symbolic references. CaMeL’s controller passes variable placeholders, so the privileged component never encounters actual untrusted content, only variable names representing it. Translated: a tool intent should carry an evidence ID, not the evidence text. Then the untrusted bytes never need to pass through the decision path at all.

Where the evidence is weak

Don’t oversell this. A preprint evaluating the out-of-band defence family under adaptive attack (arXiv 2606.26479) reports mean attack success dropping from 25.8% to 4.2%, and a hand-crafted adaptive attack failed to raise it. Sounds great. But the authors’ own caveat is that it is one small-scale data point on a weak model with a single black-box attack template, and a stronger optimised attack remains open.

So my read: the architectural direction is right, and the specific systems are not yet independently validated. Take the pattern, not the product.


The capability-budget rule

This is the one I’d put on a wall. Within a single session, an agent should satisfy no more than two of these three:

  • A. It processes untrustworthy inputs.
  • B. It accesses sensitive systems or private data.
  • C. It changes state or communicates externally.

If all three are needed, the agent should not operate autonomously and at a minimum requires supervision, via human-in-the-loop approval or another reliable means of validation.

An investigating agent keeps A and B because they are the job. Leg C is the one you can actually remove.

An investigating agent keeps A and B because they are the job. Leg C is the one you can actually remove.

Think about an agent that investigates and diagnoses problems. It reads logs, traces, change records, customer documents. That’s A, and it’s unavoidable because that is the work. It touches sensitive telemetry. That’s B, also unavoidable. But C, changing state or talking to the outside world, is removable. Make the agent read and recommend only, no automated remediation, and you’ve removed leg C.

I’d argue that’s not a product limitation. It’s probably the single most effective prompt-injection control you have, and it should be defended as such the next time someone says “can’t it just fix the thing automatically?” Adding an automated action to a session that already has A and B removes your strongest control.

One caveat on the human-approval clause: a human reviewer is a control, but overwhelming a reviewer with plausible-looking evidence is itself a named threat. Approval isn’t a free pass.


Layers, and where the strength sits

The practitioner frameworks list controls in layers: instruction hierarchies at the model level; sandboxing, filesystem scope, network egress policy, granular tool permissions and capability-style patterns at the deterministic system level; input/output classifiers and activation-level detection at the probabilistic system level; memory-poisoning containment, reasoning inspection, link-safety blocking and audit logging at the harness level; vetted tool directories at the tool level.

Given the adaptive-attack result, the strength sits in the deterministic controls. The probabilistic ones are defence in depth only. Two from that list deserve a closer look.

Memory poisoning. If an item enters long-term memory from untrusted content with no applicability conditions, you get a persistent injection. It gets retrieved in future cases, possibly in a different tenant’s context if scoping is weak, with no trace of where it came from. So the memory write gate is a security boundary, not only a quality gate.

Network egress policy. Once leg C is gone, this is the highest leverage. And it’s easy to get wrong:

  • A TLS-server-name-only allowlist is documented as bypassable by the vendors who publish it.
  • DNS resolution outside the deny policy turns the resolver into an exfiltration channel.
  • A broad address-range allow sitting next to your domain rules silently defeats them.

One permissive rule undoes the careful ones.


The retrieval path is its own injection surface

Retrieved text is untrusted input. It must not override system instructions, tool policy, tenant scope or workflow controls. Saying so isn’t enough, you need a design rule that enforces it.

The rule I like: trusted code builds the query (from the incident summary, a capability id, the evidence gaps), runs one retrieval round, and that’s it. The model never writes the query. There’s no second round.

Then walk the injection paths one by one:

  • Malicious content changes the agent’s next action: closed, because the engine chooses the next action.
  • Malicious content causes a different retrieval: closed, because code builds the query and there’s one round.
  • Malicious content causes a tool call: closed, because the agent suggests an intent, the engine dispatches, and a gateway holds the credential.
  • Malicious content is cited as evidence: only partly closed.

That last row is the live one. Method rules can forbid citing knowledge-base text as evidence. But historical precedent (past cases pulled back in) is a third source class, and if no record forbids citing it, you have a hole. It’s a knowledge-base defect with a security consequence, and I’d close it before memory enters the context window, not after.


What I’d take away

Filters lose because the attacker moves second and the filter is a model too. Architecture wins because it doesn’t need to recognise an attack, it only needs the attack to have nowhere to go.

Measure your own utility cost, though. The published comparison is 77% versus 84%. If you’ve built architectural constraints and haven’t measured what they cost you, you can’t defend them well. Knowing the number makes the argument stronger, not weaker.


References

  • “The Attacker Moves Second”, arXiv 2510.09023
  • Design patterns for securing LLM agents against prompt injection, arXiv 2506.08837
  • “Defeating Prompt Injections by Design” (CaMeL), arXiv 2503.18813
  • Adaptive evaluation of the out-of-band defence family, arXiv 2606.26479