Layout is a cost decision

Most people I talk to treat prompt layout as a readability question. Where does the system prompt go, where do the examples go, should the date be at the top. It feels like formatting.

It’s not formatting. If you run an agent loop, the order and exact bytes of your prompt decide a large part of your bill and your latency. And the failure mode is silent: no error, no wrong answer, just a bigger invoice.

So here’s the mechanism first, then the rules, then the places where the rules fight with other things you probably care about.


Prefill, decode, and why loops are the worst case

One model call has two phases. Prefill processes every prompt position at once, in parallel, and is bottlenecked by compute. Its latency grows with prompt length. Decode produces one new position per step, sequentially, bottlenecked by memory bandwidth. Its latency grows with output length. (This split is from a CMU lecture on LLM serving.)

Now think about an agent loop. Call 1 sends system prompt, tools, user request. Call 2 sends all of that again plus the tool result and the assistant turn. Call 3 sends all of that again plus more. Without caching, every call pays full prefill on everything that came before. The cost per call grows with every step.

With a stable prefix it flips. The previous call’s whole sequence is already in the KV cache, so only the new tail is paid for. The loop goes from the worst case to a pretty good one.

But only if the repeated part is byte-identical. That’s the catch. And it means prompt layout belongs to whoever assembles the context, not to whoever runs the GPUs.


The mechanism, just enough to design against

Two implementation ideas make this practical, and both are public:

  • Paged allocation (PagedAttention, the vLLM paper from SOSP 2023). The KV cache lives in fixed-size blocks, and two requests can point at the same physical block. That’s what makes sharing cheap at all.
  • Radix-tree prefix sharing (RadixAttention, SGLang). Requests that diverge split a node in the tree instead of invalidating it. A shared 18K-token system-plus-tools prefix stays shared while tenant A and tenant B go their own ways.

There’s also cache-aware routing, where the scheduler scores a worker by reusable prefix length minus queue penalty. An idle worker with a 2K match can lose to a busy worker with an 18K match. I mention it because it shows the prefix isn’t just a billing detail, it’s an input to scheduling.

What you need to remember: the matching is by prefix. First byte that differs, and everything after it is a miss.


The price gap is the design input

Published per-1M-token prices, which the source checked on 30 Aug 2026, across six models: the ratio of regular input price to cached input price is about 31x for one provider, 10x for four of them, and about 5x for the sixth. So roughly an order of magnitude.

Two caveats matter more than the numbers. First, “cached” means a cache read; cache creation may be billed separately, which changes the break-even on prefix size. Second, one of those ranges is off-peak to peak. Prices move, so treat the table as a snapshot that establishes the ratio, nothing more.

Also, don’t use physical cost to predict price. Research measures compute and memory; the price list is a commercial decision. Different quantities.

Still, do the arithmetic once. A 20K-token stable prefix re-sent on 30 calls is 600K input tokens. Paid at regular price that’s 1.0x. Paid at cached price it’s about 0.1x. For a lot of agent workloads, getting the prefix stable is worth more than most prompt tweaks you’d spend a week on. I’d go further and say it’s the first optimisation to check, before touching the wording of anything.


The layout, as three zones

Prompt layout: stable prefix, tenant or run scope, volatile suffix. Everything from the first changed byte onward is paid at the regular input price.

Prompt layout: stable prefix, tenant or run scope, volatile suffix. Everything from the first changed byte onward is paid at the regular input price.

Top is stable: system instructions, tool definitions in a fixed order, mandatory records, examples. Middle is tenant or run scope: still cacheable, but only shared within that tenant or run. Bottom is appended: new turn, observations, tool results.


The four rules

Stabilise the prefix. System instructions, tool definitions, mandatory records and examples go first, in a fixed order.

Append changing state. New turn, observations and tool results go at the end. Never interleave them into the stable region.

Avoid needless churn. No timestamps, no request IDs, no non-deterministic schema ordering anywhere in the prefix.

Measure reuse. Cached tokens, time-to-first-token and eviction rate are first-class metrics, not infrastructure trivia.

The third one is the rule that gets broken by accident, because nobody breaks it on purpose. Two examples that I think happen all the time:

  1. Tool definitions serialised from a dict where key order isn’t deterministic. Call 1 has name then desc, call 2 has desc then name. Different bytes, zero reuse. Same tools, same meaning.
  2. A system prompt that interpolates the current time down to the second. That’s a new prefix every second.

In both cases the model answers correctly. The only symptom is the bill.


Where it collides with other requirements

These are real conflicts, not oversights, so worth being honest about them.

Pinned knowledge

Mandatory records belong in the stable prefix: they’re read on every invocation and they’re the cheapest thing to cache. But if they’re tied to a pinned snapshot version, then every knowledge release changes the prefix and your hit rate drops.

That drop is correct behaviour. Don’t read it as a regression. If you alert on hit rate, annotate releases or you’ll chase a bug that isn’t there.

Tenant isolation

A shared prefix is the ideal case for caching and the worst case for isolation reasoning. The rule is simple: the shared region is tenant-free, tenant content goes after it, never inside it.

Easy to state, easy to break. A tool description that says “…for Acme’s payments endpoint” has just moved tenant data into the shared prefix. It hurts twice: reuse breaks across tenants, and you have a tenant name sitting in a region that’s supposed to be common. (Caches as a timing or hit-rate side channel is its own topic. The simple rule above is the one you need for layout.)

Compaction

Compaction rewrites the middle of the history, which invalidates the cache from that point on. Unavoidable, and acceptable, because the alternative is full prefill on an ever-growing history.

What’s avoidable is compacting more often than necessary, or compacting a region that was about to be reused. So the trigger threshold is a cost parameter as well as a capacity one. I think this is the least appreciated of the three. People tune compaction thresholds for context-window fit and never look at what the rewrite does to the cache.


What to measure

The headline number is cached input tokens divided by total input tokens. Next to it: time-to-first-token by prompt length (that’s where prefill cost shows up for the user), eviction rate (a high one means your working set exceeds capacity and the reuse is illusory), and cost per run split into cached input, regular input and output. That last view is the only one where a prefix fix is visibly worth doing, and the one that makes this work legible to people outside it.

The diagnostic I like most is the prefix hit length distribution. Healthy looks like one mode, high. If you see a second mode near zero, one code path is churning the prefix. That’s a bug, not a workload.


What I’d actually do

Hash the prefix across two consecutive calls in a test. It’s cheap, and it’s much stronger than a review rule that says “don’t put timestamps in the system prompt”. Reviews miss the dict-ordering case every time, because nobody reads serialisation order in a diff.

Then look at your cached-token ratio. If you can’t see it, every statement you make about inference cost is an estimate, and the rules above can’t be verified. Check whether your serving path even exposes per-request cache metrics, and whether cache creation is billed separately, before you size anything.

My honest caveat: the price table is a snapshot and the ratios are the only part I’d rely on. The layout rules don’t depend on any particular price. They only depend on prefix matching, and that’s not going away soon.


Further reading

  • PagedAttention (vLLM), SOSP 2023
  • SGLang RadixAttention, and the SGLang v0.4 cache-aware load balancing
  • CMU 11-768 Lecture 3, on prefill and decode and prefix caching economics