Two Things With One Name

I keep seeing “routing” used for two things that have almost nothing in common, and the conflation makes every conversation about it muddier than it needs to be.

Model routing decides which model answers a request. The signal is a classifier, an embedding, or a cascade with a confidence check. The goal is cost or quality.

Infrastructure routing decides which worker answers a request. The signal is cache utilisation, queue depth, which adapters are loaded. The goal is throughput and latency.

Same word, different problem, different evidence. Basically almost all the enthusiasm I see is about the first one, and almost all the measured gain I can find is in the second. So let me go thru both, starting with the one that’s more fun to be skeptical about.


The Result Everyone Quotes

There’s a widely cited paper (arXiv 2406.18665) on a learned router that sits between a strong model and a weak model, trained on preference data. The numbers are good. Cost reduction of more than 85% on one benchmark, 45% on another, 35% on a third, while retaining 95% of the strong model’s performance. One variant gets 95% of strong-model quality while calling the strong model for only 26% of requests.

I want to be fair here: this is real, it comes from the authors, and it’s on their setup. Nothing wrong with that. It’s just the authors measuring their own router on their own benchmarks, which is the weakest kind of evidence for “this will work for you”.

The Independent Evaluation

Then there’s a more recent preprint, “Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks” (arXiv 2608.14641). They put four routers behind one common interface and measured them on four benchmarks. A preprint, so treat it as such, but look at what it found.

First, what the routers actually did. The paper says “three routers emit constant or near-constant tier assignments”. Three out of four. One router picked the middle tier basically all the time, another nearly all the time. Only one router distributed its choices across tiers in anything like a varied way.

Second, the control. A baseline that always picks the middle tier, no model involved at all, matched one router’s performance exactly on three of four benchmarks. The paper’s own wording is that gains “track selected-tier composition more closely than demonstrated task-specific targeting”.

So the savings come from using cheaper models more often. Not from semantic cleverness about which request needs which model. That’s my reading, and I think it’s the right one, though one preprint on four open-source routers is not the final word. (Router robustness is also separately contested in another preprint, arXiv 2504.07113. I haven’t read that one in full, so I’m just flagging it.)

Why this matters for money, not just for science

If the mechanism is tier mix, you don’t need a router to get the saving. Pick the tier per phase of your workflow. That’s a field on the model binding. Compare:

  • Build a router: a classifier to train, a quality argument to defend, ongoing calibration, drift to monitor.
  • Choose the tier per phase: a config value. The model doesn’t change per request, so there’s no quality argument and no drift.

And on managed APIs there’s a stronger version of the same lever, the service tier. Priority, standard, flex and batch tiers price the identical model very differently. That’s a cost knob that needs no router and no quality argument at all.

The Managed Router, Briefly

Cloud providers sell this as a feature too. The one I looked at (Bedrock prompt routing, per its docs and pricing page) routes between exactly two models from one family, predicts response quality per model, and takes a fallback model plus a quality-difference threshold.

Three things count against it for me. It’s English-only. You can’t tune it on your own performance data (the docs say you “can’t adjust routing decisions… based on application-specific performance data”). And it’s priced at $1 per 1,000 requests, which on cheap traffic can cost more than it saves. A router that saves a fraction of a cheap call but charges more than that fraction is a net loss. I also noticed the supported-model table looked stale and still carried preview language despite the feature being generally available, which is a small thing but doesn’t build confidence.


The Half That Works

Infrastructure routing is where I’d actually put effort. Three independent systems converged on the same three signals: a Kubernetes gateway extension, a distributed-serving stack, and a routing gateway (recently rebranded under a foundation). All three look at KV-cache utilisation, queue depth, and active adapters. The gateway adds endpoint health on top.

The three signals an endpoint picker reads before choosing a worker

The three signals an endpoint picker reads before choosing a worker

That kind of convergence is the strongest consensus I can find in this whole area. The serving stack also has explicit anti-hot-spotting, including a credit-decay knob that damps a cache-rich worker’s advantage when it’s overloaded. Which makes sense: the worker with the best cache hit is also the one everybody wants to hit.

The measured gain, and its caveat

The only configured benchmark I found: on a prefill-bound workload, the configured router got 2-3x over round-robin. On a decode-bound workload it reached “parity or better”. And the same source says plainly: “no single configuration wins across all three workloads.”

So the gain is real, but it’s workload-shaped. If someone quotes you a cache-aware routing number without saying what the workload was, it’s not interpretable. Also worth noting: the cache-aware router’s own documentation provides no measured performance data. It tells you to compare random and round-robin against cache-aware yourself. That’s the right instruction and it’s also kind of an admission.

One architectural thing to track

Components recently moved between projects. The gateway specification project keeps the resource definitions, conformance, and a minimal reference implementation (the endpoint-picker reference is now optional), while the endpoint-picker, body-based routing and latency prediction went to the serving project. My takeaway: the interface is stabilising as a specification while the implementation consolidates elsewhere. Depend on the interface, expect the implementation to move.


Cascades Are Worth Measuring

If I had to do model routing at all, I’d do a cascade, not a classifier. The difference is when the decision happens.

A classifier commits before it sees any output; a cascade decides after

A classifier commits before it sees any output; a cascade decides after

A classifier-based router looks at the request, guesses which model is needed, and commits before any output exists. A cascade runs the cheap model, checks the output, and escalates only on failure. It decides after seeing the output, and that’s the decisive difference.

The escalation trigger matters too. “The cheap model says it’s confident” is a weak trigger, because the model is grading itself. A better one is a hard, mechanical check computed by the runtime: does every claim carry an evidence reference, is the required structure complete, does the output cite what it’s supposed to cite. If the cheap output fails a hard gate, escalate. That’s a principled trigger, not a confidence threshold.

Then the arithmetic decides whether it’s worth it. At a 10% escalation rate the saving is large. At 30% it’s moderate. At 60% it’s marginal. At 85% it’s negative, because you paid for both calls most of the time. So the number you need before anything else is the escalation rate, and you can measure it cheaply on an existing case set before building anything.

I’d say it plainly: a cascade is an experiment, not an assumed saving. Report the escalation rate and the quality delta together. A saving claimed without an escalation rate is incomplete.


One Boring Argument That Might Matter Most

A router turns the model into a per-request variable. If your evaluation records “model id and version” once per batch, a router breaks that: run 1, run 2, run 3 might each have hit a different model, and your manifest says “router v1”. That batch isn’t reproducible.

So if you adopt any routing, record the selected model per invocation. This is an argument for explicit, versioned model binding that’s independent of cost or quality, and it’s easy to lose if routing gets introduced as an optimisation instead of a change to the contract.


What I’d Actually Do

  • Choose the tier per workflow phase instead of building a semantic router. That’s the mechanism the independent evaluation says produces the gain.
  • Keep the capability-to-model binding explicit and versioned. Don’t let routing become implicit.
  • If you self-host, record cache utilisation, queue depth and active adapters, then test routing on your workload type, since no single configuration wins everywhere.
  • Treat a cascade as a measured experiment with hard gates as the trigger. Don’t assume it preserves quality.
  • Depend on the routing interface specification, not on components that just moved repositories.
  • Before adopting a router on the strength of its authors’ own savings, check it against an independent evaluation. And if a gain is reported without the tier composition, assume it might be entirely tier mix.

One open question I don’t have an answer to: which workflow phases tolerate a cheaper model, as opposed to just a slower tier? Those are two different levers and only the second one is free.


References

  • arXiv 2406.18665, learned router trained on preference data (the widely cited result)
  • arXiv 2608.14641, “Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks”
  • arXiv 2504.07113, router robustness (not read in full)
  • AWS Bedrock documentation on prompt routing, and Bedrock pricing