Home › What Happens After You Hit Enter › Model Routing & Cascades (Small-First Selection)
Pipeline stage · Operate

Model Routing & Cascades (Small-First Selection)

A router predicts whether this particular query needs the expensive model or whether a cheap one will produce an equally good answer, and sends it accordingly.

In one line

Most production traffic is easy, so a router that sends only the hard queries to the frontier model can halve spend while barely moving average quality.

Why you'd careThe thing you have already noticed

Two users ask your product near-identical questions an hour apart and get answers of visibly different quality — one crisp and specific, one generic. Nothing changed in your prompt, your data or your model configuration. What changed is which model actually answered. A router sits in front of the model you nominally requested and decides, per query, whether a cheaper one will do. Calibrated well, nobody notices and your bill halves. Drifted, the errors concentrate exactly where they hurt: on the hard queries that were the reason you paid for a frontier model in the first place. This stage is also why reports that the model got worse are so hard to reproduce — you and the reporter may not have been talking to the same model.

In and outWhat goes in, what comes out

InThe prompt, optionally the conversation so far, a candidate pool of models with their per-token prices, and a target: either a cost budget or a tolerated quality drop, often expressed as the fraction of calls allowed to reach the strong model.
ProcessA small scorer — typically a fine-tuned encoder or a matrix-factorization model fitted to human preference data — estimates the probability that the weak model's answer would be judged at least as good, then compares it against a threshold. In a cascade, run the cheap model first and score its actual output instead.
OutA concrete model id that may differ from the one the caller named, a routing decision record for logging, and in a cascade a discarded draft answer whose tokens you have already paid for.

The prompt survives intact; routing rewrites nothing. What is lost is the caller's assumption that the model field is a request rather than a hint, and with it the comparability of your logs: aggregate quality metrics now blend two distributions, so an average that looks stable can hide a widening gap on the hard tail. In a cascade you also lose money that produced nothing, because an escalated request pays for the draft, the judge and the final answer.

ConceptThe idea underneath

The router is a learned classifier over queries, and the only unusual thing about it is what it predicts: not the answer, but whether a cheaper model's answer would be good enough.

The training signal is pairwise human preference data, the same kind collected by public model-comparison arenas. From votes between two models on the same prompt you can fit P_strong_wins = sigma(f(q)), where q is the query, f is a small learned scorer, and sigma is the logistic function squashing its output into a probability. Route to the strong model when P_strong_wins > tau. The threshold tau is the cost-quality dial; lowering it sends more traffic to the expensive model.

RouteLLM's most effective variant does not score the query in isolation. It uses matrix factorization: learn an embedding for each model and each query, and score a pairing by their inner product. This is the machinery of a recommender system, and it generalizes better than a single difficulty score because it captures which model is good at what, not merely which query is hard.

Cascades invert the order. Run the cheap model, score its answer with a verifier g(q, a), and escalate only when the score falls below a bar. You get a much better signal — judging a real answer rather than predicting one — and you pay for it with the draft's tokens and an extra round trip.

Both designs live or die on one economic constraint: the router must be far cheaper and faster than the difference it saves. A 100 ms router in front of a 400 ms model is not a saving.

At a glanceSee it

Model Routing & Cascades (Small-First Selection) diagram

A router predicting difficulty, and a cascade paying for a draft before it escalates.

The knobsHyperparameters and nuance

  • threshold / tauthe cost-quality dial. RouteLLM parameterizes it by the fraction of calls sent to the strong model, which is far more interpretable than a raw probability. Push it down and quality erodes first on hard queries; push it up and you have added latency and a maintenance burden for nothing.
  • router typeRouteLLM ships several: matrix factorization, a BERT-style classifier, and a causal-LLM classifier. The cheap ones cost single-digit milliseconds; an LLM-based router can cost more than the call it was meant to avoid.
  • escalation threshold on the verifierin a cascade, the score below which the cheap answer is rejected. Set it strict and the escalation rate approaches 100 percent, so you pay for both models plus the judge on nearly every call.
  • cascade depthtwo stages is normal. At three or more, the expected cost of discarded drafts starts to exceed the saving from the calls that stop early.
  • fallback policywhether an availability error from the preferred model may silently swap models. Failing over on a 429 and routing on predicted difficulty are different decisions and should be logged as different decisions.

EffectHow this stage moves the answer

Quality loss from routing is never uniform, which is both the point and the danger. The queries sent to the cheap model are the ones predicted to be easy, so the errors concentrate on queries that look easy and are not — ambiguous phrasing, an unusual language, a domain the router's training data never covered. Users experience this as a product that is reliable except when it matters. Two second-order effects are more visible. Formatting habits differ between models, so a routed multi-turn conversation can change voice, heading style or list conventions between turns for no reason the user can see. And structured output is where cheap models degrade first: if a request carries a strict JSON schema or a tool definition and the router ignores those features, you get a spike in parse failures and malformed tool calls that looks like a bug in your parser.

EvalsWhat it does to your measurements

Evaluate the routed system, never either model alone — a single accuracy number cannot express a trade whose whole shape is a curve. The right artifact is a cost-quality Pareto plot: quality against the fraction of calls routed to the strong model, with the two single-model points as endpoints. The summary metric is the call fraction needed to reach some percentage of strong-model performance. The trap is distributional. Benchmark suites are curated to be hard; production traffic mostly is not. A router tuned where half the queries need the strong model will route a few percent of real traffic there, and your offline numbers predicted none of it. Cascades have their own trap: if you do not charge the discarded draft and the judge call against the cascade's cost, the Pareto curve is fiction.

Failure modesWhen it goes wrong

  • Quality is fine on your eval set and mediocre in productionthe router was calibrated on a benchmark difficulty distribution that does not match real traffic.
  • Enabling the cascade increased total costthe escalation rate is high enough that most requests pay the cheap model, the verifier and the strong model instead of just the strong model.
  • The assistant's voice or formatting changes between turns of one conversationthe router decides per turn with no session stickiness, so consecutive turns come from different models.
  • A sudden spike in JSON parse errors or malformed tool callsrequests carrying strict schemas are being routed to a model with weaker constrained-output adherence.
  • A capability regression confined to one language or one topicthe router's features do not represent that slice, so it scores every such query as easy.

PapersWhere this comes from

  • RouteLLM: Learning to Route LLMs with Preference DataIsaac Ong et al., 2024 (arXiv:2406.18665). Trains routers on pairwise human preference data and shows a matrix-factorization router recovering most of the strong model's quality while sending only a minority of calls to it; this is the reference design for the stage.
  • FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving PerformanceLingjiao Chen, Matei Zaharia and James Zou, 2023 (arXiv:2305.05176). Introduced the LLM cascade with a learned scorer deciding when to escalate, and argued that on some workloads a cascade beats the strong model on both cost and accuracy.
  • Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding et al., ICLR 2024. Routes on the predicted quality gap between models rather than on query difficulty alone, and shows the threshold behaving as a smooth cost-quality dial rather than a cliff.
A living map of modern AI — kept current every morning