Model Families & Architectures · 6 min
Model Routing and the Real Selection Tradeoffs
How to route between models at runtime across capability, latency, cost, and context — with cascades, difficulty routing, and the ways routers break.
You almost never have one model. You have a pool: a cheap small model, a mid-tier one, a frontier model, maybe a long-context specialist. And you have a stream of requests whose difficulty varies wildly. Routing is the decision layer that sits in front of that pool and picks, per request, which model does the work. Get it right and frontier-level answers go to the hard 10% of traffic while the easy 90% pays small-model prices. Get it wrong in one direction and you burn frontier money answering "what's 2+2." Get it wrong in the other and a genuinely hard query lands on a small model that quietly ships a wrong answer.
This is a different router from the one in the MoE lesson. A mixture-of-experts router steers tokens to experts inside a single model's forward pass, learned end to end and invisible to you. The router here steers whole requests to separate models. It runs in your infrastructure, and you own every one of its failures. Same word, different altitude.
The four axes, and why they fight
Every routing decision trades off four things that do not move together.
- Capability: will this model actually get the answer right? The hardest axis to measure, because it is a property of the query, not the model. A 7B model nails routine extraction and falls apart on multi-step math.
- Latency: a small model on a warm GPU replies in a few hundred milliseconds. A frontier model with long output and a cold start can take tens of seconds. For an autocomplete box that gap is the whole product; for an overnight batch job it is noise.
- Cost: the spread between tiers runs one to two orders of magnitude per token. A frontier model costs roughly $3 to $15 per million tokens; a small hosted model runs about $0.10 to $0.60. At scale that ratio is your margin.
- Context window: a hard ceiling, not a gradient. If the request plus retrieved documents is 400K tokens and your cheap model tops out at 128K, capability and cost stop mattering. The window already decided.
Here is what makes routing hard. These four axes barely correlate with each other, but every one of them correlates with query difficulty, and difficulty is the one thing you cannot read off a request without spending the compute to answer it. That circularity is the whole problem.
Two ways to route
Predictive routing decides before generating anything. A lightweight classifier, often a fine-tuned BERT-class model or an embedding-similarity lookup, scores the incoming query and picks a tier. RouteLLM (LMSYS, 2024) trained routers like this on Chatbot Arena preference data and reported over 85% cost reduction on MT-Bench and 45% on MMLU while holding 95% of GPT-4's quality. One decision, no wasted generation. The price you pay is that the router is guessing difficulty from surface features, and it guesses wrong sometimes.
Cascades decide after. Run the cheap model, judge whether its answer is good enough, and escalate only on failure. FrugalGPT (Chen, Zaharia and Zou, 2023) is the canonical version: a chain of models plus a scorer that decides accept-or-escalate. Their headline was matching the best single model's accuracy at up to 98% lower cost, or beating that model by 4% at equal cost, on their benchmarks.
CASCADE (decide after generating)
query
│
▼
┌────────┐ score ≥ τ ? yes
│ small │──────────────────► return (cheap, ~90% of traffic)
└────────┘
│ no (low confidence)
▼
┌────────┐ score ≥ τ ? yes
│ mid │──────────────────► return
└────────┘
│ no
▼
┌────────┐
│frontier│──────────────────► return (expensive, the hard tail)
└────────┘The scorer is where cascades get expensive. On multiple-choice or code-with-tests you have a cheap, honest signal: token logprobs, or does-it-compile-and-pass. On open-ended generation you have neither, so people reach for a separate judge model, and now your "cheap" path pays for two forward passes plus a judge on every escalation. Cascades also serialize latency. A query that escalates twice waits for three models in sequence. Predictive routing is one hop; a deep cascade can be three.
cheap ◄─────────── cost/capability ───────────► frontier latency-critical │ small only, no cascade (autocomplete) │ interactive chat │ predictive route, small↔mid, 1 hop batch / offline │ deep cascade, small→mid→frontier, latency is free huge context │ window forces the tier regardless of the above
What actually goes wrong
Misclassification is asymmetric, and you have to pick which way you would rather fail. A false "easy" ships a wrong answer to a user: silent, and the expensive kind of failure. A false "hard" just overpays. Tune the threshold toward over-escalation, then go attack the resulting cost.
Router calibration drifts. A router trained on last quarter's model pair degrades the moment you swap in a new model whose real capability boundary sits somewhere else. RouteLLM found its routers transferred across model pairs better than expected, but "better than expected" is not "recalibrate never." Treat the threshold as a live parameter with a dashboard behind it, not a constant you set once.
The router's own cost and latency are not free. A judge model that costs a third of the frontier model quietly eats the savings you routed to capture. Measure end-to-end cost per resolved query, including escalations and judge calls, not the sticker price of the tier you hoped to hit. RouterBench (Martian, 2024) exists precisely because "we route now" claims need an apples-to-apples cost/quality curve; it ships 405K precomputed inferences across 11 models so you can plot your policy against a fixed baseline.
Context is a cliff, not a slope. Put a hard pre-check on token count before any capability logic runs. A request that overflows the small model's window should never reach the difficulty classifier at all, because the window already made the call.
Builder: Ship one boring rule before any ML: if input_tokens > small_ctx or has_tools or lang != "en": go big. That static gate captures a surprising share of the real wins and gives you the baseline a learned router has to beat before it earns its extra cost and failure surface.Researcher: The open frontier is difficulty estimation without generating: cheaper, better-calibrated signals than "embed the prompt and classify." RouterBench is the standard harness. Report the full cost/quality Pareto curve, not one operating point, because a router that wins at the 90%-quality target can lose at 99%.
The security angle
A router is an attacker-influenced control-flow decision, which makes it a fresh attack surface. "Rerouting LLM Routers" (Shafran et al., 2025) demonstrates confounder gadgets: short, query-independent token strings a user appends to a prompt that flip the router's decision. Their upgrade attack forces queries to the expensive frontier model, a cost-inflation or denial-of-wallet attack. It works (upgrade rates of 79% to 91% against four open-source routers) and it transfers to black-box commercial routers the attacker cannot inspect. The mirror risk is a downgrade: steering traffic to the weak model to degrade quality or slip past safety tuning that only the frontier model has. The gadget paper concentrates on upgrades, but the trust boundary cuts both ways, so guard both.
The takeaway is that the router is a trust boundary. Do not let raw user text be its only input, cap the fraction of traffic any single account can push to the expensive tier, and alert on tier-mix anomalies the way you already alert on a traffic spike.
Defender: Rate-limit by resolved tier, not just by request count. An account whose queries suddenly route 90% to the frontier model is either a genuinely hard workload or someone running an upgrade attack on your bill. Both are worth a page.
A worked cut
Say 100K support queries a day. A frontier-only baseline costs, illustratively, $2,000/day. Sampled, 82% turn out to be simple lookups the small model answers correctly, and a logprob threshold catches most of the rest for escalation. Route them: 82K at small-model rates plus 18K escalating to frontier lands near $400 to $500/day, roughly a 75% cut. That holds only if your escalation scorer's false-easy rate stays low enough that ticket quality does not slide. That one clause is the entire engineering problem, and it is why you instrument resolution quality per tier from day one, not the week after finance notices the savings.