In plain terms
A hospital does not send every patient to the senior consultant. A triage nurse looks at each case for a few seconds and decides: the pharmacist can handle this one, a general practitioner that one, the specialist the third. A model router is that triage desk for AI requests. It looks at each request and picks the cheapest model that is likely to answer it well.
Why it matters
In most applications the bulk of requests are easy, and a small model answers them just as well, in less time and at a fraction of the price. Routing can therefore cut the bill substantially while quality on the test set holds. The risk sits in the router's mistakes: a hard question sent to a weak model produces a confident, poor answer that nobody flags. A router is only as trustworthy as the evaluation behind it, and it is one more component to test whenever a model or a prompt changes.
Example
A bank's support assistant handles 1 million requests a month, all on a frontier model, for 40,000 dollars. A router is added. It sends 70% of requests, such as balance and card questions, to a small model, 22% to a mid-tier model, and the remaining 8%, complaints and complex cases, to the frontier model. The monthly bill falls to 13,000 dollars; accuracy on the 2,000-question test set moves from 94% to 93%.
Most often confused with
Model Router vs. AI gateway
The two are neighbours and are often sold in one product. The router answers one question: which model for this request? The gateway is the checkpoint every request passes through, where keys, limits, budgets, logging and fallback are enforced. Routing is frequently a feature inside a gateway. A company can run a gateway with no intelligent routing at all, and that is usually the right place to start.
Under the hood
Three families. Rule-based routing uses what is already known: the product feature, the user tier, the length of the input. Learned routing uses a small classifier, embedding similarity or a small language model to predict which model will do; RouteLLM is a well-known open-source framework, trained on preference data. A cascade tries the cheap model first and escalates when a confidence signal or a checking step says the answer is weak; FrugalGPT is the best-known paper on this approach. Available signals: the query, each model's price and latency, response confidence and accumulated feedback. Requirements: a labelled evaluation set for every route, monitoring of the share of traffic sent to each model, and the routing decision recorded in every trace. Pitfalls: the router's own latency and cost; prompts tuned for one model that behave differently on another; prompt caches that split across models; drift when a provider updates a model. Sending traffic to a second provider when the first one fails is fallback, and usually lives in the gateway.