In plain terms
Think of the reception desk of a large office building. Every visitor comes through one door, shows a badge, is logged and is directed to the right floor. An AI gateway is that desk for model traffic. Applications and agents stop calling model providers directly, each with its own key; they call the gateway, and the gateway calls the providers on their behalf.
Why it matters
Without one, every team holds its own provider keys, nobody can say what the organisation spends per product, and a provider outage takes down each application separately. A gateway gives one place to set budgets, see usage, apply data rules and add or switch providers without touching application code. The costs: it is a new piece of critical infrastructure, since everything stops when it stops; it adds a little latency; and it sees every prompt, so its own logs become sensitive data that need access control and retention rules.
Example
A software company has 14 teams using three model providers through 60 separate API keys. After the move to a gateway, each team has one virtual key with a monthly budget. In the first month the dashboard shows that a single internal tool accounts for 38% of all spend. When one provider has a two-hour outage, the gateway sends that traffic to a second provider and no application goes down.
Most often confused with
AI Gateway vs. API gateway
A classic API gateway counts requests and cannot see inside them. For model traffic that is too coarse: one request can cost a thousand times more than another, the content may hold personal data, and the response arrives as a stream. An AI gateway meters tokens and cost, inspects prompts and responses, and can swap one model for another. Several AI gateways are extensions of API gateway products.
Under the hood
Typical functions: one API format in front of many providers, often the OpenAI-compatible format; virtual keys that map to teams or applications while the real provider keys stay inside the gateway; rate limits and budgets counted in tokens and in money; retries, load balancing and fallback to another model or provider; response caching, exact or semantic; guardrails on input and output, such as masking personal data; and logs, traces and cost attribution for every call. Deployment: a self-hosted proxy, a managed service, or a feature of an existing API-management or cloud platform. Products include LiteLLM, Kong AI Gateway, Cloudflare AI Gateway and OpenRouter. Increasingly the same layer governs the tool traffic of agents, as a gateway in front of MCP servers. Design points: measure the added latency; run the gateway redundantly, because it is a single point of failure; decide whether prompts are logged in full, masked or left out; and block direct calls that bypass it, because otherwise it governs only the teams that volunteer.