Infrastructure & economics

Rate Limit

A cap set by a model provider on how many requests and tokens a customer may send within a period, usually a minute; requests above it are rejected until capacity frees up.

REQUESTS PER MINUTE · illustrativelimit: 80 requests a minutethe excess is rejectedminutes →HTTP 429limit exceeded, request rejectedWait, then retry1 s → 2 s → 4 s (backoff)Lasting fixqueue, batch, a higher tierThe account limit is 400,000 tokens a minute; at 5,000 tokens per request it is reached at 80 requests.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/rate-limit

In plain terms

A rate limit works like the turnstiles at a busy station: people are let through at a set pace so that the platform never overcrowds, however long the queue outside. A provider's GPUs are shared by all its customers, so each account gets an allowance, for example so many requests and so many tokens per minute. Stay within it and nothing happens. Go over it and the extra requests are turned away with an error and have to be sent again a moment later.

Why it matters

A rate limit is a capacity ceiling that does not appear on the price list. A pilot never notices it. A launch to the whole company, a month-end peak or a single agent working through ten thousand documents can reach it within minutes, and users then see errors or long waits. Limits rise with spending history, on request or by contract, and that can take time. Compare your limits with the expected peak load before launch, as you would check any other supplier's capacity, and decide what the product does when it is told to wait.

Example

An insurer's account allows 1,000 requests and 400,000 input tokens per minute. Its claims assistant sends about 5,000 tokens per request, so the token limit is reached at 80 requests a minute, long before the request limit. On the Monday after a storm, traffic climbs to 130 requests a minute. About 50 a minute are rejected with error 429 and retried after a short wait, and adjusters see delays until the limit is raised.

Most often confused with

Rate Limit vs. Spend limit

Rate LimitCaps the pace: requests and tokens per minute
Spend limitCaps the total: money or usage per month

A rate limit concerns pace and clears by itself within seconds or minutes; retrying after a short wait is the right response. A spend limit, or quota, concerns the total for the month: once it is reached, every request fails until the period resets or someone raises the cap, and retrying achieves nothing. The two often return the same error code, so read the error message before deciding how to react.

Under the hood

Limits are usually expressed as requests per minute (RPM) and tokens per minute (TPM), sometimes with separate limits for input and output tokens and with daily caps. They apply per model or model family, at the level of the organisation or project. Providers place accounts in usage tiers that rise with payment history; large customers negotiate limits or buy reserved capacity. Many use a token-bucket algorithm: the allowance refills continuously, so a burst can exhaust a per-minute limit within seconds. An exceeded limit returns HTTP 429, usually with a retry-after header and headers showing the remaining allowance; a separate error signals that the provider itself is overloaded. Standard client behaviour: exponential backoff with random jitter, a cap on the number of retries, and respect for retry-after. Design measures: queue and smooth the traffic, set a realistic maximum output length, use prompt caching (some providers leave cached tokens out of the count), send work that can wait through the batch interface, which has its own limits, and spread load across models or providers through a gateway. Gateways also set their own limits per team to protect a shared budget.

Written by Mehmet Erkek · Last updated: