In plain terms
Latency is how long you wait. With a language model the wait has two parts, as it does in a restaurant: the time until the first plate arrives, and the time until everything is on the table. The first depends on how busy the kitchen is and how long your order was. The second depends on how much you ordered, because the model produces its answer piece by piece, at a steady speed.
Why it matters
Latency decides what a model can be used for. A chat assistant seems broken after a few seconds of silence; a voice assistant, after a fraction of a second; a nightly batch job can wait for hours. Latency is traded against quality and cost: larger models and reasoning modes give better answers and slower ones, and an agent that takes twenty steps multiplies every delay by twenty. Set the acceptable wait for each use case before choosing the model, and measure the slow requests; averages hide the ones that make users give up.
Example
A bank tests two models for its call-centre assistant. The larger one produces 50 tokens a second and the smaller one 150. After a wait of about one second for the first token, a typical 300-token answer takes 7 seconds in total on the larger model and 3 on the smaller. Agents on live calls will not wait 7 seconds. The bank uses the smaller model in the call centre and keeps the larger one for written complaints.
Most often confused with
Latency vs. Throughput
Latency is measured per request, in seconds. Throughput is measured across the system, in requests or tokens per second. The two pull against each other: a server that processes many requests together produces more tokens per second in total, and each user waits a little longer. Interactive products are tuned for latency; bulk jobs are tuned for throughput and cost.
Under the hood
Total latency ≈ time to first token + output tokens ÷ generation speed (tokens per second). Time to first token covers the network, the queue, the processing of the input and any thinking the model does before it answers. Generation time grows linearly with output length, which makes output length the largest lever for most requests. Other levers: a smaller or faster model, shorter prompts, prompt caching, a lower reasoning effort, running independent steps in parallel, hosting in a region close to the users, and priority capacity tiers where providers offer them. Streaming does not shorten the total; it lets the user start reading sooner. Report percentiles (p50, p95, p99): the slowest few percent of requests shape user experience and timeouts. Measure under realistic load, because latency rises as the provider, or your own server, gets busy. In agents and pipelines, latency is the sum of every model call and tool call in the chain; budget it per step. Voice interfaces need the first audio within a few hundred milliseconds, which usually calls for small models or dedicated speech-to-speech models.