Infrastructure & economics

Time to First Token

TTFT

The time from sending a request until the first token of the model's response arrives; the main measure of how fast an AI application feels.

WHAT MOVES TIME TO FIRST TOKEN · illustrative12,000-token prompt, distant region2.4 s+ fixed 10,000 tokens cached1.1 s+ region close to the users0.6 swith reasoning switched on6.0 s≈ 1 s: the answer seems to start at onceShortens TTFTshort prompt, caching, small model, nearby regionLengthens TTFTlong prompt, queuing, reasoning before the answerOnce the first word appears the wait is over; the rest arrives while the user reads.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/time-to-first-token

In plain terms

You ask a question in a chat and nothing happens; the seconds feel long. Once the first words appear, the wait is over: you are reading, and the rest can take its time. Time to first token measures exactly that silent gap. It corresponds to how quickly someone picks up the phone; how long the call then lasts is a separate matter.

Why it matters

For anything a person watches, TTFT matters more than total response time: an answer that starts within a second and takes ten to finish feels faster than one that arrives complete after six. That makes it the number to write into the requirements for chat, voice and coding assistants. Its limits matter just as much. A good TTFT says nothing about when the answer will be complete; it helps only if the response is streamed; and reasoning models lengthen it by design, because they think before the first visible word.

Example

A retailer's shopping assistant sends a 12,000-token prompt (instructions, a catalogue extract, the conversation), and the first word appears after 2.4 seconds. The team caches the 10,000 tokens that never change and moves the service to a region closer to its customers. TTFT falls to 0.6 seconds. The total time for a full answer falls only from 9 seconds to 7, yet the complaints about a “slow assistant” stop.

Most often confused with

TTFT vs. Latency (total response time)

TTFTThe wait until the answer starts
Latency (total response time)The wait until the answer is complete

TTFT is one component of latency. Total latency adds the time to generate every remaining token. A person reading a streamed answer mostly experiences TTFT; a program waiting for a complete JSON object, or the next step of an agent, experiences only the total. Improve TTFT for conversations and total latency for pipelines, and report the two separately.

Under the hood

TTFT is the sum of: the network round trip and connection set-up; time in the provider's queue; processing of the input tokens (prefill), which grows with prompt length; any reasoning before the visible answer; and, in a full application, everything that runs before the model call, such as retrieval, guardrail checks and tool calls. Levers: shorter prompts; prompt caching, which lets the provider skip reprocessing a repeated prefix; a smaller model; a lower reasoning effort; keeping connections open; a region near the users; and capacity tiers with latency commitments. Measure from the user's device, since server-side figures leave out the network, and track p95 as well as the median, because queuing makes TTFT spike at busy hours. With reasoning models, separate the first thinking token from the first answer token; many products fill the gap with a progress indicator or a summary of the thinking. Related measures: time per output token (inter-token latency) and tokens per second, which describe the speed after the first token.

Written by Mehmet Erkek · Last updated: