Infrastructure & economics

Streaming

Response streaming

Delivering a model's response piece by piece as it is generated, so that the user sees the first words within moments and does not have to wait for the complete answer.

Without streamingthe answer arrives whole at the endSecond 1 · blank screen, spinnerSecond 6 · still waitingSecond 12 · the whole answer appearsWith streamingthe answer arrives as it is writtenSecond 1 · first sentence on screenSecond 6 · half the answer being readSecond 12 · answer completeThe total time is the same: 12 seconds. What changes is the moment the user starts reading.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/streaming

In plain terms

A model writes its answer one token at a time. Without streaming, the application holds everything back until the last token is written and then shows the whole text at once. With streaming, each piece is passed on as soon as it exists, which is why answers in chat assistants seem to be typed out in front of you. The model works at the same speed either way; you simply start reading earlier.

Why it matters

Streaming is the cheapest way to improve perceived speed: the total time stays the same, the wait before something appears drops from many seconds to about one, and people can stop a bad answer early. It is the default for chat. It has a price as well. Text reaches the user before anyone can check the whole of it, so output filters have to work on a moving stream or take back words already shown. Errors can arrive in the middle of an answer. And programs that need a complete result, such as a JSON object, gain nothing from it.

Example

A legal-research assistant needs 12 seconds to write a 600-token answer. Without streaming, the lawyer watches a spinner for 12 seconds. With streaming, the first sentence appears after 0.8 seconds and the text keeps flowing; after 3 seconds the lawyer sees that the question was misunderstood, stops the answer and rephrases. The full answer would still have taken 12 seconds.

Most often confused with

Streaming vs. Batch processing

StreamingOne answer, delivered live, piece by piece
Batch processingMany requests, processed later at lower cost

The two sit at opposite ends of urgency. Streaming serves a person who is waiting now and wants to watch the answer form. Batch processing serves work nobody is waiting for: thousands of requests handed over together, completed within hours, commonly at half the price. In data engineering “streaming versus batch” means something else, namely how data moves through pipelines; here the subject is how a model's response is delivered.

Under the hood

Most model APIs stream over server-sent events (SSE): the client sets a streaming flag, the server keeps the HTTP connection open and sends a series of small events, each carrying a fragment of text, followed by a final event with the stop reason and the token usage. Two-way, real-time interfaces such as voice use WebSockets or WebRTC. What the developer has to handle: assembling the fragments, including partial tool calls and partial JSON that stays invalid until the last piece arrives; connections that drop in the middle of a response; proxies and gateways that buffer the response and quietly undo the streaming; and cancellation. With most providers, cancelling stops generation, and only the tokens produced up to that point are billed. Moderation and guardrails either inspect the stream in chunks or hold back a short buffer. Billing is the same as for a call without streaming. Streaming improves time to first token as the user experiences it and leaves total latency unchanged. Providers recommend streaming for requests with long outputs, because a connection that sits idle while the model works may time out.

Written by Mehmet Erkek · Last updated: