Streaming changes no tokens and saves no compute; it only moves when the user first sees progress, which is most of what fast means.
Why you'd careThe thing you have already noticed
You built the thing, it worked on your laptop, and in production the whole answer appears at once after nine seconds. Or the stream survives short answers and dies at exactly sixty seconds on long ones. Or a user reports a reply that stops mid-sentence, and your logs show HTTP 200 and no error at all. All three are this stage. Before generation starts, the response channel has to be opened, buffering has to be disabled along the entire path, and both sides have to agree on an event protocol and — crucially — on how failure is signalled after the status code has already been sent. Get any of it wrong and the model's output is fine while the experience is not.
In and outWhat goes in, what comes out
| In | The caller's stream flag and Accept header, the protocols supported along the path (HTTP/1.1 chunked encoding, HTTP/2 DATA frames, WebSocket, gRPC), and the buffering configuration of every proxy, load balancer and CDN in between. |
|---|---|
| Process | Write response headers immediately with content-type: text/event-stream and no content-length, disable proxy buffering and Nagle so small writes leave at once, and commit to an event schema. Anthropic's Messages API sends message_start, then per-block content_block_start, content_block_delta and content_block_stop, then message_delta and message_stop. |
| Out | An open, flushing byte stream carrying small JSON events; a terminal contract, since message_delta carries stop_reason and output token usage; and an in-band error channel, because the 200 is already on the wire and cannot be withdrawn. |
The token sequence is preserved exactly: streaming and non-streaming produce identical text for identical sampling, at identical cost. What is lost is the ability to fail cleanly. Once the status line is sent, an error can only arrive as an event, so any client that maps HTTP status to success records a failed generation as a success. Content-length is lost too, so nothing downstream can distinguish a complete response from a truncated one by size.
ConceptThe idea underneath
The one piece of model behaviour that makes streaming honest is autoregression. The model computes p(x_t | x_<t) — the distribution over the next token given everything before it — samples one token, appends it, and repeats. The token emitted at step 5 is never revised at step 40. A prefix of the answer is therefore a real, final prefix, not a preview. This is why streaming a language model is not the same trick as showing a progressively sharpening image, where partial output genuinely is provisional.
That splits latency into two metrics that behave differently. Time-to-first-token is dominated by prefill, which processes the whole prompt in parallel and is compute-bound, so it grows with prompt length. Inter-token latency is dominated by decode, which reads the entire weight matrix from memory to produce one token and is bandwidth-bound, so it barely depends on prompt length. Total time is TTFT + (N_out - 1) * ITL.
The useful consequence is that the two have very different value. Reading speed is roughly 250 words per minute, call it five or six tokens per second. Once a stream exceeds about 15 tokens per second the user is reading behind it and further inter-token improvements are invisible. Time-to-first-token improvements are never invisible.
The rest is plumbing, and the plumbing is where the failures live. A flush must traverse the whole path: the server's socket buffer, Nagle's algorithm coalescing small writes, the reverse proxy's response buffer, the CDN. Any one of them can hold your deltas until the response completes, at which point you have paid for streaming and shipped a non-streaming product.
At a glanceSee it
The response channel opening, flushing deltas, and the terminal event that proves the answer finished.
The knobsHyperparameters and nuance
- streamthe request flag. Anthropic's SDKs recommend streaming above roughly 16k
max_tokens, because a non-streaming request that long can exceed the client's HTTP timeout; the TypeScript SDK scales its default timeout upward for large non-streaming requests specifically to paper over this. Provider- and SDK-specific behaviour. - proxy_buffering / X-Accel-Buffering(nginx) — the single most common way streaming breaks in production. Left on, nginx accumulates the response and delivers it in one piece;
proxy_buffering offor anX-Accel-Buffering: noresponse header disables it. - TCP_NODELAYdisables Nagle's algorithm. With Nagle on, a small delta can wait for an acknowledgement or a roughly 40 ms timer before leaving the host, turning a smooth stream into visible stutter.
- idle timeouton every load balancer and gateway in the path — commonly 60 seconds by default. A long prefill or a long reasoning block produces no bytes for longer than that. Keepalive events exist to reset the timer (Anthropic's Messages API stream can interleave
pingevents), but only if every intermediary forwards them promptly. - flush granularityhow many tokens the server batches into one event. One event per token maximizes responsiveness and multiplies per-event framing overhead; a few tokens per flush is usually indistinguishable to a reader and materially cheaper.
EffectHow this stage moves the answer
Streaming forces you to publish output you cannot retract, and everything downstream inherits that. Moderation, schema validation and stop-sequence handling now run on a prefix, so a product either shows text it may have to delete — the visible retraction users find alarming — or buffers until it is safe and gives up the latency win it opened the channel for. Streaming also changes what the answer is when things go wrong. Many stacks abort generation on client disconnect, so a user who navigates away gets nothing while you have paid for everything produced. And because errors arrive in-band after a 200, a naive client renders a half-finished sentence as the final answer with no error surfaced anywhere. The discipline that fixes this is to treat the terminal event as the only proof of completion: a response without message_stop and a stop_reason is not an answer, it is a fragment.
EvalsWhat it does to your measurements
Time-to-first-token and inter-token latency only exist as measurements if you stream — a non-streaming harness can report end-to-end time and nothing else, and end-to-end time is the number users feel least. The traps are all in the reading loop. Mix cold and warm clients and TTFT includes connection setup on some samples but not others, going bimodal, so the mean describes nothing. Compute inter-token latency as total minus TTFT divided by token count and you fold in every scheduler pause; percentile inter-token latency is what correlates with perceived jank. Consume the stream slower than it arrives and TCP backpressure makes the server look slow. And a harness that stops reading at the last text delta never sees message_delta, so it records zero output tokens and no stop_reason — the run looks clean and its cost accounting is empty.
Failure modesWhen it goes wrong
- Streams token by token locally, arrives as one block in productionan intermediary is buffering the response body; nginx
proxy_bufferingand CDN response buffering are the usual causes. - The connection dies at exactly 30 or 60 seconds, and only on hard promptsa load-balancer idle timeout firing during a long prefill or reasoning block, with no keepalive events reaching the client.
- The client reports success and the text stops mid-sentencethe error arrived in-band after HTTP 200, or the connection dropped; the client checked the status code instead of waiting for the terminal event.
- Token usage logs are zero or missing for streamed requeststhe reader stopped at the last content delta and never consumed
message_delta, which is where usage andstop_reasonlive. - Fast first token, then long stalls mid-answerserver-side scheduling: another request's large prefill is occupying scheduler steps, or your sequence was preempted and recomputed.
PapersWhere this comes from
There is no research literature for the channel itself; it is a protocol question, and the authoritative sources are specifications. Server-sent events and the text/event-stream format are defined in the WHATWG HTML Living Standard, HTTP/2 framing in RFC 9113, and the small-write stall you disable with TCP_NODELAY in RFC 896 (Nagle, 1984). The event schema itself is vendor-specific and documented per provider, not standardized. What does have literature is the pair of metrics streaming exposes:
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong et al., 2024. Separates the two phases precisely because time-to-first-token and time-per-output-token have different bottlenecks and different targets, which is the argument for treating them as distinct metrics rather than averaging them into end-to-end latency.
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal et al., 2024. Shows that a long prefill admitted whole stalls every decode in the batch and that chunking it keeps inter-token latency smooth; this is the mechanism behind mid-stream stalls that have nothing to do with your network.