Home › What Happens After You Hit Enter › Prefill (processing the prompt)
Pipeline stage · Operate

Prefill (processing the prompt)

Every token of your prompt goes through the whole model at once, filling the cache — this is the pause before the first word appears.

In one line

Prefill is the only compute-bound phase of inference, which is why time-to-first-token responds to prompt length and raw FLOPs while decode speed responds to neither.

Why you'd careThe thing you have already noticed

You paste a long document and nothing happens for several seconds, then text streams out at full speed and never slows down. That gap is prefill. Every token of the prompt must pass through every layer before the first output token can exist, because the model needs a key and a value for each position in every layer's cache. It is one forward pass over the whole prompt and it is the one part of serving that genuinely saturates a GPU. It is also where your bill goes. Input tokens are cheaper per token, but a retrieval-augmented prompt with tool schemas, a system prompt and a dozen retrieved chunks can be fifty times longer than the answer, so most of the arithmetic in your deployment happens before generation starts.

In and outWhat goes in, what comes out

InThe full prompt token sequence minus any prefix already present in the cache, plus sampling parameters. Length S ranges from a handful of tokens to hundreds of thousands; this is the only stage where S is ever large.
ProcessOne forward pass over all S positions at once. Every layer computes Q, K and V for all tokens, runs causal attention across the whole sequence, and applies the feed-forward block. K and V for every position are written into the cache as they are produced.
OutA fully populated KV cache covering S positions, and logits for the final position only. Prefill produces exactly one token of output no matter how long the prompt was.

Everything except the KV cache is discarded: activations, intermediate attention outputs, and the hidden states of every position but the last. That is why asking for logprobs across the prompt costs extra. The LM head is normally applied only to the final token to avoid building an S-by-vocabulary logits tensor, which at 100,000 tokens and a 128,000-token vocabulary would be about 25 GB in bf16.

ConceptThe idea underneath

Prefill is teacher forcing at inference time. During training the model sees a whole sequence and predicts every next token in parallel, because the causal mask stops any position seeing its own future. Prefill exploits the same property for a different purpose: the prompt is already known, so all S positions can be evaluated in one pass rather than S sequential ones.

Cost has two terms. The dense part — every projection and every feed-forward — is about 2 * n_params * S FLOPs, linear in prompt length. The attention part is roughly 4 * n_layers * S^2 * d_model, quadratic. Which dominates depends on the model. For an 8B model with 32 layers and d_model 4096 the two terms are equal at around 30,000 tokens; below that, prefill FLOPs are close to linear in S.

That is worth stating plainly, because the superlinear time-to-first-token people observe at 2,000 or 8,000 tokens is usually not the S-squared term. It is queueing behind other requests, chunk scheduling, memory-allocation pressure, or a kernel falling off its fast path. Genuine quadratic behaviour does arrive, but much later than the folklore suggests.

Because prefill is compute-dense it can push a GPU to 40 to 60 percent of peak FLOPs, where decode sits in the low single digits. That asymmetry causes the scheduling problem this stage is famous for: one long prefill occupies the entire device for a step, and every user currently streaming stalls. Chunked prefill splits the prompt into slices of a few thousand tokens and interleaves decode steps between them, trading a little TTFT for a stable stream. Disaggregated designs go further and run prefill on separate hardware entirely.

At a glanceSee it

Prefill (processing the prompt) diagram

The whole prompt goes through the model once, filling the cache and producing one token.

The knobsHyperparameters and nuance

  • max_num_batched_tokensvLLM's per-step token budget, which with chunked prefill is effectively the chunk size. Small values protect streaming latency for existing users and raise TTFT for the new request; large values do exactly the opposite.
  • enable_chunked_prefillon by default in recent vLLM. Turned off, a long prompt monopolises a scheduler step and every concurrent stream visibly stutters.
  • long_prefill_token_threshold and max_num_partial_prefillshow many long prompts may be in flight partially at once. These bound head-of-line blocking when several large documents arrive together.
  • prompt caching breakpointsAnthropic's cache_control markers, automatic prefix caching on other APIs, vLLM's --enable-prefix-caching. The largest available TTFT lever, since a hit skips prefill for the matched prefix entirely. Provider-specific in both syntax and minimum cacheable length.
  • max_model_lenthe longest prompt the server accepts. Set above what your KV memory can support and you convert a clean rejection into an out-of-memory failure under load.

EffectHow this stage moves the answer

Prefill does not change what the model says. It changes what the model is allowed to be given. Every content-losing decision in a production prompt exists because of this stage's cost: truncating chat history, cutting retrieval from twenty chunks to five, dropping few-shot examples, summarising earlier turns instead of including them. Those are the levers that actually move time-to-first-token and per-request cost, and each removes evidence the model would have used. The observable result is an answer that is fluent, reasonable and missing the specific thing that was cut. A subtler effect follows chunked prefill: a prompt is processed in slices whose boundaries depend on what else was running, and floating-point accumulation differs between a prompt processed in one chunk and the same prompt processed in four. At temperature 0 that is occasionally enough to flip a near-tied token and send the completion somewhere else entirely.

EvalsWhat it does to your measurements

Prefill sets time-to-first-token, the SLO most deployments miss first under load, and it dominates total cost for every workload whose prompts dwarf their outputs — classification, extraction, retrieval question answering, LLM-as-judge. The failure that ruins benchmark runs is prefix caching. Harnesses reuse the same system prompt and few-shot block across thousands of examples, so every request after the first hits a warm cache and reported TTFT can be an order of magnitude better than any real user will see. Report cold and warm numbers separately, or disable prefix caching for the latency run. A related trap is measuring at concurrency 1, where prefill never queues. TTFT under load is a scheduling number rather than a model number, and the two diverge sharply the moment one long prompt arrives.

Failure modesWhen it goes wrong

  • Time-to-first-token excellent in development and terrible in productiondevelopment repeats one prompt and hits the prefix cache, while production prompts are unique.
  • Every active stream stutters when one large request arrivesunchunked prefill occupying a whole scheduler step while decodes wait.
  • Out-of-memory on a single long prompt when the same total tokens spread over many requests is finepeak activation memory scales with the tokens in one forward pass, not with the total served.
  • Cost far above the estimate despite short answersper-request prompt overhead from system prompts, tool schemas and retrieved context, multiplied by request volume.
  • Intermittent context-length rejections on prompts that usually fitvariable-length retrieved context or tool definitions pushing the rendered prompt over max_model_len only some of the time.

PapersWhere this comes from

  • Efficiently Scaling Transformer InferencePope et al., 2022. arXiv:2211.05102. Separated prefill and decode analytically, showed prefill is compute-bound and decode memory-bound, and derived the partitioning strategies serving stacks still follow.
  • SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked PrefillsAgrawal et al., 2023. arXiv:2308.16369. Introduced chunked prefill, the direct fix for a long prompt stalling every ongoing generation, now default in the major serving stacks.
  • DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingZhong et al., 2024. Argued the two phases have incompatible resource profiles and should run on separate hardware, which is increasingly how large deployments are actually built.
A living map of modern AI — kept current every morning