At a glanceWhat Happens After You Hit Enter
The eight phases in execution order; a tool call sends the request back through cache and forward pass rather than starting over.
Where the time goesTwo of the eight phases hold most of the clock
The figure above is the order. This one is the cost. The bar is share of wall clock, not number of stages — Meters is a single stage and costs nothing, Forward pass is nine and costs a quarter of the wait. Click any bar to filter the table below to that phase.
A shape, not a measurement. These proportions are typical of a retrieval-backed call on a hosted model with a warm connection and a cold prompt cache. Your split moves with prompt length, cache hit rate and output length. Nothing here was timed — the claim is the ordering and the concentration, not the numbers. Safety and shaping streams alongside generation rather than after it, so it is drawn faint and carries no percentage.
CompareAll forty-two stages, in the order they run
Arrival, intake, cache, the forward pass, generation, the loops that fetch and think, safety, and what you can measure. Unlike every other table on this site these are not alternatives you choose between — it is one sequence everything passes through. Filter by phase or search by symptom: the pause before the first token, the same prompt giving different answers, a long chat losing the plot.
Every row has a page — what it does, what it costs you, and how to tell when it is the thing biting you.
LandscapeTypes & approaches
Click a highlighted type to open its own page — concept, use case, and diagram.
Why it mattersThree symptoms, and not one of them is the model thinking
You run a support agent. It reads a ticket thread, calls two tools — one for the account record, one for order status — and streams an answer back. Three complaints have arrived this month and none of them has an obvious cause.
The first token takes four seconds, then the rest arrives fast. That gap is not the model deliberating. It is some combination of four named stages: admission control holding your request while the fleet is saturated, prefix-aware routing sending you to a replica that has never seen your system prompt, prompt cache key resolution returning a zero-length match because somebody put the current timestamp in the first line of the system prompt, and then prefill pushing all 9,000 tokens of preamble and ticket history through every layer before token one exists. Each of those leaves a separate number behind — queue wait, replica id, cache hit length, prompt token count — so the four seconds can be attributed rather than argued about. None of them is a model problem.
The same ticket gives a different answer on Tuesday. The obvious suspect is sampling — temperature and top-p turning a distribution into a die roll. The less obvious one is determinism and batch invariance: at temperature 0 with a fixed seed, floating-point addition is not associative, and the reduction order inside a kernel depends on what else happened to be in the batch alongside you. Your greedy regression test can flake because another customer's request arrived. And if a router sits in front of the model, the two calls may not have hit the same model at all.
A long session forgets something the customer said an hour ago. Three candidates, in this order: context compaction replaced the early turns with a summary and the detail was not in it; context assembly and ordering put the fact in the middle of a long prompt, where retrieval accuracy is measurably weakest; or the model uses sliding-window attention in most of its layers, so the token is inside the context window but outside what those layers can see.
The thesis: what sits between your call and the first character is a fixed sequence of forty-two mechanical stages, and most of them contain no intelligence at all — a load balancer picking a replica, a hash table lookup, a memory allocator handing out fixed-size blocks, a string matcher watching for a stop token. Every one of them can be the thing making your output bad. And unlike every other comparison on this site, these are not alternatives you choose between. It is one sequence, and every request goes through all of it.
The eight phasesForty-two stages, and the order is the whole argument
Each phase consumes what the last one produced, which means a decision made early constrains everything after it and cannot be repaired later. Put a volatile field at the top of your prompt during Intake and Cache & reuse has nothing to reuse, so Forward pass must prefill the entire thing, so your time-to-first-token is fixed before the model has computed anything. Fix the retrieval order at Intake and you have already decided what attention is able to find. There is no stage downstream that can undo it. This is why a stage list is more useful than a component diagram: components can be swapped, but a sequence tells you where a symptom can and cannot have originated.
| Phase | What it does | What goes wrong here |
|---|---|---|
| Arrival · 5 | Connection, admission, model routing, replica selection, stream setup — everything before your text is even read. | Queued or shed under load; routed to a replica with none of your prefix resident. Cost lands entirely on TTFT. |
| Intake · 6 | Normalization, system-prompt assembly, context ordering, evidence formatting, chat templating, tokenization. | The wrong template, or an order the model reads badly. Nothing downstream can repair either, and almost nobody logs the rendered string that would show it. |
| Cache & reuse · 5 | Prefix matching, KV cache write, paged block allocation, eviction, host-memory offload. | One volatile field near the top zeroes the hit length. Blocks you counted on were evicted under someone else's load. |
| Forward pass · 9 | Embedding, RoPE, attention and its variants, MoE routing, prefill, continuous batching. | Recall collapses past the trained length or outside the attention window; the wrong kernel backend silently takes a slow path. |
| Generation · 9 | LM head, sampling, penalties, constrained decoding, reasoning budgets, speculation, stops, detokenization. | This is where the words are chosen. Almost every knob that changes what the model says lives in this phase. |
| Loops & tools · 4 | Tool dispatch, parallel fan-out, step and spend budgets, context compaction. | An agent circles; a tool runs with too wide a blast radius; compaction deletes the one detail that mattered. |
| Safety & shaping · 3 | Injection detection on untrusted spans, streaming moderation, stop-reason handling. | A retrieved page gives the model orders; a stream aborts mid-answer; an unread stop reason ships truncated JSON. |
| Meters · 1 | Cache hit rate and cached-token accounting, annotated with deploys — the one stage that watches the pipeline rather than running inside it. | Without it, the prompt edit that doubled your input bill is invisible until the invoice. |
Your support agent does not run this once. It runs Intake through Generation to decide on the tool calls, then again to write the answer. If both tool calls go out in parallel that is two passes; if they go out one after the other it is three, and each extra pass re-prefills everything appended since the last one.
LatencyTwo kinds of compute, bound by different hardware
Prefill pushes your whole prompt through every layer at once. It is parallel and compute-bound, it fills the KV cache, and it produces the first token's logits. Time-to-first-token is prefill, plus whatever Arrival spent getting there. Attention cost grows with the square of prompt length while everything else grows linearly, so doubling the prompt more than doubles TTFT.
Decode then produces one token per step, and each step reads the entire KV cache to produce that one token. It is memory-bandwidth-bound, not compute-bound. Decode sets your tokens-per-second.
Because the two are limited by different resources, they have different fixes, and the fixes frequently oppose each other. TTFT improves when you shorten the prompt, hit the prefix cache, land on a replica that already holds it, or stop queueing behind a saturated fleet. Throughput improves when the KV cache per sequence is smaller — grouped-query attention has several query heads share one key/value head, so eight query heads per KV group is an eight-times smaller cache, and you can read that ratio straight out of a model's config — and when the batch is bigger, and when a mixture-of-experts model activates only a few billion of its parameters per token. But a bigger batch raises tokens-per-second per GPU and lengthens per-token latency for everyone in it. That single dial is the central throughput-versus-latency trade in any deployment, and it is a business decision about which number you sell.
Two decode costs surprise people. The LM head — the matrix that turns the final hidden state into a score for every token in the vocabulary — is hidden size times vocabulary size, so a 256k vocabulary on a 2,048-wide model is roughly half a billion parameters on its own. That is 5–10% of a large model and can exceed 25% of a small one, and it is read in full for every token produced, which at decode can cost as much as several transformer blocks. A “small” model with a huge vocabulary is slower than its parameter count suggests. And speculative decoding trades compute for memory-bandwidth steps: a draft model proposes several tokens and the target model verifies them in one pass. An idle GPU has spare compute and the trade is nearly free; a GPU already saturated by a full batch has none to spare and pays for every rejected draft. Teams enable it, get a beautiful single-user demo, then watch aggregate throughput drop when real traffic arrives.
Now the part that is most often misunderstood. Prompt caching removes prefill for the matched prefix. It does not make decode faster by a single token. On your support agent — four thousand tokens of system prompt and tool definitions, a long ticket thread, a two-hundred-token answer — caching is transformative. On a short prompt with a long answer it does nothing at all. It is also fragile in one specific way: the match is a prefix match, so a single early volatile field takes the hit length to zero with no error and no log line.
QualityMost of the stack cannot change your answer, by construction
Start by ruling out the phase that cannot be responsible. Paged block allocation of the KV cache is bit-exact. FlashAttention computes the same numbers in a different memory order. Continuous batching reorders scheduling. KV cache reuse is mathematically exact — a token's keys and values depend only on the tokens before it, so an identical prefix produces identical state. All of these move latency, cost and concurrency, and leave the text alone. One caveat, and it is the same one from the first section: batch composition can change floating-point reduction order, so “leaves the text alone” means no systematic change, not bit-identical output. Beyond that, if your output got worse and nothing changed in Intake or Generation, the serving layer is usually not your culprit.
The stages that change what the model says are fewer, and most of them are things you own rather than things the provider owns.
- Context assembly and ordering.Where you put a fact decides whether the model uses it. Retrieval accuracy across a long context is U-shaped: strongest at the beginning and the end, weakest in the middle. Stuffing in more retrieved chunks raises recall, lowers precision, and past a point degrades answers as near-miss passages crowd out the right one. Test this before you test models: hold the retrieved set fixed, permute the order, re-run the eval. The spread that produces is frequently wider than the spread between two candidate models on the same prompt.
- Retrieved-context formatting.Chunks get deduplicated, labelled with ids and sources, and wrapped in delimiters. This is the difference between clean citations and the right fact attributed to the wrong document. To find out how much of a retrieval win was really a formatting win, freeze the retrieved set and change only the ids, labels and delimiters. Whatever moves was never retrieval.
- Sampling.Temperature, top-k, top-p, min-p. Raise temperature for variety and you buy more hallucination and more format breakage; drop to greedy for reproducibility and you buy repetitive, occasionally degenerate text. Benchmark gaps of several points can come from sampling settings rather than weights — run the same checkpoint greedy and at temperature 0.7 on the same eval and read the difference — which is why an eval that does not record temperature, top-p and seed is not comparable to anything.
- Repetition, frequency and presence penalties.Turn them up and you get less repetition, then unnatural vocabulary, then dropped articles and broken grammar. Typical
repetition_penaltylives around 1.05–1.2. Never apply them to code or structured output, where correct answers legitimately repeat tokens. - Constrained decoding.Compile the format into a state machine and mask out every token that would break it before sampling, so format validity is 100% by construction. The cost is real: a tight constraint can corner the model until every remaining legal token is wrong and it emits filler to satisfy the shape, and a forced choice list leaves no token available for “I don't know.” Constraint also competes with reasoning — a schema that demands the answer field first makes the model commit before it has spent any tokens getting there — so let it reason in free text and constrain only the final block.
- Reasoning token budgets.The dial between a query that costs a cent and takes two seconds and one that costs a dollar and takes two minutes. On easy queries, high effort produces overthinking and can lower accuracy. Always report accuracy against tokens spent: a model that wins at 30k thinking tokens has not beaten one that ties at 3k.
- Stop criteria.Not a quality lever but a correctness one. A generation that hits
max_tokensstops wherever it happens to be, which leaves a prefix: prose that reads finished and JSON that does not parse. The stop reason is the only signal that separates the two — a capped answer and a completed answer look identical in a chat window and are completely different in a parser. Logfinish_reasonor you are guessing.
Blind spotsThree stages that bite people who have never heard of them
Unicode normalization and text hygiene sits at the very front of Intake and appears in almost no eval suite. It decides the exact byte sequence the tokenizer sees. A ticket thread pasted out of a PDF carries non-breaking spaces, smart quotes, soft hyphens and zero-width joiners, and three things follow. Token counts inflate, because one word becomes three tokens. Cache hit rates fall, because two logically identical prompts differ by an invisible character. And it is a documented injection channel: invisible characters and lookalike letters can carry instructions that no human reviewer will see. The knob is NFC versus NFKC versus nothing, plus a strip policy — and aggressive normalization destroys meaningful formatting in code, poetry and non-Latin scripts. There is no safe default here, only a decision you have or have not made.
KV block eviction policy is why your mental model of prompt caching is wrong. You think of the cache as a property of your prompt. It is a property of the fleet's memory pressure. When the block pool crosses a watermark, blocks are recycled by a policy — least-recently-used, least-frequently-used, leaf-first, priority — and nothing tells you it happened, which is why the same request can be cached at 09:00 and cold at 09:05 with nothing in your code having changed. Both directions cost: evict eagerly and you re-run prefill you already paid for; evict lazily and running sequences starve and get preempted instead. The practical consequence is that cache hit rate is a curve against load, not a number you measure once after a good test.
Incremental detokenization is the last stage before a human sees anything, and there is no model in it at all. Token ids come back one at a time, but a single character can be split across several tokens, so the decoder must hold bytes back until they form something valid. Get the window wrong and emoji and non-Latin scripts appear as a replacement character mid-stream and then silently correct themselves — which shows up as mysteriously low multilingual eval scores and nowhere else. It is also where stop-sequence matching happens, against the decoded text rather than token ids. That mismatch is exactly why a stop string sometimes fails to fire and why the last few characters of an output occasionally look chopped.
What these three share: they sit at phase boundaries, they have no owner on most teams, and every one of them fails without raising an error.
The gridRead it left to right, not as a shortlist
The grid below is the same forty-two stages in pipeline order, and there is no row you opt out of. Start on the “Why you'd care” row: every stage has one line there naming a symptom you have already seen in production, which is the fastest way to find the three or four stages that concern you. Filter by family to read one phase end to end. Every column header opens a full page — what the stage does, what goes in and what comes out, the concept underneath it, the knobs, what turning each knob costs you, and how the stage shows up in evals.
DebuggingWhat this changes: a symptom tells you which phase to open
The reason to hold forty-two stages in your head is that it converts a vague complaint into a short list of places to look. Use this backwards — start at the symptom, not the stage.
| Symptom | Phase to open | First thing to check |
|---|---|---|
| First token is slow, the rest streams fine | Arrival → Cache & reuse → Forward pass | Prefix cache hit length, then where the first volatile field sits in your prompt, then raw prompt length. |
| First token is fast, the answer crawls | Forward pass → Generation | Batch size and concurrency, KV cache per sequence, whether speculation is on under load, output length. |
| Same prompt, different answer | Generation | temperature and top_p first; then batch invariance at temperature 0; then whether a router sent the two calls to different models. |
| Broken or half-written JSON | Safety & shaping | The stop reason. It is a max_tokens cap, a stop sequence, or a content filter — and you are almost certainly not reading it. |
| Long session forgets an earlier detail | Loops & tools → Intake | Compaction trigger threshold and what is pinned, then where the fact lands in the assembled order. |
| Right fact, wrong source cited | Intake | Chunk labelling and delimiter schema. This is a formatting bug, not a retrieval bug. |
| Cost jumped after a deploy, traffic unchanged | Cache & reuse → Meters | Cache hit rate with deploy annotations. Look for a timestamp, user id or shuffled example that moved up the prompt. |
| An agent looped overnight and ran up a bill | Loops & tools | Step, token, spend and wall-clock caps, and whether loop detection exists at all. |
| Text appeared, then vanished and became a refusal | Safety & shaping | Streaming moderation abort — the guard caught up with the generator. |
| Fine in dev, degrades under real traffic | Arrival | Admission control thresholds. Measure goodput — requests finished inside the SLO — not raw throughput. |
| Emoji or non-Latin text mangles mid-stream | Generation | Detokenizer window size and skip_special_tokens. |
| Open model scores below its published numbers | Intake | The chat template. Somebody rendered the turns differently from how the model was post-trained. |
Two habits fall out of this. First, locate before you tune: most wasted debugging
is a Generation knob being turned to fix an Intake problem, or a bigger model being bought to fix a
cache miss. Second, log the boundary values — cache hit length, prompt token
count, finish_reason, tool step count, sampling parameters. Those five fields settle the
third column above for half the rows in the table before you run a single experiment.
GlossaryKey terms
CheckCheck your understanding
The first token takes four seconds and the rest streams fast. What is that?
Not the model deliberating. Some mix of admission control queueing you, routing sending you to a replica with none of your prefix cached, a prompt-cache miss, and prefill pushing the whole prompt through every layer. Each leaves its own number behind — queue wait, cache hit length, prompt token count — so it can be attributed rather than argued about.
Temperature is 0 and the seed is fixed. Why is the output still not identical?
Floating-point addition is not associative, and the reduction order inside a kernel depends on what else was in the batch alongside you. Another customer's request can change your result. If a router sits in front, the two calls may not even have hit the same model.
A long conversation forgets something said an hour ago. Which stage?
Three candidates in order: context compaction replaced the early turns with a summary that dropped the detail; context ordering put the fact mid-prompt where retrieval is measurably weakest; or the model uses sliding-window attention in most layers, so the token is inside the context window but outside what those layers can see.
Why does putting a timestamp at the top of a system prompt cost so much?
Prompt caching matches on an exact prefix. One volatile field in the first line makes the hit length zero, so every request prefills the entire prompt from scratch. The fix is ordering, not spend: stable content first, volatile content last.
What changedWhat changed here
Nothing in the daily brief has touched this page since 2026-09-25. The sweep runs every morning and checks every page on this site; when it finds something for this one, it lands here.
Three kinds of claim, strongest first. Signal runs every morning.