⚙️ · Operate

KV cache

The reuse trick that turns wasteful re-computation into fast token-by-token LLM generation

In one line

KV cache stores the attention keys and values already computed for earlier tokens so each new decode step processes just one token instead of re-running the whole sequence.

ConceptWhat it is

A transformer generates text one token at a time, and at every step each new token must attend back over all the tokens before it. That attention is built from two projections of the earlier tokens — the keys and the values. Under causal masking these never change once computed: token 40's key is the same whether you are generating token 41 or token 400. Recomputing them at every step means re-running the entire prefix through every layer again and again, which is pure waste.

The KV cache is the fix: during generation the model keeps every layer's keys and values in GPU memory and simply appends the new token's pair each step. This converts decode from re-processing the whole sequence per token into processing a single token per step, which is why it is the default in every autoregressive inference stack. The price is memory — the cache grows linearly with context length and batch size, and at long context it can rival or exceed the model weights themselves.

How it worksThe mechanics

On the first pass, called prefill, the model runs the full prompt through every layer and computes the key and value projections for all prompt tokens, writing them into the cache. Generation then enters the decode loop: for each new token the model computes query, key, and value for that single token only, appends the new key and value to the cache, and runs attention where the lone query attends over every cached key and value across all layers and heads. Because the past keys and values are read rather than recomputed, each step forwards one token instead of the entire sequence, and per-token latency stays roughly flat. The loop repeats until a stop token or a length limit, with the cache growing by exactly one position per generated token.

At a glanceSee it

KV cache diagram
KV cache diagram 1

The cache pays off — and stays correct — only because causal masking freezes every past key, cutting per-step cost from growing-with-length to near constant.

KV cache diagram 2

When the cache outgrows memory, four standard levers shrink or manage it — and sliding-window eviction is the one that permanently forgets.

When to use itWhere it fits

  • Any autoregressive, token-by-token generation — chat, completion, code, agents — where you decode more than a single token.
  • Interactive or streaming responses where per-token latency must stay low and flat as the output grows longer.
  • High-throughput serving, where not recomputing the prefix frees compute so larger batches fit in the same time budget.
  • Long-prefix workloads such as RAG or full-document prompts, where the same large context is attended to on every decode step.

When NOT to use itLimits & anti-patterns

  • Encoder-only or embedding workloads (BERT-style classification, retrieval embeddings) that run one forward pass with no decode loop — there is nothing to reuse.
  • Single-shot scoring or reranking where you never generate a sequence token by token.
  • When the real bottleneck is already the KV memory itself — reaching for a bigger per-request cache is the wrong move; you need GQA/MQA, KV quantization, or paging instead.
  • Ultra-long context on tight VRAM where a full cache would dwarf memory — sliding-window attention, cache eviction, or offloading may be required rather than a naive full cache.

Trade-offsAdvantages & costs

Advantages
  • Eliminates redundant recomputation: each decode step forwards one token, not the whole prefix, giving near-constant per-token cost.
  • Effectively mandatory and on by default in every serious inference stack, so the speedup is essentially free to obtain.
  • Provides the substrate for further wins — prefix and prompt caching, paged allocation, and streaming all build on the cached K and V.
  • Amortizes the expensive prefill across every generated token, making long-context decoding practical.
Trade-offs & costs
  • Memory footprint grows linearly with sequence length and batch size, and at long context can rival or exceed the model weights.
  • KV memory, not compute, usually becomes the hard ceiling on batch size, concurrency, and maximum context length.
  • Naive contiguous allocation fragments GPU memory and strands capacity, requiring paging schemes to recover it.
  • Adds serving complexity: per-request allocation, eviction, and correct sizing across a dynamic batch.

ExampleIn the real world

Consider an assistant answering questions over an eight-thousand-token document. Prefill runs that document through the model once, writing keys and values for all eight thousand positions into the cache. The answer then streams out token by token: each decode step computes projections for just the one new token, appends them, and attends over the cached context, so the tenth output token and the four-hundredth cost about the same to produce instead of latency climbing with position. Under batching, the server tracks that every concurrent conversation carries its own growing cache, and it is total KV memory — not arithmetic throughput — that decides how many simultaneous sessions fit on the GPU. When a long chat's cache threatens to exhaust VRAM, the serving layer is what must page, evict, or reject, not the model math.

ToolsHow to implement it

  • vLLMPagedAttention stores the KV cache in non-contiguous pages to cut fragmentation and raise concurrent-batch throughput.
  • Hugging Face Transformersthe use_cache flag plus the DynamicCache and StaticCache classes implement per-request KV caching for generate().
  • NVIDIA TensorRT-LLMproduction kernels with paged KV cache and KV quantization (for example FP8 or INT8) for high-throughput serving.
  • SGLangRadixAttention shares KV-cache prefixes across requests so common system prompts are computed once and reused.

Cost & effortWhat it takes

KV cache is not a product you buy — it is default behavior, and its cost is GPU memory. The footprint scales as roughly two (keys and values) times the number of layers, times the number of KV heads, times head dimension, times tokens, times batch, times bytes per element, so long contexts and large batches make the cache a major or even dominant memory consumer that, more than FLOPs, sets the ceiling on concurrency and maximum context. Turning it on takes essentially zero effort since every serving stack enables it by default; the real engineering investment is in optimizing it — GQA/MQA to shrink the KV heads, quantized caches, paged allocation, and cross-request prefix reuse — once you are pushing throughput or context length.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Proposes chunk-level KV-cache reuse with deviation-guided recomputation to cut prefill latency when prompts lack a shared prefix.

    arXiv cs.AI · 24 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning