PagedAttention manages the KV cache in fixed-size, non-contiguous blocks like operating-system memory pages, cutting reservation waste to near zero and raising serving concurrency.
ConceptWhat it is
Every token an LLM has already seen leaves behind key and value tensors that later tokens attend to; this KV cache grows with each generated token and dominates GPU memory during serving. Naive engines reserve one contiguous slab per request sized to the maximum sequence length, so a request that stops early or never reaches its limit strands most of that slab. Across many requests this produces heavy internal and external fragmentation and forces small batch sizes, wasting the very GPU memory that caps how many users you can serve at once.
PagedAttention borrows the operating system's idea of virtual memory paging. The KV cache for each sequence is chopped into fixed-size blocks (pages), each holding the KV for a handful of tokens. Blocks live anywhere in GPU memory, and a per-sequence block table maps logical positions to physical blocks, exactly like a page table. Memory is handed out one block at a time on demand, so waste shrinks to at most one partially filled block per sequence, and freed blocks return to a shared pool for the next request.
How it worksThe mechanics
At each decode step the engine writes the new token's key and value into the sequence's current block; when that block fills, it grabs a fresh block from a global free pool and extends the block table, so allocation tracks actual length rather than a worst-case reservation. The attention kernel is rewritten to read KV through the block table, gathering scattered physical blocks instead of assuming one contiguous run, which is what makes non-contiguous storage transparent to the math. Because blocks are addressable units, identical content can be shared across sequences: a common system prompt or parallel samples of one request point their tables at the same physical blocks and only copy-on-write when they diverge. If the free pool empties under load, the scheduler preempts a sequence and either recomputes or swaps its blocks to host memory, then restores them later.
At a glanceSee it
A naive contiguous slab sized to the max length strands memory two ways — internally in unfilled reserved slots and externally as gaps too small to reuse — while paged blocks cap waste at a single partial block.
When many samples share a prompt, PagedAttention maps them to one set of physical blocks tracked by reference count and copies a block only when a sequence writes to it — copy-on-write that avoids duplicating the prompt per beam.
When to use itWhere it fits
- Serving many concurrent sessions on shared GPUs where memory, not compute, is the ceiling on throughput and batch size.
- Workloads with highly variable or unpredictable output lengths, where fixed max-length reservations would strand most of the cache.
- Patterns that share tokens across requests, such as a large shared system prompt, few-shot prefix, parallel sampling, or beam search that benefit from copy-on-write block sharing.
- Any high-QPS chat or API endpoint where you want to push batch size and GPU utilization without buying more hardware.
When NOT to use itLimits & anti-patterns
- Single-user, single-stream, or batch-of-one inference where there is no concurrency to reclaim and the paging bookkeeping earns nothing.
- Very short, fixed-length generations where the KV cache is tiny and fragmentation was never the bottleneck.
- Deployments pinned to a runtime or custom kernel stack that does not implement paged KV, unless you are willing to migrate serving engines.
- When your true limit is prompt-processing compute or network, not KV-cache memory, so freeing memory does not raise throughput.
Trade-offsAdvantages & costs
Advantages
- Near-elimination of KV-cache waste; wasted memory drops from large max-length reservations to at most one partial block per sequence.
- Substantially higher concurrency and throughput at the same VRAM, because reclaimed memory lets far more requests batch together.
- Cheap sharing of identical KV blocks via copy-on-write, avoiding duplicate storage for shared prefixes and parallel samples.
- Graceful behavior under memory pressure through preemption, recompute, or swap instead of hard out-of-memory failures.
Trade-offs & costs
- Requires a custom attention kernel that reads through the block table, so it is engine-specific rather than a drop-in library call.
- Block-table lookups and non-contiguous gathers add indirection overhead, though in practice it is small relative to the memory won.
- Block size is a tuning knob: too large reclaims less waste, too small adds table and management overhead.
- Swapping or recomputing preempted sequences under heavy load adds latency spikes and scheduler complexity.
ExampleIn the real world
A support-chat API runs a mid-size open model on a single 80GB GPU. With contiguous allocation the engine reserves the full 4,096-token context per request even though most conversations end after a few hundred tokens, so only a dozen or so requests fit in memory at once and the GPU sits underused waiting on memory rather than compute. Switching to a serving engine with PagedAttention, the cache is handed out block by block as each conversation actually grows, and the shared system prompt is stored once and referenced by every session's block table. The engine now keeps far more sequences resident and batched together, so effective throughput climbs several-fold on the same hardware; the team tunes the block size and the fraction of GPU memory given to the cache, then adds preemption headroom so a burst of long conversations degrades gracefully instead of triggering out-of-memory errors.
ToolsHow to implement it
- vLLMthe engine that introduced PagedAttention and the most common way to adopt it.
- NVIDIA TensorRT-LLMoffers a paged KV cache implementation for optimized GPU serving.
- Hugging Face Text Generation Inference (TGI)production serving stack that incorporates paged-attention-style KV management.
- SGLangserving framework whose RadixAttention builds on paged, shareable KV blocks for prefix reuse.
Cost & effortWhat it takes
PagedAttention is an engine feature you adopt, not infrastructure you build: the heavy lifting lives inside vLLM and similar runtimes, so the realistic effort is deploying that serving stack and tuning a few knobs (block size, the fraction of GPU memory allotted to the cache, and preemption or swap policy) rather than writing kernels. The payoff is pure efficiency — more concurrent users and higher throughput per GPU, which directly lowers cost-per-token — with no change to model weights or output quality. Ongoing cost is mostly load testing to set batch and memory limits, plus watching for latency spikes when the system preempts sequences under saturation.