Home › What Happens After You Hit Enter › CPU/host-memory KV offload and tiering
Pipeline stage · Operate

CPU/host-memory KV offload and tiering

Move cold parts of the cache down into ordinary system RAM and pull them back over PCIe just before the GPU needs them.

In one line

Host RAM on a serving node is ten to twenty-five times larger than GPU memory and reachable at a few percent of its bandwidth, so offload buys sessions and spends first-token latency.

Why you'd careThe thing you have already noticed

You come back to a chat you left twenty minutes ago, send a one-line follow-up, and wait noticeably longer than for any turn before it — then everything is fast again. The request is short, the model is the same, nothing in the payload explains it. Your session's cached state was demoted to host RAM while you were idle, and the first request back had to drag it home. An 8,000-token session on a 70B model is about 2.5 GB of keys and values; across a PCIe Gen4 x16 link at a realistic 25 GB/s that is roughly 100 ms of pure transfer before the first new token can be computed, and more when the link is busy restoring somebody else. This is the trade the tier exists to make: many more live sessions per machine, paid for in wake-up latency.

In and outWhat goes in, what comes out

InBlocks the manager has marked cold — finished requests whose prefix is still worth keeping, idle sessions, or a preempted sequence's entire cache — plus a pinned host buffer and the sequence's block table.
ProcessCopy the blocks GPU → host asynchronously on a side stream into page-locked memory, mark those logical blocks host-resident in the block table, and return the physical GPU blocks to the pool. On wake, allocate fresh GPU blocks, issue the transfer in chunks, and overlap it with the first layers' compute where the engine supports prefetch.
OutAn effective cache many times GPU capacity, a residency tag per block, and a latency distribution whose tail is set by link bandwidth and contention rather than by anything about the prompt or the model.

The bytes survive exactly — DMA is lossless and a restored block computes identical attention. What is lost is predictability: the same request now has two very different latencies depending on where its state was living, and nothing in the response tells the caller which one they got. Pinned host memory is not free either. It cannot be paged out, so a large staging pool permanently reduces what the operating system can reclaim for everything else on the box.

ConceptThe idea underneath

Most of this is systems engineering rather than machine learning, but one model-side property is doing real work. Unlike weights, which every token needs, and unlike activations, which live for microseconds, KV state is read-mostly, append-only, and accessed in an order known before you need it. At decode step t you know with certainty that layer 1 will read every block of that sequence, then layer 2 will, and so on down the stack. A demand-paged operating system has to guess what comes next; a KV tier does not, so it can prefetch layer 3's blocks while layer 1 is still computing and hide most of the transfer behind work that was happening anyway.

The rest is the memory hierarchy. HBM offers a few terabytes per second across tens of gigabytes of capacity; host DRAM offers hundreds of gigabytes to terabytes reachable at roughly 25 GB/s over a PCIe Gen4 x16 link; NVMe sits below that again. Each step down is about an order of magnitude more capacity and an order of magnitude less bandwidth.

Which reduces the stage to one arithmetic comparison, evaluated per sequence. Restoring costs kv_bytes / link_bandwidth. Recomputing costs prompt_tokens / prefill_throughput. For an 8,000-token session on a 70B model that is roughly 2.5 GB and 100 ms against about a second of prefill, so restore wins by an order of magnitude. Shrink the cache with grouped-query attention on a smaller model, or saturate the link with other traffic, and the comparison flips — which is exactly why serving engines expose both swap and recompute as configurable preemption modes instead of picking one for you.

At a glanceSee it

CPU/host-memory KV offload and tiering diagram

Cold session state is demoted to pinned host memory, then restored or recomputed when the session wakes.

The knobsHyperparameters and nuance

  • kv_offloading_sizeand kv_offloading_backend (vLLM: native or lmcache) — the host-memory KV tier itself. Offloading stays off until you give it a size in GiB; the backend then decides how blocks are staged and fetched back. The older swap_space flag that pinned a CPU pool for preempted sequences belonged to the removed V0 engine and no longer exists.
  • preemption(engine behaviour, not a parameter you set) — the crossover this stage is about, minus the switch. vLLM V1 always preempts by recomputation, so swap-versus-recompute is no longer yours to choose in-engine. Whether shipping bytes over PCIe beats a full re-prefill still depends on your KV bytes per token and your prefill throughput, and that arithmetic is exactly what a KV-offload tier is betting on.
  • cpu_offload_gb(vLLM, default 0) — offloads model weights to host memory, not KV. It is routinely confused with KV offload because of the name and behaves oppositely: it costs bandwidth on every forward pass rather than only on wake.
  • enable_hierarchical_cacheand hicache_ratio (SGLang, ratio default 2.0) — adds a host-memory tier beneath the radix cache; the ratio sizes the host pool as a multiple of the device pool, and hicache_size overrides it with an absolute figure. The companion flags (hicache_write_policy, hicache_io_backend) are implementation-specific and still moving; check them on the version you actually run.
  • pinned versus pageable host allocationpage-locked buffers roughly double achievable PCIe throughput and permit genuinely asynchronous DMA; pageable memory is staged through a driver bounce buffer and serialises against compute. Engines pin their host-side staging buffers by default, which is why a host KV tier is charged against unswappable RAM.

EffectHow this stage moves the answer

Restored bytes are the same bytes, so a woken session answers exactly as it would have. The decision around the restore is where answers change. When the engine picks recompute over swap, the prefix is prefilled again, possibly chunked at different boundaries and certainly in a different batch, and the resumed generation can diverge from where it was heading — from outside that reads as the model repeating a sentence at the seam or changing direction mid-answer. Bandwidth contention spreads the cost sideways: a saturated PCIe link makes restores steal from active decode, so your wake-up slows other users' tokens rather than only your own. And some stacks compress or quantize keys and values on the way out to host memory to halve transfer time. That is a real quality change, visible only on restored sessions and only in long-range recall, which makes it unusually hard to attribute.

EvalsWhat it does to your measurements

Split time to first token by session state before you report it. Offload produces a bimodal distribution — resident sessions and woken sessions — and a single p50 averages two populations that differ by an order of magnitude. Alongside it, track PCIe utilisation, restore bytes per request, and the swap-versus-recompute ratio, which tells you whether the engine thinks the tier is paying for itself. The silent invalidation is a benchmark with no idle time: fire requests back to back and nothing is ever cold, so offload looks free. Insert a 60-second gap between turns in a multi-turn trace and the same suite reports a p99 several times higher. The second trap is hardware generalisation — a result measured over a coherent high-bandwidth host link such as NVLink-C2C does not transfer to a PCIe-attached node, where the restore-versus-recompute crossover sits an order of magnitude away.

Failure modesWhen it goes wrong

  • The first turn after a pause is slow and every turn after it is fastrestore latency on wake; once the state is back in HBM the rest of the session is resident.
  • Enabling offload made p99 worse for everyonerestores contend with active decode for PCIe bandwidth, so the tier buys capacity at the cost of tail latency under load.
  • Host memory climbs until the OOM killer takes the serverpinned staging buffers cannot be swapped out, and a pool sized per GPU multiplies across a multi-GPU node.
  • Offload is configured and nothing is ever offloadedblock reuse or prefix caching is off, so blocks are freed at request end and there is nothing cold left to demote.
  • Long-context accuracy drops only for reactivated sessionsthe stack compresses or quantizes keys and values on demotion, so restored state is lower precision than state that never left the GPU.

PapersWhere this comes from

  • FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUSheng et al., 2023 (arXiv 2303.06865). Formalised offloading across GPU, CPU and disk as a placement and scheduling search, and showed workloads can run far beyond GPU capacity — the framing every later KV tier inherits, though its target was batch throughput rather than interactive wake-up latency.
  • InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache ManagementLee et al., 2024. Showed that speculatively prefetching only the KV entries a layer will actually attend to makes host-memory offload viable at long context, because the binding constraint on this tier is bandwidth rather than capacity.
  • CacheGen: KV Cache Compression and Streaming for Fast Large Language Model ServingLiu et al., 2024. Demonstrated encoding KV state into a compact bitstream that can be shipped and decoded faster than it can be recomputed, which is the direct evidence behind the warning above that some stacks make the transfer cheaper by making the restored state lossy.
A living map of modern AI — kept current every morning