Cache hit rate is the rare metric where latency and cost move together, and a single timestamp early in your prompt can silently drive it to zero.
Why you'd careThe thing you have already noticed
Your app was fast and cheap yesterday. This morning you shipped a one-line change to the system prompt — added the current date, bumped a build ID, reordered the tool list — and now time-to-first-token has tripled and the input line on the bill has jumped while traffic, output length and answer quality are unchanged. Nothing throws. Nothing logs an error. Your latency traces just show prefill dominating where it used to be nearly free. That is the prompt cache going cold. Every request is now recomputing attention over a prefix that used to be served from an existing KV cache, and those tokens bill at full input rate instead of roughly a tenth of it. Cache hit rate is the meter that would have told you inside one deploy window instead of at the end of the month.
In and outWhat goes in, what comes out
| In | Per-request usage records. On Anthropic's Messages API that is usage.cache_read_input_tokens, usage.cache_creation_input_tokens and usage.input_tokens. Self-hosted, it is the serving engine's prefix-cache hit and query counters from its metrics endpoint. Alongside each record: a timestamp, a route label, and a deploy or prompt-template version tag. |
|---|---|
| Process | Bucket requests by window and route. Hit rate is cache_read / (cache_read + cache_creation + input_tokens) — that denominator is the whole prompt, because input_tokens reports only the uncached remainder, not the total. Reprice each bucket at roughly 0.1× for reads and 1.25× or 2× for writes, then join the series against the deploy log. |
| Out | A hit-rate time series with deploy markers, a TTFT distribution split into warm and cold requests, and a three-way input-cost breakdown per route — full price, write premium, read discount. Plus an alarm that fires when hit rate falls by more than a set fraction inside a single deploy window. |
Preserved: exact token counts, so both cost and skipped prefill are attributable per request with no sampling error. Lost: everything about why. The usage record is a scalar. It never tells you which breakpoint matched, at which block the prefix diverged, or what the previous request's rendered bytes actually were. Unless you logged the fully rendered prompt yourself, that diff is unrecoverable — and cache residency cannot be enumerated from outside the server.
ConceptThe idea underneath
Attention is causal, and everything here follows from that. Write one head as softmax(QK^T / sqrt(d_k) + M) V. Q, K and V are the query, key and value projections of the hidden states; d_k is the per-head dimension the dot products are divided by so the scores do not blow up; M is the causal mask that sets every future position to negative infinity. Because of M, the K and V rows for position i are a function of positions 0 through i and nothing else. Two requests that begin with the same tokens therefore produce mathematically identical keys and values across that shared span, at every layer, no matter how differently they continue. That is the whole basis of the cache: prefill can be skipped for any prefix already computed.
The catch is the word prefix. Positional information is baked into K — rotary embeddings rotate each key by an angle proportional to its absolute position — so a cached span is only valid at the exact offset it was computed at. Change one byte at position 40 and every key and value after it is not stale, it is wrong. That is why a single early timestamp invalidates fifty thousand downstream tokens rather than forty.
Serving systems index this with a hash chain over fixed-size token blocks: block n's key is a hash of its token IDs combined with block n−1's key. Lookup walks the chain until it misses, so hits are quantised to block boundaries — with a 16-token block, a 100-token shared prefix yields 96 cached tokens, not 100. There is no learning and no approximation anywhere in this. The metric measures arithmetic you avoided, not anything the model did.
At a glanceSee it
How a prompt's block hashes resolve to a cache read or a full recompute, and what the usage record reports.
The knobsHyperparameters and nuance
- cache_controlthe Anthropic Messages API breakpoint,
{"type": "ephemeral"}, placed on a content block. Maximum four per request. Placed at the end of the whole prompt it writes a unique entry every request and reads none; it belongs at the end of the shared span, before the varying question. - ttl(inside
cache_control) —5mby default,1hoptional. The write premium is 1.25× at five minutes and 2× at one hour, so break-even is two requests versus three. The hour TTL only pays for bursty traffic with gaps; on continuous traffic it is a doubled write cost buying nothing. - modela caching knob whether you meant it to be. Caches are model-scoped, so switching models mid-conversation zeroes the hit rate outright, and the minimum cacheable prefix is model-specific and non-monotonic: 512 tokens on Claude Opus 5, 1024 on Opus 4.8 and Sonnet 5, 2048 on Opus 4.7, 4096 on Opus 4.6 and Haiku 4.5. Below the floor nothing caches and nothing errors.
- enable_prefix_caching(vLLM) — when off, hit rate is structurally zero no matter how well-shaped your prompts are. The default flipped to enabled in the V1 engine, so check the version rather than assuming; disabling it is now the explicit flag.
- block_size(vLLM, default 16 on CUDA) — the granularity of a hit. Larger blocks mean less metadata and fewer lookups but coarser matching, so short shared prefixes round down harder; smaller blocks catch more partial prefixes at the cost of more hashing and bookkeeping per request.
- gpu_memory_utilization(vLLM, default 0.9) — sets how much of the card becomes KV pool. Lower it and blocks are evicted sooner, so hit rate falls under unchanged traffic; raise it too far and you trade OOM headroom for cache residency.
EffectHow this stage moves the answer
A warm cache does not change the answer, which is exactly why this stage is dangerous: it changes answers through the design decisions the hit-rate number pushes you toward. Teams chasing hit rate freeze the system prompt, which usually means removing the current date and the user's context from it — and the model then starts answering time-sensitive questions from its training prior instead of from today. They stop swapping tool sets per mode and leave every tool loaded at once, which raises wrong-tool selection. They move retrieved documents after the breakpoint, changing where evidence sits relative to instructions, which changes how strongly the model weights it. The failure runs the other way too: at low hit rate under load, prefill queues, the scheduler starves decode, and users read the stall as the model thinking. The usual reaction is to trim max_tokens or drop reasoning effort — and that genuinely degrades the answer.
EvalsWhat it does to your measurements
The metric that moves is time-to-first-token, and it moves so much that it will quietly invalidate your run. Eval suites are the ideal cache shape — one large shared preamble, one varying question — so the first pass writes the cache and every pass after reads it. Measured TTFT drops several-fold between run one and run two on byte-identical inputs, which looks exactly like an optimisation landing and is only cache warmth. Report cold and warm numbers separately, or discard the first pass on purpose. Concurrency compounds this: an entry only becomes readable once the first response begins streaming, so a harness that fires N identical-prefix requests in parallel has every one of them miss, and reports worse hit rate and higher cost than the same suite run serially. Quality scores should be invariant to cache state. If accuracy shifts between warm and cold runs, suspect the kernel path rather than the cache concept — cached-prefix and full-prefill attention can differ in low-order bits, and at high temperature that is enough to change a sampled token.
Failure modesWhen it goes wrong
- Hit rate is exactly zero and never recoversa per-request value is rendered ahead of the breakpoint:
datetime.now()in the system prompt, a request UUID, a non-deterministically serialised JSON blob, or a tool list built per user. - cache_creation_input_tokens grows every request while cache_read_input_tokens stays at zerothe breakpoint sits at the end of the whole prompt, after the varying question, so each request writes a unique entry that nothing will ever read. You are paying the 1.25× write premium for pure waste.
- Caching works for two turns then stops mid-conversationa single turn added more than twenty content blocks, typically tool_use and tool_result pairs in an agent loop, pushing the previously cached block outside the twenty-block lookback window that each breakpoint searches.
- Hit rate collapses at peak traffic with no code changethe KV pool is full and blocks are evicted before the next request in that session arrives. Higher concurrency or a lower
gpu_memory_utilizationshrinks the residency window below your inter-request gap. - cache_creation_input_tokens is zero, with no hit and no errorthe prefix is shorter than the model's minimum cacheable length. This is silent by design, and the floor differs enough between models that the same prompt caches on one and not on another.
PapersWhere this comes from
- Prompt Cache: Modular Attention Reuse for Low-Latency InferenceGim et al., 2024. Showed that attention states for reusable prompt segments can be precomputed once and reused across requests, and worked out the position-ID handling needed to reuse a segment that is not at the front of the prompt. It is the paper that separates the general idea of reuse from the strict prefix constraint most production caches actually enforce — useful here because it explains what your hit-rate number is not measuring.
- SGLang: Efficient Execution of Structured Language Model ProgramsZheng et al., 2024. Introduced RadixAttention, an LRU radix tree over cached KV prefixes that makes cross-request reuse automatic rather than something the caller declares. This is where hit rate becomes a serving-level metric at all, and where eviction policy starts determining it.
- Efficient Memory Management for Large Language Model Serving with PagedAttentionKwon et al., 2023. Established block-granular KV memory with an indirection table, the substrate every prefix cache is built on. It explains the quantisation you see in practice: hits land on block boundaries, so hit rate is a step function of shared prefix length rather than a smooth one.
What changedWhat changed here
Updated this page OpenAI added higher cache hit rates, new diagnostics and explicit breakpoints for GPT-6 prompt caching, which cuts latency and cost without changing the model.
Three kinds of claim, strongest first. Signal runs every morning.