Home › What Happens After You Hit Enter › Prompt Cache Key Resolution & Breakpoints
Pipeline stage · Operate

Prompt Cache Key Resolution & Breakpoints

The request is split at declared or inferred boundaries into a reusable prefix and a fresh suffix, and the system looks up whether that prefix's computed state already exists.

In one line

Caching is a prefix match on exact bytes, so the cheapest optimisation available to you is ordering the prompt static-first and never touching the front of it.

Why you'd careThe thing you have already noticed

You built an agent with a 30,000-token system prompt and a large tool schema. The first call cost what you expected. The second cost roughly a tenth as much and started streaming far sooner. Then you added Current time: {now} near the top of the system prompt while chasing an unrelated bug, and every call went back to full price — no error, no warning, no field that obviously said why. That is cache key resolution. Before a single token is processed, the server splits your request at declared or inferred boundaries, hashes the prefix, and asks whether that exact byte sequence has been computed before. One changed character at position 40 makes all 30,000 tokens after it a miss.

In and outWhat goes in, what comes out

InThe fully rendered prompt as an ordered byte string — tool definitions first, then system blocks, then messages — plus up to four cache_control markers, the model ID, and the request-level parameters that participate in one or more cache tiers.
ProcessTokenize, walk the block list computing a rolling hash over cumulative token chunks, and look each chunk hash up in a table or radix tree keyed by the chain of preceding chunks. Take the longest matching chain, searching back at most 20 content blocks.
OutA split of the prompt into N cached tokens whose KV state is reattached and M fresh tokens queued for prefill, plus the billing counters cache_read_input_tokens, cache_creation_input_tokens and input_tokens, and a pending write for the new suffix.

Nothing about the computation is lost: a cache read reattaches tensors identical to what recomputation would produce, so the model sees the same state either way. What is lost is diagnosis. The response reports how many tokens hit and missed, never at which position the prefix diverged, so a one-character invalidator looks exactly like a genuinely new prompt. Anthropic ships a beta cache-diagnostics surface that closes some of this gap; it is provider-specific and not the general case.

ConceptThe idea underneath

Every cached token is really a pair of tensors. Attention computes softmax(Q K^T / sqrt(d_k)) V: each query vector Q is scored against the key vectors K of every earlier position, the scores are normalised, and the result weights the value vectors V. The keys and values at position i are linear projections of the hidden state at position i, and because attention is causally masked that hidden state depends on tokens 1 through i only. Nothing later can change it.

That is the whole basis of prompt caching. If two requests begin with the same tokens, every key and value vector over that span is identical, so the second request can skip the arithmetic and attach the stored tensors. The property is one-directional: it holds for prefixes, not suffixes. Move a paragraph from position 5,000 to position 100 and every vector after position 100 changes, because each was computed while attending to a different history and, under rotary embeddings, at a different position.

Resolution is therefore a longest-common-prefix search. Servers do not hash the whole prompt; they hash cumulative chunks — OpenAI documents 128-token increments — and look each chunk hash up in a structure keyed by the chain of chunks before it. The longest chain that hits is your cache read; everything past it is a cache write. Declared breakpoints tell the server where you want entries stored. The matching itself happens whether or not you declare anything.

At a glanceSee it

Prompt Cache Key Resolution & Breakpoints diagram

How a request is split at breakpoints, hashed, matched against stored prefixes, and billed.

The knobsHyperparameters and nuance

  • cache_control(Anthropic) — {"type": "ephemeral"} on a content block, maximum 4 per request. Placement is the whole decision: put it on the last block of the stable span. Too early and you cache less than you could; on a per-request block and every call writes a fresh entry and reads none.
  • ttl(Anthropic, inside cache_control) — "5m" by default or "1h". Writes cost 1.25x base input at 5m and 2x at 1h; reads cost about 0.1x. Two reads break even at 5m, three at 1h, so pick the hour only for traffic with genuine multi-minute gaps.
  • minimum cacheable prefix(a model constant, not a parameter you set) — 512 tokens on Claude Opus 5, 1024 on Opus 4.8 and Sonnet 5, 4096 on Opus 4.6 and Haiku 4.5. Below it caching silently does nothing: no error, cache_creation_input_tokens just stays 0. It is not monotonic across generations, so it moves when you switch models.
  • prompt_cache_key(OpenAI) — a routing hint, not a cache key. It steers same-prefix requests toward the same machine so they can find each other's entries; it does not create or extend a cache. High-cardinality values fragment your traffic and lower the hit rate.
  • enable_prefix_caching(vLLM, on by default in V1) — the self-hosted equivalent. It hashes every filled block so unrelated requests can share prefixes. Off, blocks are freed at request end and only forks of one request share. On with no shared prefixes in your workload, you pay hashing overhead for nothing.

EffectHow this stage moves the answer

Caching does not change a single logit — a read reattaches tensors the model would otherwise have recomputed identically. It changes the answer by changing what you are willing to put in the prompt, and by distorting where you put it. When a 30,000-token prefix costs a tenth of list price on every call after the first, teams keep the twenty few-shot examples and the full tool documentation they would otherwise have trimmed, and output quality rises for reasons that have nothing to do with caching. The distortion runs the other way too: engineering for hit rate pushes retrieved documents to the top of the prompt, far from the question they support, and freezes per-request instructions into a bottom band. Both placements measurably change how well long-input instructions are followed, so a cache-optimised prompt is not the same prompt.

EvalsWhat it does to your measurements

The metrics that move are time to first token and cost per request, both by roughly an order of magnitude on a hit. The way this quietly ruins an eval run is placement. A suite that varies only the final question shares its entire prefix, so every sample after the first reads from cache: your reported p50 latency is a warm number diverse production traffic will never reproduce, and your cost line understates reality by up to 10x. A suite that shuffles few-shot order per sample for robustness does the opposite, guaranteeing a miss every time and making the same cost line meaningless in the other direction. Accounting is its own trap: on Anthropic's API input_tokens is only the uncached remainder, so an agent that ran for an hour can report 4,000 input tokens. Real prompt size is input_tokens + cache_creation_input_tokens + cache_read_input_tokens.

Failure modesWhen it goes wrong

  • Cache reads are zero on every request despite an identical-looking prompta timestamp, a UUID, a session ID, or a dict serialised without sorted keys sits somewhere inside the prefix.
  • Hit rate collapses only on turns with many tool callsthat turn appended more than 20 content blocks, and the breakpoint's 20-block lookback can no longer reach the previous entry.
  • A burst of parallel requests all pay full pricean entry only becomes readable once the first response starts streaming, so fan-out before that point writes N copies instead of reading one.
  • Caching works in development and never in productionthe production prefix is under the model's minimum cacheable length, which fails silently rather than erroring.
  • Hit rate drops the day you add a mode flag or a per-tenant toolconditional system sections make every flag combination a distinct prefix, and tool definitions render at position zero, so changing the tool set invalidates everything.

PapersWhere this comes from

  • Prompt Cache: Modular Attention Reuse for Low-Latency InferenceGim et al., 2024 (arXiv 2311.04934). Showed attention state can be reused for non-prefix reusable segments by declaring a prompt schema with position placeholders, which matters here because it establishes that the prefix-only restriction is an implementation choice rather than a mathematical necessity.
  • SGLang: Efficient Execution of Structured Language Model ProgramsZheng et al., 2024 (arXiv 2312.07104). Introduced RadixAttention, a radix tree over token sequences that finds the longest cached prefix automatically — the mechanism behind caching that works without you declaring any breakpoints.
  • Efficient Memory Management for Large Language Model Serving with PagedAttentionKwon et al., 2023 (arXiv 2309.06180). Established block-level KV storage with reference counting, which is what makes one stored copy of a shared prefix serve thousands of concurrent requests. Note that vendor breakpoint semantics and TTLs are documented product behaviour, not research results, and change without notice.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page OpenAI added higher cache hit rates, new diagnostics and explicit breakpoints for GPT-6 prompt caching, which cuts latency and cost without changing the model.

    OpenAI · 22 Sep 2026 · source

  • Updated this page Proposes chunk-level KV-cache reuse with deviation-guided recomputation to cut prefill latency when prompts lack a shared prefix.

    arXiv cs.AI · 24 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning