Caching means never paying twice to compute the same answer or the same prompt prefix.
ConceptWhat it is
Caching in LLM systems reuses previously computed results, whether that is an entire response to a repeated query or the internal key-value state for a shared prompt prefix, instead of recomputing from scratch. It exists because LLM inference is expensive and many real workloads have high redundancy, such as repeated system prompts, common questions, or shared document context across users.
Two layers matter most in practice: response-level caching of full answers, and prompt-prefix caching of the attention key-value state shared across requests.
How it worksThe mechanics
A response cache hashes the incoming prompt, or a semantic embedding of it, checks for a match above a similarity threshold, and returns the stored answer if found; a prompt-prefix cache, offered natively by several model providers, stores the computed key-value tensors for a repeated prefix so only the new suffix needs full computation on subsequent calls.
At a glanceSee it
The two caching layers side by side — a response cache reuses the finished answer, while a prefix KV cache reuses the attention state of tokens that many requests share.
How prefix KV caching actually pays off — a cache hit still runs the model but skips re-encoding the shared prefix, and one edit to that prefix drops you back to the slow path.
When to use itWhere it fits
- High-traffic applications with repeated or templated queries, like FAQ bots.
- Long, shared system prompts or document context reused across many requests.
- Cost-sensitive workloads where reducing redundant token processing directly cuts spend.
- Latency-sensitive applications where a cache hit returns instantly versus seconds of generation.
When NOT to use itLimits & anti-patterns
- Highly personalized or unique-per-request generation where cache hit rates would be near zero.
- Applications requiring strictly fresh, non-deterministic output on every call, like creative brainstorming.
Trade-offsAdvantages & costs
Advantages
- Cuts cost significantly on workloads with repeated content.
- Reduces latency to near-zero on cache hits.
- Prompt-prefix caching requires no application logic change with providers that support it natively.
Trade-offs & costs
- Stale cached answers can serve outdated information if underlying data changes.
- Semantic caching risks false-positive matches that return a wrong but similar answer.
- Cache infrastructure adds its own operational and memory cost.
ExampleIn the real world
Anthropic's prompt caching feature is used heavily by customer-support copilots that reuse the same lengthy knowledge-base context across thousands of daily conversations, cutting repeated-context cost by up to 90 percent.
ToolsHow to implement it
- Anthropic prompt cachingnative provider caching for repeated prompt prefixes.
- GPTCacheopen-source semantic caching layer for LLM responses.
- Rediscommon backing store for response and embedding caches.
- vLLM automatic prefix cachingbuilt-in KV-cache reuse for self-hosted serving.
Cost & effortWhat it takes
Cached responses cost a small fraction, often 10 to 25 percent, of a fresh generation call, with moderate engineering effort to build cache-key and invalidation logic.
What changedWhat changed here
Updated this page OpenAI added higher cache hit rates, new diagnostics and explicit breakpoints for GPT-6 prompt caching, which cuts latency and cost without changing the model.
Three kinds of claim, strongest first. Signal runs every morning.