⚙️ · Operate

Caching

Reusing prior computation or responses to avoid redundant, costly model calls.

In one line

Caching means never paying twice to compute the same answer or the same prompt prefix.

ConceptWhat it is

Caching in LLM systems reuses previously computed results, whether that is an entire response to a repeated query or the internal key-value state for a shared prompt prefix, instead of recomputing from scratch. It exists because LLM inference is expensive and many real workloads have high redundancy, such as repeated system prompts, common questions, or shared document context across users.

Two layers matter most in practice: response-level caching of full answers, and prompt-prefix caching of the attention key-value state shared across requests.

How it worksThe mechanics

A response cache hashes the incoming prompt, or a semantic embedding of it, checks for a match above a similarity threshold, and returns the stored answer if found; a prompt-prefix cache, offered natively by several model providers, stores the computed key-value tensors for a repeated prefix so only the new suffix needs full computation on subsequent calls.

At a glanceSee it

Caching diagram
Caching diagram 1

The two caching layers side by side — a response cache reuses the finished answer, while a prefix KV cache reuses the attention state of tokens that many requests share.

Caching diagram 2

How prefix KV caching actually pays off — a cache hit still runs the model but skips re-encoding the shared prefix, and one edit to that prefix drops you back to the slow path.

When to use itWhere it fits

  • High-traffic applications with repeated or templated queries, like FAQ bots.
  • Long, shared system prompts or document context reused across many requests.
  • Cost-sensitive workloads where reducing redundant token processing directly cuts spend.
  • Latency-sensitive applications where a cache hit returns instantly versus seconds of generation.

When NOT to use itLimits & anti-patterns

  • Highly personalized or unique-per-request generation where cache hit rates would be near zero.
  • Applications requiring strictly fresh, non-deterministic output on every call, like creative brainstorming.

Trade-offsAdvantages & costs

Advantages
  • Cuts cost significantly on workloads with repeated content.
  • Reduces latency to near-zero on cache hits.
  • Prompt-prefix caching requires no application logic change with providers that support it natively.
Trade-offs & costs
  • Stale cached answers can serve outdated information if underlying data changes.
  • Semantic caching risks false-positive matches that return a wrong but similar answer.
  • Cache infrastructure adds its own operational and memory cost.

ExampleIn the real world

Anthropic's prompt caching feature is used heavily by customer-support copilots that reuse the same lengthy knowledge-base context across thousands of daily conversations, cutting repeated-context cost by up to 90 percent.

ToolsHow to implement it

  • Anthropic prompt cachingnative provider caching for repeated prompt prefixes.
  • GPTCacheopen-source semantic caching layer for LLM responses.
  • Rediscommon backing store for response and embedding caches.
  • vLLM automatic prefix cachingbuilt-in KV-cache reuse for self-hosted serving.

Cost & effortWhat it takes

Cached responses cost a small fraction, often 10 to 25 percent, of a fresh generation call, with moderate engineering effort to build cache-key and invalidation logic.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page OpenAI added higher cache hit rates, new diagnostics and explicit breakpoints for GPT-6 prompt caching, which cuts latency and cost without changing the model.

    OpenAI · 22 Sep 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning