Home › Deployment, Inference & LLMOps › Key term › Caching
Key term · Operate

Caching

Reusing prior results for repeated or similar requests to cut cost and latency.

In one line

Caching skips a full model call by reusing an answer you already computed.

DefinitionWhat it means

Caching in an LLM system stores the result of a previous prompt, tool call, or prompt prefix so that an identical or sufficiently similar future request can be served instantly from the cache instead of paying for a fresh model call.

Why it mattersWhy you should care

Prompt-prefix and semantic caching can cut inference cost and latency dramatically for high-traffic products with repetitive queries, such as customer support or coding assistants, making caching strategy a core part of any production LLMOps cost-optimization plan.

At a glanceSee it

Caching diagram
Caching diagram 1

A miss is not a dead end — it calls the model, writes the result back, then the next identical query returns instantly.

Caching diagram 2

Three ways to judge two prompts as the same — exact string, semantic similarity, or a shared prompt prefix — each trading cost against a different failure mode.

Where you see itIn the wild

  • Prompt caching features offered by major model providers to cut repeated-context cost.
  • Semantic caches in customer support bots serving common questions instantly.
  • Architecture discussions on designing a cache invalidation strategy for an AI feature.
A living map of modern AI — kept current every morning