Home › RAG - Retrieval-Augmented Generation › Contextual retrieval
📚 · Ground

Contextual retrieval

Prepend an LLM-written context blurb to every chunk before indexing so retrieval keeps document meaning.

In one line

Contextual retrieval adds a short LLM-generated note about where each chunk sits in its document before embedding, sharply cutting retrieval misses on long, ambiguous corpora.

ConceptWhat it is

Contextual retrieval is a RAG indexing pattern that fixes a specific failure of naive chunking: when a document is split into passages, each chunk loses the surrounding context that gave it meaning. A chunk that says "the company grew 30 percent" no longer knows which company or which quarter, so neither semantic nor keyword search can reliably find it. The technique, introduced by Anthropic, runs an index-time LLM pass that writes a short chunk-specific blurb situating the passage within the whole document, then prepends that blurb to the chunk before it is indexed.

Because the context is baked in before indexing, it improves both retrieval arms at once: the augmented text produces better vectors (Contextual Embeddings) and better keyword matches (Contextual BM25). Each chunk becomes self-describing, so a query no longer has to guess at the entity, timeframe, or section a bare chunk silently assumed. It composes cleanly with hybrid search and reranking for compounding gains.

How it worksThe mechanics

Split each document into chunks as usual. Then, for every chunk, prompt an LLM with the full document plus that chunk and ask for a one-to-two-sentence context that locates the passage in the whole; prepend that blurb to the chunk text, embed the augmented chunk into the vector store, and add it to the BM25 index. At query time, run hybrid search over the contextualized index, fuse the dense and sparse hits, optionally rerank, and pass the top chunks to the generator. Prompt caching keeps the per-chunk cost low because the source document is cached once and reused across all of its chunks.

At a glanceSee it

Contextual retrieval diagram
Contextual retrieval diagram 1

Prompt caching pays to load the document once—every chunk after the first reads it from the warm cache at roughly a tenth the cost, which is what makes contextualizing each chunk affordable.

Contextual retrieval diagram 2

Each retrieval layer stacks—contextual embeddings, then contextual BM25, then reranking—driving the share of missed chunks down step by step against the naive baseline.

When to use itWhere it fits

  • Long documents where a chunk's meaning depends on distant context, such as which entity, quarter, or section it belongs to.
  • Corpora full of ambiguous references, like pronouns, "the company", or "this policy", that leave bare chunks stranded.
  • High-value retrieval where a missed chunk is costly and an extra index-time pass is easily affordable.
  • Pipelines already running hybrid dense plus BM25 search, since both arms improve without changing the retriever.

When NOT to use itLimits & anti-patterns

  • Short, self-contained records such as product rows, FAQs, or tickets, where each chunk already stands alone.
  • Rapidly changing corpora where re-running the LLM pass on every edit is too slow or expensive.
  • Tight indexing budgets or latency-sensitive ingestion where an LLM call per chunk is prohibitive.
  • Before trying cheaper fixes first, like larger chunks, overlap, better metadata, or a reranker alone.

Trade-offsAdvantages & costs

Advantages
  • Sharply reduces failed retrievals on long or ambiguous documents by making each chunk self-describing.
  • Improves dense and keyword search simultaneously as a purely index-time change; the query path and retriever stay the same.
  • Composes with hybrid search and reranking for compounding accuracy gains.
  • Cheap per chunk when prompt caching reuses the cached source document across all of its chunks.
Trade-offs & costs
  • Adds an LLM pass over the entire corpus at index time, which is real money and wall-clock time at scale.
  • Blurbs must be regenerated and the index rebuilt when documents change, or stale context will mislead.
  • Storage and index size grow because every chunk now carries extra prepended text.
  • Quality depends on the context-generation prompt and model; a weak or wrong blurb injects noise.

ExampleIn the real world

A support team indexes a 90-page enterprise contract into a RAG assistant. A raw chunk reads "either party may terminate with 30 days notice," but the query "how much notice to cancel the premium tier?" fails to retrieve it, because the chunk never names a tier or contract. With contextual retrieval, an index-time LLM pass prepends "this clause is from the premium-tier termination section of the customer's master services agreement," so the augmented chunk now matches both the embedding and the keyword index, and the assistant surfaces the correct 30-day clause instead of a generic or empty answer.

ToolsHow to implement it

  • Claude with prompt caching for the index-time context-generation pass, the technique's original reference implementation.
  • LlamaIndex and LangChain, which provide the chunking, hybrid retrieval, and ingestion pipelines you slot a contextual step into.
  • Embedding models from providers such as OpenAI, Cohere, or Voyage for the dense arm, paired with a BM25 store like Elasticsearch or OpenSearch for the sparse arm.
  • Rerankers such as Cohere Rerank layered on top for additional retrieval gains.

Cost & effortWhat it takes

The dominant cost is a one-time-per-document-version LLM pass over the whole corpus at index time, one short generation per chunk. Prompt caching makes this far cheaper than naive per-chunk calls, because the full document is cached once and reused across all its chunks, but at large scale it is still a meaningful indexing bill and adds ingestion latency. Query-time cost is unchanged. Effort is moderate: it is an add-on step in an existing pipeline, mostly prompt tuning and a re-index rather than a new retriever. Ongoing cost scales with how often documents change, since every edit forces regeneration.

A living map of modern AI — kept current every morning