The model cannot tell retrieved text from anything else you typed, which is why "it ignored my documents" is almost always a retrieval bug rather than a model one.
Why you'd careThe problem it solves
The instinct that gets people is putting the knowledge base in a skill. It feels right: the skill is the folder of reference material, so the four hundred product manuals go in the folder. Then the first matching question loads a file the size of a novel, the context window fills, and the answer gets worse rather than better. Skills load whole named files on a trigger; that works for a nine-hundred-line playbook and collapses at nine hundred documents. Retrieval is the construct for the collapse case. You search the corpus yourself, pick a handful of passages, and paste them into the prompt. Nothing about the model changes. What changes is that the relevant two thousand tokens are present and the irrelevant two million are not.
ConceptWhat it is
Retrieval-augmented generation is a pattern, not a product. At request time your code runs a query against an index — vector, keyword, hybrid, whatever you built — ranks the results, takes the top few, and injects them into the prompt as content the model reads like any other text. Then it asks the question.
Every interesting decision is yours and none of them are in the API: how you chunk, what you embed, how you build the query from a conversational turn, how many passages you keep, whether you rerank, where in the prompt they land, and what you do when retrieval returns nothing useful. The model's only contribution is reading what arrived.
That is also where the boundaries sit. Against a skill: a skill names files and loads them whole on a trigger; retrieval selects fragments from something unloadable. Against long context: a million-token window genuinely removes the need to retrieve for mid-sized corpora, and “just put it all in” is now a real answer for a few hundred pages — but it is not an answer for a document store that grows weekly, and it costs full price on every request unless it caches. Against memory: retrieval pulls from a corpus you curate; memory is state the agent itself writes and reads back across sessions.
Anthropic's surface for the injection step is ordinary content blocks, plus document blocks with citations when you want the answer to point at sources. There is no hosted retrieval service in that surface; the search half is yours.
How it worksThe mechanics
The injection is unremarkable — text or document content blocks in a user turn. Turning on citations is what makes the result auditable, and it changes the response shape: cited output splits into multiple text blocks, each carrying a citations array.
{
"type": "document",
"source": {"type": "base64",
"media_type": "application/pdf",
"data": "..."},
"citations": {"enabled": true}
}Each citation carries cited_text, document_index, document_title and a location object typed by source kind — character offsets for plain text, page numbers for PDFs, block indices for custom content. Citations must be enabled on all documents in a request or none, and they are incompatible with structured output formats, which returns a 400.
The expensive mistake is positional, not conceptual. Prompt caching is a prefix match: the cache key is the exact bytes up to each breakpoint, and any byte that changes at position N invalidates everything at or after N. Render order is tools, then system, then messages. Retrieved passages change on every request by definition. Put them ahead of your stable instructions — a tempting layout, since “here is the context, now here are the rules” reads well — and you invalidate the whole prefix downstream and pay full price every single time. Put them after the last breakpoint and the stable part keeps hitting.
The economics make that worth care. Cache reads cost roughly a tenth of base input; writes cost about 1.25× for the five-minute TTL and 2× for the one-hour. The minimum cacheable prefix is model-dependent and not monotonic across generations — 512 tokens on Claude Opus 5, 1024 on Opus 4.8 and Sonnet 5, 2048 on Opus 4.7, 4096 on Opus 4.6 and Haiku 4.5. Below the minimum nothing caches and you get no error, just cache_creation_input_tokens of zero. Verify with cache_read_input_tokens across repeated requests; a persistent zero means something in your prefix is moving.
At a glanceSee it
Where retrieved passages sit relative to the cache breakpoint decides what every later request costs.
Where it runsSurfaces and availability
| Surface | Status | Notes |
|---|---|---|
| Claude Code | Yes | In the sense that any harness can inject retrieved text; in practice file search and grep often replace an index entirely for a repository. |
| Claude API / Messages API | Yes | It is just content blocks. Citations are GA; document blocks accept base64, plain text, and Files API references. Confirmed against the “Features overview” table. |
| Managed Agents | Yes | Files and repositories mount into the session container, so the agent can search them with ordinary file tools instead of a pre-built index. |
| Claude Desktop / claude.ai | Unverified | Project knowledge and connected files behave like retrieval from the user’s side, but the selection mechanism is not documented as a developer surface. |
| Claude Agent SDK | Yes | Same as any harness; its built-in grep and file tools are frequently a better fit than embeddings for code. |
| Amazon Bedrock | Yes | Injection and citations work — citations are GA here. Automatic prompt caching and the Files API do not: Bedrock is the only platform absent from the automatic-caching row of the “Features overview”, and the Files API is listed under “Features not supported” on “Claude in Amazon Bedrock”. So you manage cache breakpoints explicitly and pass content inline. Explicit 5-minute and 1-hour prompt caching are both available. |
| Google Vertex AI | Yes | Injection and citations work. Automatic prompt caching is available here — the Features overview lists it on the Claude API, Claude Platform on AWS, Google Cloud and Microsoft Foundry — so only the Files API caveat carries over from Bedrock (listed under “Features not supported” on “Claude on Google Cloud”). |
| Microsoft Foundry | Yes | Core injection is available and citations are GA, not beta (Features overview lists Citations on all five platforms; Foundry’s own not-supported list omits them). The Files API is the beta piece, and it additionally requires a Hosted on Anthropic deployment. Prompt caching, including automatic caching, is available. |
| Anthropic-hosted vector store | Unverified | Nothing in the Messages API surface provides one, and the canonical “Features overview” page lists no managed retrieval or vector-store product in its Files and assets section. If such a product exists it is not part of the request shape, so plan on owning the index. |
| Other vendors | Yes | Every chat-completion API supports it, because it is prompt content. Some vendors also sell a hosted retrieval layer; that part is not portable. |
The uniform Yes column is the point: retrieval is the most portable thing in this family precisely because no vendor feature is involved. What that buys you is a design that survives a provider migration untouched. What it costs you is that nobody is going to hand you the hard part — chunking, ranking, and knowing when retrieval returned junk are yours on every platform. If you are choosing where to build and retrieval is central, citation support is no longer a differentiator (it is GA everywhere); choose on the Files API, which only the Anthropic-operated surfaces have, and on automatic prompt caching, which everything except Amazon Bedrock has.
ExampleIn the real world
A support tool answers questions over 40,000 help-centre articles and release notes. Loading them is not an option at any context size, and they change daily.
A ticket arrives: “export to CSV is truncating at 1,000 rows on the team plan.” The service embeds the question, searches the index, reranks the top fifty hits down to six passages — two from the export documentation, one release note about a limit change, three from prior tickets — and assembles a request. The system prompt is the frozen support persona and output format with a cache breakpoint on it. The six passages go into the user turn, after that breakpoint, as document blocks with citations enabled.
The model answers in four sentences and each factual claim carries a citation the UI renders as a footnote linking back to the source article. The reviewer can see that the limit claim came from the release note rather than from the model's memory.
Two weeks later an engineer moves the passages into the system prompt “so they get more weight”. Answer quality does not change. Cost triples, because the cached prefix now begins with text that differs on every request, and cache_read_input_tokens drops to zero on every call.
Not thisWhat it is often confused with
- Not a skilla skill loads named files whole when the task matches. Retrieval picks fragments from a corpus that could never be loaded. If you are chunking and embedding the contents of a skill folder, you crossed the line and should build an index.
- Not memoryretrieval reads a corpus you curate and the model never writes to. Memory is state the agent records for itself and reads back in later sessions. Different write path, different governance, different failure modes.
- Not fine-tuningretrieval changes what is in the prompt on this request; training changes the weights. Retrieved facts are current and removable; trained facts are neither.
- Not long contexta large window lets you skip retrieval for a few hundred pages, and often should. It does not scale to a growing document store, and it bills for everything you pasted whether it mattered or not.
- Not a toolthough it can be wrapped in one. Injected retrieval happens before the model speaks; a search tool lets the model decide to search mid-turn. The second is often better when the query needs the model's own reformulation.
LimitsWhen not to reach for it
- The corpus fits.A playbook, a style guide, a schema reference. Load it whole in a skill or the system prompt and skip the entire retrieval stack, including its failure modes.
- The answer requires an action.Retrieval will produce a well-cited description of how to issue the refund. Only a tool issues it.
- The material is public and current.Building an index of the open web is a large project with a maintained alternative; a web search tool covers it, with the caveat that it is unavailable on Bedrock.
- You need facts about this user across sessions.That is memory, not retrieval — a store the agent writes to, with versioning and deletion, not a corpus you index.
- You have not measured retrieval quality.If you cannot say what fraction of questions get the right passage in the top five, adding a bigger model or a longer prompt is fixing the wrong layer.
Verified 2026-09-12. Stable — the shape of this is unlikely to move. Provider: Cross-vendor.