CRAG scores retrieved documents for relevance, refines the good ones, and falls back to web search when the corpus comes up short.
ConceptWhat it is
Corrective Retrieval-Augmented Generation (CRAG) inserts a quality gate between retrieval and generation. Plain RAG trusts whatever the retriever returns and hands it straight to the model; CRAG does not. A lightweight retrieval evaluator (a grader) scores how relevant the retrieved passages are to the query and sorts the result into three states: correct (confidently relevant), incorrect (confidently nothing useful), and ambiguous (uncertain).
Each state triggers a different corrective action. Correct hits are kept but run through knowledge refinement — decomposed into small knowledge strips and recomposed so noise is filtered out. Incorrect retrievals are discarded and the system does a web search fallback for fresh external knowledge. Ambiguous cases blend both. CRAG exists because retrieval is the weakest link in most RAG stacks: a stale, sparse, or off-topic corpus silently produces confident hallucinations, and CRAG makes the pipeline robust to that.
How it worksThe mechanics
The query first hits the normal retriever to pull top-k passages from the vector store. The grader scores each passage (or the set) for relevance and assigns a confidence label. On correct, CRAG refines the passages into filtered knowledge strips and passes them forward. On incorrect, it drops the retrieved context entirely and issues a web search, treating those results as the new evidence. On ambiguous, it combines refined internal context with web results. The chosen knowledge is then handed to the generator, which produces the grounded answer only from vetted evidence rather than from raw, untriaged retrieval.
At a glanceSee it
Inside the refine step CRAG does not trust a whole document — it decomposes each into fine-grained knowledge strips, grades and drops the noisy ones strip by strip, then recomposes only the survivors.
The grader is not a plain three-way classifier — a continuous relevance score is split by two tunable thresholds into correct, ambiguous, and incorrect, so operators trade precision against recall and cost by moving the cut points.
When to use itWhere it fits
- The corpus is incomplete, stale, or uneven in coverage, so retrieval often misses or returns partially relevant passages.
- The domain changes faster than the index is refreshed and a live web fallback can supply what the corpus lacks.
- Wrong answers are costly and you would rather escalate to search than let the model guess from thin evidence.
- You already run standard RAG and see hallucinations that trace back to bad retrievals rather than to the generator.
When NOT to use itLimits & anti-patterns
- The corpus is curated, complete, and reliably relevant — the grader and fallback add cost with little to correct.
- Latency and per-query budget are tight; the extra grading call and conditional web hop widen the tail sharply.
- Web search is unavailable, disallowed, or the domain is closed and confidential, removing the main corrective path.
- The failure mode is generation quality or reasoning, not retrieval — CRAG fixes evidence, not the model.
Trade-offsAdvantages & costs
Advantages
- Directly attacks the biggest RAG failure mode by catching bad retrievals before they reach the model.
- The web fallback lets answers stay fresh even when the internal index lags behind reality.
- Knowledge refinement strips out noise, so the generator sees denser, higher-signal context.
- The grader is a modular, swappable component that bolts onto an existing RAG pipeline.
Trade-offs & costs
- Every query pays for at least one extra grading step, and triggered fallbacks add a full search-and-regenerate cycle.
- The system is only as good as the grader — a miscalibrated evaluator either lets weak docs through or wastes searches.
- Web fallback introduces an external dependency with its own latency, cost, rate limits, and trust and provenance concerns.
- More branches mean more to tune, monitor, and debug than a single-path RAG flow.
ExampleIn the real world
An internal support assistant answers over a product documentation corpus that is re-indexed weekly. A user asks about a configuration flag shipped three days ago. The retriever returns only older pages that never mention it; the grader scores them as incorrect because none address the flag. Rather than confidently citing the outdated page, CRAG discards those passages and runs a web search against the public release notes, refines the fresh result into a few knowledge strips, and generates an accurate answer that names the new flag and links its source. Standard RAG, by contrast, would have paraphrased the stale docs and quietly gotten it wrong.
ToolsHow to implement it
- LangGraphits documented Corrective RAG example wires the grade-then-branch flow as an explicit state machine.
- LlamaIndexships a Corrective RAG pack that implements the evaluate, refine, and web-search-fallback loop.
- Tavily(or a comparable search API) — a common web-search fallback for the incorrect and ambiguous branches.
- Cohere Rerankor a cross-encoder reranker — supplies relevance scores the grader can threshold on instead of a bespoke evaluator.
Cost & effortWhat it takes
CRAG adds a grader on the critical path of every query plus a conditional fallback that only fires on weak retrievals. The grader is cheap when it is a small fine-tuned evaluator, a reranker score, or a single LLM-as-judge call; the expensive tail is the fallback, where a web search followed by re-generation roughly doubles work and widens latency for the fraction of queries that trigger it. Budget for variable, bimodal latency rather than a flat number, and expect ongoing effort to calibrate grader thresholds and monitor how often each branch fires.