📚 · Ground

Self-RAG

Self-RAG: let the model decide when to retrieve, then grade its own evidence and answer.

In one line

A RAG variant where the model itself chooses whether to retrieve and emits reflection tokens that critique each retrieved passage and its own draft, so retrieval happens only when it actually helps.

ConceptWhat it is

Self-RAG (self-reflective retrieval-augmented generation) is a RAG pattern in which a single model learns to control retrieval and to grade evidence on the fly, rather than retrieving a fixed number of passages for every query. It exists to fix a structural weakness of naive RAG: retrieval is not always needed, and forcing it can inject irrelevant or distracting context that degrades an answer the model already knew.

The mechanism is a small vocabulary of reflection tokens the model emits alongside normal text. A Retrieve token decides whether to call the retriever at all; relevance tokens judge whether a fetched passage is on-topic; and support and utility tokens judge whether the draft is actually grounded in that passage and worth keeping. In the original work this grading ability is fine-tuned into the generator itself so one model both answers and critiques; in practice it can also be approximated with a separately prompted grader, which is cheaper to stand up but easier to miscalibrate.

How it worksThe mechanics

Given a query, the model first emits a Retrieve decision. If retrieval is skipped, it answers directly from parametric memory. If retrieval fires, the retriever returns candidate passages and the model tags each with a relevance judgment, dropping the off-topic ones. For the surviving passages it drafts a candidate answer segment, then emits support and utility judgments scoring how well that segment is grounded in the passage and how useful it is. These reflection scores rank competing segments, and a poorly supported draft can trigger another retrieval pass. The best-scoring, best-grounded segment is stitched into the final answer, with the reflection tokens left behind as an audit trail.

At a glanceSee it

Self-RAG diagram
Self-RAG diagram 1

The four reflection token families Self-RAG interleaves with text — one control token that gates retrieval and three critique tokens that grade relevance, support, and utility.

Self-RAG diagram 2

How the skill is learned — an offline critic annotates a corpus with reflection tokens, the generator is trained to emit them itself, and the critic is discarded at inference.

When to use itWhere it fits

  • Traffic mixes questions that need fresh or private documents with questions the model can answer from its own weights, and paying for retrieval on every call is wasteful.
  • Grounding and faithfulness matter enough that you want per-passage evidence checks, not a single blind top-k dump into the prompt.
  • You need an inspectable trace of why a passage was used or rejected, for review or compliance.
  • Retrieval quality is uneven and you want the model to discard irrelevant hits instead of being derailed by them.

When NOT to use itLimits & anti-patterns

  • Every query genuinely needs retrieval, so the retrieve-or-not decision adds cost without ever saving a call.
  • Latency and token budgets are tight, and the extra reflection tokens plus possible re-retrieval loops are unaffordable overhead.
  • You cannot fine-tune a critic and cannot absorb the extra prompted grading calls that approximate one.
  • A simpler reranker or a fixed retrieve-then-read pipeline already clears your grounding bar.

Trade-offsAdvantages & costs

Advantages
  • Retrieves only when it helps, cutting needless retriever calls and context bloat on questions the model already knows.
  • Per-passage relevance and support checks raise faithfulness and reduce hallucination from irrelevant context.
  • Reflection tokens give an inspectable trace of the retrieve, keep, and reject decisions.
  • Adapts to uneven corpora by discarding bad hits rather than forcing them into the answer.
Trade-offs & costs
  • Reflection tokens and extra critic calls add token cost and latency to every query.
  • Reliable behavior usually needs a fine-tuned critic; a purely prompted critic is easy to miscalibrate.
  • Re-retrieval loops can stall or run long on hard queries unless capped.
  • More moving parts to build, tune, and evaluate than a straight retrieve-then-read chain.

ExampleIn the real world

Consider a support assistant over an enterprise wiki. A user asks what OKR stands for; the model's Retrieve token says no, and it answers from parametric memory instantly, skipping a vector search. The next user asks for the Q3 refund SLA for enterprise accounts; the Retrieve token fires, the retriever returns five passages, and relevance tokens keep the two that mention refund SLAs while dropping three about billing. The model drafts an answer, its support token flags that the draft cites a figure absent from either passage, and a second retrieval pulls the SLA table. The regrounded draft scores well and is returned along with the passage IDs it relied on.

ToolsHow to implement it

  • The reference Self-RAG implementation and critic models from Asai et al., which fine-tune a Llama-2 base to emit reflection tokens.
  • LangGraph, whose Self-RAG tutorial wires the retrieve, grade-documents, generate, and grade-generation steps into a stateful loop.
  • LlamaIndex, which ships a Self-RAG LlamaPack that reproduces the reflection-token flow over its retrievers.
  • Faithfulness and context-relevance evaluators such as Ragas or TruLens to measure whether the critic actually improves grounding.

Cost & effortWhat it takes

Cost lives in two places: the reflection tokens the model emits on every step, and the extra model calls when a weak support score triggers another retrieval pass. On top of inference, the upfront effort is the critic itself, either fine-tuning a base model to emit the tokens reliably (data plus training) or engineering and evaluating prompted grading calls, which are cheaper to start but easier to miscalibrate. Expect Self-RAG to cost more per query than fixed retrieve-then-read, and to earn that back only where the retrieve-versus-skip split is real and grounding quality is worth paying for.

A living map of modern AI — kept current every morning