📚 · Ground

Adaptive RAG

Adaptive RAG routes each query to the cheapest retrieval strategy that will actually answer it.

In one line

A router classifies each query's difficulty and dispatches it to a matched pipeline — from no retrieval to iterative multi-step search — instead of forcing every question through one fixed path.

ConceptWhat it is

Adaptive RAG is a retrieval-augmented generation pattern that puts a router in front of retrieval, so the pipeline chosen depends on the query rather than being fixed in advance. A naive RAG system runs the same embed, retrieve, generate steps for every question, which over-serves trivial lookups and under-serves genuinely hard, multi-hop ones. Adaptive RAG instead classifies incoming query difficulty and dispatches each query to a matched strategy — commonly no retrieval for questions the model already knows, single-shot retrieval for straightforward factual lookups, and iterative multi-step retrieval for questions that need several passes of evidence gathering.

The idea, formalized in the Adaptive-RAG research line, is that difficulty is not uniform, so a single pipeline is either too expensive for easy queries or too shallow for hard ones. By spending one small routing decision up front, the system puts cost and latency where they earn their keep.

How it worksThe mechanics

A query arrives and a lightweight router — a trained classifier, a small language model, or a few-shot prompt — scores its difficulty. That verdict selects a branch: the simplest queries are answered directly from the model's parameters with no retrieval; moderate queries trigger one round of vector search and generation; complex queries enter a loop that retrieves, reasons about what evidence is still missing, and retrieves again until it has enough. Each branch then composes a grounded answer, and because the router is cheap relative to a full multi-step run, the average query resolves faster and cheaper than if every query took the heaviest path.

At a glanceSee it

Adaptive RAG diagram
Adaptive RAG diagram 1

How the router learns difficulty offline — each training query is labeled with the cheapest pipeline tier that still answers it correctly, and those labels are distilled into the lightweight scorer.

Adaptive RAG diagram 2

A runtime verify-and-escalate loop that fixes the misroute the base diagram only warns about — a groundedness check on every answer decides whether to return it or bump the query up to a heavier tier.

When to use itWhere it fits

  • Traffic mixes trivial lookups with genuinely hard multi-hop questions, so no single pipeline can be tuned for both.
  • Latency and cost per query matter, and you want to avoid running expensive multi-step retrieval on every request.
  • You can label or cheaply estimate query difficulty, whether by a trained router or a small prompt-based classifier.
  • You already run several retrieval strategies and need a principled way to choose between them per query.

When NOT to use itLimits & anti-patterns

  • Queries are uniform in difficulty — if they all look alike, one well-tuned pipeline is simpler and just as good.
  • Volume is low or the domain is narrow, where the router's extra moving parts outweigh the savings.
  • You cannot reliably classify difficulty, since a weak router silently sends hard queries down cheap paths.
  • A strict accuracy floor on every query makes running the strongest pipeline for all traffic the safest default, regardless of cost.

Trade-offsAdvantages & costs

Advantages
  • Spends compute where it helps: cheap queries stay cheap, hard queries get the depth they need.
  • Lower average latency and cost than routing all traffic through the heaviest multi-step pipeline.
  • Modular — new retrieval strategies can be added as branches behind the same router.
  • Often lifts answer quality on hard queries versus naive RAG without penalizing the easy ones.
Trade-offs & costs
  • The router is a new component to build, evaluate, and maintain, and it can misroute.
  • A wrong route degrades quality invisibly: a hard query sent to no-retrieval yields a confident but ungrounded answer.
  • Difficulty labels for training or evaluating the router are often expensive to produce.
  • More branches mean more code paths to test, monitor, and keep behaving consistently.

ExampleIn the real world

A customer-support assistant over a product knowledge base sees three kinds of questions. "What is your refund window" is answered from the model's own knowledge with no retrieval. "Does the enterprise tier include single sign-on" triggers one vector search over the docs and a single generation pass. "Compare the data-retention terms across the enterprise, team, and free tiers and tell me which fits a regulated healthcare buyer" is routed to the multi-step branch, which retrieves each tier's terms, reasons about the regulatory constraint, and retrieves again before composing the answer. A front-end router — a small classifier trained on labeled past tickets — makes the call in a few milliseconds, so the common cheap questions never pay for the expensive path.

ToolsHow to implement it

  • LangGraphits documented adaptive-RAG example wires a router node to no-retrieval, single-shot, and multi-step branches as a graph.
  • LlamaIndexRouterQueryEngine and router retrievers select among query engines per query.
  • Haystackconditional routers and branching pipelines dispatch queries to different retrieval flows.
  • semantic-router(Aurelio Labs) — fast embedding-based routing that picks a strategy without a full LLM call.

Cost & effortWhat it takes

The marginal cost of Adaptive RAG is one routing decision per query — cheap when the router is an embedding classifier or small model, a little less cheap when it is a full LLM call. The real investment is up front: building the router, gathering difficulty labels, and measuring routing accuracy against the pipelines it feeds. Run cheaply, it pays for itself by keeping most traffic off the expensive multi-step path; run carelessly, a weak router can cost more in wrong answers than it saves in compute. Budget for ongoing evaluation as query patterns drift.

A living map of modern AI — kept current every morning