📚 · Ground

Agentic RAG

Let the model decide what to retrieve, when to retrieve again, and when it has enough.

In one line

Retrieval becomes a tool the agent calls repeatedly instead of a fixed one-shot step.

ConceptWhat it is

Agentic RAG replaces the fixed retrieve-once pipeline with an agent loop: the model plans what information it needs, calls search or other tools, inspects what comes back, and decides whether to search again, search differently, or answer. It exists because many real questions cannot be resolved by a single retrieval pass, and rigid pipelines fail silently when the first guess at relevant context is wrong.

The retrieval step becomes just one tool among several, alongside things like a calculator, a SQL query, or a live API call.

How it worksThe mechanics

An orchestration loop, typically built with a graph or agent framework, gives the model a set of tools including a retriever; the model reasons about the question, calls a tool, reads the result, and either issues another tool call with a refined query or produces a final answer once it judges the evidence sufficient.

At a glanceSee it

Agentic RAG diagram
Agentic RAG diagram 1

The single ‘search differently’ move unfolds into a classifier that names the gap — wrong words, too broad, or wrong source — and picks the matching corrective action.

Agentic RAG diagram 2

A worked multi-hop question where hop 2 can only be built from hop 1’s result — the dependency chain a one-shot retrieval pass cannot express.

When to use itWhere it fits

  • Open-ended research questions that need iterative digging across sources.
  • Tasks that mix retrieval with other actions, like looking up a customer record then querying a knowledge base.
  • Situations where the first retrieval attempt is often insufficient and needs follow-up queries.
  • Customer support agents that must decide between searching docs, calling an API, or escalating.

When NOT to use itLimits & anti-patterns

  • Simple factual lookups where a single retrieval pass already answers the question, since agentic looping just adds latency.
  • Cost-sensitive, high-volume endpoints where multiple LLM calls per query blow the budget.
  • Latency-critical paths, like live chat first-response, where iterative tool calling is too slow.

Trade-offsAdvantages & costs

Advantages
  • Handles multi-part and ambiguous questions that fixed pipelines miss.
  • Can recover from a bad first retrieval by rephrasing and searching again.
  • Composable with other tools beyond just document search.
  • Naturally supports self-correction and citation checking as extra steps.
Trade-offs & costs
  • Multiple LLM calls per question drive up cost and latency.
  • Harder to debug and evaluate than a linear pipeline.
  • Risk of infinite or wasteful tool-call loops without guardrails.
  • Requires careful prompt and framework engineering to keep the agent on task.

ExampleIn the real world

Perplexity's answer engine plans multiple search queries, reads back snippets, and issues follow-up searches before synthesizing a cited answer, rather than answering off a single retrieval pass.

ToolsHow to implement it

  • LangGraphgraph-based orchestration for building controllable retrieval-then-reason loops.
  • LlamaIndex Agentsagent abstractions with built-in retriever tools and query engines.
  • Vertex AI Searchmanaged retrieval backend that plugs into agent tool-calling setups.
  • LangSmithtracing to debug multi-step agentic retrieval loops in production.

Cost & effortWhat it takes

Higher cost and latency than single-pass RAG since each query can trigger several LLM and tool calls; meaningful engineering effort to add loop guardrails and evaluation.

A living map of modern AI — kept current every morning