Answer questions that need chained facts by retrieving once per reasoning step.
ConceptWhat it is
Multi-hop RAG handles questions that cannot be answered from a single retrieved chunk because the answer requires chaining facts across documents, such as "which company acquired the startup founded by the person who invented X." It exists because single-pass top-k retrieval returns evidence for one aspect of a question and silently misses the rest.
The system decomposes the question into sub-questions, retrieves for each in sequence, and carries intermediate answers forward until enough evidence is assembled.
How it worksThe mechanics
The question is decomposed, either by an LLM prompt or a planning module, into an ordered set of sub-questions; each sub-question triggers its own retrieval step, and the answer or evidence from one hop is fed into the retrieval query for the next hop until the chain resolves into a final synthesized answer.
At a glanceSee it
Before any retrieval, a planner decides whether the sub-questions are independent enough to fan out in parallel or must be chained through a bridge entity — a routing choice the purely sequential loop never surfaces.
A single hop can fail three distinct ways and each bad answer is fed forward, so multi-hop accuracy degrades multiplicatively with hop count unless every hop is verified before the next retrieval.
When to use itWhere it fits
- Questions that explicitly chain facts, like comparative or transitive questions across documents.
- Research or due-diligence tasks where evidence must be assembled from several independent sources.
- Knowledge bases spanning many separate documents with no single one holding the full answer.
- QA benchmarks or products where accuracy on compound questions is a measured requirement.
When NOT to use itLimits & anti-patterns
- Single-fact lookups where one retrieval pass is already sufficient, since decomposition just adds latency.
- Latency-sensitive interactive use cases where multiple sequential retrieval hops are too slow.
- Corpora too small or too shallow to contain genuinely chained facts worth decomposing for.
Trade-offsAdvantages & costs
Advantages
- Correctly answers compound questions that single-pass RAG gets wrong or incomplete.
- Each hop's retrieval is more targeted than one broad query trying to cover everything.
- Decomposition traces are inspectable, aiding debugging and trust.
- Composable with reranking or agentic loops for further precision.
Trade-offs & costs
- Multiple sequential retrieval and generation calls increase latency and cost.
- Errors in early hops propagate and compound into the final answer.
- Decomposition quality depends heavily on prompt engineering.
- Harder to evaluate end-to-end than single-hop retrieval accuracy.
ExampleIn the real world
A financial research assistant answers "what regulatory risk does the supplier of Company X's main chip vendor face" by first identifying the vendor, then the supplier, then retrieving regulatory filings for that entity.ToolsHow to implement it
- LlamaIndex Query Decompositionbuilt-in sub-question query engine for multi-hop retrieval.
- LangGraphstate machine orchestration for chaining retrieval steps with intermediate results.
- Ragasevaluation metrics for multi-hop faithfulness and answer correctness.
- HotpotQA-style benchmarksstandard test sets for validating multi-hop retrieval pipelines.
Cost & effortWhat it takes
Cost and latency scale with the number of hops, since each adds a retrieval and often an LLM call; moderate to high engineering effort to build reliable decomposition and stopping logic.