📚 · Ground

RAG-Fusion

Rewrite one question into several, retrieve for each, then fuse the ranked lists into one.

In one line

Generate several phrasings of a query, retrieve for each in parallel, and merge the results with reciprocal rank fusion so relevant chunks any single phrasing missed still rise to the top.

ConceptWhat it is

RAG-Fusion attacks the fragility of single-query retrieval. A user's exact wording is only one of many ways to ask a question, and dense retrieval returns a different set of chunks for each phrasing, so a slightly awkward query can miss the passage that actually holds the answer. RAG-Fusion uses an LLM to expand the original query into several alternative phrasings (multi-query), runs an independent retrieval for each, and then combines the separate ranked lists with Reciprocal Rank Fusion (RRF) into one consolidated ranking.

The insight behind RRF is that a chunk ranking reasonably well across many of the query variants is more trustworthy than one that ranks first for a single lucky phrasing. RRF scores each document by summing 1 divided by k plus its rank across every list it appears in, where k is a small constant, commonly 60. This rewards consensus across queries and dampens the noise of any one retrieval, lifting recall without needing a trained reranker.

How it worksThe mechanics

Given the incoming question, an LLM prompt produces a handful of reworded and reframed queries that cover synonyms, sub-aspects and different levels of specificity. Each generated query is sent to the vector store independently, producing its own top-k ranked list of chunks. Every document is then scored by reciprocal rank fusion, which adds up 1 divided by k plus rank for each list the document appears in, so items ranked highly by multiple variants float to the top. The fused list is truncated to the best few chunks, and only those are packed into the final prompt the generator sees.

At a glanceSee it

RAG-Fusion diagram
RAG-Fusion diagram 1

How reciprocal rank fusion actually scores a chunk — each list contributes the reciprocal of k plus rank, so broad moderate consensus outscores a single top hit.

RAG-Fusion diagram 2

Diagnosing why the first retrieval missed decides the payoff — RAG-Fusion recovers phrasing-driven misses but only adds cost when the answer is absent, and drifting variants can boost a wrong chunk on false consensus.

When to use itWhere it fits

  • Recall-critical search where missing the right passage is expensive — support knowledge bases, research, discovery over large corpora.
  • Users phrase questions inconsistently or use vocabulary that differs from the language of the source documents.
  • Corpora where the same concept is described many ways, so no single query reliably surfaces the best chunk.
  • When you want a recall boost without training, hosting or paying for a cross-encoder reranker.

When NOT to use itLimits & anti-patterns

  • Tight latency budgets — an extra LLM call to write the variants plus N parallel retrievals add real round trips.
  • Cost-sensitive, high-volume endpoints where N-times retrieval and an extra generation per request is prohibitive.
  • Narrow, well-curated corpora where a single query already returns the correct chunk and fusion adds cost for nothing.
  • When the failure is precision, not recall — RRF widens the net; sharpening what comes back is a reranker's job.

Trade-offsAdvantages & costs

Advantages
  • Higher recall than single-query RAG — relevant chunks one phrasing misses are recovered by another and fused back in.
  • Simple and model-agnostic — RRF needs no training, weights or labelled data, only the ranks each retrieval already produces.
  • Robust to messy user wording, since consensus across variants smooths out any one bad or ambiguous query.
  • Composes cleanly as a front stage ahead of hybrid search and reranking in a larger pipeline.
Trade-offs & costs
  • Multiplies retrieval cost — N queries per request, plus one extra LLM call just to generate them.
  • Added latency from the query-generation step and the fan-out of parallel searches before any answer starts.
  • Generated variants can drift off-topic, pulling in irrelevant chunks that fusion then has to absorb.
  • Improves recall but not precision — bad chunking or a genuinely thin corpus is not fixed by fusing more lists.

ExampleIn the real world

An internal engineering help desk answers questions over years of runbooks and incident write-ups. A query like "why is the checkout service timing out" retrieves a thin set of chunks, because past incidents described the same failure as "payment gateway latency" or "order API 504s". With RAG-Fusion, an LLM expands the question into four variants covering those alternate phrasings, each retrieves its own passages, and reciprocal rank fusion merges them so a runbook that never uses the word "checkout" still surfaces near the top. The answer now cites the incident that single-query retrieval had buried several pages down.

ToolsHow to implement it

  • LangChainMultiQueryRetriever generates the query variants, and EnsembleRetriever combines the result sets with weighted reciprocal rank fusion.
  • LlamaIndexQueryFusionRetriever implements RAG-Fusion directly, with a reciprocal-rerank fusion mode over multiple generated queries.
  • Elasticsearch / OpenSearchnative RRF to combine several query result sets, such as keyword and vector, into a single ranking.
  • Qdrantits Query API exposes native RRF fusion to merge multiple prefetch result lists server-side.

Cost & effortWhat it takes

Low engineering effort but a real per-query cost multiplier: expect one extra LLM call to write the variants plus N retrievals instead of one, so runtime cost and latency scale with the number of variants you choose. Fusion itself is nearly free — RRF is a handful of additions over integer ranks with no model to host. The main knobs are the number of variants, the top-k per variant and the RRF constant, all of which trade recall against cost and are best set against a retrieval-eval set rather than by feel.

A living map of modern AI — kept current every morning