Cast a wider retrieval net, then use a cross-encoder to keep only what actually matters.
ConceptWhat it is
Advanced RAG layers two fixes onto naive retrieval: hybrid search, which combines sparse keyword matching like BM25 with dense vector similarity so exact terms and semantic meaning both count, and reranking, which passes the wider candidate set through a cross-encoder that scores query-chunk pairs jointly rather than by cosine distance alone.
It exists because pure embedding similarity often ranks a superficially similar chunk above the actually relevant one, and reranking recovers precision that a single vector pass throws away.
How it worksThe mechanics
The system first retrieves a broad candidate set, say top-50, using both a keyword index and a vector index in parallel, merges and deduplicates the results, then runs a cross-encoder reranker over the candidates to score true relevance, and finally passes only the top-3 to top-5 reranked chunks into the LLM prompt.
At a glanceSee it
Why reranking beats cosine alone — a bi-encoder scores query and chunk in isolation, while a cross-encoder reads them jointly for a sharper but costlier relevance signal.
The two-stage funnel — a wide high-recall net first, then a narrow high-precision rerank, with the reminder that stage two can only reorder what stage one retrieved.
When to use itWhere it fits
- Large, heterogeneous corpora where naive top-k pulls in near-miss chunks.
- Domains with precise terminology, like legal or medical, where exact keyword hits matter as much as semantic meaning.
- Production systems where hallucination cost is high enough to justify an extra latency hop.
- Any case where naive RAG evaluation scores are visibly mediocre on relevance.
When NOT to use itLimits & anti-patterns
- Tiny, curated corpora where naive top-k already returns near-perfect chunks and reranking adds cost with no benefit.
- Ultra-low-latency use cases like autocomplete, where an extra cross-encoder pass blows the latency budget.
- Very early prototypes where the priority is proving the concept, not squeezing out precision.
Trade-offsAdvantages & costs
Advantages
- Meaningfully higher retrieval precision than naive vector-only search.
- Hybrid search catches exact-match terms, like part numbers, that embeddings blur.
- Reranking is a bolt-on that does not require re-architecting the index.
- Reduces downstream hallucination by feeding the model cleaner context.
Trade-offs & costs
- Extra latency and compute cost from the reranking pass.
- More moving parts to tune: fusion weights, candidate count, reranker choice.
- Cross-encoder rerankers do not scale to huge candidate sets as cheaply as vector search.
- Requires its own evaluation harness to prove the added complexity is worth it.
ExampleIn the real world
A legal research tool retrieves contract clauses with hybrid BM25 plus vector search, then reranks with Cohere Rerank so a query about "indemnification cap" surfaces the exact clause instead of a generally related one.ToolsHow to implement it
- Cohere Rerankmanaged cross-encoder reranking API purpose-built for this second-pass step.
- Elasticsearch or OpenSearchbattle-tested BM25 keyword search to fuse with vector results.
- Weaviatenative hybrid search combining sparse and dense retrieval in one query.
- Ragasevaluation framework to measure whether reranking actually improved context precision.
Cost & effortWhat it takes
Moderate cost and latency: adds a reranker call per query on top of retrieval, meaningful engineering effort to tune fusion and reranking thresholds, worth it once relevance failures show up in evaluation.