Home › RAG - Retrieval-Augmented Generation › Advanced (rerank, hybrid)
📚 · Ground

Advanced (rerank, hybrid)

Blend keyword and vector search, then rerank before the model ever sees the context.

In one line

Cast a wider retrieval net, then use a cross-encoder to keep only what actually matters.

ConceptWhat it is

Advanced RAG layers two fixes onto naive retrieval: hybrid search, which combines sparse keyword matching like BM25 with dense vector similarity so exact terms and semantic meaning both count, and reranking, which passes the wider candidate set through a cross-encoder that scores query-chunk pairs jointly rather than by cosine distance alone.

It exists because pure embedding similarity often ranks a superficially similar chunk above the actually relevant one, and reranking recovers precision that a single vector pass throws away.

How it worksThe mechanics

The system first retrieves a broad candidate set, say top-50, using both a keyword index and a vector index in parallel, merges and deduplicates the results, then runs a cross-encoder reranker over the candidates to score true relevance, and finally passes only the top-3 to top-5 reranked chunks into the LLM prompt.

At a glanceSee it

Advanced (rerank, hybrid) diagram
Advanced (rerank, hybrid) diagram 1

Why reranking beats cosine alone — a bi-encoder scores query and chunk in isolation, while a cross-encoder reads them jointly for a sharper but costlier relevance signal.

Advanced (rerank, hybrid) diagram 2

The two-stage funnel — a wide high-recall net first, then a narrow high-precision rerank, with the reminder that stage two can only reorder what stage one retrieved.

When to use itWhere it fits

  • Large, heterogeneous corpora where naive top-k pulls in near-miss chunks.
  • Domains with precise terminology, like legal or medical, where exact keyword hits matter as much as semantic meaning.
  • Production systems where hallucination cost is high enough to justify an extra latency hop.
  • Any case where naive RAG evaluation scores are visibly mediocre on relevance.

When NOT to use itLimits & anti-patterns

  • Tiny, curated corpora where naive top-k already returns near-perfect chunks and reranking adds cost with no benefit.
  • Ultra-low-latency use cases like autocomplete, where an extra cross-encoder pass blows the latency budget.
  • Very early prototypes where the priority is proving the concept, not squeezing out precision.

Trade-offsAdvantages & costs

Advantages
  • Meaningfully higher retrieval precision than naive vector-only search.
  • Hybrid search catches exact-match terms, like part numbers, that embeddings blur.
  • Reranking is a bolt-on that does not require re-architecting the index.
  • Reduces downstream hallucination by feeding the model cleaner context.
Trade-offs & costs
  • Extra latency and compute cost from the reranking pass.
  • More moving parts to tune: fusion weights, candidate count, reranker choice.
  • Cross-encoder rerankers do not scale to huge candidate sets as cheaply as vector search.
  • Requires its own evaluation harness to prove the added complexity is worth it.

ExampleIn the real world

A legal research tool retrieves contract clauses with hybrid BM25 plus vector search, then reranks with Cohere Rerank so a query about "indemnification cap" surfaces the exact clause instead of a generally related one.

ToolsHow to implement it

  • Cohere Rerankmanaged cross-encoder reranking API purpose-built for this second-pass step.
  • Elasticsearch or OpenSearchbattle-tested BM25 keyword search to fuse with vector results.
  • Weaviatenative hybrid search combining sparse and dense retrieval in one query.
  • Ragasevaluation framework to measure whether reranking actually improved context precision.

Cost & effortWhat it takes

Moderate cost and latency: adds a reranker call per query on top of retrieval, meaningful engineering effort to tune fusion and reranking thresholds, worth it once relevance failures show up in evaluation.

A living map of modern AI — kept current every morning