🧭 · Models

Rerankers

A second-pass model that re-scores retrieved results for precision before they reach the LLM.

In one line

Rerankers take a rough shortlist and put the truly best answer at the top.

ConceptWhat it is

A reranker is a model, usually a cross-encoder, that re-scores a shortlist of retrieved candidates by jointly reading the query and each candidate together, producing a more accurate relevance ranking than the initial vector search. It exists because embedding-based retrieval, which encodes query and document separately, is fast but less precise than models that compare them directly.

Rerankers are applied after an initial retrieval step narrows millions of documents down to a few dozen, since cross-encoders are too slow to run over an entire corpus.

How it worksThe mechanics

The initial retrieval stage, often ANN vector search, returns the top 50 to 100 candidates; each candidate is then paired with the original query and fed through a cross-encoder that outputs a single relevance score, and the candidates are re-sorted by that score before the final top few are passed to the LLM.

At a glanceSee it

Rerankers diagram
Rerankers diagram 1

The mechanism behind the accuracy gain — a bi-encoder scores two vectors it built in isolation, while a cross-encoder lets query and document tokens attend to each other, trading cacheability for precision.

Rerankers diagram 2

Not all rerankers are cross-encoders — listwise LLM rerankers reason across the whole shortlist and late-interaction models like ColBERT precompute document vectors, each landing at a different speed, cost, and reliability point.

When to use itWhere it fits

  • RAG pipelines where retrieval precision directly affects answer quality.
  • Search products where the top result matters more than average recall.
  • Any pipeline where initial retrieval returns noisy or loosely relevant candidates.
  • Multi-source retrieval that needs a unified relevance ranking.

When NOT to use itLimits & anti-patterns

  • Latency-critical paths where the extra model call is too slow.
  • Cases where initial vector search is already highly precise.
  • Very small candidate sets where reordering makes negligible difference.

Trade-offsAdvantages & costs

Advantages
  • Significantly improves precision of the final retrieved context.
  • Reduces irrelevant context reaching the LLM, cutting hallucination risk.
  • Works as a drop-in second stage after any retrieval method.
  • Can incorporate more nuanced signals than embedding similarity alone.
Trade-offs & costs
  • Adds latency, since it scores each candidate individually.
  • Only as good as the shortlist it receives from initial retrieval.
  • Extra cost per query compared to vector search alone.
  • Requires maintaining a second model in the pipeline.

ExampleIn the real world

Cohere's Rerank API is commonly bolted onto Pinecone or Weaviate retrieval pipelines to boost RAG answer accuracy by re-scoring the top 50 vector search hits before generation.

ToolsHow to implement it

  • Cohere Rerankmanaged cross-encoder reranking API, widely used in RAG stacks.
  • BGE Rerankeropen-source cross-encoder from BAAI, self-hostable.
  • Sentence-Transformers CrossEncoderopen-source cross-encoder models for reranking.
  • Ragasevaluation framework to measure reranking impact on retrieval quality.

Cost & effortWhat it takes

Adds tens to low hundreds of milliseconds per query and a per-call fee on managed APIs; worthwhile when answer quality matters more than raw latency.

A living map of modern AI — kept current every morning