Home › RAG - Retrieval-Augmented Generation › Hybrid (keyword+vector)
📚 · Ground

Hybrid (keyword+vector)

Run BM25 and vector search together, fuse the scores, and catch what either alone would miss.

In one line

Pair keyword retrieval with vector retrieval and merge their rankings so exact terms and meaning both make it into the context.

ConceptWhat it is

Hybrid retrieval runs two retrievers over the same corpus and blends their output: a keyword index scored with BM25 for exact lexical matches, and a dense vector index scored by embedding similarity for semantic matches. It exists because vector-only RAG has one specific blind spot — embeddings compress meaning, so a rare literal token like a product SKU, an API error code, or an unusual acronym often fails to pull back the one chunk that actually contains it, even when a plain keyword search would find it instantly.

Rather than choosing between precision on exact terms and recall on paraphrased questions, hybrid keeps both by fusing the two ranked lists into a single list before any chunk reaches the model. The dense side handles "how do I reset my password" phrased fifty different ways; the sparse side guarantees that a string like "ORA-01017" or "clause 7.3.1" is never silently dropped.

How it worksThe mechanics

The corpus is indexed twice up front — once into an inverted keyword index and once into a vector store built with the same embedding model you query with. At request time the query is sent to both indexes in parallel, each returning its own ranked candidate list on an incompatible scoring scale. A fusion step reconciles them: most commonly reciprocal rank fusion, which scores each document by its position in each list so no cross-scale normalization is needed, or a weighted score blend with a tunable alpha that leans the result toward keyword or vector. The fused top-k chunks are concatenated into the prompt, and the model generates a grounded answer from context that carries both exact and semantic hits.

At a glanceSee it

Hybrid (keyword+vector) diagram
Hybrid (keyword+vector) diagram 1

Keyword and vector fail on opposite query types, so hybrid’s real value is keeping both candidate sets alive instead of betting the retrieval on one retriever.

Hybrid (keyword+vector) diagram 2

Reciprocal rank fusion turns each list position into 1 over k-plus-rank and adds them, which is why it merges BM25 and cosine results without ever putting their raw scores on the same scale.

When to use itWhere it fits

  • Corpora where exact identifiers matter — SKUs, error codes, ticket numbers, legal clause references, drug or gene names.
  • Retrieval that must serve both natural-language questions and precise term lookups from the same endpoint.
  • An existing vector RAG pipeline that mostly works but embarrassingly misses on rare proper nouns and acronyms.
  • Enterprise, legal, medical, or e-commerce search where a single missed exact match is a visible, reportable failure.

When NOT to use itLimits & anti-patterns

  • Purely conversational question answering over prose, where semantic similarity alone already covers recall.
  • Small or homogeneous corpora where either retriever alone tests just as well, leaving the second index as dead weight.
  • Hard latency or infrastructure budgets that cannot absorb a second index, a second query path, and fusion overhead.
  • When a cross-encoder reranker on vector-only results already fixes the precision gap more cheaply than a whole second index.

Trade-offsAdvantages & costs

Advantages
  • Recovers the exact-match precision that vector-only RAG quietly loses on names, codes, and acronyms.
  • Keeps the semantic recall that keyword-only search never had, so you get both strengths rather than a compromise.
  • Reciprocal rank fusion needs no score normalization, making it easy to add without deep tuning.
  • Supported natively in most modern vector databases and retriever libraries, so adoption is incremental.
Trade-offs & costs
  • Two indexes to build, store, and keep in sync — reindex one and forget the other and the halves drift apart.
  • A second retrieval path adds latency and compute to every query.
  • Weighted-blend fusion introduces an alpha weight that must be tuned per corpus and can bury signal if set wrong.
  • Harder to debug: when the two retrievers disagree, tracing why a given chunk ranked where it did takes more work.

ExampleIn the real world

A support assistant sits over an internal engineering knowledge base, and users paste literal failures like "deploy fails with ORA-01017." Vector-only retrieval drifts toward semantically similar authentication articles and buries the single runbook page that names that exact code, so the answer comes back vague. Switching to hybrid, a BM25 index catches the literal "ORA-01017" token while the dense index still handles paraphrased questions like "login keeps getting rejected," and reciprocal rank fusion merges both lists so the exact-match runbook and the semantically relevant background both land in the top chunks the model reads.

ToolsHow to implement it

  • Weaviatenative hybrid search with a single alpha parameter blending BM25 and vector scores.
  • Elasticsearch / OpenSearchBM25 plus dense-vector kNN in one engine, with reciprocal rank fusion built in.
  • Qdrantfuses sparse and dense vectors natively, the sparse side standing in for the keyword signal.
  • LlamaIndex / LangChainretriever-fusion abstractions that wrap a keyword retriever and a vector retriever and merge them inside the RAG pipeline.

Cost & effortWhat it takes

Hybrid roughly doubles retrieval-side storage and compute versus vector-only, since every document lives in two indexes and every query hits both; with proper indexing the added latency usually stays modest. Engineering effort is low if your database supports hybrid natively and you lean on rank fusion, and moderate if you hand-roll a weighted blend and tune its alpha per corpus. The real recurring cost is operational rather than per-query: keeping the keyword and vector indexes rebuilt together so they never fall out of sync.

A living map of modern AI — kept current every morning