Hybrid search catches the exact matches keywords need and the meaning matches vectors need.
ConceptWhat it is
Hybrid search combines traditional keyword search, like BM25, with vector similarity search, then merges the two rankings into one result list. It exists because embeddings alone often miss exact terms, like product SKUs or rare proper nouns, that keyword search catches instantly.
By running both retrieval methods and fusing their scores, hybrid search gets the recall benefits of semantic matching without losing the precision of exact-term matching.
How it worksThe mechanics
A query is run simultaneously through a keyword index, like BM25 or Elasticsearch, and a vector index; both produce ranked candidate lists with different scoring scales, which are then combined using a fusion method such as reciprocal rank fusion or a weighted score blend, producing a single re-ranked result set.
At a glanceSee it
Reciprocal rank fusion scores a doc by its position in each list, not its raw score, so anything ranked well in both rises to the top without cross-scale tuning.
A query-type decision showing when hybrid actually earns its overhead — and how a mistuned blend weight can bury the very signal it was meant to protect.
When to use itWhere it fits
- Enterprise search where exact codes, names, or IDs must be found reliably.
- RAG systems needing both semantic recall and precise term matching.
- E-commerce search mixing brand names with descriptive queries.
- Legal or medical search where terminology precision matters.
When NOT to use itLimits & anti-patterns
- Pure natural-language question answering where semantic match alone suffices.
- Very small corpora where either method alone performs adequately.
- Systems where the added complexity of score fusion is not worth the marginal recall gain.
Trade-offsAdvantages & costs
Advantages
- Captures both exact-term precision and semantic recall.
- Reduces embarrassing misses on names, codes, and acronyms.
- Improves overall retrieval quality for RAG pipelines.
- Supported natively by most modern vector databases.
Trade-offs & costs
- Requires tuning fusion weights between keyword and vector scores.
- Adds latency from running two retrieval paths.
- More moving parts to maintain and monitor.
- Harder to debug when rankings disagree.
ExampleIn the real world
Elastic's hybrid search combines its native BM25 engine with dense vector kNN so customer support search finds both exact ticket IDs and semantically related past tickets.
ToolsHow to implement it
- Elasticsearch / OpenSearchnative BM25 plus dense vector fields in one engine.
- Weaviatebuilt-in hybrid search with tunable alpha fusion parameter.
- Vespaproduction-grade hybrid ranking at large scale.
- Qdrantsupports sparse and dense vector fusion natively.
Cost & effortWhat it takes
Roughly doubles retrieval compute versus vector-only search; latency stays sub-200ms with proper indexing, engineering effort is moderate for fusion tuning.