Home › Embeddings & Vector Search › Elasticsearch
🧭 · Models

Elasticsearch

Elasticsearch: bolt kNN vector search onto a battle-tested BM25 search engine for hybrid retrieval at scale

In one line

Elasticsearch lets you add a dense_vector field and HNSW-based kNN to a mature Lucene search engine, giving you semantic search, keyword search, filtering, and rank fusion inside one cluster.

ConceptWhat it is

Elasticsearch is a distributed search and analytics engine built on Apache Lucene, best known for fast full-text keyword search (BM25) over JSON documents. It is not a purpose-built vector database; instead it takes an extension approach — you add a dense_vector field to an existing index and gain approximate nearest neighbor (kNN) search backed by an HNSW graph, running alongside the inverted index you already operate.

It exists because most teams adopting semantic search already run Elasticsearch for logs, catalogs, or site search and would rather not stand up a separate datastore. Co-locating vectors with text unlocks hybrid retrieval — combining BM25 and kNN scores, typically fused with Reciprocal Rank Fusion (RRF) — which usually beats either method alone, while metadata filtering, role-based access, aggregations, and operational tooling all carry over for free.

How it worksThe mechanics

You define a mapping with a dense_vector field specifying its dimensions and similarity metric (cosine, dot product, or L2), then embed each document with an external or built-in model and index the vectors, which Lucene stores as a per-segment HNSW graph next to the normal inverted index. At query time you embed the incoming query and issue a kNN query — optionally pre-filtered on metadata so the graph search only traverses matching candidates — and you can pair it with a BM25 match query, fusing the two rankings via RRF. Elasticsearch returns scored, filtered hits; you tune recall versus speed with the num_candidates knob, or drop to exact brute-force scoring for small result sets, and re-embedding into a new mapping requires a reindex.

At a glanceSee it

Elasticsearch diagram

When to use itWhere it fits

  • You already run Elasticsearch and want to add semantic search without introducing a new datastore or ops surface.
  • You need hybrid keyword-plus-vector retrieval with rich metadata filtering expressed in a single query.
  • You want production maturity out of the box: RBAC, snapshots, monitoring, and horizontal multi-tenant scale.
  • Your workload mixes search, logs/analytics, and RAG retrieval that can share the same cluster.

When NOT to use itLimits & anti-patterns

  • You want a lightweight, purpose-built vector store with minimal operations for a prototype or small project.
  • You are chasing extreme-scale pure-ANN with tight latency and memory budgets, where a dedicated vector DB is more efficient.
  • Your team has no appetite for cluster operations like heap sizing, shard and segment management, and provisioning RAM for the vector index.
  • You expect to swap embedding models frequently, since each change forces a costly reindex.

Trade-offsAdvantages & costs

Advantages
  • Reuses existing infrastructure: one system covers keyword, vector, filtering, aggregations, and security.
  • Strong hybrid retrieval with BM25, kNN, and RRF fusion, plus mature filtered search that constrains the graph traversal.
  • Scales horizontally through sharding and replication, with proven production ops, monitoring, and snapshots.
  • Deep ecosystem: Kibana, official clients, LangChain and LlamaIndex integrations, and the ELSER sparse model option.
Trade-offs & costs
  • Heavier cluster operations: shard and segment management, JVM heap sizing, and substantial off-heap RAM so the HNSW graph and vectors stay resident in the filesystem page cache.
  • Not purpose-built for vectors, so it can trail dedicated vector databases on raw ANN latency and memory efficiency at very large scale.
  • Changing the embedding model or vector dimensions requires a full reindex.
  • Licensing and cost picture is layered: Elastic License 2.0, an AGPL open-source option, or managed Elastic Cloud pricing.

ExampleIn the real world

A retailer already uses Elasticsearch for product-catalog keyword search. To add semantic relevance, they extend the product mapping with a dense_vector field, embed each item's title and description using a sentence-transformer model, and index the vectors alongside the existing text fields. When a shopper searches "cozy autumn jacket," the query is embedded and run as a filtered kNN (in stock, within a price band) combined with a BM25 match, and the two rankings are fused with RRF — so semantically similar items surface even without exact wording, while precise SKU and brand matches still rank strongly. No new datastore is introduced, and the team keeps its existing Kibana dashboards and access controls.

ToolsHow to implement it

  • Elasticsearchthe dense_vector field type and kNN query, backed by Lucene HNSW.
  • Kibanaand the official clients (elasticsearch-py, the Elasticsearch JS client) for indexing, queries, and observability.
  • LangChainand LlamaIndex ElasticsearchStore integrations for wiring retrieval into RAG pipelines.
  • ELSER(Elastic Learned Sparse EncodeR) for built-in sparse semantic retrieval without an external embedding model.

Cost & effortWhat it takes

The open-source core is free to self-host, so the dominant cost is operational rather than licensing: enough off-heap RAM to keep the HNSW graph and vectors resident in the filesystem page cache, careful JVM heap, shard, and segment sizing, and reindexing effort whenever the embedding model changes. Elastic Cloud and its serverless option offload much of that ops burden for a per-resource fee. For teams already operating a cluster the marginal cost of adding vectors is modest; standing up Elasticsearch purely to serve vectors is heavier than adopting a dedicated, managed vector service.

A living map of modern AI — kept current every morning