OpenSearch is an open-source search engine whose k-NN plugin turns a familiar Lucene index into a filterable, hybrid vector store you can self-host at scale.
ConceptWhat it is
OpenSearch is a community-driven, Apache-2.0 licensed fork of Elasticsearch and Kibana, created in 2021 after Elastic relicensed its code. It kept the full lexical search stack — inverted indexes, BM25 scoring, aggregations, sharding — and layered on a k-NN plugin that stores dense embedding vectors alongside your normal documents and metadata. That combination is the whole point: you get a mature, self-hostable search engine that is also a vector store, rather than bolting a separate specialist database onto your stack.
Vector search runs as approximate nearest-neighbor over an HNSW graph, backed by one of three interchangeable engines — Lucene (pure Java, in-process), Faiss, or nmslib — so you trade recall, memory, and native-dependency footprint against each other. Because vectors live in the same index as keyword fields, OpenSearch does hybrid search (blending BM25 and vector scores through a normalization search pipeline) and filtered k-NN (restricting the neighbor search by structured predicates) natively, which is why teams reach for it when they already speak Elasticsearch.
How it worksThe mechanics
At ingest, each document is passed through an embedding model — either externally or via the built-in ml-commons/neural-search plugin — and the resulting vector is written to a knn_vector field while the raw text and metadata are indexed normally. You choose an engine (Lucene, Faiss, or nmslib) and HNSW parameters (ef_construction, m) that shape the proximity graph as segments are built and merged. At query time a request embeds the user's text, runs an approximate k-NN traversal of the HNSW graph to fetch candidate neighbors, optionally applies metadata filters during the search, and — for hybrid queries — a search pipeline normalizes and combines the vector score with a BM25 keyword score before returning the ranked hits, all spread across shards for horizontal scale.
At a glanceSee it
When to use itWhere it fits
- You already run OpenSearch or Elasticsearch for logs, text search, or observability and want to add semantic retrieval without introducing a second datastore.
- Your retrieval needs true hybrid search — strong keyword matching plus vector similarity — with per-query metadata filtering over tenant, date, permission, or category fields.
- You need an OSS, self-hostable option with a permissive license, or a managed equivalent, and want to avoid vendor lock-in or per-vector SaaS pricing.
- You are operating at large document and query volumes where mature sharding, replication, and cluster ops matter as much as raw vector recall.
When NOT to use itLimits & anti-patterns
- You want a low-effort, plug-and-play vector store for a small prototype — a purpose-built or embedded option gets you to first results with far less setup.
- Your team has no appetite for JVM cluster operations, heap and shard tuning, or capacity planning; OpenSearch rewards operational depth it also demands.
- You need cutting-edge ANN features (e.g. the newest quantization, disk-based indexes, or filtered-recall tricks) the instant they ship — specialist vector databases often lead here.
- Your corpus is tiny and static, where an in-process library index would be cheaper and simpler than standing up a cluster.
Trade-offsAdvantages & costs
Advantages
- One engine for keyword and vector: hybrid ranking and filtered k-NN work over the same documents, avoiding a separate system and a sync pipeline.
- Permissive Apache-2.0 license with no source-available strings, plus a broad ecosystem inherited from the Elasticsearch lineage.
- Flexible ANN backend — swap between Lucene, Faiss, and nmslib to trade recall, memory, and native dependencies per index.
- Battle-tested horizontal scale: sharding, replicas, rolling upgrades, snapshots, and role-based security are already part of the platform.
Trade-offs & costs
- Tuning complexity: getting good recall and latency means tuning HNSW
ef/m, shard counts, refresh, and JVM heap — a real operational skill. - HNSW is memory-hungry; large vector indexes push RAM and cost, and force decisions about quantization or engine choice.
- Segment merges and reindexing on graph parameters can be expensive, making some changes slow or disruptive on big clusters.
- As a general-purpose engine, its newest vector features can trail dedicated vector databases, and the API surface is heavier to learn.
ExampleIn the real world
A company building an internal knowledge assistant over policy documents, support tickets, and wiki pages already uses OpenSearch for full-text search. Rather than adding a separate vector database, they enable the k-NN plugin, add a knn_vector field to the existing document index, and backfill embeddings from a sentence-embedding model. Retrieval for the assistant issues a hybrid query: BM25 catches exact product names and error codes while vector similarity handles paraphrased questions, and a filter clause restricts results to the requesting user's department and to non-archived documents. They start on the Lucene engine to avoid native libraries, then move the largest index to Faiss when recall at scale matters, tuning ef_search until latency and answer quality both meet target.
ToolsHow to implement it
- OpenSearch k-NN pluginthe core vector field type and approximate/exact HNSW search over Lucene, Faiss, or nmslib.
- ml-commons and the neural-search pluginin-cluster model hosting and embedding generation plus hybrid search pipelines.
- LangChain(
OpenSearchVectorSearch) and LlamaIndex (OpensearchVectorStore) — RAG integrations that read and write the k-NN index. - Amazon OpenSearch Serviceand OpenSearch Dashboards — a managed hosting option and the Kibana-derived visualization and admin UI.
Cost & effortWhat it takes
The software is free under Apache-2.0, so there are no license or per-vector fees; the real cost is infrastructure and operations. HNSW indexes are memory-intensive, so RAM tends to dominate self-hosted spend, and you carry the ongoing effort of cluster sizing, shard and heap tuning, upgrades, and monitoring. A managed offering such as Amazon OpenSearch Service converts much of that operational effort into a predictable service bill while still leaving index and query tuning to you. Budget for a moderate-to-high engineering learning curve up front, and expect vector quality to depend on deliberate parameter tuning rather than out-of-the-box defaults.