Home › Embeddings & Vector Search › MongoDB Atlas Vector Search
🧭 · Models

MongoDB Atlas Vector Search

Store embeddings inside your MongoDB documents and query them with native aggregation stages

In one line

Atlas Vector Search adds HNSW approximate-nearest-neighbor search to MongoDB so vectors, metadata, and operational data live and are queried together in one place.

ConceptWhat it is

MongoDB Atlas Vector Search is a feature of the managed Atlas cloud database that lets you store vector embeddings in the same documents as your normal application data and search them by semantic similarity. Instead of running a separate vector store beside your primary database, you add an array field of floats to existing documents and build a vector search index over it. Queries run through the standard MongoDB aggregation pipeline via the $vectorSearch stage, so a similarity search can be chained with $match, $lookup, and $project like any other query.

It exists to collapse the operational gap between an app's system of record and its retrieval layer. For teams already on MongoDB, embeddings, source text, and business metadata stay in one engine with one connection string, one access-control model, and one backup, avoiding the dual-write and sync problems of keeping a standalone vector database consistent with a document store. Under the hood the index is HNSW built on Apache Lucene, and it supports metadata pre-filtering and hybrid search that fuses full-text $search with vector results.

How it worksThe mechanics

You generate an embedding for each piece of content and write it as a numeric array into a document field, then define a vector search index declaring the field, its dimension count, and a similarity metric (cosine, dot product, or Euclidean). At query time you embed the incoming query with the same model and issue a $vectorSearch aggregation stage, passing numCandidates (how wide the HNSW graph is explored), limit, and an optional filter on indexed metadata fields so only matching documents are considered. Lucene's HNSW graph returns approximate nearest neighbors ranked by similarity score, which you can expose via $meta, and subsequent pipeline stages join related collections or reshape the output; for hybrid retrieval you run a parallel $search text query and combine the two rankings with reciprocal-rank fusion.

At a glanceSee it

MongoDB Atlas Vector Search diagram

When to use itWhere it fits

  • Your operational data already lives in MongoDB Atlas and you want retrieval without standing up and syncing a separate vector database.
  • You need to combine semantic similarity with tight metadata filtering (tenant, product line, language, ACLs) enforced in the same query.
  • You want hybrid retrieval that blends keyword precision (exact error codes, SKUs, names) with semantic recall in one pipeline.
  • Retrieval results must be joined to live records via $lookup so answers reflect current prices, status, or inventory rather than a stale copy.

When NOT to use itLimits & anti-patterns

  • You are not on Atlas and do not want to be — the feature is delivered as part of the Atlas managed cloud platform, not as a drop-in for an arbitrary self-hosted deployment.
  • You need multi-cloud or on-prem portability, or want to avoid coupling your vector layer to a single managed service.
  • You are at extreme vector scale with specialized needs (billion-scale, custom quantization tuning, GPU indexing) where a dedicated engine gives finer control.
  • Your data has no home yet and lightweight local prototyping matters more than integration — an embedded or open-source store is faster to start.

Trade-offsAdvantages & costs

Advantages
  • One system for documents, metadata, and vectors: a single connection, security model, and backup, with no dual-write consistency problem.
  • Vector search is just an aggregation stage, so it composes with $match, $lookup, and $project and reuses your existing MongoDB skills and drivers.
  • Strong pre-filtering and hybrid full-text plus vector fusion out of the box, backed by mature Lucene indexing.
  • Managed operations, plus optional dedicated Search Nodes to isolate search workload from primary database traffic, and scalar or binary quantization to shrink memory cost.
Trade-offs & costs
  • Atlas-managed and cloud-tied — choosing it couples your retrieval layer to the managed platform and its hosting model rather than a portable, self-run stack.
  • Cost scales with cluster tier and, for isolation and headroom, with separate Search Nodes; large float vectors are RAM-hungry unless quantized.
  • ANN indexing adds resource pressure to a database also serving transactional load if you do not separate the workload.
  • Less specialized tuning surface than purpose-built vector engines for the highest-scale or most latency-sensitive cases.

ExampleIn the real world

A support team runs its ticketing and knowledge base on Atlas. Each KB article is embedded and the vector is written into the same document that already holds the article's product line, language, and visibility flags. When an agent opens a ticket, the app embeds the customer's question and runs a $vectorSearch that pre-filters to the ticket's product line and the agent's language, returning the ten most similar articles; a parallel $search catches exact error codes the customer pasted, and the two rankings are fused so both the semantically related how-to and the article naming that precise code surface together. A $lookup then attaches each article's current published status so retired guidance is dropped before the results reach the answer-drafting step, all without a second datastore to keep in sync.

ToolsHow to implement it

  • LangChainthe MongoDBAtlasVectorSearch integration for retrieval and RAG chains.
  • LlamaIndexan Atlas vector store connector for indexing and query engines.
  • Voyage AIembedding and rerank models (now part of MongoDB), alongside OpenAI or Cohere embeddings.
  • mongoshthe Atlas UI, and official drivers for defining vector indexes and writing $vectorSearch pipelines; Spring AI for JVM apps.

Cost & effortWhat it takes

Cost is driven by your Atlas cluster tier plus, optionally, dedicated Search Nodes that isolate search from primary traffic, so pricing tracks the database you are already paying for rather than a standalone vector service. Raw float vectors consume significant RAM at scale; scalar or binary quantization materially cuts that footprint and cost with a modest recall trade-off. Integration effort is low for teams already on Atlas: add a vector field, define an index, and write one aggregation stage. The larger, recurring cost is operational discipline around embeddings — re-embedding when you change models and managing index build time as data grows — not new infrastructure to run.

A living map of modern AI — kept current every morning