Vector databases store millions of embeddings and answer nearest-neighbor queries in milliseconds.
ConceptWhat it is
A vector database is a data store optimized to index and query high-dimensional embedding vectors, returning the closest matches to a query vector instead of exact-match rows. It exists because relational databases scan linearly and cannot efficiently answer "what is semantically closest to this" across millions of vectors.
These systems combine an approximate nearest neighbor index with metadata filtering, replication, and updates, so applications can search by meaning while still filtering by attributes like date or category.
How it worksThe mechanics
Embeddings are inserted along with metadata; the database builds an index, commonly HNSW or IVF, that organizes vectors so a query vector can find near neighbors without comparing against every stored vector, then it returns the top-k closest items ranked by distance, optionally filtered by metadata predicates.
At a glanceSee it
Inside the HNSW index box — a greedy hop repeats until no neighbor is closer, then the search steps down layer by layer to the dense bottom graph.
Choosing an index is a workload decision — trading exactness, recall, and memory, with over-compression as the recall-killing failure mode.
When to use itWhere it fits
- RAG pipelines needing fast retrieval over large document sets.
- Recommendation systems matching users to items by similarity.
- Deduplication or anomaly detection across large embedding sets.
- Semantic search products serving low-latency queries at scale.
When NOT to use itLimits & anti-patterns
- Small datasets under a few thousand vectors, where a flat in-memory search or pgvector is simpler.
- Workloads that are purely relational, where a normal SQL database is a better fit.
- Strict transactional consistency requirements that vector-first databases do not prioritize.
Trade-offsAdvantages & costs
Advantages
- Sub-second search across millions or billions of vectors.
- Supports metadata filtering alongside similarity search.
- Scales horizontally with managed cloud offerings.
- Integrates directly with RAG and LangChain-style pipelines.
Trade-offs & costs
- Adds a new piece of infrastructure to operate and monitor.
- Approximate indexes trade some recall for speed.
- Re-indexing needed after embedding model changes.
- Cost grows with vector count and dimensionality.
ExampleIn the real world
Notion and Shopify both use Pinecone-backed vector search to power semantic search and recommendations across millions of documents and products.
ToolsHow to implement it
- Pineconefully managed vector database with simple scaling.
- pgvectorPostgres extension for teams wanting vectors alongside relational data.
- Weaviateopen-source, supports hybrid search and multimodal out of the box.
- Milvusopen-source, built for very large-scale deployments.
Cost & effortWhat it takes
Managed options run tens to hundreds of dollars monthly for small workloads, scaling with vector count; self-hosted pgvector is cheapest but needs ops effort.