Weaviate is a vector-native database that fuses HNSW vector search with BM25 keyword search and lets you plug in embedding, reranking, and generative modules, running either self-hosted or as managed cloud.
ConceptWhat it is
Weaviate is an open-source (BSD-3), purpose-built vector database written in Go. It stores each object together with its vector embedding, indexes those vectors with HNSW for fast approximate nearest-neighbor search, and keeps a separate inverted index for keyword lookups and metadata filters. It exists because general-purpose databases bolt vectors on as an afterthought; Weaviate is designed vector-first, so semantic retrieval, keyword ranking, and structured filtering all live in one engine.
Its defining trait is a modular architecture. Vectorizer modules can embed your data at import time by calling a model provider, generative modules wire retrieved context straight into an LLM for RAG, and reranker modules re-score candidates. Hybrid search combines BM25 and vector scores through a fusion step, and features like multi-tenancy and replication push it toward high scale.
How it worksThe mechanics
On import, each object optionally passes through a vectorizer module that produces an embedding; the vector goes into the HNSW graph while the object's properties populate the inverted index. At query time you choose a mode: nearVector/nearText for pure semantic search, BM25 for keyword, or hybrid, which runs both and fuses the two score lists. Any filters are resolved through the inverted index first (pre-filtering), so the HNSW traversal only walks candidates that already satisfy the constraints. The top-k results can then be handed to a generative module, which templates them into an LLM prompt and returns a grounded answer in the same round trip.
At a glanceSee it
When to use itWhere it fits
- You need hybrid retrieval where exact keyword matches (product codes, names, acronyms) and semantic similarity both matter for the same query.
- You want embedding, reranking, or RAG generation handled inside the database via modules instead of stitching separate services together.
- You are building multi-tenant applications and want per-tenant isolation with the option to scale out through sharding and replication.
- You value an open-source core with a clear path to a managed offering if you later want to shed operational work.
When NOT to use itLimits & anti-patterns
- Your corpus is small or your team wants zero infrastructure — a lightweight embedded store or a managed serverless index is less to run.
- You already run a relational or document database and just need modest vector support; a vector extension there avoids adding a new system.
- You lack the operations capacity to size RAM, tune HNSW, and manage upgrades and backups for a self-hosted cluster.
- Your workload is pure keyword search with no semantic component, where a dedicated lexical engine is a simpler fit.
Trade-offsAdvantages & costs
Advantages
- First-class hybrid search with configurable BM25-plus-vector fusion, avoiding a separate keyword engine.
- Modular design that integrates vectorization, reranking, and generative RAG directly, cutting glue code.
- Efficient pre-filtering, so metadata constraints narrow candidates before the ANN traversal rather than after.
- Open-source and self-hostable, with a compatible managed cloud for teams that want to offload ops.
Trade-offs & costs
- Self-hosting carries real operational burden: HNSW is RAM-resident, so memory sizing, sharding, and upgrades demand attention.
- Module calls to external model providers add latency and separate, metered costs on top of the database.
- The breadth of GraphQL, REST, and gRPC APIs plus module configuration is a learning curve for newcomers.
- High-recall HNSW settings raise memory and index-build cost, forcing an accuracy-versus-resource trade-off.
ExampleIn the real world
A support team builds an assistant over its knowledge base. Articles are imported with a vectorizer module that embeds each one, while fields like product line and locale go into the inverted index. A user asks about an error code for a specific product in French. The app issues a hybrid query filtered to that product and locale: BM25 pins the exact error code while the vector side pulls semantically related troubleshooting steps, and the two lists are fused into one ranking. The filtered top passages then flow through a generative module that prompts an LLM to write a grounded answer citing those articles, all in a single round trip.
ToolsHow to implement it
- Weaviate Cloudthe managed, hosted version of the same engine.
- Official Python and TypeScript/JavaScript clients (with gRPC transport in current versions).
- LangChainand LlamaIndex integrations that expose Weaviate as a vector store or retriever.
- Built-in vectorizer, reranker, and generative modules (for example text2vec and generative connectors to model providers).
Cost & effortWhat it takes
The open-source core is free to run, but self-hosting means you carry the cost: HNSW keeps vectors in RAM, so memory is the dominant line item, plus effort for sharding, backups, and version upgrades. Weaviate Cloud shifts that operational load to a managed service, typically priced by the volume of vector dimensions stored and the SLA tier. Either way, calls made by vectorizer and generative modules are billed separately by the model provider. The effort profile is friendly at prototype scale — a client plus a module gets you retrieving quickly — but production self-hosting requires genuine capacity planning because the index is memory-bound.