Home › Embeddings & Vector Search › Milvus / Zilliz
🧭 · Models

Milvus / Zilliz

Open-source, cloud-native vector database built to index billions of embeddings for similarity search at scale.

In one line

Milvus is the heavy-duty, open-source vector database for teams that need similarity search across billions of vectors.

ConceptWhat it is

Milvus is a purpose-built, open-source vector database (Apache-2 licensed) engineered for similarity search at massive scale — comfortably into the billions of vectors. Where a bolt-on like pgvector adds vector columns to an existing store, Milvus is a distributed, cloud-native system whose entire architecture is organized around approximate nearest-neighbor retrieval.

It separates compute from storage and splits work across specialized nodes, which is what lets it index and query far more vectors than a single machine could hold in memory. Zilliz Cloud is the fully managed service from Milvus's original creators, offering the same engine without the operational burden of running the cluster yourself.

How it worksThe mechanics

Vectors and their scalar fields are inserted and first buffered into growing segments; once a segment seals, Milvus builds a per-segment index — HNSW or IVF for in-memory speed, or DiskANN to keep billions of vectors on SSD — and persists it to object storage through a log broker. A query is routed by a proxy to query nodes, each searching its own segments in parallel with optional metadata filters and hybrid dense-plus-sparse scoring, and the partial top-k results are merged and ranked before returning.

At a glanceSee it

Milvus / Zilliz diagram

When to use itWhere it fits

  • Similarity search over hundreds of millions to billions of vectors that outgrow a single node.
  • RAG or semantic-search products needing horizontal scale with metadata filtering and hybrid retrieval.
  • Recommendation, deduplication, or visual-search workloads where memory cost forces on-disk indexing like DiskANN.
  • Teams that want an open-source core with the option to offload operations to Zilliz Cloud.

When NOT to use itLimits & anti-patterns

  • Small or mid-size corpora under a few million vectors, where pgvector or a lighter store is far simpler.
  • Teams without the platform capacity to run a distributed system with etcd, object storage, and a message queue.
  • Workloads that are mostly relational or transactional rather than similarity-first.
  • Quick prototypes where standing up a cluster outweighs the value of scale you do not yet need.

Trade-offsAdvantages & costs

Advantages
  • Scales to billions of vectors while holding sub-second query latency.
  • Broad index choice — HNSW, IVF, and DiskANN — so you can trade memory, recall, and cost per workload.
  • Native hybrid search and rich metadata filtering built into the engine.
  • Apache-2 open source with a drop-in managed path via Zilliz Cloud.
Trade-offs & costs
  • Heavier to operate than most vector stores — many moving parts and external dependencies.
  • Overkill for small datasets, where its distributed design adds cost without benefit.
  • Index tuning and re-indexing after embedding-model changes demand real expertise.
  • Self-hosted infrastructure and ops effort grow with vector count and cluster size.

ExampleIn the real world

A media-monitoring platform needs to search roughly two billion image and text embeddings to flag near-duplicate and visually similar content in near real time. The team deploys Milvus with DiskANN indexing so the bulk of vectors live on SSD rather than RAM, partitions collections by source and date for metadata filtering, and runs hybrid dense-plus-keyword queries to catch both semantic and exact matches. As traffic grows they migrate to Zilliz Cloud to hand off cluster operations while keeping the same query API.

ToolsHow to implement it

  • Zilliz Cloudfully managed Milvus from its creators, removing cluster operations.
  • pymilvusthe official Python SDK for schema, insert, and search operations.
  • Attuthe open-source GUI for managing and inspecting Milvus collections.
  • LangChain and LlamaIndexRAG frameworks with built-in Milvus vector-store integrations.

Cost & effortWhat it takes

The Milvus core is free and open source, so self-hosting cost is mostly infrastructure — compute nodes, object storage, and a message queue — plus meaningful ops effort to run and tune the cluster. Zilliz Cloud shifts that to a managed bill that scales with vector count, dimensionality, and query volume, trading dollars for far lower operational load. Either way the effort profile is higher than a single-node store, and it only pays off at scale.

A living map of modern AI — kept current every morning