🧭 · Models

Qdrant

An Apache-2 licensed vector database written in Rust, built for fast filtered similarity search at scale.

In one line

Qdrant is an open-source, Rust-built vector store whose filterable HNSW index keeps similarity search fast even under heavy metadata filtering.

ConceptWhat it is

Qdrant is a purpose-built vector database written in Rust and released under the permissive Apache-2 license, available both as self-hosted open source and as a managed Qdrant Cloud service. It stores points — each an id, one or more vectors, and a JSON payload of metadata — and answers nearest-neighbor queries over them in milliseconds.

It exists because generic databases index scalars and text, not high-dimensional embeddings, and because naive filtering bolted onto an approximate index quietly wrecks recall. Qdrant's answer is a filterable HNSW index that applies payload filters during graph traversal rather than before or after it, so tight, high-selectivity filters stay both fast and accurate. It also supports hybrid search over sparse and dense vectors and several quantization modes that trade a little accuracy for large memory savings.

How it worksThe mechanics

You create a collection with a chosen distance metric and vector size, then upsert points carrying their embeddings and payload; Qdrant builds an HNSW graph incrementally and can apply scalar, product, or binary quantization to shrink the in-memory footprint. A query sends a vector plus an optional payload filter, and the engine walks the HNSW graph while enforcing the filter conditions inline, returning the top-k points with their scores and payload. For hybrid retrieval the Query API runs a sparse and a dense search and fuses the two result lists with a rank-fusion step.

At a glanceSee it

Qdrant diagram

When to use itWhere it fits

  • RAG or semantic search over large corpora where queries almost always carry metadata filters like tenant, language, or date.
  • Teams that want to self-host on their own infrastructure without a per-vector SaaS bill.
  • Latency-sensitive services needing fast approximate search combined with strict, high-selectivity filtering.
  • Hybrid retrieval that blends keyword-style sparse vectors with dense embeddings for better recall.

When NOT to use itLimits & anti-patterns

  • Small collections of a few thousand vectors, where a flat scan or pgvector inside your existing Postgres is simpler.
  • Teams wanting the widest catalog of turnkey connectors and reference integrations, where a more established managed service may ship faster.
  • Workloads needing strong transactional guarantees across mixed relational and vector data in one store.
  • Cases where you would rather not run and scale any stateful infrastructure at all.

Trade-offsAdvantages & costs

Advantages
  • Fast Rust engine with low, predictable latency and no garbage-collection pauses.
  • Filterable HNSW keeps recall high even under tight payload filters.
  • Permissive Apache-2 license with a genuinely usable single-binary self-host path.
  • Built-in hybrid search and quantization to control memory cost at scale.
Trade-offs & costs
  • Fewer turnkey integrations than the largest managed incumbents, so more glue code to write.
  • Self-hosting distributed clusters brings sharding, replication, and backup ops burden.
  • HNSW is memory-hungry; large collections force quantization or careful capacity planning.
  • Rich payload filtering is powerful but adds schema and index-tuning decisions to get right.

ExampleIn the real world

A team building support-documentation search chunks a few million help-center passages, embeds each into a dense vector, and also generates a sparse vector for keyword overlap. Each chunk becomes a Qdrant point whose payload records product, version, and language. At query time the app searches with the user's embedded question plus a filter of product equals current and language equals the user's locale; the filterable HNSW returns the top passages in milliseconds without scanning irrelevant tenants. As traffic grows the team enables scalar quantization to cut memory and moves from a self-hosted Docker container to Qdrant Cloud without changing client code.

ToolsHow to implement it

  • Qdrant Cloudthe managed offering with hosted clusters, scaling, and backups.
  • FastEmbedQdrant's lightweight library for generating dense and sparse embeddings.
  • LangChain and LlamaIndexboth ship Qdrant vector-store integrations for RAG pipelines.
  • Official client SDKsPython, Rust, JavaScript, and Go clients over REST and gRPC.

Cost & effortWhat it takes

The open-source engine is free to run from a single Docker container, so a proof of concept costs only the compute and RAM the HNSW index needs; memory, not license, is the real bill at scale, which is exactly why quantization matters. Qdrant Cloud adds managed tiers priced by cluster size, trading dollars for the ops effort of sharding, replication, and backups. Overall effort is moderate: easy to start, more involved once you run a distributed cluster yourself.

A living map of modern AI — kept current every morning