🧭 · Models

Vespa

Vespa: one open-source engine unifying vector search, filtering, and machine-learned ranking at massive scale.

In one line

Vespa is a purpose-built engine that stores vectors alongside structured data and runs hybrid retrieval plus multi-phase ML ranking over billions of documents in a single system.

ConceptWhat it is

Vespa is an open-source (Apache-2) big-data serving engine, originally built at Yahoo, that combines a vector database, a lexical search engine, and a machine-learned ranking layer in one distributed system. Where most vector stores handle only approximate nearest-neighbor lookup, Vespa treats embeddings as first-class tensor fields that live next to structured attributes and full-text fields, so a single query can filter, match, and score against all of them at once.

It exists because production relevance is rarely pure similarity. Real ranking needs hybrid retrieval — blending HNSW vector recall with BM25 text matching and metadata filters — followed by expensive re-scoring with gradient-boosted trees or neural models. Vespa runs that whole pipeline server-side, close to the data, avoiding the round-trips and re-ranking glue that a separate vector store, search engine, and model server would require.

How it worksThe mechanics

Documents are fed into a content cluster where each node indexes vector fields with HNSW graphs, text fields with inverted indexes, and attributes for filtering; a schema declares the fields and one or more rank profiles. At query time a stateless container node parses the request, pushes matching and filtering down to every content node in parallel, and each node runs a cheap first-phase score (for example approximate nearest neighbor plus BM25) to gather top candidates. Those survivors go through a more expensive second-phase re-rank — often an ONNX or GBDT model over tensor features — and the merged, globally-ordered results return to the caller. Tuning relevance means editing the rank profile and re-deploying, not re-indexing the data.

At a glanceSee it

Vespa diagram

When to use itWhere it fits

  • You need hybrid search — vector plus keyword plus structured filters — evaluated together rather than stitched from separate systems.
  • Relevance depends on learned ranking: you want to run GBDT or neural re-rankers server-side over hundreds of candidate features.
  • Scale is large and latency-sensitive — tens of millions to billions of documents with high query throughput and low tail latency.
  • Your data changes continuously and you need real-time indexing alongside serving, not batch rebuilds.

When NOT to use itLimits & anti-patterns

  • You have a small corpus and simple similarity search — a lightweight library or managed store gets you there with far less setup.
  • The team wants a plug-in-and-go API and has no appetite for schemas, rank profiles, and cluster tuning.
  • You only need embeddings bolted onto an existing app and lack the operations capacity to run a distributed serving system.
  • Your workload is analytical or transactional rather than search and ranking — a general database fits better.

Trade-offsAdvantages & costs

Advantages
  • Genuine native hybrid ranking: vectors, text, and filters scored in one pass, no external re-ranker required.
  • Proven scale and low latency for very large, high-traffic corpora with real-time writes.
  • Rich ranking expressiveness — multi-phase scoring, tensor math, and embedded ONNX or GBDT models.
  • Open source with a managed Vespa Cloud option, so you can self-host or offload operations.
Trade-offs & costs
  • Steep learning curve: schemas, rank profiles, and its query and ranking language take real investment to master.
  • Operationally heavy to self-host — distributed content and container clusters need capacity planning and monitoring.
  • Overkill for small or simple use cases, where the setup cost dwarfs the benefit.
  • Smaller ecosystem and community than the most popular vector databases, so fewer tutorials and off-the-shelf recipes.

ExampleIn the real world

A retailer wants product search that respects both meaning and hard constraints. Each product is fed to Vespa with a text embedding, a title and description for BM25, and attributes like price, category, and in-stock status. A shopper query is embedded and issued as a hybrid request: an approximate-nearest-neighbor operator over the embedding, combined with keyword matching, and gated by a filter for in-stock items under a price cap. The first-phase rank profile mixes vector closeness with BM25; the second phase runs a gradient-boosted model over features such as margin, popularity, and recency to re-order the top few hundred candidates. The team improves conversion by editing the rank profile and re-deploying — the index and embeddings stay untouched.

ToolsHow to implement it

  • pyvespathe Python client for defining schemas, feeding documents, and querying from notebooks and pipelines.
  • Vespa CLIcommand-line tool for deploying applications and managing clusters locally or on Vespa Cloud.
  • ONNX Runtimeembedded to run neural embedding and ranking models inside second-phase scoring.
  • LangChain and LlamaIndexboth offer Vespa retriever integrations for RAG applications.

Cost & effortWhat it takes

The software is free under Apache-2, so the real cost is effort and infrastructure. Self-hosting means running and tuning a distributed cluster — sized to your document count, vector dimensions, and query volume — plus the engineering time to climb the schema and ranking learning curve. Vespa Cloud trades that operational burden for a usage-based bill covering compute and storage. Expect a meaningful upfront investment to reach production relevance, repaid at scale where its hybrid ranking and low-latency serving replace several separate systems.

A living map of modern AI — kept current every morning