🧭 · Models

Chroma

An open-source embedding database built for fast prototyping and local retrieval-augmented development.

In one line

Chroma is a developer-friendly, Apache-2 vector store that runs embedded on your laptop or as a managed cloud service, made for prototyping RAG quickly.

ConceptWhat it is

Chroma is an open-source embedding database released under Apache-2 that lets developers store documents, their embeddings, and metadata in named collections and query them by semantic similarity. It exists to make vector search feel like an ordinary library import rather than a piece of infrastructure: you spin one up in-process in a few lines of Python or JavaScript with no separate server to run.

Its design bias is developer experience and iteration speed. Chroma can auto-embed raw text through a pluggable embedding function, runs embedded on a laptop or as a client-server process, and offers a managed Chroma Cloud for teams that want a hosted endpoint. The trade-off is scale: it targets low-to-medium workloads and prototyping, not billion-vector production fleets.

How it worksThe mechanics

You create a collection and add records as raw text or precomputed vectors, each with an id and optional metadata; when you pass text, Chroma runs the collection's embedding function to produce vectors, then indexes them with HNSW (via hnswlib) for approximate nearest-neighbor search. A query embeds the query text the same way, walks the HNSW graph to find the closest vectors, and returns the top-k results ranked by distance, optionally narrowed by a metadata where filter or a document-substring where_document filter. Persistence writes the collection to local storage so it survives restarts.

At a glanceSee it

Chroma diagram

When to use itWhere it fits

  • Prototyping a RAG pipeline locally, where you want retrieval working in minutes rather than a provisioning ticket.
  • Small internal tools and demos over thousands to low millions of chunks that fit comfortably on one node.
  • Teaching, notebooks, and experiments where an embedded, zero-config store keeps the feedback loop tight.
  • Projects already using LangChain or LlamaIndex that want a sensible default vector store.

When NOT to use itLimits & anti-patterns

  • Very large-scale or high-QPS production retrieval, where a distributed store like Milvus, Qdrant, or managed Pinecone fits better.
  • Workloads that need first-class hybrid keyword-plus-vector fusion or reranking out of the box.
  • Teams already running Postgres who would rather keep vectors beside relational data with pgvector.
  • Strict multi-tenant isolation, heavy horizontal sharding, or strong transactional guarantees.

Trade-offsAdvantages & costs

Advantages
  • Fastest path from zero to working similarity search; embedded, with no server to operate.
  • Apache-2 open source with Python and JavaScript clients and tight LangChain and LlamaIndex integration.
  • Auto-embedding through pluggable embedding functions, so you can add raw text directly.
  • Supports metadata and document-substring filtering alongside vector search.
Trade-offs & costs
  • The open-source engine is historically single-node, so it is not built for very large-scale production traffic.
  • Hybrid keyword-plus-vector search is limited compared to Weaviate or Qdrant.
  • Scaling up often means moving to Chroma Cloud or migrating to a heavier store once data or QPS grows.
  • Fewer enterprise operational features such as replication, sharding, and fine-grained access control.

ExampleIn the real world

A two-person team building an internal documentation assistant embeds roughly forty thousand help-center chunks into a persistent Chroma collection running on a laptop, wiring it to a RAG chain in LangChain. During development they add and re-embed docs and requery in seconds, filtering by a product-area metadata field to keep answers scoped. When the tool graduates to company-wide use and the corpus grows toward tens of millions of chunks, they lift the same embedding-and-collection pattern onto a managed, horizontally scaling vector service, having used Chroma to move fast through the prototype phase.

ToolsHow to implement it

  • chromadbthe core Python package with embedded, persistent, and client-server modes.
  • Chroma Cloudthe managed hosted service for a shared team endpoint.
  • LangChain and LlamaIndexframeworks with built-in Chroma vector-store integrations.
  • hnswlib and sentence-transformersthe HNSW index and default local embedding model under the hood.

Cost & effortWhat it takes

The software is free and Apache-2 licensed; embedded local use costs only your own machine or a small VM, which is why prototyping on Chroma is nearly free. Managed Chroma Cloud adds usage-based hosting fees, and the larger real cost is usually the eventual migration effort when a prototype outgrows a single node.

A living map of modern AI — kept current every morning