sqlite-vec is a tiny, dependency-free SQLite extension that adds exact brute-force vector search to a single database file, with no server and no ops.
ConceptWhat it is
sqlite-vec is a small, single-file SQLite extension that adds vector storage and similarity search to the world's most widely deployed database. Written in dependency-free C by the author of the earlier sqlite-vss, it loads into any SQLite build and exposes a virtual table (the vec0 module) where you store embedding vectors alongside ordinary columns. Because SQLite runs in-process with no server, sqlite-vec inherits that model: your vector search executes inside the same library call as the rest of your queries, on the same file, with zero network hops and zero infrastructure to operate.
It exists to make semantic search a feature you ship inside an application rather than a service you stand up. The trade is deliberate: it performs exact brute-force KNN — comparing the query vector against every stored vector — rather than building an approximate-nearest-neighbor graph. That keeps the code tiny and the results exact, at the cost of search time that grows linearly with the collection. It shines on-device and at the edge, where a full vector database would be overkill.
How it worksThe mechanics
You load the extension, then create a vec0 virtual table declaring an embedding column of a fixed dimension (float32, int8, or binary vectors are supported), plus any auxiliary or metadata columns you want to filter on. Your application computes embeddings with a separate model and inserts them as rows. At query time you embed the user's text into a query vector and run a normal SQL SELECT that matches against the embedding column and asks for the k nearest by distance; sqlite-vec scans the stored vectors, computes distances (cosine, L2, or Hamming for binary), and returns the top-k rows. Because it is all SQL, you can add WHERE clauses for metadata pre-filtering and join against an FTS5 full-text table to blend keyword and vector scores into hybrid search.
At a glanceSee it
When to use itWhere it fits
- On-device or edge apps — mobile, desktop, browser via WASM — that need offline semantic search with no backend.
- Small-to-moderate collections (thousands to low millions of vectors) where exact results matter more than sub-millisecond ANN latency.
- Local-first and privacy-sensitive products where embeddings must never leave the user's machine.
- Prototypes and RAG demos where you want vector search that ships as one file and one dependency.
When NOT to use itLimits & anti-patterns
- Large corpora (many millions to billions of vectors) where linear brute-force scans become too slow.
- High-QPS multi-tenant services that need a managed, horizontally scalable vector backend.
- Workloads demanding tunable ANN recall and latency trade-offs, like HNSW or IVF graphs, out of the box.
- Heavy concurrent write throughput, where SQLite's single-writer model becomes the bottleneck.
Trade-offsAdvantages & costs
Advantages
- Zero servers, zero ops: it is a library call on a single file that goes anywhere SQLite goes.
- Exact results — brute-force KNN has no recall loss to tune or worry about.
- Full SQL power: metadata filtering, joins, transactions, and hybrid search with FTS5 in one query.
- Tiny, permissively licensed (Apache-2.0 or MIT), dependency-free C that runs on mobile, WASM, and edge runtimes.
Trade-offs & costs
- Search time grows linearly with collection size — no ANN graph, so large sets get slow.
- Scales to modest data only; it is not a fit for very large or very high-throughput deployments.
- Single-writer concurrency inherited from SQLite limits write-heavy multi-user workloads.
- You bring your own embedding model and pipeline; it stores and searches vectors but does not generate them.
ExampleIn the real world
A field-service mobile app lets technicians search thousands of equipment manuals and past repair notes while working in basements with no signal. During the nightly sync, each document chunk is embedded and written into a vec0 table in the app's local SQLite file, alongside columns for equipment type and site. Offline, a technician types "compressor won't hold pressure"; the app embeds that phrase on-device, runs a single SELECT that pre-filters to the relevant equipment type via a WHERE clause and returns the ten nearest chunks by cosine distance, then joins an FTS5 table so exact part numbers still rank. Results appear instantly with no network call, and the brute-force scan stays fast because each device only holds that technician's slice of the corpus.
ToolsHow to implement it
- sqlite-vec — the core extension, with official bindings for Python, Node.js, Ruby, and the browser via WASM.
- SQLite FTS5 — the built-in full-text module you join against for hybrid keyword-plus-vector search.
- better-sqlite3 or Python's sqlite3 module — common host drivers used to load the extension and run queries.
- sentence-transformers or a hosted embeddings API — to generate the vectors sqlite-vec stores and searches.
Cost & effortWhat it takes
Cost is close to zero at rest: no service to run, no license fee, and the vectors live inside a file you already back up. The real budget line is engineering — you own the embedding pipeline, choose vector precision (int8 or binary quantization to shrink storage and speed scans), and must benchmark brute-force latency against your collection size to know when you have outgrown it. Effort to get a first working search is very low; effort to scale it past a few million vectors is effectively a migration to a dedicated ANN store.