🧭 · Models

FAISS

A library for fast similarity search over vectors — not a database, which is the whole distinction.

In one line

FAISS is the fastest way to search vectors in one process and gives you none of what a database provides, which is fine until you need any of it.

ConceptWhat it is

FAISS is a library for similarity search over dense vectors. It is not a database: there is no server, no query language, no permissions, no transactions and no durability beyond writing the index to a file yourself.

That is the trade, and it is a good one for a narrower set of cases than it gets used for. In-process search over a fixed corpus is extremely fast and has no operational surface at all. The moment the corpus changes often, or a second process needs the same index, or a filter has to be applied per user, the missing pieces have to be built.

How it worksThe mechanics

Vectors are added to an index whose type encodes the trade-off being made. A flat index compares against everything and is exact but linear in corpus size. IVF partitions the space into cells and searches a few, trading a little recall for a large speed gain. HNSW builds a navigable graph and is fast with good recall at higher memory cost. Product quantization compresses vectors so far more fit in memory, at some accuracy cost.

The index is built, then searched, and typically written to disk and loaded by the serving process. Updates are the weak point: many index types are not designed for incremental addition and deletion, so the common pattern is periodic rebuild rather than live mutation.

At a glanceSee it

FAISS diagram

The index type is the trade-off, chosen once. Most types are built rather than mutated, which is why a changing corpus wants a database instead.

When to use itWhere it fits

  • A fixed or slowly changing corpus that fits in the memory of one machine.
  • Batch and offline work — evaluation, clustering, deduplication, analysis.
  • Embedded use where adding a database dependency is disproportionate.
  • As the fast baseline to measure a managed vector store against, since it sets the speed ceiling.

When NOT to use itLimits & anti-patterns

  • When you need metadata filtering inside the search, which is where a real vector database earns its cost.
  • For frequently updated corpora, since most index types prefer rebuild to mutation.
  • Multi-tenant products needing per-user scoping, which has to be built entirely by hand.
  • When durability, backup, replication and access control matter, none of which it provides.

Trade-offsAdvantages & costs

Advantages
  • Extremely fast, and the reference implementation many alternatives are benchmarked against.
  • No server, no operational surface, no network hop on the search path.
  • Fine-grained control over the accuracy, speed and memory trade-off.
  • Mature, well-documented, and free.
Trade-offs & costs
  • Not a database — no persistence guarantees, permissions, transactions or replication.
  • Metadata filtering is not built in and hand-rolled post-filtering silently shrinks result sets.
  • Index types that support deletion do so awkwardly, so churn means rebuilds.
  • Scaling past one machine is your problem, not the library's.

ExampleIn the real world

A prototype runs on a flat FAISS index over a hundred thousand chunks and is fast and simple. Going multi-tenant requires filtering results to the asking customer, done by over-fetching and filtering afterwards. Tenants with small corpora start receiving fewer results than requested, and quality degrades in proportion to how little data a customer has — which is exactly backwards.

ToolsHow to implement it

  • faiss-cpu or faiss-gputhe library itself; GPU builds matter mostly for index construction at scale.
  • HNSW for most serving casesthe best default balance of recall and speed when memory allows.
  • pgvector or Qdrantthe answer when filtering, updates or permissions are required.
  • A recall measurement against a flat indexthe exact baseline that tells you what an approximate index is costing.

Cost & effortWhat it takes

Free, and the cost is memory. Vector count times dimensions times four bytes is the floor, before index overhead; quantization is the lever when that number is too large. Engineering effort is low to start and grows sharply the moment filtering or updates are needed, which is the point at which a database is cheaper overall.

A living map of modern AI — kept current every morning