Filtering before the search keeps k results; filtering after it silently returns fewer, and quality degrades exactly where the filter is most selective.
ConceptWhat it is
A metadata filter restricts retrieval to records matching structured criteria — this tenant, this department, documents after a date, this document type. It is how a general index serves a specific question, and it is the mechanism behind both tenant isolation and retrieval-time access control.
The design question is where the filter runs relative to the vector search, and the two orders are not equivalent. That distinction is the single most consequential thing on this page, because one of them degrades quietly.
How it worksThe mechanics
Post-filtering runs the vector search first, takes the top k, then discards non-matching results. It works with any index and needs nothing special. Its flaw is that it returns fewer than k whenever matches are sparse in the top results, so a narrow filter can return almost nothing while the index contains plenty of good matches further down.
Pre-filtering constrains the candidate set before scoring, so the search returns k results from the permitted subset. It needs an index that supports predicates during traversal, which is precisely what a vector database provides and a bare library does not. Some stores over-fetch adaptively as a middle path, widening the search until enough matches survive.
At a glanceSee it
Two orders, not equivalent. Post-filtering degrades exactly where the filter is most selective, and returns a short result set with no error.
When to use itWhere it fits
- Multi-tenant retrieval, where isolation is a correctness requirement rather than a preference.
- Access control at retrieval time, which is the only place it can be enforced properly.
- Time-bounded questions, where superseded documents would otherwise retrieve as readily as current ones.
- Any corpus mixing document types where the question implies one of them.
When NOT to use itLimits & anti-patterns
- As post-filtering on a selective predicate, which is the failure mode this page exists to name.
- For a criterion better expressed semantically, where a filter makes the system brittle.
- On a field whose values are unreliable, since the filter will faithfully enforce bad metadata.
- As the only isolation mechanism in a multi-tenant product, without a test that proves it holds.
Trade-offsAdvantages & costs
Advantages
- Turns one index into many logical corpora with no duplication.
- The mechanism that makes retrieval-time access control possible at all.
- Cheap on the query path when the store supports predicates natively.
- Improves precision by removing whole categories of irrelevant match before scoring.
Trade-offs & costs
- Post-filtering silently shrinks results, and nothing in the response says so.
- Requires metadata attached at ingest; adding a filterable field later means re-indexing.
- Highly selective filters can force a wider scan and slow the search rather than speeding it.
- Correctness depends entirely on metadata quality, which is usually assumed rather than checked.
ExampleIn the real world
A multi-tenant assistant filters retrieval to the asking customer after the search. Large customers work well. A customer with a few dozen documents gets almost nothing, because none of their chunks reach the global top fifty. The system appeared to work in testing, which was done with the largest account.
ToolsHow to implement it
- Qdrant, Weaviate or pgvectorstores that apply predicates during search rather than after it.
- Adaptive over-fetchwidening k until enough matches survive, when only post-filtering is available.
- A result-count assertionalerting when a filtered search returns fewer than k, which is the silent failure made loud.
- A tenant-isolation testproving one tenant's query cannot surface another's chunk, run on every build.
Cost & effortWhat it takes
Near zero at query time when supported natively. The cost sits at ingest, in attaching and maintaining the metadata, and in the index type chosen to support predicates. Engineering effort is low if the store supports it and significant if you are building over-fetch logic to compensate for a store that does not.