Access control in a RAG system has to happen at retrieval time, because a chunk the model has already read cannot be un-read.
ConceptWhat it is
Access control is the rule set that decides which parts of a corpus a particular person may receive an answer from. In a conventional application it is a filter on rows returned to a screen. In a retrieval system it is a filter on what enters the model's context in the first place.
The distinction is the whole subject. A permission check applied to the model's output is not access control — it is redaction after disclosure, and the model has already conditioned its answer on text the person asking was never entitled to.
How it worksThe mechanics
Every chunk carries the access metadata of the document it came from — owner, group, classification, and whatever the source system uses. The retriever is handed the identity of the person asking along with the query, and the vector search is constrained by that metadata before scoring, not after it.
The ordering matters and is easy to get wrong. Filtering after the top-k selection returns fewer than k results and quietly degrades answer quality as permissions tighten; filtering inside the search keeps k intact but needs an index that supports metadata predicates. Which one you have is a property of the vector store, so it is a choice made at the architecture stage, not at the end.
At a glanceSee it
The filter sits before retrieval. Anything it removes never reaches the model, so the answer is in scope by construction rather than by review.
When to use itWhere it fits
- Any corpus where two readers are entitled to different subsets — which is nearly every corpus inside a company.
- HR, legal, finance and clinical material, where the classification already exists on the source document.
- Multi-tenant products, where one customer's data reaching another is the failure that ends the product.
- Anywhere an auditor will ask you to prove that a given answer could only have used permitted sources.
When NOT to use itLimits & anti-patterns
- A single-classification public corpus, where every reader is entitled to everything and the filter is pure cost.
- As a substitute for redaction — access control decides who sees a document, not which sentences inside it are safe.
- As the only control on an agent that can call tools; a permitted document read by an agent that can send email is still an exfiltration path.
- When the access metadata on the source is known to be wrong; the filter will faithfully enforce a wrong answer.
Trade-offsAdvantages & costs
Advantages
- Makes the safety property structural — the model cannot leak what it was never given.
- Reuses permissions that already exist in the source system rather than inventing a second scheme.
- Produces an answerable question for a security review, with a demonstrable filter rather than a promise.
- Scales to multi-tenant without a separate index per tenant, when the store supports predicates.
Trade-offs & costs
- Post-filtering silently shrinks the result set, so answer quality degrades as permissions tighten and nothing reports it.
- Access metadata has to be carried through ingestion, chunking and embedding without being dropped — four places it can be lost.
- Re-indexing is required when permissions change, unless the filter reads them live.
- A permission model that is complex on paper becomes a query predicate that is slow in practice.
ExampleIn the real world
An internal policy assistant indexes the whole HR handbook alongside the compensation review files. Both live in the same store, and the compensation chunks carry a group tag. A manager asking about parental leave and a compensation partner asking about band midpoints hit the same index and the same model; the filter is what makes the first question safe to answer at all.
ToolsHow to implement it
- Vector stores with metadata filteringpgvector, Qdrant, Weaviate and Pinecone all support predicates inside the search rather than after it.
- The source system's own permissionsSharePoint, Drive or Confluence ACLs, mirrored into chunk metadata at ingest so there is one scheme, not two.
- OPA or Cedara policy engine, when the rule is richer than a group tag and you want it expressed once and testable.
- Row-level security in Postgresif the corpus already lives there, the database can enforce the filter it is best placed to enforce.
Cost & effortWhat it takes
Little runtime cost — a metadata predicate is cheap next to the embedding search around it, though a highly selective filter over a large index can force a wider scan. The real cost is ingestion engineering: carrying access metadata intact through every transform, and keeping it fresh when the source's permissions change. Budget that as a first-class part of the pipeline rather than a field added late.
What changedWhat changed here
- Your AI agents can now control your Google Home devices
Google launched early access to an MCP server for Google Home, letting AI agents like Claude and ChatGPT control connected devices, review camera summaries, and query smart home activity in natural language. This is a concrete template for how a consumer hardware platform exposes itself to third-party agents — useful if you're designing tool surfaces for your own product.
Three kinds of claim, strongest first. Signal runs every morning.