Why post-filtering leaks
The obvious design is to retrieve the top k documents by similarity, then drop the ones the asker cannot see. It is wrong in two ways, and the second one is the dangerous one.
- It degrades badly. If eight of the top ten are filtered out, the model answers from the remaining two and sounds just as confident.
- It leaks through the answer even when it drops the document. Rankings, counts, and the model’s own summary of what it nearly said all carry information out. "I found several documents about the Rahman acquisition but cannot show them" has already told the reader that the acquisition exists.
The test that finds this: ask as a user with no entitlements at all. If any answer differs from what an empty corpus would produce, you are leaking.
Filter at query time instead
The permission set has to be part of the query, so an entry the asker cannot see is never a candidate and never influences the ranking.
- At ingest, store the document’s access control list alongside its vector: the group ids, not the user ids, or you will re-index on every staff change.
- At query time, resolve the asker to their group ids once, from your identity provider, not from anything the model can influence.
- Pass those ids as a hard pre-filter to the vector store, so the search space itself is restricted.
- Retrieve k from the filtered space. The ranking is now computed only over things the asker may see.
Every serious vector store supports this: a metadata pre-filter, a partition key, or a per-tenant namespace. If the one you have chosen does not, that is a reason to change it, not a reason to post-filter.
The parts people forget
- Chunks inherit permissions from their document, and derived artefacts inherit from their source. A summary index built over restricted documents is restricted.
- Permissions change. A revocation has to invalidate the index entry, not wait for the next full rebuild. Treat it as an event, not a batch.
- Deleted means deleted. A document removed from the source system must leave the index in the same transaction, or your assistant will be the only place it still exists.
- Citations are part of the answer. A cited title or path is a disclosure even when the body is withheld.
- Logs are a corpus too. Prompt and response logs contain retrieved content, and inherit its sensitivity.
What to test, and with whom
Permission bugs do not show up in accuracy metrics. They need their own tests, and the tests need to be adversarial.
- A user with no entitlements: every answer must be indistinguishable from an empty corpus.
- A user entitled to one document among many similar ones: the answer must not reflect the others, including in what it declines to say.
- A user whose access was revoked one minute ago.
- Direct attempts: ask the assistant to ignore its restrictions, to summarise what it cannot show, to say how many results it found.
- A shared document that later becomes private, and the reverse.
Run these on every change to retrieval, in the same gate as your accuracy evals. They fail loudly and they fail rarely, which is exactly what you want from a security test.
Want us to run this with you?
The Audit is this method pointed at your systems, with a costed build plan at the end of it.
Schedule call
