Thinklytics

Knowledge Retrieval · 10 min read · October 2026

Permission-aware retrieval: the design decision that decides approval

By Sean Majidi, Founder, Thinklytics

One design decision separates a knowledge assistant that can be approved from one that cannot: permissions resolved per user at query time, or a permission graph copied into the index. One test tells them apart, and it is the test most often skipped.

Two knowledge assistant proposals. Both index the same sources, both answer in the same format, both demo identically.

One can be approved and one cannot, and the difference is a single design decision that is usually made by default rather than deliberately.

The two models

Two ways to handle permissions, and only one of them holds

This is the single design decision that decides whether a knowledge assistant can be approved. It is usually made by default rather than deliberately.

Resolve per user at query timeCopy the permission graph at index time
What it checks againstThe live permission model in the source systemA snapshot taken when the document was indexed
When it is wrongWithin the propagation lag of a permission change, which you can measure and stateFrom the moment anyone changes a share, with nothing to signal it
The leaver caseAccess revoked, next query returns nothing from that contentKeeps answering from content the account can no longer open
CostA permission check per query, so latency and source-system loadCheap at query time, expensive the first time it is wrong
What a reviewer does with itCan bound the worst case, so can signCannot bound it, so the safe answer is no

Test it the same way every time: revoke one account's access to one known document, wait your stated propagation lag, and query again. That single test separates the two models, and it is the one most often skipped.

Source: Thinklytics engagement pattern across the retrieval and governance engagements in the case library.

Resolve per user at query time. Every query checks entitlements against the live permission model in the source system, then retrieves only from what that person may read.

Copy the permission graph at index time. When a document is ingested, the system records who could read it then, and uses that snapshot afterwards.

In a demo these are indistinguishable. They separate the first time anyone changes a share. Someone leaves a project, a folder is restricted, a contractor's access is revoked, and the snapshot keeps answering from the old model with nothing to signal that it is wrong.

The second model is cheaper, faster at query time, and easier to build. It is chosen far more often than it is argued for.

The test that tells them apart

Revoke one account's access to one known document. Wait whatever propagation lag the vendor states. Query again as that account.

A live per-user check returns nothing. An index-time copy keeps answering.

That is the whole test, it takes an afternoon, and it is the one most often skipped because it requires asking for a real restricted document and a real account rather than a sandbox.

If the vendor cannot state a propagation lag, that is itself the answer. An unmeasured lag means nobody has tested this, and a lag you cannot state is one a reviewer cannot accept.

Why this is worse than leaking a file

A file that gets shared leaves a trace. Someone opened it, the access log records it, and an investigation has somewhere to start.

An assistant that answers from a document the asker could not open has disclosed the contents without anyone opening the file. Nothing in the access log records it, because no access to the file occurred under that person's identity.

The disclosure is real and the audit trail is silent. That combination is why a security reviewer treats this differently from an ordinary permissions question, and why the honest version of the answer is an architecture decision rather than a policy statement. The broader version of that review conversation is in why security will not sign off your AI pilot.

Filter before retrieval, not after

Filter before retrieval, not after

A post-filter means the restricted content was retrieved and passed through the model before being removed. The boundary then depends on the filter rather than on the retrieval.

  • Resolve who is asking
  • Reduce the candidate set to what they may read
  • Retrieve only from that set
  • Answer, citing sources the asker can open
  • Log what was withheld, and why

The last step is the one nobody asks for until an investigation. A log of what was withheld is the only evidence that the boundary was enforced rather than merely configured.

Source: Thinklytics engagement pattern across the retrieval and governance engagements in the case library.

Given a per-user model, there is still a second decision inside it.

Reduce the candidate set to what the asker may read, then retrieve from that set. Or retrieve first, then remove results they may not see.

The second is weaker and a reviewer will notice. A post-filter means the restricted content was retrieved and passed through the model before being removed, so the boundary depends on the filter holding rather than on the retrieval being bounded. There are implementations where the post-filter is the only option available, and that is survivable if you state it and explain why rather than leaving it to be discovered.

Two more choices worth making on purpose.

What the assistant says when the answer exists and is not permitted. Silence, a generic refusal, or naming that inaccessible material exists. The third is most useful to a legitimate user and it also confirms the document exists, which is itself disclosure during an investigation or an acquisition. All three are defensible. A default is not.

Logging what was withheld. Who asked, what was retrieved, what was held back and why. The withheld entry is the one that matters in an investigation and the one usually missing, and it is the only evidence that the boundary was enforced rather than merely configured.

What a permission model does not fix

What a permission model does not solve

Four failure modes that survive a correct permission model, because they arrive inside content the assistant was entitled to read.

  • Indirect prompt injection through indexed content. An instruction planted in a document the assistant retrieves legitimately. Prompt injection stayed first in the 2026 edition of the OWASP Top 10 for LLM Applications, published 4 August 2026.
  • A superseded version answered as current. Both copies are permitted. Nothing in a permission model knows which one is right, which is why an owner per document set is part of the fix.
  • A confident answer assembled from two contradicting sources. Both readable, both retrieved, one answer. The citation is what lets a reader catch this, and it is the reason citations are not optional.
  • Inference across documents the asker may each read separately. Two permitted facts can combine into a conclusion nobody intended to publish. Rare, and the one to raise with legal rather than engineering.
  • Surfacing a document the asker could not open. This one the permission model does solve, and it is the only one of the five it solves.

Worth being explicit with a reviewer about which of these you have addressed. A proposal claiming permissions solve retrieval safety is making a claim a reviewer can disprove in one question.

Source: OWASP Top 10 for LLM Applications, 2026 edition, 4 August 2026; Thinklytics engagement pattern across the retrieval engagements in the case library.

Four failure modes survive a correct permission model, because they arrive inside content the assistant was entitled to read.

Indirect prompt injection through indexed content. An instruction planted in a document the assistant retrieves legitimately. The permission model is irrelevant to it. Prompt injection stayed first in the 2026 edition of the OWASP Top 10 for LLM Applications, published 4 August 2026.

A superseded version answered as current. Both copies are permitted. Nothing in a permission model knows which is right, which is why an owner per document set belongs in the same project.

A confident answer assembled from two contradicting permitted sources. Both readable, both retrieved, one answer, no indication of the conflict. Citations are what let a reader catch this, which is why they are not a nice-to-have.

Inference across separately permitted documents. Two facts a person may each read can combine into a conclusion nobody intended to publish. Rare, and a conversation for legal rather than engineering, but worth naming before someone else does.

Being explicit about which of these you have addressed is the difference between a proposal a reviewer trusts and one they start testing.

What the per-query check actually costs

Worth stating, because it is the honest reason the weaker design gets built.

Every query resolves entitlements against the live source, so you pay latency and load on systems that were not designed for that traffic pattern. On a large estate it is a real engineering problem rather than a configuration flag.

The usual answer is short-lived caching, and the distinction matters. Caching the resolved answer to "what may this user see" for a few minutes is defensible and still fails closed within minutes of a revocation. Caching the whole permission graph for a week is the index-time model wearing a different name.

Which data has to be ready underneath any of this is covered in the honest guide to LLM grounding data architecture.

What we would do first

Pick one restricted document and one account that should not see it. Ask whoever is proposing the build to demonstrate the revoked-access case on your data, with your propagation lag.

If they can, you are choosing between implementations. If they cannot, you have found the decision that was never made, and you have found it before the index was built rather than after.

What to put in the scope and the contract is in what to ask before a knowledge assistant build, and the retrieval access and permissions checklist is the worksheet we hand to whoever has to approve the access, which is usually a different person from the one approving the budget.

Delivery sits in RAG consulting for the retrieval and the citation layer, AI security consulting for the access model and the adversarial testing, and data governance consulting for the source ownership that makes currency answerable. The full set of work in this area sits under nobody can find what we already know.

Frequently asked questions

What is permission-aware retrieval?

An assistant that can only use documents the person asking is already entitled to read, enforced at query time against the live permission model in the source system. The alternative is copying the permission graph into the index when the document is ingested. Both look identical in a demo. They differ the first time someone's access changes, because the copy keeps answering from the old model and nothing signals it.

How do I tell which model a vendor has built?

One test. Revoke a single account's access to one known document, wait whatever propagation lag the vendor states, and query again as that account. A live per-user check returns nothing. An index-time copy keeps answering. If the vendor cannot state a propagation lag, that is the answer: nobody has measured it, which means nobody has tested this.

Should permissions filter before or after retrieval?

Before. Filtering the candidate set to what the asker may read, then retrieving only from that set, means restricted content is never retrieved. A post-filter on results means the restricted content was retrieved and passed through the model before being removed, so the boundary depends on the filter holding rather than on the retrieval being bounded. A reviewer will ask which one you have, and the weaker answer is survivable if you state it and say why.

Why is this worse than leaking a file?

Because there is no event to find afterwards. A file that gets shared leaves a trace in an access log. An assistant that answers from a document the asker could not open has disclosed the contents without anyone opening the file, so nothing in the access log records it. The disclosure is real and the audit trail is silent, which is the combination that makes a reviewer refuse.

What should the assistant say when the answer exists but is not permitted?

Choose deliberately between silence, a generic refusal, and an explicit statement that relevant material exists but is not accessible to you. The third is the most useful to a legitimate user and it also confirms the document exists, which is itself disclosure in some settings such as an investigation or an acquisition. All three are defensible. A default is not, because nobody decided it.

Does permission-aware retrieval make the assistant safe?

No, and claiming it does is a claim a reviewer can disprove in one question. Four failure modes survive a correct permission model, because they arrive inside content the assistant was entitled to read: indirect prompt injection through indexed content, a superseded version answered as current, a confident answer assembled from two contradicting permitted sources, and inference across documents the asker may each read separately. Prompt injection stayed first in the 2026 edition of the OWASP Top 10 for LLM Applications.

What does the per-query permission check cost?

Latency and load on the source system, since every query resolves entitlements against the live model. That is a real cost and the honest reason the cheaper design gets chosen. It is usually manageable with short-lived caching of the entitlement result rather than of the permission graph, and the distinction matters: caching what this user may see for a few minutes is defensible, caching the whole graph for a week is the failure mode.

What do citations have to do with permissions?

They turn an invisible boundary into a visible one. An answer that cites its sources lets the reader check which version they got, and a citation the reader cannot open is a correct, checkable signal that material exists outside their access, provided the answer itself did not use it. Without citations, a reader cannot tell a current source from a superseded one or a permitted one from a leaked one.

The work behind this

The retrieval work sits on top of the eighteen governance, privacy and security engagements in the case library, which are what settle who owns a source and who may read it. A federal agency across six regional offices took statutory information requests from 34 days to 8 on exactly that foundation.

Governance, privacy and security, 18 engagements.

Topics covered

  • permission aware retrieval
  • RAG security
  • document level permissions
  • knowledge assistant access control
  • query time filtering
  • retrieval architecture

Frequently asked questions

What is permission-aware retrieval?

An assistant that can only use documents the person asking is already entitled to read, enforced at query time against the live permission model in the source system. The alternative is copying the permission graph into the index when the document is ingested. Both look identical in a demo. They differ the first time someone's access changes, because the copy keeps answering from the old model and nothing signals it.

How do I tell which model a vendor has built?

One test. Revoke a single account's access to one known document, wait whatever propagation lag the vendor states, and query again as that account. A live per-user check returns nothing. An index-time copy keeps answering. If the vendor cannot state a propagation lag, that is the answer: nobody has measured it, which means nobody has tested this.

Should permissions filter before or after retrieval?

Before. Filtering the candidate set to what the asker may read, then retrieving only from that set, means restricted content is never retrieved. A post-filter on results means the restricted content was retrieved and passed through the model before being removed, so the boundary depends on the filter holding rather than on the retrieval being bounded. A reviewer will ask which one you have, and the weaker answer is survivable if you state it and say why.

Why is this worse than leaking a file?

Because there is no event to find afterwards. A file that gets shared leaves a trace in an access log. An assistant that answers from a document the asker could not open has disclosed the contents without anyone opening the file, so nothing in the access log records it. The disclosure is real and the audit trail is silent, which is the combination that makes a reviewer refuse.

What should the assistant say when the answer exists but is not permitted?

Choose deliberately between silence, a generic refusal, and an explicit statement that relevant material exists but is not accessible to you. The third is the most useful to a legitimate user and it also confirms the document exists, which is itself disclosure in some settings such as an investigation or an acquisition. All three are defensible. A default is not, because nobody decided it.

Does permission-aware retrieval make the assistant safe?

No, and claiming it does is a claim a reviewer can disprove in one question. Four failure modes survive a correct permission model, because they arrive inside content the assistant was entitled to read: indirect prompt injection through indexed content, a superseded version answered as current, a confident answer assembled from two contradicting permitted sources, and inference across documents the asker may each read separately. Prompt injection stayed first in the 2026 edition of the OWASP Top 10 for LLM Applications.

What does the per-query permission check cost?

Latency and load on the source system, since every query resolves entitlements against the live model. That is a real cost and the honest reason the cheaper design gets chosen. It is usually manageable with short-lived caching of the entitlement result rather than of the permission graph, and the distinction matters: caching what this user may see for a few minutes is defensible, caching the whole graph for a week is the failure mode.

What do citations have to do with permissions?

They turn an invisible boundary into a visible one. An answer that cites its sources lets the reader check which version they got, and a citation the reader cannot open is a correct, checkable signal that material exists outside their access, provided the answer itself did not use it. Without citations, a reader cannot tell a current source from a superseded one or a permitted one from a leaked one.

Related reading

If this is the problem you have