1. Inventory what will actually be indexed
2. Establish the permission model you are inheriting
3. Decide what happens at query time
4. Prove it with a test matrix, before launch
5. Agree what gets logged, and who reads it
What this looks like when the access question is the project
Questions to put to any firm, including us
A knowledge assistant is an index plus a permission model, and the permission model is the part that gets deferred. An assistant that answers from a document the asker could not open has not leaked a file, it has leaked the contents, which is worse because nothing in an access log shows it. Fill this in before the index is built. Every row is a decision somebody has to make, and making them afterwards means rebuilding.
A worksheet for deciding who may see what a knowledge assistant answers: source inventory, the permission model you inherit, query-time enforcement, a test matrix, and what gets logged.
Permission-aware retrieval means an assistant can only use documents the person asking is already entitled to read, enforced at query time against the live permission model rather than at index time against a copy of it. Without it, an assistant answers from whatever was indexed, so a correct answer can disclose the contents of a file the asker could never open.
Not the systems. The content inside them, because permissions live at the item level and that is where the exposure sits. Most teams discover here that the index is wider than the use case.
Shared drives, wikis, ticketing, email, chat, the CRM, recorded calls. Name a person per source. A source with no owner cannot have its access questions answered, so it either gets an owner or it comes out of scope.
HR files, legal privilege, M&A folders, salary data, anything under a separate retention obligation. Write the exclusion list before the inclusion list; it is the shorter document and the more useful one.
This is the uncomfortable one. Most shared drives have accumulated access nobody has reviewed. An assistant inherits that drift and makes it queryable, which is how a permission problem that existed quietly becomes one that answers questions.
How much of the corpus is misfiled, duplicated, or contradictory
Sample 50 documents and count. The share matters because retrieval surfaces the wrong-but-plausible version as confidently as the right one, and a reader cannot tell which they got.
You are not designing a permission model. You are inheriting several and reconciling them, which is a different and harder job.
If the answer is a service account, the assistant can read everything that account can and no per-user restriction is possible downstream. Write the answer down even if it is uncomfortable, because this single field decides whether the rest of the checklist is achievable.
Per-user resolution against the live source is the only model that stays correct. A copy of the permission graph taken at index time is wrong the first time someone changes a share, and nothing alerts you.
Someone leaves a project, a folder is restricted, a person changes role. State the mechanism and the lag. A lag of days is usually acceptable if it is known; a lag nobody has measured is not.
Nested groups, external collaborators, and shared links are where inherited models break. Name the cases your implementation does not handle, so they can be excluded rather than discovered.
Four behaviours to choose deliberately. Defaults are rarely what a reviewer would have picked.
Filtering the candidate set before retrieval is the safer design, because a post-filter on results means the restricted content was retrieved and passed through the model before being removed. State which one you have.
Silence, a generic no-answer, or an explicit 'there is material you cannot access'. The third is more useful and it also confirms the document exists, which is itself disclosure in some contexts. Pick on purpose.
Whether every answer cites its sources, and whether the citation is clickable
Citations are the only way a reader can check an answer, and a link the reader cannot open is a visible, checkable permission boundary rather than a silent one.
Retention, whether it leaves your tenancy, whether it trains anything. This is the section every enterprise security questionnaire reuses, so write it once and properly.
The only part of this checklist that produces evidence rather than intent. Build the matrix with real accounts and real restricted documents.
An executive, a line manager, an individual contributor, a contractor, and a leaver whose access was revoked. Five accounts covers most models.
A known-restricted document per tier, with the expected result
The expected result for each account and document pair, written down before the test. A test without a pre-stated expectation grades itself.
An instruction planted inside a document the assistant will retrieve. This is the vector that does not care about your permission model, because it arrives inside content the assistant was entitled to read.
Revoke an account's access and query again after your stated propagation lag. This is the test that catches an index-time permission copy, and it is the one most often skipped.
An assistant that cannot be audited cannot be approved twice. The second approval is the one this section buys.
Who asked, what was retrieved, what was withheld and why. The withheld entry is the one that matters in an investigation and the one usually missing.
A log nobody is permitted to read is not a control. Name the role.
What pattern should notify someone: repeated withheld results, queries against the exclusion list, a spike from one account. Pick thresholds now rather than after an incident.
A new source added to the index, a change to the authentication model, a new permission tier. One line each. Without it the original approval quietly stops covering what is running.
A federal agency across six regional offices had the same problem in its statutory form: requests arriving for information that existed somewhere, with a legal obligation to release what was releasable and withhold what was not, and no reliable way to tell the difference quickly.
Governance over what existed and who could see it, before any retrieval was built on top
These separate a firm that has shipped permission-aware retrieval from one that has shipped a demo over a shared drive.
Do you filter the candidate set before retrieval, or filter the results afterwards, and why did you choose that?
Are permissions resolved per user against the live source at query time, or copied into the index?
Show me the test matrix from a previous build, including the revoked-access case and its propagation lag.
What does the assistant do when the answer exists and the asker is not entitled to it?
What is logged when content is withheld, and who is permitted to read that log?
How much of our corpus do you expect to exclude, and who decides the exclusion list?
An assistant that can only use documents the person asking is already entitled to read, enforced at query time against the live permission model rather than at index time against a copy. Without it the assistant answers from whatever was indexed, so a correct answer can disclose the contents of a file the asker could never open. Nothing in a file access log records that, which is what makes it hard to detect after the fact.
It is weaker, and a reviewer will notice. Filtering afterwards means the restricted content was retrieved and passed through the model before being removed, so the boundary depends on the filter rather than on the retrieval. Filtering the candidate set before retrieval is the design that holds. If you have the weaker version, state it and say why rather than leaving it to be discovered.
Because it is wrong the first time anyone changes a share, and nothing alerts you. Someone leaves a project, a folder is restricted, a contractor's access is revoked, and the index keeps answering from the old model. The test that catches this is the revoked-access case: revoke an account, wait your stated propagation lag, and query again.
What should the assistant say when the answer exists but is not permitted?
Choose deliberately between silence, a generic no-answer, and an explicit statement that relevant material exists but is not accessible. The third is the most useful to a legitimate user and it also confirms the document exists, which is itself disclosure in some settings. There is no universally right answer, which is exactly why it should be a written decision rather than a default.
No, and the two should not be conflated. A permission model controls what the assistant may retrieve. Indirect prompt injection arrives inside content the assistant was entitled to retrieve, so it bypasses the permission question entirely. Both belong in the test matrix, and prompt injection remained first on the 2026 edition of the OWASP Top 10 for LLM Applications.
More than feels comfortable at first, and that is the right direction. HR files, legal privilege, M&A material, anything under a separate retention obligation. A narrow index with a clear boundary gets approved; a wide index with a filter in front of it is the proposal that stalls, because a reviewer cannot bound the worst case.
Then that is the finding, and it is better to have it now. An assistant inherits whatever drift has accumulated in your shared drives and makes it queryable, which turns a quiet permission problem into one that answers questions. Sample the sources, measure the drift, and decide whether to fix it or exclude the source. Both are legitimate; proceeding without doing either is not.
Retrieval that cites what it used, with permissions enforced at query time.
The written scope and the adversarial testing behind a sign-off.
FOIA response from 34 days to 8 across six regional offices.