Knowledge Retrieval · 10 min read · October 2026
What to ask before a knowledge assistant build
By Sean Majidi, Founder, Thinklytics
Eight questions, six of them about the permission model and the evaluation set, because those decide whether it gets approved and whether it works. Plus the acceptance criteria split into the four your security reviewer owns and the four your business owner does.
A proposal for a knowledge assistant is in front of you. The architecture diagram looks fine and the demo answered three questions correctly.
What decides this is narrower than the architecture: whether the permission question has been answered before, and whether anyone has written down what a right answer looks like.
What belongs in the scope
What belongs in the scope, and what gets added
The deferred column is where these builds overrun. Each item is reasonable and none of them makes the first answer correct sooner.
| In scope | Why | Deferred |
|---|---|---|
| One question type, from one audience, over a named source list | Narrow enough that a wrong answer is diagnosable. Support answering from the help centre, not everything from everywhere | Every document the company owns |
| The exclusion list, written before the inclusion list | Shorter, more useful, and it is what a reviewer reads first | A filter in front of a wide index |
| Permissions resolved per user at query time, with a stated propagation lag | The decision that makes the build approvable | Index-time permission copies, revisited later |
| Citations on every answer, linking to something the asker can open | The only way a reader can check an answer, and a visible permission boundary | A confidence score, which readers trust more than it deserves |
| An owner per document set, so currency has an answer | Governance, not tooling. Two permitted copies cannot be ranked without it | A company-wide taxonomy project |
| An evaluation set of real questions with known answers, built before the system | Without it there is no way to say whether it works, only whether it responds | Model selection, which matters far less than the corpus |
A 2026 benchmark over roughly 500,000 documents across nine real platforms found BM25 keyword ranking beating vector search by more than 17 points on correctness. Scope the corpus and the evaluation, not the embedding model.
Source: Enterprise retrieval benchmark as cited in our internal knowledge assistant analysis; Thinklytics engagement pattern across the retrieval engagements in the case library.
Six items. Two of them are the project.
One question type, from one audience, over a named source list. Support answering from the help centre. Not everything, from everywhere, for everyone. Narrow enough that a wrong answer is diagnosable rather than merely disappointing.
The exclusion list, written before the inclusion list. It is shorter, more useful, and it is what a reviewer reads first.
Permissions resolved per user at query time, with a stated propagation lag. The decision that makes the build approvable at all. Why, and how to test it, is in retrieval that respects who is asking.
Citations on every answer, linking to something the asker can open. The only mechanism by which a reader can check which version they got.
An owner per document set. Currency has no answer without it, and two permitted copies of the same policy cannot be ranked by any amount of retrieval sophistication.
An evaluation set of real questions with known answers, built before the system exists. The section below.
Deferred: every document the company owns, a company-wide taxonomy project, a confidence score that readers will trust more than it deserves, and model selection.
The evaluation set is the project
Without one you can establish that the system responds. You cannot establish that it is right.
The set should be real questions from the real audience, with answers agreed by whoever owns the content. Twenty to fifty is usually enough to be informative. Two rules about where it comes from.
A vendor-supplied evaluation set measures the vendor. It will be representative of what their system handles well, not of what your people ask.
A set written after the build is written around what the build already does. The questions it fails get softened into questions it passes, usually without anyone intending to.
So it is written first, by you, and the cost of writing it is the cheapest honest thing in the project.
This is also where the model question resolves itself. In a 2026 benchmark over roughly 500,000 documents across nine real enterprise platforms, deliberately seeded with misfiled documents, near-duplicates and conflicting information, BM25 keyword ranking beat vector search by more than 17 points on correctness. If the retrieval approach is that contested, the model on top of it is not where this is won. The detail is in internal knowledge assistants, what breaks.
Eight questions
Questions to ask before signing
Six on the permission model and the evaluation, because those are the two things that decide whether this gets approved and whether it works.
- Are permissions resolved per user at query time, or copied into the index?. If the answer is a copy, ask what the propagation lag is and whether it has been measured. An unmeasured lag is the answer.
- Do you filter the candidate set before retrieval, or filter the results?. Before is the design that holds. After means restricted content passed through the model before being removed.
- Show me the test matrix from a previous build, including the revoked-access case. Five test accounts across permission tiers, a known-restricted document each, expected results written before the test ran.
- What does the assistant do when the answer exists and the asker is not entitled to it?. A deliberate choice between silence, a generic refusal, and naming that inaccessible material exists. All three are defensible; a default is not.
- What is your evaluation set, and who wrote the known answers?. Real questions from the real audience, with answers agreed by someone who owns the content. A vendor-supplied evaluation set measures the vendor.
- What share of our corpus do you expect to exclude, and who decides?. A firm that expects to exclude nothing has not looked at the corpus.
- What is logged when content is withheld, and who may read that log?. The entry that matters in an investigation and the one usually missing.
- What is the residual annual run cost, and who owns the index after handover?. Re-indexing, permission sync, evaluation re-runs, and the owner's time. A proposal with no ongoing cost is hiding one.
The federal agency engagement is the shape to ask about: 2,400 statutory requests a year, response time from 34 days to 8, and $1.2M of annual penalties eliminated, built on governance over what existed and who could see it.
Source: Thinklytics case library, published engagement scopes and outcomes.
Six concern the permission model and the evaluation. That ratio is deliberate.
Are permissions resolved per user at query time, or copied into the index. Do you filter the candidate set before retrieval, or filter the results. Show me a test matrix from a previous build, including the revoked-access case. What does the assistant say when the answer exists and the asker is not entitled to it. What is your evaluation set, and who wrote the known answers. What share of our corpus do you expect to exclude, and who decides. What is logged when content is withheld, and who may read that log. What is the residual annual run cost, and who owns the index after handover.
The two warning signs, if the answers are thin: no stated propagation lag, and no test matrix from a previous engagement. Both mean the permission question has not been answered before, so it will be answered during your project, on your timeline, in front of your reviewer.
If the answers come from documents rather than records
One number to settle before signing, because it changes the economics rather than the design.
Vendors quote field-level extraction accuracy, which lands in the mid 90s. Whole-document accuracy in a 2026 benchmark of 10,000 filings ran 63% to 76%. Getting 19 of 20 fields right means getting the document wrong, and only the per-document figure decides whether a person still has to read everything that comes through.
Ask which figure the proposal is quoting. The full breakdown is in what AI document processing costs, and the design consequence, which is that the escalation path matters more than the accuracy rate, is in document processing: design the escalation path.
Acceptance criteria, in two groups
Acceptance criteria, written before the build
Tested on the last day by someone who was not in the project. Four are permission tests and four are answer-quality tests.
- Correctness on the agreed evaluation set, at or above a stated figure. Stated as a number on a set written before the build. Not a demo.
- Every answer carries citations the asker can open. Checkable by clicking. A citation to a document the reader cannot open is a feature, not a fault, provided the answer did not use it.
- The known-restricted document returns nothing, for every test tier. Five accounts, five documents, expected results pre-stated.
- The revoked-access case passes after the stated propagation lag. The test that distinguishes a live permission check from an index-time copy.
- Indirect prompt injection through indexed content is attempted and the result documented. A permission model does not address this, so it needs its own result rather than being assumed covered.
- Withheld-content logging live, with a stated retention and a named reader. Demonstrated by producing a real log entry, not by showing a settings page.
- A named owner per document set, recorded. Currency has no answer without it, and two permitted copies cannot be ranked.
- A re-review trigger: a new source, a new audience, a change to the auth model. One line each. Without it the approval stops covering what is running the first time the index widens.
The permission criteria are the ones to hand to security and the answer-quality criteria are the ones to hand to the business owner. They are usually two different people with two different definitions of working.
Source: Thinklytics case library, published delivery outcomes and acceptance measures per engagement.
Eight, and they split cleanly because they go to two different people with two different definitions of working.
Four for the security reviewer. The known-restricted document returns nothing for every test tier. The revoked-access case passes after the stated propagation lag. Indirect prompt injection through indexed content is attempted and the result documented, because a permission model does not address it and a reviewer will not assume it is covered. And withheld-content logging is live with a stated retention and a named reader.
Four for the business owner. Correctness on the agreed evaluation set at a stated figure. Citations on every answer, linking to something the asker can open. A named owner per document set. And a re-review trigger: a new source, a new audience, a change to the authentication model.
Handing each group to the right person is worth doing explicitly. A security reviewer asked to assess answer quality will defer, and a business owner asked to assess a permission model will approve it.
What the ongoing cost covers
Re-indexing as content changes. Permission synchronisation. Periodic evaluation re-runs, because correctness regresses quietly when the corpus moves. And the named owner's time.
Ask for a range with the assumptions stated rather than a single number. The pattern is the same one that catches out AI builds generally, and it is set out in what an AI system costs in year two.
What we would do first
Write twenty evaluation questions with known answers, before reading the proposal again. Real questions, from the people who will use this, with answers someone will stand behind.
Then ask whoever is proposing the build to run those twenty against whatever they would deliver, and to demonstrate the revoked-access case on one of your own restricted documents.
Those two artefacts, a scored evaluation set and a passed permission test, are the entire basis on which this decision should be made. Everything else in the proposal is description.
The retrieval access and permissions checklist is the worksheet we hand to whoever approves the access, which is usually not the person approving the budget.
Delivery sits in RAG consulting for retrieval and citation, AI security consulting for the access model and the adversarial testing, and data governance consulting for the source ownership that makes currency answerable. The full set of work in this area sits under nobody can find what we already know.
Frequently asked questions
What should I ask before commissioning a knowledge assistant?
Eight questions. Are permissions resolved per user at query time or copied into the index. Do you filter the candidate set before retrieval or filter the results. Show me a test matrix from a previous build including the revoked-access case. What does it say when the answer exists and the asker is not entitled to it. What is the evaluation set and who wrote the known answers. What share of our corpus do you expect to exclude and who decides. What is logged when content is withheld and who may read it. And what is the residual annual run cost and who owns the index afterwards.
What belongs in the scope?
One question type from one audience over a named source list. The exclusion list, written before the inclusion list. Permissions resolved per user at query time with a stated propagation lag. Citations on every answer, linking to something the asker can open. An owner per document set so currency has an answer. And an evaluation set of real questions with known answers, built before the system exists. Six items, and the last one is the one that makes the result checkable rather than merely demonstrable.
Why does the evaluation set have to exist first?
Because without it you can only establish whether the system responds, not whether it is right. The set should be real questions from the real audience with answers agreed by whoever owns the content. A vendor-supplied evaluation set measures the vendor, and a set written after the build is written around what the build already does. Twenty to fifty questions is usually enough to be informative.
What acceptance criteria should go in the contract?
Eight, splitting into two groups. Four for the security reviewer: the known-restricted document returns nothing for every test tier, the revoked-access case passes after the stated lag, indirect prompt injection through indexed content is attempted and the result documented, and withheld-content logging is live with a named reader. Four for the business owner: correctness on the agreed evaluation set at a stated figure, citations on every answer the asker can open, a named owner per document set, and a re-review trigger.
How much of our content should be excluded?
More than feels comfortable, and a firm that expects to exclude nothing has not looked at your corpus. HR files, legal privilege, M&A material, anything under a separate retention obligation. A narrow index with a clear boundary gets approved. A wide index with a filter in front of it is the proposal that stalls, because a reviewer cannot bound the worst case and declining is the defensible position.
Does the model choice matter?
Far less than the corpus and the evaluation. In a 2026 benchmark over roughly 500,000 documents across nine real enterprise platforms, BM25 keyword ranking beat vector search by more than 17 points on correctness. If the retrieval approach is that contested, the model sitting on top of it is not where the project is won. Scope the corpus, the permissions and the evaluation set, and treat the model as replaceable.
What is the ongoing cost?
Re-indexing as content changes, permission synchronisation, periodic evaluation re-runs to catch regression, and the named owner's time. Ask for a range with the assumptions stated rather than a number. A proposal with no year-two figure is quoting half the project, and the second half arrives as an operating expense somebody else has to find.
What is the warning sign a build will not get approved?
No stated propagation lag on the permission model, and no test matrix from a previous engagement. Both mean the permission question has not been answered before, which means it will be answered during your project, on your timeline, in front of your reviewer. The second warning sign is a scope defined by systems rather than by question types, because that produces an index nobody can bound.
The work behind this
Six engagements in the case library carry document processing and eighteen carry governance, privacy and security. The published outcomes include adverse event intake from 18 days to 4 hours across six unified sources, and statutory information requests from 34 days to 8 across six regional offices.
Document processing, 6 engagements.
Topics covered
- knowledge assistant build
- RAG project scope
- retrieval evaluation set
- AI assistant acceptance criteria
- internal chatbot procurement
- document assistant questions
Frequently asked questions
What should I ask before commissioning a knowledge assistant?
Eight questions. Are permissions resolved per user at query time or copied into the index. Do you filter the candidate set before retrieval or filter the results. Show me a test matrix from a previous build including the revoked-access case. What does it say when the answer exists and the asker is not entitled to it. What is the evaluation set and who wrote the known answers. What share of our corpus do you expect to exclude and who decides. What is logged when content is withheld and who may read it. And what is the residual annual run cost and who owns the index afterwards.
What belongs in the scope?
One question type from one audience over a named source list. The exclusion list, written before the inclusion list. Permissions resolved per user at query time with a stated propagation lag. Citations on every answer, linking to something the asker can open. An owner per document set so currency has an answer. And an evaluation set of real questions with known answers, built before the system exists. Six items, and the last one is the one that makes the result checkable rather than merely demonstrable.
Why does the evaluation set have to exist first?
Because without it you can only establish whether the system responds, not whether it is right. The set should be real questions from the real audience with answers agreed by whoever owns the content. A vendor-supplied evaluation set measures the vendor, and a set written after the build is written around what the build already does. Twenty to fifty questions is usually enough to be informative.
What acceptance criteria should go in the contract?
Eight, splitting into two groups. Four for the security reviewer: the known-restricted document returns nothing for every test tier, the revoked-access case passes after the stated lag, indirect prompt injection through indexed content is attempted and the result documented, and withheld-content logging is live with a named reader. Four for the business owner: correctness on the agreed evaluation set at a stated figure, citations on every answer the asker can open, a named owner per document set, and a re-review trigger.
How much of our content should be excluded?
More than feels comfortable, and a firm that expects to exclude nothing has not looked at your corpus. HR files, legal privilege, M&A material, anything under a separate retention obligation. A narrow index with a clear boundary gets approved. A wide index with a filter in front of it is the proposal that stalls, because a reviewer cannot bound the worst case and declining is the defensible position.
Does the model choice matter?
Far less than the corpus and the evaluation. In a 2026 benchmark over roughly 500,000 documents across nine real enterprise platforms, BM25 keyword ranking beat vector search by more than 17 points on correctness. If the retrieval approach is that contested, the model sitting on top of it is not where the project is won. Scope the corpus, the permissions and the evaluation set, and treat the model as replaceable.
What is the ongoing cost?
Re-indexing as content changes, permission synchronisation, periodic evaluation re-runs to catch regression, and the named owner's time. Ask for a range with the assumptions stated rather than a number. A proposal with no year-two figure is quoting half the project, and the second half arrives as an operating expense somebody else has to find.
What is the warning sign a build will not get approved?
No stated propagation lag on the permission model, and no test matrix from a previous engagement. Both mean the permission question has not been answered before, which means it will be answered during your project, on your timeline, in front of your reviewer. The second warning sign is a scope defined by systems rather than by question types, because that produces an index nobody can bound.
Related reading
If this is the problem you have
- Nobody can find what we already know, resolved by 2 services.
- Retrieval Access and Permissions Checklist, the worksheet for whoever has to approve the spend.
- The 30 day Corporate Drag and Risk Diagnostic, findings yours either way.