AI Security · 10 min read · October 2026
Scope the permissions or add guardrails? The order that gets an AI system approved
By Sean Majidi, Founder, Thinklytics
They answer different reviewer questions and most approvals need both, in a specific order. Scope first, because it decides what the guardrails have to cover and because a guardrail over unbounded access bounds nothing. OWASP moved excessive agency from sixth to third for the same reason.
Two proposals arrive after the security refusal. One scopes what the system is allowed to reach and do. The other adds guardrails: an injection classifier, a content filter, an output check.
The second one demos better and is easier to buy. The first one is what gets signed.
They answer different questions
Two ways to make an AI system reviewable
They answer different questions and most approvals need both. Choosing only the second is the common mistake, because it is the one a vendor can sell you.
| Scope the permissions | Add runtime guardrails | |
|---|---|---|
| What it establishes | What the system can reach and do, by design | What the system did, and whether anything was blocked |
| The reviewer question it answers | What is the worst case if this is fully compromised? | How would we know if it went wrong? |
| How it fails | Scope creeps after approval and nobody re-reviews | A filter is treated as a boundary, and filters are probabilistic |
| What it costs | Mostly argument. Deciding which data the agent does not get | Mostly build and ongoing tuning |
| Can it be signed alone | Usually yes, if the scope is narrow and written | Rarely. A guardrail with unbounded access bounds nothing |
| Order | First. It determines what the guardrails have to cover | Second, sized to the scope you settled on |
This is why excessive agency moved to third in the 2026 OWASP list while output handling fell to tenth. The reviewable question moved from what the model says to what the model is allowed to touch.
Source: OWASP Top 10 for LLM Applications, 2026 edition, 4 August 2026; Thinklytics engagement pattern across the governance, privacy and security engagements in the case library.
A reviewer is asking two things, and only one of them is about runtime behaviour.
What is the worst case if this system is fully compromised? Scoping answers that. It is a statement about design: these sources, these actions, not those.
How would we know if it went wrong? Guardrails answer that. Detection, blocking, logging, alerting.
The first has to be answerable for an approval to exist at all, because a reviewer cannot accept an unbounded worst case on the strength of a detection rate. The second is then sized to whatever the first settled.
Which is why the order matters more than the choice. A guardrail layered over unbounded access does not bound anything. It reports on an unbounded system more attentively.
Why the industry moved in the same direction
OWASP published the 2026 edition of its Top 10 for LLM Applications on 4 August 2026. Excessive agency went from sixth to third. Improper output handling went from fifth to tenth.
That is the same shift in one artefact: away from what the model says, toward what the model is permitted to touch. Prompt injection stayed first and sensitive information disclosure second, which is consistent rather than contradictory, because disclosure is what excessive agency makes possible.
One caution on quoting it. The ranking was weighted 75% community vote and 25% real-world incident data, and OWASP notes that prompt injection held first place despite a thin incident record, which it attributes to heavy investment keeping successful attacks out of public databases. A list position is a judgement about risk. It is not a frequency, and presenting it as one to a security reviewer will cost you credibility in the room where you most need it.
What scoping actually involves
The written scope, which is the actual deliverable
One page. A reviewer with this can assess; a reviewer without it has nothing to assess and the safe answer is no.
- Every data source the system may read, named. Named systems and named objects, not 'the data warehouse'. The list is the scope.
- Every action the system may take, named, split into read and write. Query a table, draft a message, send a message, update a record, call an external API. The read and write boundary is the one reviewers care about most.
- What it may NOT reach, stated explicitly. An exclusion list is more persuasive than an inclusion list, because it shows the question was asked.
- Identity: what the agent authenticates as, and whose permissions it inherits. 21% of organisations govern agent access with shared credentials or broad-permission service accounts. If that is the answer, it is the finding.
- What happens to the data it reads, including retention and whether it leaves the tenancy. The question every security questionnaire asks, and the one that stalls enterprise deals when unanswered.
- The human review step, with who signs and on what. Named role, not 'the team'.
- An adversarial test against the written scope, with results. Indirect prompt injection via retrieved content, attempted access outside the list, attempted action outside the list. Prompt injection is still first on the 2026 OWASP list.
- A model card or vendor security questionnaire. Useful, and not this. It describes the model, not what you connected it to, and the connection is what is being reviewed.
A university system closed 31 identified infrastructure gaps across 14 departments in 14 weeks by doing exactly this work: cataloguing, access controls, lineage and reproducibility documentation.
Source: Okta Global CISO Insights 2026, n=306, data finalised June 2026; Thinklytics case library, published engagement scopes and outcomes.
Mostly argument, not engineering. That is why it gets deferred.
Deciding which data the agent does not get. Splitting its actions into read and write, because that boundary is the one reviewers care about most. Writing an explicit exclusion list, which is more persuasive than an inclusion list because it shows the question was asked. And giving the agent an identity of its own.
That last item is where most of the work is. Okta's Global CISO Insights 2026, fielded by Apprize360 Intelligence with data finalised in June 2026 across 306 CISOs, cybersecurity heads and senior security executives in six markets, found 21% of organisations governing AI agent access with shared credentials or broad-permission service accounts. It is vendor-commissioned, and the finding is specific enough to be useful: in a fifth of organisations the agent's permissions are whatever a shared integration account happens to have.
An agent in that position cannot be scoped, because its reach is not its own. Fixing the identity is a prerequisite for the document rather than a line in it.
What the guardrails then cover
Whatever the scope leaves live, and now you can say what that is.
Indirect prompt injection through retrieved content, which is the dominant vector and the reason prompt injection remains first on the OWASP list. Content the system fetches, a document it summarises, a record it reads: the instruction arrives in the data rather than from the user.
Attempted access outside the named list, which should fail loudly and alert rather than fail silently. Attempted actions outside the named list, treated the same way. A reviewer will test exactly this.
Output review with a named signer where the output reaches a customer or a filing. And logging detailed enough to produce a post-incident account, which is what makes approval something other than a one-way bet.
The order
The order that converts a refusal into a signature
Each step produces a document the reviewer asked for. Running the guardrail build first produces a system nobody can scope.
- Write what it may read and do, and what it may not
- Fix the identity so each agent authenticates as itself
- Test adversarially against that written scope
- Add guardrails sized to the scope, not to the model
- Name the signer and the re-review trigger
The last step is the one that gets skipped and the one that prevents the second refusal. Scope changes after approval, and without a re-review trigger the original signature silently stops covering the system.
Source: Thinklytics engagement pattern across the governance, privacy and security engagements in the case library.
Write the scope. Fix the identity. Test adversarially against the written scope. Add guardrails sized to that scope. Name the signer and the re-review trigger.
The last step is the one that gets dropped and the one that prevents the second refusal. Scope changes after an approval: a new data source gets connected, a read action becomes a write action, the model is swapped. Without an agreed trigger for coming back, the original signature quietly stops covering the system that is actually running, and the next review starts from a worse place because trust has been spent.
Agree the trigger in one line while everyone is still in the room.
Where frameworks fit
A framework is an index into your evidence rather than a substitute for it.
NIST AI RMF, ISO 42001 and the OWASP LLM Top 10 each tell you which categories of control to have an answer for. None of them answers what your specific agent may read. Map the written scope and the test results to whichever one your organisation has adopted and present the mapping as a table of pointers into the pack.
A reviewer handed a framework mapping with nothing behind it will ask for the scope anyway, and the mapping will have cost you a week. Our view on the operating model around this is in the 2026 AI governance operating model, and where the obligation is regulatory rather than internal, the EU AI Act in 2026 covers what actually applies and when.
What we would do first
Pick the single use case closest to approval and write its scope on one page this week. Sources, actions split into read and write, exclusions, identity.
Then try to answer the identity question out loud. If the answer is the name of a shared integration account, stop there and fix that, because no amount of guardrail work will make the system scopeable while its permissions belong to something else.
What the reviewer needs in front of them, and the acceptance criteria to put in the engagement, are in what an AI security review needs from you. The AI deployment approval pack carries the business case, the security review and the responsibilities in the form approvers usually ask for.
Delivery sits in AI security consulting for the scope and the adversarial testing, AI governance managed operations for operating the controls afterwards, MLOps consulting for logging and monitoring, and EU AI Act compliance on the regulatory side. The full set of work in this area sits under AI is too risky to approve.
Frequently asked questions
Should we scope the permissions or add guardrails first?
Scope first, in almost every case. The two answer different reviewer questions: scoping answers what the worst case is if the system is fully compromised, guardrails answer how you would know if something went wrong. A reviewer needs the first to sign at all, and the second is sized to whatever the first settled. A guardrail layered over unbounded access does not bound anything, it just reports on an unbounded system more attentively.
What does scoping the permissions actually involve?
Mostly argument rather than engineering. Deciding which data the agent does not get, splitting its actions into read and write, writing an explicit exclusion list, and giving it an identity of its own so it stops inheriting a service account's reach. Okta's Global CISO Insights 2026, fielded by Apprize360 Intelligence with data finalised June 2026 across 306 security executives, found 21% of organisations governing AI agent access with shared credentials or broad-permission service accounts, which is the identity half of this work in one number.
Why did OWASP move excessive agency up the list?
Because that is where the reviewable risk sits. The 2026 edition of the OWASP Top 10 for LLM Applications, published 4 August 2026, promoted excessive agency from sixth to third while dropping improper output handling from fifth to tenth. Attention moved from what the model says to what the model is permitted to touch. The ranking was weighted 75% community vote and 25% real-world incident data, so read a position as a judgement about risk rather than as a count of incidents.
Are guardrails ever enough on their own?
Rarely, and the reason is structural rather than a matter of quality. A content filter or an injection classifier is probabilistic: it reduces the rate of a bad outcome without bounding the worst case. A reviewer asked to approve a system whose worst case is unbounded has nothing to approve, however good the filter is. The exception is a strictly read-only system over data the requester already has access to, where the worst case is already bounded by the user's own permissions.
What should the guardrails cover once the scope is settled?
What the scope leaves live. Indirect prompt injection through content the system retrieves, since prompt injection is still first on the OWASP list and retrieved content is the dominant vector. Attempted access outside the named list, which should fail and alert rather than fail silently. Attempted actions outside the named list, with the same treatment. Output review where a person signs. And logging detailed enough to produce a post-incident account.
What is the most common mistake?
Buying the guardrail first, because it is the part a vendor can sell you and the part that demos well. Scoping produces a document and an argument with colleagues about who loses access to what, which nobody is selling and nobody enjoys. So the purchasable half gets funded, the reviewer still cannot bound the worst case, and the approval does not move.
What happens after the approval?
Scope changes, and without a written re-review trigger the original signature silently stops covering the system. Agree in advance what requires coming back: a new data source, a new write action, a model change, a new integration. That one line prevents the second refusal, which is usually harder than the first because trust has been spent.
How does this interact with a framework like NIST AI RMF or ISO 42001?
A framework is an index into your evidence, not a substitute for it. It tells you which categories of control to have an answer for; it does not answer what your specific agent may read. Map the written scope and the test results to whichever framework your organisation has adopted, and present the mapping as a table of pointers. A reviewer handed a framework mapping with nothing behind it will ask for the scope anyway.
The work behind this
Eighteen engagements in the case library carry governance, privacy and security, running 12 to 22 weeks. Six published a zero-findings outcome at the next audit, exam or review, which is the criterion a reviewer sets rather than one a vendor claims.
Governance, privacy and security, 18 engagements.
Topics covered
- AI agent permissions
- AI guardrails
- least privilege AI
- prompt injection defence
- excessive agency
- AI security scope
- agent identity
Frequently asked questions
Should we scope the permissions or add guardrails first?
Scope first, in almost every case. The two answer different reviewer questions: scoping answers what the worst case is if the system is fully compromised, guardrails answer how you would know if something went wrong. A reviewer needs the first to sign at all, and the second is sized to whatever the first settled. A guardrail layered over unbounded access does not bound anything, it just reports on an unbounded system more attentively.
What does scoping the permissions actually involve?
Mostly argument rather than engineering. Deciding which data the agent does not get, splitting its actions into read and write, writing an explicit exclusion list, and giving it an identity of its own so it stops inheriting a service account's reach. Okta's Global CISO Insights 2026, fielded by Apprize360 Intelligence with data finalised June 2026 across 306 security executives, found 21% of organisations governing AI agent access with shared credentials or broad-permission service accounts, which is the identity half of this work in one number.
Why did OWASP move excessive agency up the list?
Because that is where the reviewable risk sits. The 2026 edition of the OWASP Top 10 for LLM Applications, published 4 August 2026, promoted excessive agency from sixth to third while dropping improper output handling from fifth to tenth. Attention moved from what the model says to what the model is permitted to touch. The ranking was weighted 75% community vote and 25% real-world incident data, so read a position as a judgement about risk rather than as a count of incidents.
Are guardrails ever enough on their own?
Rarely, and the reason is structural rather than a matter of quality. A content filter or an injection classifier is probabilistic: it reduces the rate of a bad outcome without bounding the worst case. A reviewer asked to approve a system whose worst case is unbounded has nothing to approve, however good the filter is. The exception is a strictly read-only system over data the requester already has access to, where the worst case is already bounded by the user's own permissions.
What should the guardrails cover once the scope is settled?
What the scope leaves live. Indirect prompt injection through content the system retrieves, since prompt injection is still first on the OWASP list and retrieved content is the dominant vector. Attempted access outside the named list, which should fail and alert rather than fail silently. Attempted actions outside the named list, with the same treatment. Output review where a person signs. And logging detailed enough to produce a post-incident account.
What is the most common mistake?
Buying the guardrail first, because it is the part a vendor can sell you and the part that demos well. Scoping produces a document and an argument with colleagues about who loses access to what, which nobody is selling and nobody enjoys. So the purchasable half gets funded, the reviewer still cannot bound the worst case, and the approval does not move.
What happens after the approval?
Scope changes, and without a written re-review trigger the original signature silently stops covering the system. Agree in advance what requires coming back: a new data source, a new write action, a model change, a new integration. That one line prevents the second refusal, which is usually harder than the first because trust has been spent.
How does this interact with a framework like NIST AI RMF or ISO 42001?
A framework is an index into your evidence, not a substitute for it. It tells you which categories of control to have an answer for; it does not answer what your specific agent may read. Map the written scope and the test results to whichever framework your organisation has adopted, and present the mapping as a table of pointers. A reviewer handed a framework mapping with nothing behind it will ask for the scope anyway.
Related reading
If this is the problem you have
- AI is too risky to approve, resolved by 4 services.
- AI Deployment Approval Pack, the worksheet for whoever has to approve the spend.
- The 30 day Corporate Drag and Risk Diagnostic, findings yours either way.