Thinklytics

AI Security · 10 min read · October 2026

What an AI security review needs from you, and what to contract for

By Sean Majidi, Founder, Thinklytics

Eight pieces of evidence, each with a named signer, assembled before the meeting rather than requested during it. Plus the acceptance criteria to contract for, including the only one the reviewer owns: zero findings at the first audit after go-live. Six engagements contracted for that and hit it.

The scope is written and the review is scheduled. What decides how that meeting goes is what arrives with you, because a reviewer who has to request evidence during a review will schedule a second one.

The evidence pack

The evidence pack, and who signs each piece

Assembled before the review meeting rather than requested during it. Every item is something a reviewer has asked us for.

  • The written scope: sources, actions, exclusions, identity. Signed by the system owner. This is the document the review is actually about.
  • Adversarial test results against that scope, with the failures included. A pack showing only passes reads as marketing. The failures and what you changed are what build confidence.
  • Data handling: retention, residency, whether anything leaves the tenancy, and sub-processors. Signed by whoever owns the vendor relationship. This is the section enterprise security questionnaires reuse.
  • The human review step and the named signer for outputs. Role and name. A reviewer will not accept a team.
  • Logging: what is recorded, for how long, and who can read it. Without this there is no post-incident account, which makes approval a one-way bet.
  • The re-review trigger: what change requires coming back. A new data source, a new write action, a model change. Agree it now, in writing.
  • A named owner for the system after launch, with an on-call path. Approval is for an operated system, not a deployed one.
  • A framework mapping, where the organisation uses one. NIST AI RMF, ISO 42001, the OWASP LLM Top 10. Useful as an index into the pack, not as a substitute for it.

Only 31% of CISOs feel fully aligned with their C-suite and board on acceptable AI risk. A pack that states the scope and its limits plainly is how that conversation gets had with a document instead of adjectives.

Source: Okta Global CISO Insights 2026, fielded by Apprize360 Intelligence, data finalised June 2026, n=306 security executives across six markets; Thinklytics case library.

Eight pieces, each with a named signer. Every one is something a reviewer has asked us for.

The written scope, naming sources, actions, exclusions and identity, signed by the system owner. This is the document the review is about. How to produce it is in scope the permissions or add guardrails.

Adversarial test results against that scope, with the failures included. Covered below, because this is the piece most often sanitised.

Data handling. Retention, residency, whether anything leaves your tenancy, and the sub-processors involved. Signed by whoever owns the vendor relationship. This section gets reused verbatim in every enterprise security questionnaire you answer afterwards, so it is worth writing properly once.

The human review step, with a named signer for outputs. Role and name. A reviewer will not accept a team.

Logging. What is recorded, for how long, and who can read it. Without it there is no post-incident account, which makes an approval a one-way bet and reviewers know it.

The re-review trigger. What change requires coming back.

A named owner and on-call path after launch. Approval is for an operated system, not a deployed one.

A framework mapping where your organisation uses one. NIST AI RMF, ISO 42001, the OWASP LLM Top 10. Present it as a table of pointers into the pack rather than as the pack.

Include the failures

The instinct is to bring a clean test report. It is the wrong instinct.

A pack showing only passes reads as marketing, and an experienced reviewer discounts it on that basis alone. Showing what the adversarial testing broke, and what you changed in response, is the strongest available evidence that the testing was real rather than ceremonial.

It also moves the meeting to the right question. Every system has residual risk. A reviewer's actual job is to decide whether to accept yours, and they cannot do that from a document that implies there is none. Handing them the residual explicitly, with the mitigations next to it, is the fastest route to a signature.

Our own engagements publish the counts rather than the adjectives for the same reason: 31 infrastructure gaps identified and 31 closed at a university system, 23 audit deficiencies resolved at a software company, six FDA deficiencies resolved at a food manufacturer. The number found is part of the evidence.

Acceptance criteria

Acceptance criteria to write into the engagement

Agreed before the work, tested on the last day, and most of them stated as a count a reviewer sets rather than as a quality you assess.

  • A signed scope document naming sources, actions, exclusions and identity. Signed, not drafted. The signature is the deliverable.
  • Adversarial test results against that scope, including every failure found and closed. State the number of findings and the number closed. 31 of 31 gaps closed at one university system, 23 deficiencies resolved at a software company.
  • Each agent authenticating as itself, with no shared credentials in scope. Checkable in the identity provider by someone who was not on the project.
  • Zero findings at the first review, exam or audit after go-live. The criterion the reviewer owns. Six engagements in the library contracted and hit a zero.
  • Logging live, with a stated retention period and a named reader. Tested by producing a real query result, not by showing a configuration screen.
  • A re-review trigger written down and agreed with the signer. Without it the approval quietly stops covering the system the first time scope changes.
  • A named owner and on-call path after handover, plus the residual run cost. A proposal with no ongoing cost is hiding one.
  • A security questionnaire response pack derived from the same evidence. One SaaS company shortened enterprise sales cycles by three weeks because the governance work answered the questionnaires. The security spend paid for itself on the revenue side.

The last item is the one to raise with a CFO. 46% of CISOs feel their board views AI security as a business enabler, and a measured sales-cycle reduction is the argument that moves that number.

Source: Okta Global CISO Insights 2026, n=306, data finalised June 2026; Thinklytics case library, published delivery outcomes and acceptance measures.

Eight, agreed before the work starts and tested on the last day. Three are worth expanding.

Each agent authenticating as itself, with no shared credentials in scope. Checkable in the identity provider by someone who was not on the project. Okta's Global CISO Insights 2026, fielded by Apprize360 Intelligence with data finalised June 2026 across 306 security executives in six markets, found 21% of organisations governing agent access with shared credentials or broad-permission service accounts, so this criterion is not theoretical.

Zero findings at the first review, exam or audit after go-live. The section below.

A security questionnaire response pack derived from the same evidence. The one criterion with a revenue argument attached, and the one to raise with a CFO.

Why zero findings is the criterion to contract for

Six engagements that went into a review and came out with nothing to fix

Published outcomes. The common factor is that the evidence existed before the reviewer asked, rather than being assembled under time pressure.

ReviewGoing inComing out
SOC 2, enterprise SaaS across 4 product lines8 weeks of senior engineer time per audit, documentation auditors questioned5 days of prep, zero auditor findings, and enterprise sales cycles 3 weeks shorter
Regulatory exam, regional bank4 data quality issues flagged at the previous exam, 6 months to fixZero matters requiring attention at the next exam
Federal grant re-audit, city governmentReporting errors across 12 departments putting $4.1M at riskZero deficiencies in the re-audit, funding preserved
Compliance review, public universityFour years of inconsistent enrollment and research dataZero compliance findings in review, 42 metrics certified across 8 units
Statutory filing, insurer6 weeks and 480 person-hours per quarterZero restatements across the first four quarters after deployment
External ESG assurance, utility14 auditor discrepancy findingsZero findings, and quarterly reporting possible for the first time

Twelve to 22 weeks across the governance engagements in the library. The outcome worth contracting for is the count at the next review, because it is the only one a reviewer sets.

Source: Thinklytics case library, published delivery outcomes per engagement.

Because the reviewer sets it and you do not.

Every other measure in a governance proposal is one the supplier defines and then grades itself against. A maturity score, a control count, a coverage percentage: all of them are marked by the person being paid. The count of findings at the next audit is set by an auditor, an examiner or a regulator who has no interest in the project's success.

Six engagements in our case library contracted for a zero and hit it. Zero auditor findings on the first SOC 2 after deployment at an enterprise SaaS company. Zero matters requiring attention at a regional bank's next regulatory exam, having gone in with four data quality issues flagged at the previous one. Zero deficiencies in a city government's re-audit, preserving $4.1M of federal grants. Zero compliance findings in a public university's review. Zero restatements across the first four quarters at an insurer. And 14 external ESG auditor discrepancy findings reduced to none at a utility.

The common factor in all six is unglamorous: the evidence existed before the reviewer asked for it, rather than being assembled under time pressure from people's memories.

The re-review trigger

One line, in the contract, and it prevents the expensive failure.

Scope always changes after an approval. A new data source gets connected. A read action becomes a write action. The model is swapped for a newer one. A new integration is added because it was easy.

Without an agreed trigger, the original signature silently stops covering the system that is actually running, and nobody notices until an incident or an audit. The second review is then harder than the first, because trust has been spent and the reviewer now has evidence that scope drifts unannounced here.

Agree the trigger while everyone is still in the room and the approval is fresh. A new data source, a new write action, a model change, a new integration: four items, one line each.

Timeline, and the revenue argument

The governance engagements in our case library ran 12 to 22 weeks. Twelve for an ESG data foundation that took external auditor findings from 14 to zero. Fourteen for a SOC 2 framework across four product lines. Sixteen for a research data certification that closed 31 gaps across 14 departments against a 16-week funding deadline, delivered in 14. Eighteen for a bank's call report and exam readiness. Twenty-two for a federal agency across six regional offices, which also cut FOIA response from 34 days to 8 and eliminated $1.2M a year in penalties.

One of those produced a measurable commercial return rather than a risk reduction. The enterprise SaaS company's framework shortened enterprise sales cycles by three weeks, because the same evidence pack answered the security questionnaires that large deals gate on, and it freed roughly seven engineer-weeks per audit. See the SOC 2 governance engagement.

That is the argument to take to a CFO, and it is worth making deliberately: only 46% of CISOs in the Okta survey feel their board views AI security as a business enabler, and a measured sales-cycle reduction is the kind of evidence that moves that number.

Five questions before signing

Show me a scope document and an adversarial test report from a previous engagement, with the failures in it. What count will you contract for at the first audit after go-live. Who on my side has to be in the room, and for how many hours a week. What is the residual annual cost after handover, and who owns the controls. And what happens if my CISO rejects the scope we write.

The last one matters most. That is a negotiation rather than a technical problem, and a firm that has been in that room answers with how they handle a rejected scope rather than with an assurance that it will not happen.

What we would do first

Assemble the pack for one use case before booking the review. Eight items, most of them a page.

Then read it as the reviewer would, looking for the question you cannot answer. There is usually one, it is usually the identity question or the logging retention, and finding it yourself a week early is worth more than any amount of preparation for the meeting.

Where the obligation is regulatory rather than internal, the EU AI Act in 2026 covers what applies and when, and the AI deployment approval pack is the worksheet we hand to whoever has to approve the spend alongside the security sign-off, because those are usually two different people with two different questions.

Delivery sits in AI security consulting for the scope and the adversarial testing, AI governance managed operations for operating the controls after handover, MLOps consulting for the logging and monitoring layer, and EU AI Act compliance on the regulatory side. The full set of work in this area sits under AI is too risky to approve.

Frequently asked questions

What evidence does an AI security review need?

Eight pieces. The written scope naming sources, actions, exclusions and identity, signed by the system owner. Adversarial test results against that scope, with the failures included. Data handling: retention, residency, whether anything leaves your tenancy, and sub-processors. The human review step with a named signer for outputs. Logging: what is recorded, for how long, who can read it. The re-review trigger. A named owner and on-call path after launch. And a framework mapping where your organisation uses one, as an index rather than a substitute.

Why include the test failures rather than only the passes?

Because a pack showing only passes reads as marketing and a reviewer discounts it accordingly. Showing what the adversarial testing broke, and what you changed in response, is the strongest evidence available that the testing was real. It also moves the conversation to the residual risk you are asking them to accept, which is the decision they are actually being asked to make.

What acceptance criteria should go in the engagement?

Eight, most of them counts rather than qualities. A signed scope document. Adversarial test results with findings found and closed stated as numbers. Each agent authenticating as itself with no shared credentials in scope. Zero findings at the first review, exam or audit after go-live. Logging live with a stated retention and a named reader. A written re-review trigger agreed with the signer. A named owner and on-call path plus the residual run cost. And a security questionnaire response pack derived from the same evidence.

Why is zero findings the criterion worth contracting for?

Because the reviewer sets it, not the vendor. Every other measure in a governance proposal is one the supplier defines and grades itself against. Six engagements in our case library contracted for and hit a zero: zero auditor findings on a first SOC 2 after deployment, zero matters requiring attention at a bank's next regulatory exam, zero deficiencies in a city government re-audit, zero compliance findings in a university review, zero restatements across four quarters at an insurer, and 14 external ESG auditor findings reduced to none.

What is the re-review trigger and why does it belong in the contract?

A written statement of what change requires coming back for approval: a new data source, a read action becoming a write action, a model change, a new integration. It belongs in the contract because scope always changes after an approval, and without the trigger the original signature silently stops covering the system that is actually running. The second refusal is harder than the first, because trust has already been spent.

How long do these take?

The governance engagements in our case library ran 12 to 22 weeks. Twelve weeks for an ESG data foundation that took external auditor findings from 14 to zero. Fourteen weeks for a SOC 2 framework across four product lines that cut preparation from eight weeks to five days. Sixteen weeks for a research data certification that closed 31 gaps across 14 departments against a funding deadline, delivered in 14. Twenty-two weeks for a federal agency across six regional offices.

Does any of this pay for itself?

In one case measurably, on the revenue side rather than the risk side. An enterprise SaaS company's governance framework shortened enterprise sales cycles by three weeks, because the same evidence pack answered the security questionnaires that large deals gate on, and it freed roughly seven engineer-weeks per audit. That is the argument to take to a CFO, and it matters because only 46% of CISOs in Okta's 2026 survey feel their board views AI security as a business enabler.

What should I ask a firm before signing?

Five questions. Show me a scope document and an adversarial test report from a previous engagement, with the failures in it. What count will you contract for at the first audit after go-live. Who on my side has to be in the room, and for how many hours. What is the residual annual cost after handover and who owns the controls. And what happens if my CISO rejects the scope we write, because that is a negotiation rather than a technical problem and the answer tells you whether they have been in that room before.

The work behind this

Eighteen engagements in the case library carry governance, privacy and security, running 12 to 22 weeks. Six published a zero-findings outcome at the next audit, exam or review, and one shortened enterprise sales cycles by three weeks because the same evidence answered the security questionnaires.

Governance, privacy and security, 18 engagements.

Topics covered

  • AI security review
  • AI deployment approval pack
  • security questionnaire
  • AI audit evidence
  • re-review trigger
  • AI governance deliverables
  • zero findings

Frequently asked questions

What evidence does an AI security review need?

Eight pieces. The written scope naming sources, actions, exclusions and identity, signed by the system owner. Adversarial test results against that scope, with the failures included. Data handling: retention, residency, whether anything leaves your tenancy, and sub-processors. The human review step with a named signer for outputs. Logging: what is recorded, for how long, who can read it. The re-review trigger. A named owner and on-call path after launch. And a framework mapping where your organisation uses one, as an index rather than a substitute.

Why include the test failures rather than only the passes?

Because a pack showing only passes reads as marketing and a reviewer discounts it accordingly. Showing what the adversarial testing broke, and what you changed in response, is the strongest evidence available that the testing was real. It also moves the conversation to the residual risk you are asking them to accept, which is the decision they are actually being asked to make.

What acceptance criteria should go in the engagement?

Eight, most of them counts rather than qualities. A signed scope document. Adversarial test results with findings found and closed stated as numbers. Each agent authenticating as itself with no shared credentials in scope. Zero findings at the first review, exam or audit after go-live. Logging live with a stated retention and a named reader. A written re-review trigger agreed with the signer. A named owner and on-call path plus the residual run cost. And a security questionnaire response pack derived from the same evidence.

Why is zero findings the criterion worth contracting for?

Because the reviewer sets it, not the vendor. Every other measure in a governance proposal is one the supplier defines and grades itself against. Six engagements in our case library contracted for and hit a zero: zero auditor findings on a first SOC 2 after deployment, zero matters requiring attention at a bank's next regulatory exam, zero deficiencies in a city government re-audit, zero compliance findings in a university review, zero restatements across four quarters at an insurer, and 14 external ESG auditor findings reduced to none.

What is the re-review trigger and why does it belong in the contract?

A written statement of what change requires coming back for approval: a new data source, a read action becoming a write action, a model change, a new integration. It belongs in the contract because scope always changes after an approval, and without the trigger the original signature silently stops covering the system that is actually running. The second refusal is harder than the first, because trust has already been spent.

How long do these take?

The governance engagements in our case library ran 12 to 22 weeks. Twelve weeks for an ESG data foundation that took external auditor findings from 14 to zero. Fourteen weeks for a SOC 2 framework across four product lines that cut preparation from eight weeks to five days. Sixteen weeks for a research data certification that closed 31 gaps across 14 departments against a funding deadline, delivered in 14. Twenty-two weeks for a federal agency across six regional offices.

Does any of this pay for itself?

In one case measurably, on the revenue side rather than the risk side. An enterprise SaaS company's governance framework shortened enterprise sales cycles by three weeks, because the same evidence pack answered the security questionnaires that large deals gate on, and it freed roughly seven engineer-weeks per audit. That is the argument to take to a CFO, and it matters because only 46% of CISOs in Okta's 2026 survey feel their board views AI security as a business enabler.

What should I ask a firm before signing?

Five questions. Show me a scope document and an adversarial test report from a previous engagement, with the failures in it. What count will you contract for at the first audit after go-live. Who on my side has to be in the room, and for how many hours. What is the residual annual cost after handover and who owns the controls. And what happens if my CISO rejects the scope we write, because that is a negotiation rather than a technical problem and the answer tells you whether they have been in that room before.

Related reading

If this is the problem you have