1. The business case
2. The security position
3. Ownership after launch
4. What each reviewer will ask
What unblocking a stalled pilot looked like
Questions to put to any firm, including us
Most AI work does not die in the build. It dies between a working pilot and production, in a review nobody scoped, with no business case, no security answer, and no agreement on who owns it once it is live. This is the pack that closes those three gaps. Fill it in and the approval meeting becomes a decision instead of a list of new questions.
The business case, security review, and ownership questions to settle before moving an AI pilot into production, so the approval meeting is a decision rather than a delay.
Moving an AI pilot into production needs three things agreed: a business case with a measurable baseline, a security position covering what the system may read and do, and a named owner accountable after launch. Pilots stall because these are discovered during the approval meeting rather than prepared before it.
A pilot proves the thing can work. The case has to show it is worth running at volume, which is a different claim.
What the process cost in time, headcount, or error rate before. If the pilot did not record this, record it now from the pre pilot period, because every later number depends on it.
Pilots run at low volume with short context, which is the cheapest possible case. Model inference, retrieval, and review time at the volume you actually expect. This number surprises people and it is better to surprise them now.
Most production AI keeps a person on the consequential decisions. That review time is an operating cost and belongs in the case, not omitted to make the number look better.
Redeployed, or reduced. Sponsors ask and a case that dodges it loses credibility. Answer it plainly.
Maintenance, re-evaluation, and the upstream changes that break things. Inference is predictable; maintenance is not. Budget for the second year, because that is where these programmes get cancelled.
The single most common reason a working pilot does not ship. Answer these before the review rather than in it.
If retrieval ignores the asking user's access, the assistant can surface documents that person could not open. Permission aware retrieval is the answer reviewers are looking for.
A scoped tool list with an approval gate on anything touching customers, money, or a record of truth. Written down, not implied.
Prompts, retrievals, tool calls, and who the request was for. If an incident happens, this is the only thing that reconstructs it.
A kill switch, who can operate it, and how quickly. Reviewers ask this and a vague answer reads as an unmanaged risk.
Prompt injection, data exfiltration through tool calls, unsafe actions. Findings with reproduction steps, and what was fixed.
Approval often stalls here, because nobody has agreed who is accountable when the system is wrong at two in the morning.
A person, not a team. Accountable for outputs and for deciding when the system is switched off. Without this, the approval has nowhere to land.
An evaluation set and a threshold, checked on a schedule. Quality drift is discovered by a customer otherwise.
Who is told, who can pause it, and what the customer remedy is if something reached them.
When the system is re-assessed against its business case. Six months is typical. Systems nobody re-evaluates keep running long after they stopped being worth it.
Cost per unit at real volume, the baseline, and year two. See section 1.
What data the system processes, where it goes, what is retained, and which obligations apply. Regulated sectors need this as evidence rather than assurance.
What changes in their day, what the exception queue looks like, and who clears it. They are rarely consulted and frequently the reason adoption fails.
Three machine learning pilots at a pharmacy benefit manager had been deadlocked for over a year. The blocker was underneath the models rather than in them.
Three separate member ID systems, so records could not be matched
These surface whether a firm has taken anything to production before.
What is the cost per unit at our expected volume, and what assumptions is that built on?
What adversarial testing will you run, and will we get the findings with reproduction steps?
Who owns this after you leave, and what does handover include?
What would make you recommend we do not put this into production?
Three reasons, usually together. No business case at real volume, because the pilot ran cheap and nobody modelled what volume costs. No security answer, because nobody wrote down what the system may read and do. No named owner, so there is nobody for the approval to land on. All three are preparable in advance and rarely are.
What the system can read and whose permissions apply, what actions it can take and which need human approval, what is logged, how it is stopped, and what happened when someone tried to break it. A policy document without controls behind it does not answer any of these, which is why reviews stall.
Take expected requests per period, average context length including retrieved content, and the model's pricing, then add the human review time that stays in the loop. Pilot figures mislead because pilots are low volume with short context. Caching and routing simpler requests to cheaper models change the number materially and are easier to design in before launch.
A named person on the business side who is accountable for outputs and can decide to switch it off, supported by whoever monitors quality. Not the vendor and not a committee. Approval frequently stalls precisely because this has not been agreed, and it is the cheapest gap to close.
Better than finding out in production. Usually the gap is the data underneath rather than the model, and that is a different and more valuable project. The Express Scripts pilots were stalled on entity resolution, not modelling.