AI Deployment · 10 min read · October 2026
What an AI pilot-to-production engagement should include
By Sean Majidi, Founder, Thinklytics
Seven deliverables, four of which concern what happens after go-live. That ratio is the tell for whether a firm has run production AI or only pilots. Plus eight acceptance criteria, the year-two run cost most proposals omit, and the handover question to ask: what does day 91 look like.
The pilot works, the blocker has been named, and a proposal to take it to production is in front of you. The useful question is how much of the scope applies after the go-live week, because that is where these engagements are won or lost.
The seven deliverables
What a pilot-to-production engagement has to hand over
Seven items. A proposal that covers the model and not these has scoped a second pilot.
- A pipeline that survives schema change, with a monitored contract per source. One churn model never shipped because the pipeline broke several times a week when CRM and product usage formats changed. Self-healing schema detection was the fix.
- Validation on a held-back period, not the training window. A six-month holdout at one client, a three-month holdout at another. The holdout is what makes the accuracy figure mean anything.
- Resolved identity for every entity the system reasons about. Member match accuracy from 75 to 94 of every 100 records unblocked three stalled pilots at one client. If records do not join, nothing downstream is reliable.
- A scheduled scoring or inference run, and alerts into a tool the owner already uses. Daily scoring with Slack alerts to named customer success managers. Not a dashboard.
- A written scope of what the system may read and do, with controls implemented and tested against it. This is what turns a security review from a refusal into a signature.
- Drift and quality monitoring with a threshold and a named owner. Accuracy degrades. Someone has to be told, at a number agreed in advance.
- Handover: documentation and the internal team operating it. One engagement delivered rebuilt feature pipelines to the internal ML team with full documentation, and that team now owns it. Ask what ownership looks like on day 91.
- A larger model. Almost never the blocker. In our assessments the failures were identity, instrumentation and pipeline stability, in that order.
Four of the seven are about what happens after go-live. That ratio is the tell for whether a firm has run production AI or only pilots.
Source: Thinklytics case library, published engagement scopes and outcomes.
A pipeline that survives schema change, with a monitored contract per source. One churn model never shipped because the pipeline broke several times a week when CRM and product usage data changed format. Rebuilding it to detect and repair schema changes automatically was the entire unlock.
Validation on a held-back period, not the training window. A six-month holdout at one client, a three-month holdout at another. The holdout is what makes the accuracy figure mean anything, and it is the first thing to check in a proposal that quotes one.
Resolved identity for every entity the system reasons about. Member match accuracy moving from 75 to 94 of every 100 records unblocked three stalled pilots at one client. If records do not join, nothing downstream is reliable, including the evaluation.
A scheduled scoring or inference run, with alerts into a tool the owner already uses. Daily scoring with Slack alerts to named customer success managers, not a dashboard somebody has to remember to open.
A written scope of what the system may read and do, with controls implemented and tested against it. This is what turns a security review from a refusal into a signature, because a reviewer with no scope has nothing to assess and the safe answer is no.
Drift and quality monitoring, with a threshold and a named owner. Accuracy degrades. Someone has to be told, at a number agreed in advance rather than after the first complaint.
Handover: documentation and the internal team operating it. One engagement delivered rebuilt feature engineering pipelines to the client's internal ML team with complete documentation, and that team now fully owns them.
Four of those seven are about life after go-live. That ratio is the tell. Count how many items in the proposal in front of you apply after the first week, and if the answer is one or none, you are buying a second pilot.
What is usually not the deliverable
A larger or better model.
Across our assessments the failures were identity resolution, instrumentation and pipeline stability, in that order. A mid-market SaaS platform had spent $2.4M on three ML initiatives that would not run for three different data reasons: customer IDs that did not match, incomplete product interaction data, and unstable pipelines. The data science team had been rebuilding the models repeatedly, which is the natural response to a problem you have misdiagnosed.
If a proposal's centre of gravity is model selection or model size, ask what evidence says the model is the constraint. The honest answer is usually that nobody has checked, and the test that checks it is in fix data, integration or governance first.
Acceptance criteria
Acceptance criteria, written before the work
Tested on the last day by someone who was not in the project. Half of them concern the period after handover.
- A stated accuracy or quality figure, measured on a held-back period. 94 of every 100 records on a validation set. 84 of 100 on a three-month holdout. State the figure and the holdout in the statement of work.
- The system runs on its schedule for a stated number of consecutive weeks without intervention. One pipeline ran 18 weeks continuously after go-live. That is the criterion that distinguishes production from a demo.
- Alerts land in a named tool, for a named owner, within a stated latency. Daily scoring, alerts to the customer success managers who own the accounts.
- A security sign-off against the written scope, with the controls tested. Not a review meeting. A signature against a document that says what the system may read and do.
- A measured business result attributable to the deployment, in a stated window. $4.8M a year of misrouted claims addressed. $2.6M of at-risk ARR flagged and 340 accounts retained in a quarter. 31 of 31 gaps closed and $6.8M released.
- Drift monitoring live, with a threshold and a named owner. Agreed in advance, because after go-live nobody wants to be the person who sets it.
- The internal team operating it, with documentation, before the engagement ends. Ask what day 91 looks like and who is on the rota.
- A stated run cost for year two. Inference, monitoring, retraining and the owner's time. A proposal with no year-two number is quoting half the project.
The fifth criterion is the one that makes this fundable. Four of our assessments ran against money already spent on a model that had no measured result, which is exactly the position a business case has to avoid repeating.
Source: Thinklytics case library, published delivery outcomes and acceptance measures per engagement.
Eight, written before the work starts and tested on the last day by someone who was not in the project. Three are worth expanding.
The system runs on its schedule for a stated number of consecutive weeks without intervention. One pipeline ran 18 weeks continuously after go-live. That criterion, more than any accuracy figure, is what separates production from a demonstration, and it is the one most often absent from a statement of work.
A measured business result attributable to the deployment, in a stated window. $4.8M a year of misrouted claims addressed once member match accuracy reached 94 of every 100. $2.6M of at-risk ARR flagged and 340 accounts retained in the first quarter. All 31 certification gaps closed and $6.8M of federal research funding released. Agree the measure and the window before the work, because afterwards it becomes a negotiation.
A security sign-off against the written scope, with the controls tested. Not a review meeting and not a risk register entry. A signature against a document that states what the system may read and what it may do. Deloitte's 2Q 2026 CFO Signals, fielded 22 May to 7 June 2026 across 200 North American CFOs above $1B revenue, found only 19% of CFOs saying they hold the greatest responsibility for AI governance against 33% pointing at the CISO, so name the signer in the contract rather than assuming one exists.
Timeline, and what was already spent
What these took, and what was already spent before they started
Published durations. The pattern worth noticing is the second column.
| Engagement | Delivery | Already spent before the engagement |
|---|---|---|
| Mid-market SaaS, 3 stalled ML initiatives diagnosed | 6 weeks assessment, then 12 weeks to production for each | $2.4M across the three, 6 months, nothing live |
| Analytics vendor, churn model into production | 10 weeks | A model at 78 of 100 that had never shipped |
| Regional P&C insurer, claims data for AI underwriting | 13 weeks | 18 months of internal attempts, failed twice |
| Pharmacy benefit manager, 3 pilots restarted | 14 weeks | Over a year of the pilots being stuck |
| Workers comp carrier, policy data to platform threshold | 16 weeks at $340K | $1.2M on the platform, plus a $2.8M vendor remediation quote |
| University system, research data certification | 16 weeks, closed in 14 | $6.8M of funding conditional on it |
| State agency, 8 programme areas scored | 20 weeks | A $12M decision with no readiness view |
In six of the seven, significant money went into the model or the platform before anyone measured the foundation underneath it. The engagements above are what that sequence costs to correct.
Source: Thinklytics case library, published delivery durations and prior-spend figures per engagement.
Six to 20 weeks in our case library, depending on how much of the foundation had to be built rather than fixed.
Ten weeks to take a churn model that had never shipped into daily production with alerts. Thirteen weeks for a claims data layer across four systems, after 18 months of internal attempts that had failed twice. Fourteen weeks for an entity resolution layer that restarted three pilots stuck for over a year. Sixteen weeks to lift policy data from 61 to 94 against a vendor-mandated threshold of 80. Sixteen weeks, closed in 14, for a research data certification with $6.8M conditional on it.
The second column of that table is the part worth reading. In six of seven, significant money had gone into the model or the platform before anyone measured the foundation: $2.4M across three initiatives, $1.2M on an underwriting platform, a $12M modernisation decision with no readiness view. The engagements are what correcting that sequence costs.
Worth sizing the alternative while reading a proposal. Across the 47 engagements in the 2026 Enterprise Data Readiness Report, the stalled projects had burned roughly $4.2M each across model build, infrastructure, licences and team time before being abandoned, over 14 to 22 months. Correcting the data layer for a mid-sized company ran 6 to 12 months and $800K to $2.4M. Those are our own engagement figures rather than an industry benchmark, and the comparison is the useful part: the remediation is the cheaper half of that pair, and it is the half that produces something.
The year-two run cost
Ask for it, and expect resistance.
It covers inference or compute, monitoring and alerting, periodic retraining, the data pipeline's ongoing maintenance, and the named owner's time. It gets omitted because it makes the first-year figure look worse and because it really is uncertain at proposal stage.
A range with stated assumptions is an acceptable answer. Silence is not. A proposal with no year-two figure is quoting half the project, and the second half arrives later as an operating expense that somebody else has to find, usually in the quarter when the system's value is still being established. The detail is in what an AI system costs in year two.
The handover question
Ask what day 91 looks like.
Who is on the rota. What they are alerted about. What they do when the drift threshold trips. Where the runbook lives. Who retrains it, on what trigger, and who signs off that the retrained version is better than the one it replaced.
A firm that has run production AI answers that specifically and without hesitation, because it has had the conversation before. A firm that has only run pilots answers with a training session and a documentation deliverable, which is not the same thing.
The engagement that handed rebuilt feature pipelines to an internal ML team with full documentation is the shape to aim for: the client's team operating it before the engagement ended, not after.
Five questions before signing
How many of your deliverables apply after the go-live week. What holdout period will the accuracy figure be measured on. What business result will we measure, and in what window. What is the year-two run cost, with assumptions. And what does day 91 look like.
A firm with production experience answers the first with a number above half and the last one specifically. Those two answers predict the rest.
What we would do first
Before re-reading the proposal, write down the measured result you will claim at the end and the window you will claim it in. One sentence.
If you cannot write it, the engagement is not ready to scope, and the thing to fix is the business case rather than the model. That is the fourth blocker in why AI pilots stall before production, and it is the one no engineering addresses.
Then check the proposal against the seven deliverables and the eight criteria. The AI deployment approval pack carries the business case, the security review and the responsibilities in the form the approver usually asks for, and the data layer requirements are in what production AI automation requires from data.
Delivery sits in AI readiness for the diagnosis, managed data readiness where remediation runs as a service, MLOps consulting for the production and monitoring layer, AI governance managed operations for the controls, and team enablement where the internal team has to own it afterwards. The full set of work in this area sits under our AI work is not delivering.
Frequently asked questions
What should a pilot-to-production engagement deliver?
Seven things. A pipeline that survives schema change with a monitored contract per source. Validation on a held-back period rather than the training window. Resolved identity for every entity the system reasons about. A scheduled scoring or inference run with alerts into a tool the owner already uses. A written scope of what the system may read and do, with controls implemented and tested against it. Drift and quality monitoring with a threshold and a named owner. And a handover with documentation to the internal team that will operate it.
Why do four of the seven concern life after go-live?
Because that is where pilots die, and the ratio is the clearest signal of whether a firm has run production AI or only pilots. A model that works on the day it is demonstrated and has no monitoring, no owner and no alerting path is a pilot with a launch date. Ask how many of the seven items in a proposal apply after the go-live week. If the answer is one or none, the engagement will end with the system in the same state the last one did.
What acceptance criteria should go in the contract?
Eight. A stated accuracy or quality figure measured on a held-back period. The system running on its schedule for a stated number of consecutive weeks without intervention. Alerts landing in a named tool for a named owner within a stated latency. A security sign-off against the written scope with controls tested. A measured business result attributable to the deployment in a stated window. Drift monitoring live with a threshold and an owner. The internal team operating it with documentation. And a stated year-two run cost.
What does the year-two run cost include and why is it omitted?
Inference or compute, monitoring and alerting, periodic retraining, the data pipeline's maintenance, and the named owner's time. It is omitted because it makes the first-year figure look worse and because it really is uncertain at proposal stage. Ask for a range with the assumptions stated rather than a number. A proposal with no year-two figure is quoting half the project, and the second half arrives as an operating expense somebody else has to find.
How long do these take?
In our case library, 6 to 20 weeks depending on how much of the foundation had to be built. Ten weeks to take a churn model that had never shipped into daily production with alerts. Thirteen weeks for a claims data layer after 18 months of failed internal attempts. Fourteen weeks for an entity resolution layer that restarted three stalled pilots. Sixteen weeks for a policy data remediation to a vendor-mandated quality threshold, and 16 for a research data certification against a funding deadline.
What does a good handover look like?
The internal team operating the system before the engagement ends, with documentation, not after it. One engagement delivered rebuilt feature engineering pipelines to the client's internal ML team with complete documentation, and that team now fully owns them. The question to ask a firm is what day 91 looks like: who is on the rota, what they are alerted about, and what they do when the drift threshold trips. A firm that has done this answers specifically.
Should the engagement include a larger or better model?
Almost never, because the model is almost never the blocker. Across our assessments the failures were identity resolution, instrumentation and pipeline stability, in that order. One client had spent $2.4M on three ML initiatives that would not run for three different data reasons. If a proposal's centre of gravity is model selection, check what evidence says the model is the constraint, because the usual answer is none.
What is the measured result criterion and why does it matter most?
A business figure attributable to the deployment inside a stated window, agreed before the work starts. $4.8M a year of misrouted claims addressed once member match accuracy went from 75 to 94 of every 100. $2.6M of at-risk ARR flagged and 340 accounts retained in one quarter. 31 of 31 gaps closed and $6.8M of funding released. It matters most because the alternative is the position most of these clients started from: money spent on a model with no measured result and no way to defend the next request.
The work behind this
Ten AI readiness assessment engagements and nine AI agent and workflow automation engagements in the case library carry this work. Each states the duration, what was already spent before it started, and the measured result at the end.
AI readiness assessment, 10 engagements.
Topics covered
- AI pilot to production
- AI deployment engagement scope
- MLOps handover
- AI acceptance criteria
- drift monitoring
- AI year two cost
- production AI deliverables
Frequently asked questions
What should a pilot-to-production engagement deliver?
Seven things. A pipeline that survives schema change with a monitored contract per source. Validation on a held-back period rather than the training window. Resolved identity for every entity the system reasons about. A scheduled scoring or inference run with alerts into a tool the owner already uses. A written scope of what the system may read and do, with controls implemented and tested against it. Drift and quality monitoring with a threshold and a named owner. And a handover with documentation to the internal team that will operate it.
Why do four of the seven concern life after go-live?
Because that is where pilots die, and the ratio is the clearest signal of whether a firm has run production AI or only pilots. A model that works on the day it is demonstrated and has no monitoring, no owner and no alerting path is a pilot with a launch date. Ask how many of the seven items in a proposal apply after the go-live week. If the answer is one or none, the engagement will end with the system in the same state the last one did.
What acceptance criteria should go in the contract?
Eight. A stated accuracy or quality figure measured on a held-back period. The system running on its schedule for a stated number of consecutive weeks without intervention. Alerts landing in a named tool for a named owner within a stated latency. A security sign-off against the written scope with controls tested. A measured business result attributable to the deployment in a stated window. Drift monitoring live with a threshold and an owner. The internal team operating it with documentation. And a stated year-two run cost.
What does the year-two run cost include and why is it omitted?
Inference or compute, monitoring and alerting, periodic retraining, the data pipeline's maintenance, and the named owner's time. It is omitted because it makes the first-year figure look worse and because it really is uncertain at proposal stage. Ask for a range with the assumptions stated rather than a number. A proposal with no year-two figure is quoting half the project, and the second half arrives as an operating expense somebody else has to find.
How long do these take?
In our case library, 6 to 20 weeks depending on how much of the foundation had to be built. Ten weeks to take a churn model that had never shipped into daily production with alerts. Thirteen weeks for a claims data layer after 18 months of failed internal attempts. Fourteen weeks for an entity resolution layer that restarted three stalled pilots. Sixteen weeks for a policy data remediation to a vendor-mandated quality threshold, and 16 for a research data certification against a funding deadline.
What does a good handover look like?
The internal team operating the system before the engagement ends, with documentation, not after it. One engagement delivered rebuilt feature engineering pipelines to the client's internal ML team with complete documentation, and that team now fully owns them. The question to ask a firm is what day 91 looks like: who is on the rota, what they are alerted about, and what they do when the drift threshold trips. A firm that has done this answers specifically.
Should the engagement include a larger or better model?
Almost never, because the model is almost never the blocker. Across our assessments the failures were identity resolution, instrumentation and pipeline stability, in that order. One client had spent $2.4M on three ML initiatives that would not run for three different data reasons. If a proposal's centre of gravity is model selection, check what evidence says the model is the constraint, because the usual answer is none.
What is the measured result criterion and why does it matter most?
A business figure attributable to the deployment inside a stated window, agreed before the work starts. $4.8M a year of misrouted claims addressed once member match accuracy went from 75 to 94 of every 100. $2.6M of at-risk ARR flagged and 340 accounts retained in one quarter. 31 of 31 gaps closed and $6.8M of funding released. It matters most because the alternative is the position most of these clients started from: money spent on a model with no measured result and no way to defend the next request.
Related reading
If this is the problem you have
- Our AI work is not delivering, resolved by 5 services.
- AI Deployment Approval Pack, the worksheet for whoever has to approve the spend.
- The 30 day Corporate Drag and Risk Diagnostic, findings yours either way.