Thinklytics

AI Deployment · 10 min read · October 2026

Why AI pilots stall before production, and what actually blocks them

By Sean Majidi, Founder, Thinklytics

88% of agent pilots never reach production, and the reported blockers are evaluation, governance and reliability rather than model capability. One client had spent $2.4M across three ML initiatives in six months with none live, for three different data reasons. Also the one figure to keep out of your business case.

The pilot worked. The demo went well. Nine months later it is still a pilot, and the quarterly review has started asking what happened.

Something specific happened, and it is almost never the model.

What the published figures say

Where agent pilots are getting stuck

Three out of four enterprises now have an agent project underway. The blockers reported are not about model capability.

  • AI agent pilots that never reach production
  • Blocked by evaluation gaps
  • Blocked by governance friction
  • Blocked by model reliability

Source: Forrester with Anaconda, The State of Agentic AI in 2026, published 9 June 2026. No sample size is published for these figures, so read them as the direction of a named study rather than a measured population.

Forrester, with Anaconda, published The State of Agentic AI in 2026 on 9 June 2026. It reports 88% of AI agent pilots never reaching production, against three out of four enterprises having an agent project underway.

The blockers it names are worth reading in order: evaluation gaps in 64% of cases, governance friction in 57%, model reliability in 51%. The most cited blocker is not the technology, it is the absence of clear criteria for what the agent working actually means.

Forrester does not publish a sample size alongside those figures, so treat them as the direction of a named study rather than a measured population. The ranking is the useful part, and it puts two organisational problems above the one technical problem.

The figure to leave out of the business case

The figure to leave out of the business case

  • The number everyone quotes. 95%. Of generative AI pilots producing no measurable financial return. From MIT Project NANDA's The GenAI Divide, State of AI in Business, published July 2025.
  • Why it will not survive review. Unsupported. A Wharton professor publicly stated he could not locate the basis for the 95% figure. The report describes interviews, surveys and analysis of 300 public implementations without publishing the derivation.

Use the Forrester and Anaconda production-rate figures instead, and your own measured pilot history. A business case built on a contested headline fails at the first finance review that checks it, and the argument it was supporting usually fails with it.

Source: MIT Project NANDA, The GenAI Divide: State of AI in Business, July 2025; public methodological criticism of the 95% figure.

Before going further, one statistic to retire.

The claim that 95% of generative AI pilots produce no measurable financial return comes from MIT Project NANDA's The GenAI Divide: State of AI in Business, published July 2025. It is the most quoted number in this category.

It will not survive a review. A Wharton professor has publicly stated he could not locate the basis for the 95% figure, and the report describes interviews, surveys and analysis of 300 public implementations without publishing the derivation.

Use the Forrester and Anaconda production-rate figures instead, and your own pilot history, which is more persuasive than either because it is about your company. A business case resting on a contested headline fails at the first finance review that checks it, and the argument it was supporting usually fails with it.

What actually blocks them

Four things. The first three are technical and organisational in turn, and the fourth is neither.

The data will not support it. Records do not join. The signal was never instrumented. A metric is defined two ways so the label is unreliable.

The integration breaks. It runs in a notebook and fails in production when a schema changes or a source moves.

Governance will not approve it. It works, and security, legal or compliance will not sign.

Nobody can say what it returned or who owns it after launch. Not a technical blocker, and no amount of engineering addresses it.

Three stalled models, three different data causes

Three stalled models, three different data causes

One mid-market SaaS platform. $2.4M spent across three initiatives, six months elapsed, none live, the data science team rebuilding models repeatedly.

InitiativeWhy it would not runCategory
Churn predictionCustomer IDs did not match consistently across sourcesIdentity resolution
Recommendation engineProduct interaction data was incomplete, so the signal was not there to learn fromInstrumentation
Anomaly detectionData pipelines were unstable, so the model could not be relied on in productionPipeline reliability

A six-week assessment across data completeness, metric consistency, pipeline reliability, feature engineering reproducibility and inference infrastructure readiness found 14 specific data layer failures. After remediation all three were live within 12 weeks.

Source: Thinklytics case library, mid-market SaaS platform AI readiness assessment, published outcome.

A mid-market SaaS platform had spent $2.4M developing three machine learning projects: churn prediction, recommendations and anomaly detection. Six months in, none were live, and the data science team was repeatedly rebuilding the models. The company believed it had a modelling problem.

A six-week assessment across data completeness, metric consistency, pipeline reliability, feature engineering reproducibility and inference infrastructure readiness found 14 specific data layer failures, and the three projects were blocked for three different reasons.

The churn model could not run because customer IDs did not match consistently across sources. The recommendation engine lacked complete product interaction data, so there was no signal to learn from. The anomaly detection system failed because the pipelines were unstable.

Identity, instrumentation, pipeline stability. Three different problems, all in the data layer, none in the models that kept being rebuilt. After remediation all three were live within 12 weeks. See the AI readiness assessment engagement.

Our own 47-engagement audit puts a distribution on this. Across 47 enterprise engagements Thinklytics audited between 2022 and 2025, about 75% never reached production, and the data-layer causes ranked: undefined metric semantics in 29 of the 47, identity resolution gaps in 25, lineage gaps at inference time in 22, feature store governance in 18, and infrastructure coupling in 15. Those counts sum to more than 47 because most stalled projects carried two or three of them at once, which is the part worth knowing before scoping a fix. The full method and the findings are in the 2026 Enterprise Data Readiness Report.

The pattern repeats. A pharmacy benefit manager had three AI projects stuck for more than a year because internal teams could not prioritise fixing the data; member match accuracy was 75 of every 100 records across three member ID systems. Lifting it to 94 on a validation set restarted all three pilots within four weeks of delivery, against $4.8M a year of misrouted claims. See the member matching engagement.

A regional insurer's data science team spent 18 months trying to build a usable training dataset across four claims systems and failed twice. The claims data layer took 13 weeks once it was treated as the actual project.

Why governance says no to something that works

Usually because there is nothing concrete to assess.

If nobody has written down what the system is allowed to read and what it is allowed to do, a reviewer has no scope to test against, and the defensible answer is no. The review is not obstructive, it is under-specified.

The ownership picture does not help. Deloitte's 2Q 2026 CFO Signals, fielded between 22 May and 7 June 2026 across 200 North American CFOs at organisations above $1B revenue, found 96% at least somewhat confident in their AI governance framework, but only 43% confident, with 53.5% only somewhat. Only 19% of CFOs said they hold the greatest responsibility for AI governance, and 33% pointed at the CISO. Ambiguity at that level becomes nobody reviewing.

On the quality side, EY's 2026 Global DNA of the CFO survey, fielded 16 February to 30 March 2026 across 1,610 finance leaders at organisations above $1B revenue in 28 countries, found 61% citing data quality and bias as the top barrier to AI investment. Deloitte's Finance Trends 2026 survey of 1,326 global finance leaders found 47% citing data issues as a barrier to AI adoption in finance.

Those are aggregate answers, which is their limitation. "Data issues" covers all three technical blockers above and the triage in fix data, integration or governance first is what separates them.

The sequence that costs the most

Spending on the model or the platform before measuring the foundation.

$2.4M across three initiatives at the SaaS platform. $1.2M on an AI underwriting platform at a workers compensation carrier, which then required a data quality score of 80 across 10 categories against policy data scoring 61. A $12M AI modernisation decision at a state agency with no view of which of its eight programme areas could actually use AI.

In six of the seven assessment and remediation engagements in our case library, significant money had already gone into the model or the platform before anyone scored the data underneath it. That sequence, not the assessment, is the expensive part.

What we would do first

Run the pilot once against production-shaped data instead of a curated extract, and watch which way it fails. That costs about a week and it names the blocker.

Cannot be trained or scored, because records will not join or the field is missing: the blocker is data. Runs, then breaks on the next schema change: integration. Works, and nobody will approve it: governance. Works, is approved, and nobody can say what it returned: the business case and the owner.

Then fix in that order, because fixing governance on a model that cannot join its own records wastes the quarter. The triage is in fix data, integration or governance first, the data layer requirements in what production AI automation requires from data, and the proof of concept that produces a decision rather than a demo in the AI proof of concept.

Delivery sits in AI readiness for the assessment, managed data readiness where the remediation needs to run as a service, and team enablement where the internal team will own it afterwards. The full set of work in this area sits under our AI work is not delivering.

Frequently asked questions

Why do AI pilots stall before production?

Four reasons, and the model is rarely one of them. The data will not support it, usually because records do not join, the signal was never instrumented, or a metric is defined two ways. The integration breaks, so it runs in a notebook and fails weekly in production on schema change. Governance will not approve it, because nobody wrote down what it may read and do. Or nobody can say what it returned and who owns it after launch, which is not a technical blocker at all and no engineering fixes it.

What do the published figures say?

Forrester, with Anaconda, in The State of Agentic AI in 2026 published 9 June 2026, reports 88% of AI agent pilots never reaching production, with three out of four enterprises having an agent project underway. The blockers it reports are evaluation gaps in 64% of cases, governance friction in 57% and model reliability in 51%. Forrester does not publish a sample size for those figures, so read them as the direction of a named study rather than a measured population.

What does this look like in practice?

A mid-market SaaS platform had spent $2.4M across three ML initiatives, churn prediction, recommendations and anomaly detection. Six months in, none were live and the data science team was repeatedly rebuilding the models. A six-week assessment found 14 data layer failures and three different causes: customer IDs did not match consistently, product interaction data was incomplete, and the pipelines were unstable. After remediation all three were live within 12 weeks.

Which figure should I keep out of the business case?

The 95% one. MIT Project NANDA's The GenAI Divide, published July 2025, is the source of the widely quoted claim that 95% of generative AI pilots produce no measurable financial return, and a Wharton professor has publicly stated he could not locate the basis for it. The report describes interviews, surveys and analysis of 300 public implementations without publishing the derivation. A case built on a contested headline fails at the first finance review that checks it.

Is the data the blocker more often than the model?

In our own assessments, consistently. The 14 failures at the SaaS platform were identity, instrumentation and pipeline stability. A pharmacy benefit manager had three AI projects stuck for over a year with member match accuracy at 75 of every 100 records; lifting it to 94 restarted all three within four weeks of delivery. A regional insurer's data science team spent 18 months trying to build a usable training dataset across four claims systems and failed twice.

Why does security block a pilot that works?

Because there is usually nothing concrete to assess. If nobody has written down what the system is allowed to read and what it is allowed to do, a reviewer has no scope to test against, and the safe answer is no. The fix is a written scope, adversarial testing against that scope, and controls implemented and evidenced. Deloitte's 2Q 2026 CFO Signals, fielded 22 May to 7 June 2026 across 200 North American CFOs above $1B revenue, found only 19% of CFOs saying they hold the greatest responsibility for AI governance, with 33% pointing at the CISO.

How much gets spent before anyone checks the foundation?

More than it should, and that is the expensive pattern rather than the assessment. $2.4M across three initiatives at one client. $1.2M on an AI underwriting platform at another, before anyone measured the policy data against the score the platform required. A $12M modernisation decision at a state agency with no view of which programme areas were ready. In six of seven engagements, significant money went into the model or the platform before the foundation was measured.

What should we do first?

Run the pilot once against production-shaped data rather than a curated extract, and watch which way it fails. If it cannot be trained or scored, the blocker is data. If it runs and then breaks on the next schema change, the blocker is integration. If it works and nobody will approve it, the blocker is governance. If it works, is approved and nobody can say what it returned, the blocker is the business case and the owner. The failure mode names the blocker, and it costs a week.

The work behind this

Ten AI readiness assessment engagements in the case library diagnosed why a deployment was blocked and what to fix first, from 14 data layer failures across three stalled ML initiatives to eight programme areas scored before a $12M modernisation decision.

AI readiness assessment, 10 engagements.

Topics covered

  • AI pilots stall
  • AI pilot to production
  • why AI projects fail
  • AI readiness
  • proof of concept to production
  • AI deployment blockers
  • AI data readiness

Frequently asked questions

Why do AI pilots stall before production?

Four reasons, and the model is rarely one of them. The data will not support it, usually because records do not join, the signal was never instrumented, or a metric is defined two ways. The integration breaks, so it runs in a notebook and fails weekly in production on schema change. Governance will not approve it, because nobody wrote down what it may read and do. Or nobody can say what it returned and who owns it after launch, which is not a technical blocker at all and no engineering fixes it.

What do the published figures say?

Forrester, with Anaconda, in The State of Agentic AI in 2026 published 9 June 2026, reports 88% of AI agent pilots never reaching production, with three out of four enterprises having an agent project underway. The blockers it reports are evaluation gaps in 64% of cases, governance friction in 57% and model reliability in 51%. Forrester does not publish a sample size for those figures, so read them as the direction of a named study rather than a measured population.

What does this look like in practice?

A mid-market SaaS platform had spent $2.4M across three ML initiatives, churn prediction, recommendations and anomaly detection. Six months in, none were live and the data science team was repeatedly rebuilding the models. A six-week assessment found 14 data layer failures and three different causes: customer IDs did not match consistently, product interaction data was incomplete, and the pipelines were unstable. After remediation all three were live within 12 weeks.

Which figure should I keep out of the business case?

The 95% one. MIT Project NANDA's The GenAI Divide, published July 2025, is the source of the widely quoted claim that 95% of generative AI pilots produce no measurable financial return, and a Wharton professor has publicly stated he could not locate the basis for it. The report describes interviews, surveys and analysis of 300 public implementations without publishing the derivation. A case built on a contested headline fails at the first finance review that checks it.

Is the data the blocker more often than the model?

In our own assessments, consistently. The 14 failures at the SaaS platform were identity, instrumentation and pipeline stability. A pharmacy benefit manager had three AI projects stuck for over a year with member match accuracy at 75 of every 100 records; lifting it to 94 restarted all three within four weeks of delivery. A regional insurer's data science team spent 18 months trying to build a usable training dataset across four claims systems and failed twice.

Why does security block a pilot that works?

Because there is usually nothing concrete to assess. If nobody has written down what the system is allowed to read and what it is allowed to do, a reviewer has no scope to test against, and the safe answer is no. The fix is a written scope, adversarial testing against that scope, and controls implemented and evidenced. Deloitte's 2Q 2026 CFO Signals, fielded 22 May to 7 June 2026 across 200 North American CFOs above $1B revenue, found only 19% of CFOs saying they hold the greatest responsibility for AI governance, with 33% pointing at the CISO.

How much gets spent before anyone checks the foundation?

More than it should, and that is the expensive pattern rather than the assessment. $2.4M across three initiatives at one client. $1.2M on an AI underwriting platform at another, before anyone measured the policy data against the score the platform required. A $12M modernisation decision at a state agency with no view of which programme areas were ready. In six of seven engagements, significant money went into the model or the platform before the foundation was measured.

What should we do first?

Run the pilot once against production-shaped data rather than a curated extract, and watch which way it fails. If it cannot be trained or scored, the blocker is data. If it runs and then breaks on the next schema change, the blocker is integration. If it works and nobody will approve it, the blocker is governance. If it works, is approved and nobody can say what it returned, the blocker is the business case and the owner. The failure mode names the blocker, and it costs a week.

Related reading

If this is the problem you have