AI Deployment · 10 min read · October 2026
Fix data, integration or governance first? An AI deployment assessment
By Sean Majidi, Founder, Thinklytics
One test decides the order: run the pilot against production-shaped data and watch which way it fails. The failure mode names the blocker. Two real assessment rubrics, five dimensions each, and the five engagements where the triage moved money before the build rather than after it.
The pilot is stuck and three groups each have an explanation. The data team says the data is not ready. The platform team says the integration is the problem. Security says it cannot be approved as described. All three are usually right, which is why the question is order rather than blame.
The test
Three symptoms, three different first moves
Run the pilot once against production-shaped data and watch which way it fails. The failure mode names the blocker.
| What you observe | The blocker | What to fix first |
|---|---|---|
| The model cannot be trained or scored because records do not join, fields are missing, or the metric is defined two ways | Data | Identity resolution, instrumentation, or one written definition per metric |
| It runs in a notebook and breaks weekly in production when a schema changes or a source moves | Integration | Pipelines that detect and survive schema change, plus a monitored contract per source |
| It works, and security, legal or compliance will not approve it | Governance | A written scope of what it may read and do, adversarial testing against that scope, and implemented controls |
| Nobody can say what it returned or who owns it after launch | None of the three | A business case at real volume and a named owner. This is not a technical blocker and no engineering fixes it |
These are not mutually exclusive and most estates have two. The point of the test is the ORDER, because fixing governance on a model that cannot join its own records wastes the quarter.
Source: Thinklytics engagement pattern across the 10 AI readiness assessment engagements in the case library.
Run the pilot once against production-shaped data rather than a curated extract, and watch which way it fails. The failure mode names the blocker, and it costs about a week.
It cannot be trained or scored. Records will not join, a field is empty across a material share of rows, or the metric it is predicting is defined two ways so the label is unreliable. The blocker is data.
It runs, then breaks. It works in a notebook and fails in production when a schema changes or a source moves. The blocker is integration.
It works, and nobody will approve it. Security, legal or compliance will not sign. The blocker is governance, and usually because nobody has written down what the system may read and do, so there is no scope to assess.
It works, is approved, and nobody can say what it returned. The blocker is the business case and the owner. This is not a technical problem and no engineering addresses it.
Most estates have two of the four. The order still matters, because governance work on a model that cannot join its own records wastes the quarter, and a reliable pipeline carrying unreliable data is a faster route to the wrong answer.
Why data usually goes first
Not as a principle. Because of what the failures actually were.
A mid-market SaaS platform had spent $2.4M across three ML initiatives with none live after six months. The assessment found 14 data layer failures, and the three projects were blocked for three different reasons: customer IDs that did not match consistently, incomplete product interaction data, and unstable pipelines. Identity, instrumentation, pipeline stability.
A pharmacy benefit manager had three AI projects stuck for over a year across three member ID systems that never aligned, with match accuracy at 75 of every 100 records. A regional insurer's data science team spent 18 months on a training dataset across four claims systems and failed twice.
The aggregate surveys point the same way without resolving it. EY's 2026 Global DNA of the CFO survey, fielded 16 February to 30 March 2026 across 1,610 finance leaders at organisations above $1B revenue in 28 countries, found 61% citing data quality and bias as the top barrier to AI investment. Deloitte's Finance Trends 2026 survey of 1,326 global finance leaders found 47% citing data issues. Those are useful for a board paper and useless for sequencing, because "data issues" covers all three technical blockers. That is the gap the test closes.
Our own data puts a number on the order. Across the 47 engagements in the 2026 Enterprise Data Readiness Report, teams that spent six to eight weeks agreeing and writing down what their metrics meant before building shipped 81% of the time. Teams that skipped that step shipped 13% of the time. Same practice, same kind of problem, a six-fold difference in whether anything reached production.
That is one firm's engagement history rather than a controlled trial, and the teams that chose to do the definition work first were probably better run in other ways too. Read it as the strongest version of the sequencing argument we can evidence, not as a causal estimate.
When integration leads
When the model and the data are both fine in a notebook and the system cannot stay up.
An analytics vendor had a churn model predicting correctly 78 of every 100 times in testing that never went live. The pipeline broke several times a week because CRM and product usage data kept changing format, and without reliable data the customer success team could not see which accounts were at risk.
The unlock was rebuilding the pipeline to detect schema changes and repair them automatically. The accuracy improvement to 84 came from validating on a three-month holdout, which was a secondary gain. Then daily scoring in production with alerts to named customer success managers. Live in week 10, $2.6M of at-risk ARR flagged in three months, 340 accounts retained, and the pipeline ran 18 weeks continuously afterwards. See the churn model deployment engagement.
Nothing in that engagement was a data quality project, and starting with one would have left the model exactly where it was.
When governance leads
When a deadline or a condition sits on it.
A university system had $6.8M of National Science Foundation research funding conditional on certifying its research data infrastructure within 16 weeks, across 14 research departments, with 31 identified gaps. There, governance was the project: data cataloguing, access controls, lineage tracking and reproducibility documentation. All 31 gaps closed by week 14, two weeks before the deadline, certification granted, funding released, and 340 researchers now on the platform. See the research data certification engagement.
The same logic applies where a security sign-off blocks a fixed launch date, or an audit cycle is immovable. An enterprise SaaS company that had grown to four product lines through acquisition cut SOC 2 audit preparation from eight weeks to five days, freed roughly seven engineer-weeks per audit, had no findings in the first audit afterwards, and shortened enterprise sales cycles by three weeks because the governance posture answered the security questionnaires.
The ownership picture is why governance stalls when nobody owns it. Deloitte's 2Q 2026 CFO Signals, fielded 22 May to 7 June 2026 across 200 North American CFOs above $1B revenue, found 96% at least somewhat confident in their AI governance framework but only 43% confident, with 53.5% only somewhat, and just 19% of CFOs saying they hold the greatest responsibility for it against 33% pointing at the CISO.
Two rubrics we have actually run
Two assessment rubrics we have actually run
Five dimensions each. The difference between them is the audience: one scores an ML initiative, the other scores a programme area deciding whether to start.
| What it asks about | Mid-market SaaS, scoring three stalled ML initiatives | State agency, scoring 8 programme areas |
|---|---|---|
| Is the data there | Data completeness | Data quality |
| Does it mean one thing | Metric consistency | Completeness |
| Does it arrive reliably | Pipeline reliability | Process documentation |
| Can the work be reproduced | Feature engineering reproducibility | Staff skills |
| Can it run and be governed in production | Inference infrastructure readiness | Governance maturity |
The SaaS rubric found 14 specific data layer failures in six weeks and all three initiatives shipped within 12 weeks of remediation. The agency rubric found 3 of 8 programme areas ready now and 5 needing 6 to 18 months of foundation work first.
Source: Thinklytics case library, published assessment approaches and outcomes per engagement.
Five dimensions each, and the difference between them is the audience.
For scoring stalled ML initiatives: data completeness, metric consistency, pipeline reliability, feature engineering reproducibility, inference infrastructure readiness. That produced 14 named failures in six weeks and all three initiatives shipped within 12 weeks of remediation starting.
For scoring programme areas deciding whether to start at all: data quality, completeness, process documentation, staff skills, governance maturity. That scored eight programme areas and found three ready immediately and five needing 6 to 18 months of foundation work first.
Name the dimensions and the scoring before the assessment starts. An assessment with undeclared dimensions produces narrative, and narrative cannot be compared across programme areas or used to move a budget.
What the triage actually changed
What the triage changed, in five engagements
In each case the assessment moved money or sequence before the build, which is the only point of running one.
| Situation before | What the assessment found | What changed |
|---|---|---|
| $12M AI modernisation decision, 8 programme areas, no view of which were ready | 3 areas ready immediately, 5 needing 6 to 18 months of data and process work | $4.2M moved from AI projects to foundation work, $7.3M of opportunities sequenced over 24 months |
| $2.4M spent on 3 ML initiatives, 6 months, none live | 14 data layer failures: identity, instrumentation, pipeline stability | All three live within 12 weeks of remediation starting |
| 3 AI projects stuck for over a year, teams unable to prioritise the data work | Member match accuracy at 75 of every 100 records | Accuracy to 94, all three pilots restarted within 4 weeks of delivery, $4.8M a year of misrouted claims addressed |
| $1.2M already spent on an AI underwriting platform requiring a data quality score of 80 | Policy data at 61, with 8 failing dimensions on the vendor's own rubric | Score to 94 in 16 weeks for $340K against a $2.8M vendor quote, platform live 3 weeks later |
| $6.8M of federal research funding conditional on certifying data infrastructure in 16 weeks | 31 infrastructure gaps across 14 departments | All 31 closed by week 14, certification granted, funding released |
Four of the five had already spent money on the model or the platform before anyone scored the foundation. That sequence is the expensive part, not the assessment.
Source: Thinklytics case library, published delivery approaches and outcomes per engagement.
An assessment that does not move money or sequence before the build was not worth running. Five engagements where it did.
A state agency facing a $12M AI modernisation decision scored eight programme areas, found three ready and five not, and moved $4.2M from AI projects to foundational data work, with $7.3M of automation opportunities sequenced over 24 months. Past technology projects there had failed by ignoring exactly this. See the programme readiness assessment.
A workers compensation carrier had spent $1.2M on an AI underwriting platform requiring a data quality score of 80 across 10 categories, with policy data scoring 61 and a $2.8M vendor remediation quote on the table. Reading the vendor's own rubric found eight failing dimensions, six addressable with automated pipelines. Score to 94 in 16 weeks for $340K, platform live three weeks later.
In four of the five, money had already gone into the model or the platform before the foundation was scored. That sequence is the expensive part.
How to keep it from becoming a report
Three requirements, and they belong in the engagement terms.
Name the dimensions and the scoring before it starts. Require a prioritised remediation plan with a duration against each item rather than a maturity score. And agree what decision the assessment is feeding, and the date that decision gets made.
An assessment with no decision attached produces a document. That is the most common failure mode in this category and it is a scoping failure rather than an analytical one.
What we would do first
Run the test this week. One pilot, production-shaped data, watch the failure.
Then write the three-column table: what failed, which blocker it names, what the first remediation item would be. If that table has entries in two columns, you have your order. If it has entries in all four, the fourth one is the real problem and it is a conversation with a sponsor rather than a project.
What the resulting engagement has to hand over is in what a pilot-to-production engagement includes, the cheap version of the diagnostic is in the 3-question AI-ready data test, and what a full assessment covers is in the 30-day AI readiness assessment.
To put the result in front of whoever has to approve it, the AI deployment approval pack carries the business case, the security review and the responsibilities in one place.
Delivery sits in AI readiness for the assessment, managed data readiness where remediation runs as a service, data governance consulting and AI governance managed operations where governance leads, and vendor-neutral system integration where the pipelines are the blocker. The full set of work in this area sits under our AI work is not delivering.
Frequently asked questions
How do I decide whether to fix data, integration or governance first?
Run the pilot once against production-shaped data rather than a curated extract, and watch which way it fails. If it cannot be trained or scored because records will not join or a field is missing, the blocker is data. If it runs and then breaks when a schema changes or a source moves, the blocker is integration. If it works and security, legal or compliance will not approve it, the blocker is governance. If it works, is approved, and nobody can say what it returned, the blocker is the business case and the owner, which no engineering fixes.
What if two or three of them are blocking at once?
That is the normal case, and it is why the question is about order rather than selection. Fix data first where it blocks, because governance work on a model that cannot join its own records wastes the quarter, and integration work on records that do not resolve produces a reliable pipeline carrying unreliable data. Governance can run in parallel once there is something concrete to describe, and in regulated or funded work it sometimes has to lead because a deadline sits on it.
What does an assessment rubric actually look like?
We have run two shapes. For scoring stalled ML initiatives: data completeness, metric consistency, pipeline reliability, feature engineering reproducibility, and inference infrastructure readiness. For scoring programme areas deciding whether to start at all: data quality, completeness, process documentation, staff skills, and governance maturity. Five dimensions each. The difference is the audience: one diagnoses a build in flight, the other decides where to spend.
How long does an assessment take and what does it produce?
Six to 20 weeks in our case library, depending on scope. A six-week assessment of three stalled ML initiatives produced 14 named data layer failures and a prioritised remediation plan; all three were live within 12 weeks of remediation starting. A 20-week assessment of eight programme areas at a state agency produced three areas ready immediately, five needing 6 to 18 months of foundation work, and a 24-month sequencing of $7.3M of opportunities.
Does an assessment actually change anything?
It should change money or sequence before the build, and if it does not, it was not worth running. A state agency facing a $12M AI modernisation decision moved $4.2M from AI projects to foundational data work after scoring eight programme areas. A workers compensation carrier facing a $2.8M vendor remediation quote found eight failing dimensions on the vendor's own rubric and reached the required score for $340K. Those are decisions, not reports.
When should governance lead rather than follow?
When a deadline or a condition sits on it. A university system had $6.8M of federal research funding conditional on certifying its research data infrastructure within 16 weeks across 14 departments, with 31 identified gaps. There, governance was the project and the data work was in service of it; all 31 were closed by week 14. The same applies where a security sign-off blocks a launch date, or where an audit or exam cycle is fixed.
Is integration ever the first thing to fix?
Yes, when the model and the data are both fine in a notebook and the system cannot stay up. One churn model tested at 78 of every 100 correct and never shipped, because the pipeline broke several times a week when CRM and product usage data changed format. Rebuilding the pipeline to detect and repair schema changes was the whole unlock, and the model's accuracy was a secondary gain from a proper holdout.
How do I keep an assessment from becoming a report nobody acts on?
Three things. Name the dimensions and the scoring before it starts, so the output is comparable rather than narrative. Require a prioritised remediation plan with durations per item rather than a maturity score. And agree in advance what decision the assessment is feeding, with the date that decision gets made. An assessment with no decision attached produces a document, which is the most common failure mode in this category.
The work behind this
Ten AI readiness assessment engagements in the case library diagnosed what was blocking a deployment and in what order to fix it. In four of the five where the spend is published, money had already gone into the model or the platform before the foundation was scored.
AI readiness assessment, 10 engagements.
Topics covered
- AI deployment assessment
- fix data or governance first
- AI readiness triage
- AI blockers
- data readiness assessment rubric
- AI pilot remediation
- sequencing AI investment
Frequently asked questions
How do I decide whether to fix data, integration or governance first?
Run the pilot once against production-shaped data rather than a curated extract, and watch which way it fails. If it cannot be trained or scored because records will not join or a field is missing, the blocker is data. If it runs and then breaks when a schema changes or a source moves, the blocker is integration. If it works and security, legal or compliance will not approve it, the blocker is governance. If it works, is approved, and nobody can say what it returned, the blocker is the business case and the owner, which no engineering fixes.
What if two or three of them are blocking at once?
That is the normal case, and it is why the question is about order rather than selection. Fix data first where it blocks, because governance work on a model that cannot join its own records wastes the quarter, and integration work on records that do not resolve produces a reliable pipeline carrying unreliable data. Governance can run in parallel once there is something concrete to describe, and in regulated or funded work it sometimes has to lead because a deadline sits on it.
What does an assessment rubric actually look like?
We have run two shapes. For scoring stalled ML initiatives: data completeness, metric consistency, pipeline reliability, feature engineering reproducibility, and inference infrastructure readiness. For scoring programme areas deciding whether to start at all: data quality, completeness, process documentation, staff skills, and governance maturity. Five dimensions each. The difference is the audience: one diagnoses a build in flight, the other decides where to spend.
How long does an assessment take and what does it produce?
Six to 20 weeks in our case library, depending on scope. A six-week assessment of three stalled ML initiatives produced 14 named data layer failures and a prioritised remediation plan; all three were live within 12 weeks of remediation starting. A 20-week assessment of eight programme areas at a state agency produced three areas ready immediately, five needing 6 to 18 months of foundation work, and a 24-month sequencing of $7.3M of opportunities.
Does an assessment actually change anything?
It should change money or sequence before the build, and if it does not, it was not worth running. A state agency facing a $12M AI modernisation decision moved $4.2M from AI projects to foundational data work after scoring eight programme areas. A workers compensation carrier facing a $2.8M vendor remediation quote found eight failing dimensions on the vendor's own rubric and reached the required score for $340K. Those are decisions, not reports.
When should governance lead rather than follow?
When a deadline or a condition sits on it. A university system had $6.8M of federal research funding conditional on certifying its research data infrastructure within 16 weeks across 14 departments, with 31 identified gaps. There, governance was the project and the data work was in service of it; all 31 were closed by week 14. The same applies where a security sign-off blocks a launch date, or where an audit or exam cycle is fixed.
Is integration ever the first thing to fix?
Yes, when the model and the data are both fine in a notebook and the system cannot stay up. One churn model tested at 78 of every 100 correct and never shipped, because the pipeline broke several times a week when CRM and product usage data changed format. Rebuilding the pipeline to detect and repair schema changes was the whole unlock, and the model's accuracy was a secondary gain from a proper holdout.
How do I keep an assessment from becoming a report nobody acts on?
Three things. Name the dimensions and the scoring before it starts, so the output is comparable rather than narrative. Require a prioritised remediation plan with durations per item rather than a maturity score. And agree in advance what decision the assessment is feeding, with the date that decision gets made. An assessment with no decision attached produces a document, which is the most common failure mode in this category.
Related reading
If this is the problem you have
- Our AI work is not delivering, resolved by 5 services.
- AI Deployment Approval Pack, the worksheet for whoever has to approve the spend.
- The 30 day Corporate Drag and Risk Diagnostic, findings yours either way.