Cost & ROI · 8 min read · September 2026
How to Verify an AI Outcome Claim Before You Budget Against It
By Thinklytics Partners, Cost & ROI
Almost every outcome number published about enterprise AI is unsourced. Not disputed, unsourced. No study, no sample, no baseline, no measurement window. Why the document with the most confident percentages is reliably the least trustworthy one, and the five questions that settle it.
Almost every outcome figure published about enterprise AI is unsourced. Not disputed, not contested, not results may vary. Unsourced. There is no study behind it, no sample, no definition of what was measured or against what baseline. Once you start checking, the market's confidence and the market's evidence turn out to have almost nothing to do with each other.
What prompted this
Six catalogues, one outlier
- Five of the six. Directional language. Faster cycle time, fewer exceptions, less rework. Reasonable, and impossible to check, which is the correct posture when you cannot check it.
- The sixth. Confident percentages. No sources and no conditions. No description of the workload, the baseline, the measurement window, or the number of clients the figure was drawn from.
We read six service catalogues from firms selling AI implementation, four of them from organisations considerably larger than this one. The document with the most confident numbers was the least reliable document in the set.
Source: Thinklytics AI readiness practice, 2026.
We read six service catalogues from firms selling AI implementation, four of them from organisations considerably larger than this one. Five described outcomes in directional language: faster cycle time, fewer exceptions, less rework. Reasonable, and impossible to check, which is the correct posture when you cannot check it.
The sixth offered numbers. Over 90% reduction in processing costs. Resolves up to 70% of inbound inquiries. Slashes cloud costs by 30% to 60%. Guaranteed system compliance.
No sources. No conditions. No description of the workload, the baseline, the measurement window, or the number of clients the figure was drawn from. The same document listed fine tuning as the cure for hallucination, which it is not.
That is the pattern worth naming. The document with the most confident numbers was the least reliable document in the set, and not by a little.
Why the numbers are constructed this way
Which of these can you check
Ticked claims can be checked by the reader. The rest cannot be, whatever confidence they arrive with.
- Over 90% reduction in processing costs. The baseline is unstated. A reduction against a fully manual baseline is a different claim from one against a partly automated baseline, and the first is much easier to achieve.
- Resolves up to 70% of inbound inquiries. Up to has no floor. One client, one quarter, one queue of password resets satisfies it, and so does zero.
- Slashes cloud costs by 30% to 60%. No workload described, no measurement window, and no count of the clients the range came from.
- Guaranteed system compliance. Compliance is a judgement made by a regulator or an auditor about your organisation. A vendor can supply controls, evidence, logging and documentation. It cannot guarantee the judgement.
- Directional outcomes, stated as directional. Five of the six catalogues described results this way. Clear about its own limits, and still not something you can check.
- A per-engagement figure attached to a named situation. The starting condition is on the page, so you can decide for yourself whether your situation resembles it. A weaker claim than a benchmark and a more useful one.
- One number from a scoped pilot on your own data. Success measure written down before it starts, baseline captured before anyone touches anything. Two to four weeks.
There is currently no credible public benchmark for enterprise AI outcomes. Anyone presenting one is presenting marketing.
Source: Thinklytics AI readiness practice, 2026.
None of this requires bad faith. It is what the incentives produce.
Up to is doing all the work. Resolves up to 70% of inbound inquiries is satisfied by one client, one quarter, one queue of password resets. It is also satisfied by zero, since up to has no floor. As a sentence it is close to unfalsifiable, which is precisely why it survives legal review.
The baseline is unstated and usually generous. A 90% reduction in processing cost against a fully manual baseline is a different claim from 90% against a partly automated one. Nobody publishes which they mean, and the first is much easier to achieve.
Pilots are reported, production is not. The impressive figure is nearly always from a narrow pilot on a curated subset. What it does across the whole queue, twelve months in, with the messy cases included, is rarely measured and never published.
Successful clients do not release their numbers. Firms with real data are usually contractually unable to publish it, and clients who could publish it have no reason to. So the loudest numbers come from the firms with the least to lose by being wrong.
Guaranteed compliance is not a thing anyone can sell you. Compliance is a judgement made by a regulator or an auditor about your organisation. A vendor can supply controls, evidence, logging and documentation that make a favourable judgement more likely. It cannot guarantee the judgement.
Why we do not publish outcome benchmarks either
We are asked for them. The honest answer is that the number would not mean anything.
Our engagements differ in starting condition more than in anything we do. Two clients asking the same question, why does the close take nine days, arrive with different systems, different data quality, different tolerance for changing a definition, and different willingness to make the decision the answer requires. A percentage averaged across those is arithmetic, not evidence.
What we do publish is per-engagement figures, attached to a named situation, with the starting condition described. A reconciliation cost that went from one figure to another, in a stated number of weeks, at an organisation whose circumstances are on the page. You can decide for yourself whether your situation resembles it. That is a weaker claim than a benchmark and a more useful one.
What to ask when someone shows you a number
Five questions to ask when someone shows you a number
In order. They are not hostile and a good firm will enjoy them.
| Ask | What it exposes |
|---|---|
| Measured against what baseline, and who established it? | If the baseline was estimated after the fact by the party being paid, the figure is an opinion. |
| Over what period, and was it still true at twelve months? | Most AI figures are pilot figures. Accuracy decays, content goes stale, upstream systems change. Year one and year two are different numbers. |
| What share of the total volume does this cover? | Resolves 70% often means 70% of the category chosen for the pilot, which may be 15% of the actual queue. |
| How many clients is this drawn from, and can I speak to one? | A single reference call tells you more than any published figure. |
| What did it cost to run in year two? | Build cost is quoted. Running cost is where programmes quietly become unaffordable. |
Source: Thinklytics AI readiness practice, 2026.
Five questions, in order. They are not hostile and a good firm will enjoy them.
1. Measured against what baseline, and who established it? If the baseline was estimated after the fact by the party being paid, the figure is an opinion. 2. Over what period, and was it still true at twelve months? Most AI figures are pilot figures. Accuracy decays, content goes stale, upstream systems change. Year one and year two are different numbers. 3. What share of the total volume does this cover? Resolves 70% often means 70% of the category chosen for the pilot, which may be 15% of the actual queue. 4. How many clients is this drawn from, and can I speak to one? A single reference call tells you more than any published figure. 5. What did it cost to run in year two? Build cost is quoted. Running cost is where programmes quietly become unaffordable.
The position this leaves you in
There is currently no credible public benchmark for enterprise AI outcomes. Anyone presenting one is presenting marketing. That is uncomfortable to sit with when a board wants a projected return, and it is better than the alternative, which is committing to a number you took from a vendor's brochure and will be asked about in twelve months.
The workable substitute is a scoped pilot on your own data, with the success measure written down before it starts and the baseline captured before anyone touches anything. Two to four weeks. It produces one number, it is about you, and it is the only number in this whole category that will survive a question.
Frequently asked questions
Are published AI ROI figures reliable?
Mostly not. The figures circulating in vendor material are unsourced: no study behind them, no sample size, no definition of what was measured or against what baseline. The pattern worth noticing is that the documents with the most confident percentages tend to be the least reliable documents in the set.
What does up to 70% actually mean in a vendor claim?
It has no floor. A claim that a system resolves up to 70% of inbound inquiries is satisfied by one client, one quarter, one queue of password resets, and it is also satisfied by zero. That is why the phrasing survives legal review, and it is why the number tells you nothing about your own queue.
Why do pilot numbers not hold in production?
The impressive figure is usually from a narrow pilot on a curated subset. The number that matters is what the system does across the whole queue, twelve months in, with the messy cases included. That number is rarely measured and almost never published, and it is lower.
Can a vendor guarantee compliance?
No. Compliance is a judgement made by a regulator or an auditor about your organisation. A vendor can supply controls, evidence, logging and documentation that make a favourable judgement more likely, but it cannot guarantee the judgement. A firm that writes guaranteed compliance has either not thought about it or is hoping you will not.
What questions should we ask when a firm shows us a number?
Five. What baseline was this measured against and who established it. Over what period, and was it still true at twelve months. What share of total volume does it cover. How many clients is it drawn from, and can we speak to one. What did it cost to run in year two. A good firm will enjoy all five.
Why does Thinklytics not publish outcome benchmarks?
Because the number would not mean anything. Our engagements differ in starting condition more than in anything we do, and a percentage averaged across those is arithmetic rather than evidence. We publish per-engagement figures with the starting condition described, so you can judge whether your situation resembles it.
What should we do instead of using a vendor benchmark?
Run a scoped pilot on your own data, with the success measure written down before it starts and the baseline captured before anyone touches anything. Two to four weeks. It produces one number, it is about you, and it is the only number in this category that will survive a question in twelve months.
Topics covered
- AI ROI
- AI outcome claims
- vendor benchmarks
- AI business case
- baseline measurement
- AI pilot
Frequently asked questions
Are published AI ROI figures reliable?
Mostly not. The figures circulating in vendor material are unsourced: no study behind them, no sample size, no definition of what was measured or against what baseline. The pattern worth noticing is that the documents with the most confident percentages tend to be the least reliable documents in the set.
What does up to 70% actually mean in a vendor claim?
It has no floor. A claim that a system resolves up to 70% of inbound inquiries is satisfied by one client, one quarter, one queue of password resets, and it is also satisfied by zero. That is why the phrasing survives legal review, and it is why the number tells you nothing about your own queue.
Why do pilot numbers not hold in production?
The impressive figure is usually from a narrow pilot on a curated subset. The number that matters is what the system does across the whole queue, twelve months in, with the messy cases included. That number is rarely measured and almost never published, and it is lower.
Can a vendor guarantee compliance?
No. Compliance is a judgement made by a regulator or an auditor about your organisation. A vendor can supply controls, evidence, logging and documentation that make a favourable judgement more likely, but it cannot guarantee the judgement. A firm that writes guaranteed compliance has either not thought about it or is hoping you will not.
What questions should we ask when a firm shows us a number?
Five. What baseline was this measured against and who established it. Over what period, and was it still true at twelve months. What share of total volume does it cover. How many clients is it drawn from, and can we speak to one. What did it cost to run in year two. A good firm will enjoy all five.
Why does Thinklytics not publish outcome benchmarks?
Because the number would not mean anything. Our engagements differ in starting condition more than in anything we do, and a percentage averaged across those is arithmetic rather than evidence. We publish per-engagement figures with the starting condition described, so you can judge whether your situation resembles it.
What should we do instead of using a vendor benchmark?
Run a scoped pilot on your own data, with the success measure written down before it starts and the baseline captured before anyone touches anything. Two to four weeks. It produces one number, it is about you, and it is the only number in this category that will survive a question in twelve months.