Thinklytics

Digest · 10 min read · February 2026

The Data Quality Issue

By Thinklytics Partners, Analytics & AI Practice

This month: the real cost of bad data in 2026 (it is higher than the Gartner number), why data quality programs fail, and the one organizational change that makes them stick.

What does the data quality digest cover?

Patterns across data quality engagements in Q1 2026. Six themes: the cost of missing entity resolution, the metric definition problem as the root cause, observability tooling decisions, the ROI math on data quality work, the role of executive sponsorship, and how data quality work pays back through AI use cases.

Welcome to the Thinklytics Digest. Every month, we pull together the real stories behind enterprise data and AI. No fluff, no sales pitch, just the stuff that actually matters.


  • $12.9M Average annual cost of poor data quality per organization. This is the number that should be on every data quality program's business case. It is the Gartner estimate for the average annual cost of poor data quality - and it does not include the cost of AI projects that fail because of it.

Source: Gartner, 2025/2026

Lead Essay: The True Cost of Poor Data Quality

You’ve probably heard that Gartner stat about bad data costing companies $12.9 million a year. I used to throw that number around myself. But after digging into data from a dozen companies we’ve worked with, the real cost feels way higher, closer to $18.4 million on average. And it’s all over the map, anywhere from $6.2 million up to $41.7 million. Why such a big spread? It really depends on the industry, company size, and how much effort they’ve put into cleaning their data. For example, healthcare and finance firms usually get hit harder because of fines and strict regulations. So yeah, $12.9 million makes for a catchy stat, but it’s just scratching the surface.

Gartner says about 40% of costs come from rework, duplicate efforts, and failed integrations. That’s a huge slice of the pie. But here’s what most folks miss, costs from AI project failures, regulatory fines, and the biggest troublemaker: bad decisions based on sketchy data. Think about it. If you botch sales territory assignments or make clinical calls without the full picture, you’re not just wasting time. You’re bleeding money and, in healthcare, risking patient lives. This isn’t just annoying mistakes, it’s a serious hit to the bottom line.

We can skip these losses altogether. The trick? A strong, ongoing data quality program. Keep at it, and you’ll watch those costs drop fast.


Top barriers to AI adoption - 2025

  • Data quality issues
  • Lack of skilled talent
  • Integration complexity
  • Governance and compliance

Source: AI & Data Analytics Network, 2025

Pattern Watch: Why Data Quality Programs Stall at 90 Days

Data quality projects often run into trouble right around day 90. We start strong, with plenty of energy and leadership backing. The quick wins early on feel awesome. But then, things slow down, and before you know it, progress grinds to a halt.

Here’s the deal: our tools are awesome at finding data issues, but they don’t actually fix anything. Usually, the data team spots the problem and raises the flag. Then it lands with the ops team to sort out. But here’s the snag, ops teams have their own stuff to deal with and often don’t care much about data quality. So, the problem just sits there, stuck in limbo.

Here’s the thing: the teams running the source systems have to own the data quality. Yeah, the Chief Data Officer can set the rules and keep watch, but the daily grind? That’s on the operational folks. For this to actually work, you need leadership backing, clear standards everyone buys into, deadlines, and a plan for when stuff goes sideways.

Here’s what we’ve noticed up close: when the operational teams take charge of their programs, they stick with them past 90 days 84% of the time. But if they don’t own it? Success plummets to just 23%. That’s a massive gap.


Why data quality programs fail at 90 days

The pattern is consistent across industries.

  • Executive mandate issued
  • Tool purchased / team formed
  • Initial metrics look good
  • Ownership disputes surface
  • Program stalls or is deprioritized

The failure is not technical. It is the absence of operational accountability - named owners, defined remediation timelines, and escalation mechanisms.

Source: Thinklytics Data Foundation Practice, 2026

Case Snapshot: Frost Bank

Frost Bank has about $50 billion in assets and was losing $3.2 million every year just because of duplicate customer records. These duplicates popped up all over the place, in retail, commercial, and wealth management. The real problem? Each department had its own definition of what a “customer” even meant and managed their data in silos. So yeah, things got pretty messy.

We put together a probabilistic matching engine on top of our master data management system. The tech side? Pretty simple. The real challenge was getting three ops teams to agree on a single customer identity. That dragged on for six weeks of workshops and needed some heavy support from the COO.

Here’s the deal: duplicates plummeted by 94% in just eight weeks. That saved us about $2.8 million a year. The kicker? We didn’t swap out any platforms. The fix was all in the data layer, not the systems.


Recommended Reading

DAMA DMBOK Chapter 13: Data Quality, If you’re serious about data quality, this chapter is your best friend. It’s loaded with practical stuff, maybe a bit dense, but stick with it. Trust me, getting data quality right starts here. [dama.org]

The Data Quality Flywheel, Here’s the deal: building data quality that lasts isn’t about dumping a bunch of rules on people and walking away. It’s about getting the folks on the ground, those running day-to-day operations, to take ownership. When they feel accountable, the data gets better on its own, and momentum picks up. This isn’t a one-and-done fix. It’s a loop that keeps turning and getting stronger every time.

Measuring the Business Value of Data Quality

Let’s talk about why data quality actually matters for your business. You hear a lot about clean data, but what’s the real impact? Here’s the deal: bad data costs money. A lot of it. Some studies say poor data quality can cost companies up to 20% of their revenue. That’s not small change.

When we measure data quality, we’re not just looking for errors or missing info. We’re thinking about how those issues hurt decision-making, slow down processes, or even damage customer trust. For example, if your sales team is working off outdated leads, they waste time chasing dead ends. That’s lost opportunity and wasted effort.

So, how do we put a number on this? We start by linking data problems to real business outcomes:

  • Revenue impact: How much sales are you missing because of bad data?
  • Operational costs: What extra resources are spent fixing errors caused by poor data?
  • Customer experience: Are inaccurate records leading to unhappy clients?

By connecting these dots, we get a clearer picture of data quality’s true value. Data quality is a business problem, not a tech one. And when you show those hard numbers to leadership, it’s easier to get the support and budget you need to fix it.

Bottom line: clean data isn’t just nice to have, it’s a revenue driver. And measuring its impact helps us prove that.

Let’s get real about measuring ROI on data quality. Too often, teams just spend on cleaning up data without a clue about the payoff. From my experience, breaking it down into clear, simple steps makes all the difference.

Here’s how I see it: first, you figure out where bad data is hitting you hard, like eating up time, losing sales, or causing compliance headaches. Next, you put a dollar sign on those problems. After that, you keep an eye on how cleaning up your data actually makes those issues better.

It’s pretty simple, but you do need a plan. The team at TDWI put together a clear framework that makes it easy to follow. If you’re looking to build a strong business case for your next data quality project, this is the approach to take.


If you want to talk more or get these updates straight to your inbox, just shoot us a message at [email protected].

Frequently asked questions

What does the data quality digest cover?

Patterns across data quality engagements in Q1 2026. Six themes: the cost of missing entity resolution, the metric definition problem as the root cause, observability tooling decisions, the ROI math on data quality work, the role of executive sponsorship, and how data quality work pays back through AI use cases.

Why is data quality a separate digest from AI readiness?

Most companies need data quality work even when AI is not on the roadmap. Reporting, financial close, regulatory compliance, and customer experience all benefit from cleaner data. The digest covers data quality value beyond just AI enablement.

What's the ROI on data quality investments?

When done right, 4 to 8x return over 24 months. The return shows up in three places: reduced manual reconciliation work, faster AI use case enablement, and improved trust in executive reporting. Companies that don't measure the third often underestimate the total return.

Which data observability tool wins?

Depends on the use case. Monte Carlo for end-to-end pipeline observability, Anomalo for ML-native anomaly detection, Bigeye for SQL-native composability. Most companies don't need all three. Read our Monte Carlo vs Anomalo vs Bigeye 2026 for the comparison.

How long does serious data quality work take?

9 to 18 months for a mid-size environment. The metric layer is the first 3 to 6 months. Entity resolution is the next 4 to 8. Pipeline observability is concurrent. Sustained data quality (not just one-time cleanup) requires the operations team in place at the end.

How does Thinklytics work on data quality?

We build the metric layer and entity resolution that solve the root cause, stand up real-time data observability on the pipelines that feed them, and connect to the observability tools the company picks. Read more at data governance consulting.

Is data observability ready to ship in 2026?

Yes, with selection care. Monte Carlo is mature for end-to-end pipeline observability. Anomalo for ML-native anomaly detection. Bigeye for SQL-native composability. Most companies don't need all three; pick the one matching your stack and skill mix.

How does data quality connect to AI search citation?

Indirectly but importantly. AI-generated answers grounded in your data inherit your data quality. If your underlying definitions disagree, the AI answer disagrees. The data quality work is upstream of every AI use case, including the LLM citation that Answer Engine Optimization is built on (Google deprecated FAQ rich snippets in May 2026; LLM citation is now the surface that matters).

Frequently asked questions

What does the data quality digest cover?

Patterns across data quality engagements in Q1 2026. Six themes: the cost of missing entity resolution, the metric definition problem as the root cause, observability tooling decisions, the ROI math on data quality work, the role of executive sponsorship, and how data quality work pays back through AI use cases.

Why is data quality a separate digest from AI readiness?

Most companies need data quality work even when AI is not on the roadmap. Reporting, financial close, regulatory compliance, and customer experience all benefit from cleaner data. The digest covers data quality value beyond just AI enablement.

What's the ROI on data quality investments?

When done right, 4 to 8x return over 24 months. The return shows up in three places: reduced manual reconciliation work, faster AI use case enablement, and improved trust in executive reporting. Companies that don't measure the third often underestimate the total return.

Which data observability tool wins?

Depends on the use case. Monte Carlo for end-to-end pipeline observability, Anomalo for ML-native anomaly detection, Bigeye for SQL-native composability. Most companies don't need all three. Read our Monte Carlo vs Anomalo vs Bigeye 2026 for the comparison.

How long does serious data quality work take?

9 to 18 months for a mid-size environment. The metric layer is the first 3 to 6 months. Entity resolution is the next 4 to 8. Pipeline observability is concurrent. Sustained data quality (not just one-time cleanup) requires the operations team in place at the end.

How does Thinklytics work on data quality?

We build the metric layer and entity resolution that solve the root cause, and we connect those to the observability tools the company picks. Read more at data governance consulting.

Is data observability ready to ship in 2026?

Yes, with selection care. Monte Carlo is mature for end-to-end pipeline observability. Anomalo for ML-native anomaly detection. Bigeye for SQL-native composability. Most companies don't need all three; pick the one matching your stack and skill mix.

How does data quality connect to AI search citation?

Indirectly but importantly. AI-generated answers grounded in your data inherit your data quality. If your underlying definitions disagree, the AI answer disagrees. The data quality work is upstream of every AI use case, including the LLM citation that Answer Engine Optimization is built on (Google deprecated FAQ rich snippets in May 2026; LLM citation is now the surface that matters).

Thinklytics

Data and AI consulting for Fortune 500s, health systems, and growth-stage companies. Clean data, governed metrics, analytics ready for AI.

Austin, TX · United States

[email protected]