Thinklytics

Metric Governance · 10 min read · October 2026

Fix your KPI definitions or rebuild your data pipelines? The test that decides

By Sean Majidi, Founder, Thinklytics

Both diagnoses look identical from the executive seat, and they cost different amounts by a factor of three. There is a half-day test that tells you which one you have, and most companies skip it and buy the expensive answer.

Two teams report the same metric and get different answers. From the executive seat the two possible causes look identical, and they cost different amounts by roughly a factor of three.

Cause one: the rules differ. Each team is computing a defensible figure under its own definition, and both are reproducible. Cause two: the plumbing is unreliable. The same rule on the same source returns a different answer depending on when you ask.

Almost nobody runs the test that separates them, because the test is unglamorous and the platform proposal arrives with a diagram.

The test

The test that separates the two diagnoses

Half a day of work, and it decides where a six-figure budget goes.

  • Pick one disputed number and one period
  • Each team writes down its rule before looking at the data
  • Run both rules against the SAME source table
  • Numbers now match: definition problem
  • Numbers still differ: pipeline problem

If both rules applied to one source produce the same figure, nothing is broken in the plumbing. The disagreement was the rules, and no amount of re-engineering will settle it.

Source: Thinklytics engagement pattern across the 11 semantic layer and metric governance engagements in the case library.

Pick one disputed number and one period. Have each team write its rule down in prose, before looking at any data, so nobody reverse-engineers a rule from the answer they want. Then run both rules against the same source table.

Two outcomes, and they point in opposite directions.

The figures now agree. The plumbing was never the problem, because one source under two stated rules produced one answer. The disagreement was the rules, and a pipeline rebuild will deliver a more modern system that still returns two numbers.

The figures still differ. Same rule, same source, different answers means something underneath is wrong, and the engineering case is real. Now you have evidence for it rather than a suspicion, which also means you can scope it.

Half a day of work decides where a six-figure budget goes. In our experience the test is skipped in most organisations that have been arguing about a number for more than a quarter.

Which one it usually is

What data teams say the obstacle is

Trust rose as a priority faster than the ownership question got answered, which is the shape of a definition problem rather than a tooling one.

  • Trust in data as a strategic priority (from 66% a year earlier)
  • Ambiguous data ownership still a challenge
  • Stakeholder data literacy still a barrier
  • Lack of stakeholder trust as a primary challenge (from 33%)

Source: dbt Labs State of Analytics Engineering 2026, fielded 5 December 2025 to 1 February 2026, n=363, 73% practitioners and 27% management, concentrated in North America and Europe. Vendor-run and practitioner-skewed, so read the direction rather than the level.

The giveaway for a definitions problem is reliability. Each team reproduces its own number cycle after cycle. A broken pipeline does not produce two consistently different answers, it produces unreliable ones, so consistency is the tell.

The dbt Labs State of Analytics Engineering 2026 survey, fielded 5 December 2025 to 1 February 2026 across 363 respondents at 73% practitioners and 27% management, concentrated in North America and Europe, found ambiguous data ownership still a challenge for 41%, effectively unchanged year over year, while trust in data rose as a strategic priority from 66% to 83%. That is a small, vendor-run, practitioner-skewed sample, so read the direction rather than the level. The direction is clear: the priority moved and the ownership question did not get answered.

Deloitte's Finance Trends 2026 survey of 1,326 global finance leaders found 47% citing data issues as a barrier to AI adoption in the finance function. EY's 2026 Global DNA of the CFO survey, fielded 16 February to 30 March 2026 across 1,610 finance leaders at organisations above $1B revenue in 28 countries, found 61% citing data quality and bias as the top barrier to AI investment.

Those are both aggregate "data" answers, and that is the problem with using them to choose a route. "Data quality" in a survey response covers both diagnoses, which is exactly why the test matters more than the benchmark.

What each route actually took

What each route actually took

Delivery durations from the case library. The definition route resolved the disputed number without touching the platform.

EngagementRoute takenDelivery
Growth-stage SaaS, five ARR definitionsDefinitions first, metric layer on the existing warehouse8 weeks
Kaiser Permanente, 14 regional encounter definitionsDefinitions first, certified layer on the existing EDW11 weeks
Enterprise SaaS, six revenue metricsDefinitions first, four reporting surfaces repointed14 weeks
Public university, 42 metrics across 8 unitsDefinitions first, certified layer in the existing warehouse20 weeks
Texas A&M System, 11 campus warehousesPlatform consolidation24 weeks
Telecom, nine reporting systems replacedPlatform consolidation26 weeks

The two routes scale on different things. Definition work scales with the number of contested metrics, which is why 42 of them took 20 weeks. Platform work scales with the number of systems. Kaiser had already failed twice internally before the definitions were written down, and neither attempt was short of engineering.

Source: Thinklytics case library, delivery durations as published per engagement.

These are delivery durations from our own case library rather than an industry average, and the overlap is the interesting part.

The definition route ran 8 to 20 weeks. Five competing ARR definitions at a growth-stage SaaS platform took 8 weeks, with the metric layer built in dbt on the warehouse already in place. Fourteen regional definitions of a patient encounter at Kaiser Permanente took 11 weeks, on the existing enterprise data warehouse, with the platform untouched throughout. Forty-two metrics across eight administrative units at a public university took 20.

The platform route ran 16 to 26 weeks. Eleven campus warehouses consolidated into one took 24. Nine reporting systems replaced took 26.

So the honest claim is not that definitions are always faster. It is that the two routes scale on different things. Definition work scales with the number of contested metrics. Platform work scales with the number of systems. If you have four contested metrics and nine warehouses, the definitions are the cheap half. If you have 42 contested metrics and one warehouse, they are not.

What is consistent is the failure mode when the order is wrong. Kaiser had already failed twice internally before the definitions were written down, and neither attempt was short of engineering. See the Kaiser Permanente metric governance engagement.

When the pipeline really is the problem

Four conditions. At least one has to hold, and if none of them do, the engineering case is not yet made.

The same query on the same table returns different answers on different days. Non-determinism in the load, usually late-arriving records or an upsert that is not idempotent.

Numbers change retroactively with no documented restatement. Someone can see last quarter moved and nobody can say why. This is a lineage and change-control failure and no definition fixes it.

A material share of rows fail a basic integrity check. Orphaned foreign keys, duplicated primary keys, nulls in required fields. Measure it before asserting it, because the share is often far smaller than the argument implies.

The data arrives too late to use. Correct and unusable is still unusable. Intuit's Future of Finance 2026 report, fielded in May 2026 by CatalystMR across 2,000 US CFOs, controllers and VPs of Finance at businesses above $2.5M revenue, found only 14% had same-day data for their most recent major business decision and 57% had missed a time-sensitive strategic action because visibility arrived too late. This is the condition most often misread as a definitions problem, and it is not one.

The sequencing mistake that costs the most

Running the platform migration first and expecting it to resolve the disagreement.

A migration carries the existing definitions forward unless somebody stops it, because the reports have to keep working through cutover. So the new warehouse arrives with the same two numbers in a more modern format, the migration gets blamed for not fixing a problem it was never scoped to touch, and the definitions work then has to be funded a second time against an organisation that has just spent six months on data and has nothing it trusts to show for it.

If both pieces are needed, write the definitions and migrate onto them. Both of the large consolidations in our case library built a shared semantic layer as part of the migration rather than after it, which is why the reporting stayed usable through cutover.

What we would do first

Run the test on one metric this week. Not the whole portfolio, one metric, the one with the largest gap between versions and the most senior person complaining about it.

Write both rules down. Run both against one source. Then you have a diagnosis instead of a debate, and whichever way it comes out you can scope the next step and defend the number in front of a finance committee.

If it comes out as definitions, the enforcement layer is semantic layer engineering and the standing rules are data governance consulting. If the dispute turns out to be about which records refer to the same entity, it is master data management or data cleaning and preparation instead, and on SAP estates SAP data quality and governance. The mechanism behind the disagreement is in why sales and finance report different revenue, and what the resulting engagement has to hand over is in what a reconciliation engagement delivers.

To put a number on either route before you ask for the budget, the reporting improvement business case is the worksheet we use with clients. The full set of work in this area sits under we cannot trust the numbers.

Frequently asked questions

How do I know whether the problem is definitions or pipelines?

Run one test before committing budget. Pick a single disputed number and one period. Have each team write its rule down before looking at the data. Then run both rules against the same source table. If the two figures now agree, the plumbing works and the disagreement was always the rules, so a pipeline rebuild buys you nothing. If they still differ when the rule and the source are identical, something underneath is broken and the engineering case is real. The test costs about half a day and it decides where a six-figure budget goes.

Which problem is it usually?

Definitions, in most cases we are called into. The giveaway is that each team can reproduce its own number reliably, cycle after cycle. Reliable production of two different answers is a rules problem by definition, because a broken pipeline produces unreliable answers rather than consistently different ones. The dbt Labs State of Analytics Engineering 2026 survey, fielded 5 December 2025 to 1 February 2026 across 363 respondents, found ambiguous data ownership still a challenge for 41%, effectively unchanged year over year, while trust in data rose as a strategic priority from 66% to 83%. Priority moved, ownership did not.

What does each route cost in time?

In our case library the definition route ran 8 to 20 weeks and the platform consolidation route ran 16 to 26. The ranges overlap, and the reason matters more than the averages: definition work scales with the number of contested metrics, platform work scales with the number of systems. Five ARR definitions took 8 weeks, 14 regional definitions took 11, and 42 metrics across eight administrative units took 20. Nine reporting systems replaced took 26.

Can we do both at once?

You can, and it is usually the most expensive order. A platform migration carries the existing definitions across unless someone stops it, so the new warehouse arrives with the same disagreement in a more modern format, and the migration is then blamed for not fixing a problem it was never scoped to touch. If both are needed, write the definitions first and migrate onto them. Texas A&M and a national telecom both built a shared semantic layer as part of the consolidation rather than after it.

When is the pipeline actually the problem?

Four conditions, and you need at least one to hold. The same query on the same table returns different answers on different days. Numbers change retroactively without a documented restatement. A material share of rows fail a basic integrity check such as orphaned keys or duplicated primary keys. Or the data arrives too late to be used at all, regardless of whether it is correct. The fourth is the one most often misdiagnosed as a definitions problem, and it is not.

Will fixing definitions improve data quality on its own?

Partly, and in a specific way. A written definition makes a test possible, and a test makes a quality failure visible. Before the definition exists there is nothing to test against, so quality problems surface as arguments rather than as failures. In practice the certified metric layer is where the automated tests get attached, which is why the definition work tends to expose pipeline problems rather than hide them. It does not repair a broken load.

Does a new BI tool help either way?

No, and it is the most common purchase made against this problem. A reporting tool renders a number it is given. Pointed at three definitions it renders three numbers faster. One retail client was running 220 store reports resting on only 40 distinct metrics; the resolution was cutting to 14 certified dashboards on a standardised semantic layer, not a different rendering engine.

How does this change if we are deploying AI on top?

It raises the cost of getting the order wrong. A natural-language assistant inherits whichever definitions it can reach, and answers with more fluency than the dashboards it replaced, so the disagreement is distributed faster and with more apparent authority. EY's 2026 Global DNA of the CFO survey, fielded 16 February to 30 March 2026 across 1,610 finance leaders at organisations above $1B revenue, found 61% citing data quality and bias as the top barrier to AI investment. The definition layer is the part of an AI programme that keeps its value when the model is replaced.

The work behind this

Eleven engagements in the case library carry semantic layer and metric governance, and 31 carry data engineering and lakehouse work. The split between them is the decision this article is about, and every engagement states which route was taken and what it delivered.

Semantic layer and metric governance, 11 engagements.

Topics covered

  • KPI definitions
  • data pipeline rebuild
  • semantic layer versus data warehouse
  • data governance versus data engineering
  • metric layer
  • data platform investment decision
  • standardize KPI definitions across departments

Frequently asked questions

How do I know whether the problem is definitions or pipelines?

Run one test before committing budget. Pick a single disputed number and one period. Have each team write its rule down before looking at the data. Then run both rules against the same source table. If the two figures now agree, the plumbing works and the disagreement was always the rules, so a pipeline rebuild buys you nothing. If they still differ when the rule and the source are identical, something underneath is broken and the engineering case is real. The test costs about half a day and it decides where a six-figure budget goes.

Which problem is it usually?

Definitions, in most cases we are called into. The giveaway is that each team can reproduce its own number reliably, cycle after cycle. Reliable production of two different answers is a rules problem by definition, because a broken pipeline produces unreliable answers rather than consistently different ones. The dbt Labs State of Analytics Engineering 2026 survey, fielded 5 December 2025 to 1 February 2026 across 363 respondents, found ambiguous data ownership still a challenge for 41%, effectively unchanged year over year, while trust in data rose as a strategic priority from 66% to 83%. Priority moved, ownership did not.

What does each route cost in time?

In our case library the definition route ran 8 to 20 weeks and the platform consolidation route ran 16 to 26. The ranges overlap, and the reason matters more than the averages: definition work scales with the number of contested metrics, platform work scales with the number of systems. Five ARR definitions took 8 weeks, 14 regional definitions took 11, and 42 metrics across eight administrative units took 20. Nine reporting systems replaced took 26.

Can we do both at once?

You can, and it is usually the most expensive order. A platform migration carries the existing definitions across unless someone stops it, so the new warehouse arrives with the same disagreement in a more modern format, and the migration is then blamed for not fixing a problem it was never scoped to touch. If both are needed, write the definitions first and migrate onto them. Texas A&M and a national telecom both built a shared semantic layer as part of the consolidation rather than after it.

When is the pipeline actually the problem?

Four conditions, and you need at least one to hold. The same query on the same table returns different answers on different days. Numbers change retroactively without a documented restatement. A material share of rows fail a basic integrity check such as orphaned keys or duplicated primary keys. Or the data arrives too late to be used at all, regardless of whether it is correct. The fourth is the one most often misdiagnosed as a definitions problem, and it is not.

Will fixing definitions improve data quality on its own?

Partly, and in a specific way. A written definition makes a test possible, and a test makes a quality failure visible. Before the definition exists there is nothing to test against, so quality problems surface as arguments rather than as failures. In practice the certified metric layer is where the automated tests get attached, which is why the definition work tends to expose pipeline problems rather than hide them. It does not repair a broken load.

Does a new BI tool help either way?

No, and it is the most common purchase made against this problem. A reporting tool renders a number it is given. Pointed at three definitions it renders three numbers faster. One retail client was running 220 store reports resting on only 40 distinct metrics; the resolution was cutting to 14 certified dashboards on a standardised semantic layer, not a different rendering engine.

How does this change if we are deploying AI on top?

It raises the cost of getting the order wrong. A natural-language assistant inherits whichever definitions it can reach, and answers with more fluency than the dashboards it replaced, so the disagreement is distributed faster and with more apparent authority. EY's 2026 Global DNA of the CFO survey, fielded 16 February to 30 March 2026 across 1,610 finance leaders at organisations above $1B revenue, found 61% citing data quality and bias as the top barrier to AI investment. The definition layer is the part of an AI programme that keeps its value when the model is replaced.

Related reading

If this is the problem you have