Thinklytics

Platform Cost · 10 min read · October 2026

Cut the bill or fix attribution first? The order that makes a saving hold

By Sean Majidi, Founder, Thinklytics

Both reduce spend. They differ in whether the reduction holds after the engagement ends. A saving nobody owns is the one that comes back, which is why attribution decides who the engineering targets go to rather than competing with them.

Two proposals. One tags the estate, builds a showback and reports cost per unit of outcome. The other rightsizes instances, tiers storage and negotiates commitments, and shows a lower invoice next month.

The second one looks better in a quarter. The first one is still working in two years.

What each route actually produces

Two routes, and why the order is not a preference

Both reduce spend. They differ in whether the reduction holds after the engagement ends.

Fix attribution firstCut unit cost first
What it producesA bill with owners, and a cost per outcomeA lower invoice next month
Who can act afterwardsEvery named owner, continuouslyWhoever ran the exercise, once
Time to first savingWeeks, often from owners switching off what they were not usingDays. Rightsizing and scheduling are fast
Why it decaysIt does not, if the owner and threshold surviveConsumption regrows, because nobody who caused it saw it
What it cannot do aloneFix infrastructure that is simply oversizedStop the next unowned thing being provisioned
The honest answerDo both. Attribution first, because it decides who the engineering targets go toRun the fast levers in parallel where they need no owner's consent

Only 49% of organisations track unit economics, up 9 points year over year. Cutting a bill you cannot attribute produces a saving nobody owns, and an unowned saving is the one that comes back.

Source: FinOps Foundation, State of FinOps 2026: 1,192 practitioners representing more than $83B in annual cloud spend; Thinklytics engagement pattern across the cost optimisation engagements in the case library.

They are usually presented as alternatives and they are not. Attribution decides who the engineering targets go to.

The attribution route produces a bill with owners and a cost per outcome. Afterwards every named owner can act, continuously, without anyone running an exercise.

The unit-cost route produces a lower invoice. Afterwards whoever ran the exercise can do it again, and nobody else can do anything.

That difference is why the order matters. A reduction target handed to someone who cannot see their own consumption is a request they cannot act on, so it produces either no action or arbitrary action.

Why the externally-driven cut decays

Because the mechanism that created the spend is untouched.

Someone outside the team resizes instances, tiers cold storage and buys commitments. The invoice falls, the engagement reports a saving, and provisioning behaviour continues exactly as before. Over the next few quarters consumption regrows, usually in different services so the comparison is awkward.

The saving was real. The cause was not addressed, and a saving nobody owns is the one that comes back.

Attribution changes behaviour rather than configuration. A team that receives its own number monthly, in a meeting that already happens, makes different provisioning decisions without being asked to. That is the durable half, and the FinOps Foundation's State of FinOps 2026, covering 1,192 practitioners representing more than $83 billion in annual cloud spend, found only 49% of organisations tracking unit economics at all, up 9 points year over year. Half are still managing a total rather than a rate.

What to run in parallel

Not everything has to wait for attribution, and saying otherwise would be wrong.

The fast levers need no owner's consent: rightsizing, auto-suspend discipline, storage tiering, commitment coverage, query result caching. They are uncontroversial, they are measurable, and they should start immediately. Our platform-specific version of that work is in Snowflake cost optimization without AI, and the broader lever set is in the 2026 FinOps playbook.

What should not run first is a percentage target handed to teams who cannot see their own line.

The four decisions that come before tagging

Tagging is the implementation. These are the decisions, and teams that tag first usually re-tag.

The dimension. Team, product, customer, or use case. Pick the one that matches how decisions are made in your organisation. If budget is held by product and you attribute by team, the reports reach people with no authority to act on them.

The shared-cost method. Split evenly, split by consumption, or leave it in a central pool reported separately. All three are defensible. An unstated method is what makes every recipient dispute the figure.

The unit of outcome. An order, a claim, a report, an active customer, a processed document. One per product area, with the calculation written down so someone else can recompute it next quarter.

Showback or chargeback, and when. Showback reports the cost. Chargeback moves the budget. Starting with chargeback before the numbers are trusted turns every month into an argument about allocation rather than a conversation about spend. Two or three cycles of showback first, until recipients stop disputing the figures.

The AI case needs its own handling

Why a pilot's AI unit cost will not survive production

Five assumptions that are all true in a pilot and all false at volume. State each one next to any AI cost figure.

  • Volume. A pilot runs on a fraction of real traffic. State the per-unit cost, then state what it becomes at ten and a hundred times the volume.
  • Context length. Pilots use short, clean inputs. Real documents, real threads and real histories are longer, and context is usually the largest cost term.
  • Retry and failure behaviour. A pilot's failures get handled by the person running it. In production they get retried, often more than once, and each retry is billed.
  • Who is in the loop. A pilot has an expert reviewing output. If production needs the same review, the labour is part of the unit cost and usually larger than the inference.
  • What happens on a bad month. The figure that matters for a budget is the ceiling, not the average. Ask for the worst observed week extrapolated, not the mean.
  • Model price falling. Often true and not a plan. A cost case resting on future price cuts is a forecast of someone else's pricing decisions.

98% of FinOps teams now manage some form of AI spend, up from 31% two years ago. Most of that arrived faster than anyone's attribution model, which is why the per-use-case key matters before the volume does.

Source: FinOps Foundation, State of FinOps 2026: 1,192 practitioners, more than $83B in represented annual cloud spend; Thinklytics engagement pattern across the AI and cost optimisation engagements.

Five assumptions are true in a pilot and false at volume, and they compound.

Volume itself. Context length, since real documents, threads and histories are longer than pilot inputs and context is usually the largest cost term. Retry behaviour, because a pilot's failures are handled by the person running it while production failures get retried and each retry is billed. Who is in the loop, since an expert reviewing output during a pilot becomes either a labour line or an accuracy risk afterwards. And what a bad month looks like, because a budget needs the ceiling rather than the mean.

State each assumption next to any AI unit cost, then state what the figure becomes at ten and a hundred times the volume. A cost case resting on model prices falling is a forecast of someone else's pricing decisions rather than a plan.

The attribution problem is sharper here too. 98% of FinOps teams now manage some form of AI spend against 31% two years ago, and most of that arrived under one organisation-wide key. A key or project per use case is cheap now and impossible to backfill, because you cannot retroactively separate what was never separated.

The order

The order that makes the saving hold

The fast engineering levers run in parallel. What cannot run in parallel is handing a reduction target to someone who cannot see their own consumption.

  • Measure the attributable share as it really is
  • Decide the dimension and the shared-cost method
  • Pick the unit of outcome, not of consumption
  • Show the number to a named owner, monthly
  • Then set engineering targets per owner

Step four usually produces a reduction on its own, before any engineering, because a team that can see what it is running switches off what it was not using. Measure that separately so it is not credited to the optimisation work.

Source: Thinklytics engagement pattern across the cost optimisation engagements in the case library.

Measure the attributable share as it really is. Decide the dimension and the shared-cost method. Pick the unit of outcome. Show the number to a named owner, monthly, in a forum that already exists. Then set engineering targets per owner.

Step four is where the first reduction usually appears, before any engineering, because a team that can see what it is running switches off what it was not using. Measure that effect separately so it is not credited to the optimisation work, since the two decay differently: an engineering cut can regrow and a behavioural change does not, provided the owner and the threshold survive.

When ownership is the whole problem

One case worth holding, because the technical answer followed from the ownership answer rather than the reverse.

A cloud provider had 14 business units each running their own data warehouse, built independently over eight years, costing $4.7M a year with no shared metrics and no cross-unit reporting. The CTO had ordered a consolidation and earlier attempts had stalled on business unit resistance.

The engagement that worked did not centralise. Each unit kept control of its own data and published through a shared catalogue with standard interfaces, migrated in groups of three over 30 weeks. All 14 integrated by week 30, cost down to $1.1M a year, cross-unit reporting for the first time, and zero business unit escalations. See the federated data mesh engagement.

Attribution and ownership were the intervention. The architecture was chosen to make them possible.

What we would do first

Pick one product area and compute its cost per unit of outcome by hand, in a spreadsheet, for last month. Attributed cost over units produced, with the method written next to it.

It will be approximate and it will still be the most useful number in the next budget conversation, because it is the first time anyone has expressed the spend as a rate rather than a total.

Then show it to the person who owns that product area and watch what they ask. That question tells you whether you have an attribution problem or an engineering one.

What belongs in the engagement, and the acceptance criteria to contract for, is in how to scope a cloud and AI cost engagement. The cost attribution and showback plan is the worksheet we fill in before anyone is asked to cut anything.

Delivery sits in cloud and AI cost optimization for both halves, system consolidation where the estate itself is the cost, and data foundation for the platform layer underneath. The full set of work in this area sits under the platform bill keeps climbing.

Frequently asked questions

Should we cut the bill or fix attribution first?

Attribution first, with the fast engineering levers running in parallel. They are not competing options: attribution decides who the engineering targets go to. A reduction target handed to someone who cannot see their own consumption is a request they cannot act on, and a saving produced by an external exercise regrows because nobody who caused the consumption ever saw it. Rightsizing, scheduling, storage tiering and commitment coverage need no owner's consent, so run those immediately and in parallel.

Why does an externally-driven cut decay?

Because the mechanism that created the spend is untouched. Someone outside the team resizes instances and tiers storage, the invoice falls, and provisioning behaviour continues unchanged, so consumption regrows over the following quarters. The reduction was real and the cause was not addressed. Attribution is what changes the behaviour, because a team that receives its own number monthly makes different decisions without being asked.

What should we attribute by?

The dimension that matches how decisions are actually made: team, product, customer, or use case. If budget is held by product and you attribute by team, the reports go to people who cannot act on them. Pick the decision-making dimension rather than the one that is easiest to tag, because re-tagging an estate a year later costs more than settling the model first. Tagging is the implementation of this decision, not a substitute for it.

Showback or chargeback?

Showback first, almost always. Showback reports attributed cost back to the owner without moving budget; chargeback moves the budget too. Starting with chargeback before the numbers are trusted turns every month into a dispute about the allocation method rather than a conversation about the spend. Move to chargeback once recipients have stopped arguing with the figures, which is usually two or three cycles.

How do we handle shared cost?

Pick a method and write it down: split evenly, split by consumption, or leave it in a central pool reported separately. All three are defensible and the unstated method is what makes recipients dispute the number instead of acting on it. Then give the unattributable remainder a named home, a target share for shrinking it, and a date, or it becomes permanent and quietly absorbs anything inconvenient.

Why do AI costs from a pilot mislead?

Five assumptions are all true in a pilot and all false at volume. Volume itself. Context length, since real documents and threads are longer than pilot inputs and context is usually the largest cost term. Retry behaviour, because production failures get retried and each retry is billed. Who is in the loop, since an expert reviewing output in a pilot becomes either a labour cost or an accuracy risk in production. And what a bad month looks like, because a budget needs the ceiling rather than the average.

What reduction comes from attribution alone?

Often a material one, before any engineering. A team that can see what it is running switches off what it was not using, and that effect arrives within a cycle or two of the first showback. Measure it separately from the engineering effect so the two are not conflated, because they decay differently: the engineering cut can regrow and the behavioural change does not, provided the owner and the threshold survive.

What if we do the engineering work anyway?

Do it, in parallel, for the levers that need nobody's consent. Rightsizing, auto-suspend discipline, storage tiering, commitment coverage and query result caching are fast and uncontroversial. What should not run first is handing a percentage target to a team that cannot see its own line, because that produces either no action or arbitrary action, and both get attributed to the cost programme.

The work behind this

Nine engagements in the case library carry cost optimization and seven published an annual before and after figure. The largest changed the ownership model rather than centralising it, and recorded zero business unit escalations across 30 weeks.

Cost optimization, 9 engagements.

Topics covered

  • cloud cost attribution
  • showback versus chargeback
  • FinOps order of work
  • cloud cost reduction
  • AI unit cost
  • tagging strategy
  • unit economics

Frequently asked questions

Should we cut the bill or fix attribution first?

Attribution first, with the fast engineering levers running in parallel. They are not competing options: attribution decides who the engineering targets go to. A reduction target handed to someone who cannot see their own consumption is a request they cannot act on, and a saving produced by an external exercise regrows because nobody who caused the consumption ever saw it. Rightsizing, scheduling, storage tiering and commitment coverage need no owner's consent, so run those immediately and in parallel.

Why does an externally-driven cut decay?

Because the mechanism that created the spend is untouched. Someone outside the team resizes instances and tiers storage, the invoice falls, and provisioning behaviour continues unchanged, so consumption regrows over the following quarters. The reduction was real and the cause was not addressed. Attribution is what changes the behaviour, because a team that receives its own number monthly makes different decisions without being asked.

What should we attribute by?

The dimension that matches how decisions are actually made: team, product, customer, or use case. If budget is held by product and you attribute by team, the reports go to people who cannot act on them. Pick the decision-making dimension rather than the one that is easiest to tag, because re-tagging an estate a year later costs more than settling the model first. Tagging is the implementation of this decision, not a substitute for it.

Showback or chargeback?

Showback first, almost always. Showback reports attributed cost back to the owner without moving budget; chargeback moves the budget too. Starting with chargeback before the numbers are trusted turns every month into a dispute about the allocation method rather than a conversation about the spend. Move to chargeback once recipients have stopped arguing with the figures, which is usually two or three cycles.

How do we handle shared cost?

Pick a method and write it down: split evenly, split by consumption, or leave it in a central pool reported separately. All three are defensible and the unstated method is what makes recipients dispute the number instead of acting on it. Then give the unattributable remainder a named home, a target share for shrinking it, and a date, or it becomes permanent and quietly absorbs anything inconvenient.

Why do AI costs from a pilot mislead?

Five assumptions are all true in a pilot and all false at volume. Volume itself. Context length, since real documents and threads are longer than pilot inputs and context is usually the largest cost term. Retry behaviour, because production failures get retried and each retry is billed. Who is in the loop, since an expert reviewing output in a pilot becomes either a labour cost or an accuracy risk in production. And what a bad month looks like, because a budget needs the ceiling rather than the average.

What reduction comes from attribution alone?

Often a material one, before any engineering. A team that can see what it is running switches off what it was not using, and that effect arrives within a cycle or two of the first showback. Measure it separately from the engineering effect so the two are not conflated, because they decay differently: the engineering cut can regrow and the behavioural change does not, provided the owner and the threshold survive.

What if we do the engineering work anyway?

Do it, in parallel, for the levers that need nobody's consent. Rightsizing, auto-suspend discipline, storage tiering, commitment coverage and query result caching are fast and uncontroversial. What should not run first is handing a percentage target to a team that cannot see its own line, because that produces either no action or arbitrary action, and both get attributed to the cost programme.

Related reading

If this is the problem you have