Thinklytics

FinOps · 9 min read · May 2026

The 2026 FinOps Playbook: Controlling Cloud and AI Spend

By Thinklytics Partners, Cloud & AI Cost Optimization Practice

Public cloud spend passed a trillion dollars and AI is piling on. Here is where the money actually hides, the four levers that recover it, and why the first two usually pay for the whole engagement.

The bill arrived. After three years of building first and counting later, cloud and AI spend is now a board-level line, and the question is no longer "how much AI can we ship" but "how do we ship it without the bill running away." That is FinOps, and in 2026 it is one of the most controllable levers a data team owns.

The money is not where everyone is looking

Most teams watch the compute line because it is the biggest number on the invoice. The recoverable money is somewhere else: capacity that sits idle, tools that duplicate each other, and AI calls that nobody is metering.

The pattern repeats across environments. The warehouse is provisioned for a peak that happens twice a quarter. Three BI tools cover the same reports because nobody retired the old one. And the new AI features bill per call with no budget or cap, so the cost grows with adoption instead of with value.

The four levers, in order

The order matters. Right-sizing the warehouse and rationalizing the BI tools are the fastest wins and usually cover the cost of the work before the AI line is even touched. That makes the engagement self-funding, which is the same logic behind data stack consolidation: the cleanup pays for itself in reduced software spend.

Metering AI and LLM spend comes next. Token budgets, the right model for each task, and caching turn an unpredictable usage-based bill into one that tracks value. Then the operating model, ownership and a monthly review, is what keeps the savings from quietly drifting back over the following two quarters.

Why this connects to data readiness

Untuned pipelines and re-runs are a data-engineering problem wearing a cost-report disguise. A pipeline that reprocesses the same data three times is both a reliability risk and a line on the cloud bill. Fixing it is where cloud and AI cost optimization overlaps with the data foundation work, and why teams that keep the foundation healthy through managed data readiness tend to have lower bills without trying.

The move this quarter

Run a cost audit before the next budget cycle. Inventory warehouse capacity against actual query patterns, list every BI tool and its real usage, and put a meter on AI spend. If you cannot answer "what would we save by right-sizing" in a week, that is the gap, and it is almost always larger than the AI line everyone is worried about. The 30-day Analytics Truth Audit is where most teams start.

Frequently asked questions

What is FinOps?

FinOps is the operating model for managing cloud and AI spend as a shared, accountable discipline: visibility into what is being spent, ownership of each cost, and a regular review that turns one-time cleanups into durable savings. It is the difference between a cost-cutting project and a cost that stays under control.

Where does most recoverable cloud spend hide?

Rarely in the compute line teams watch. It hides in idle or over-provisioned warehouse capacity, redundant BI tools and licenses, unmetered AI and LLM API calls, and untuned pipelines that re-run more than they need to. The first two usually cover the cost of the engagement before the AI line is touched.

Does cost optimization mean using less AI?

No. It means AI cost scales with value rather than with usage. Token budgets, model selection, and caching let you keep shipping AI while the bill stays predictable. The goal is control, not a freeze.

How fast does a FinOps engagement pay back?

The warehouse right-sizing and BI tool rationalization are typically the fastest, and in many environments they recover more than the cost of the engagement within the first quarter. The AI metering and the operating model are what keep the savings from drifting back.

How much can a first pass recover?

A first-pass optimization sprint typically recovers 30 to 45 percent of cloud, warehouse, and AI compute spend. The largest savings come from idle or oversized capacity, inefficient queries, duplicate pipelines, and unused licenses, with environments that have never been optimized seeing the biggest first cut.

What keeps the savings from coming back?

The operating model: cost allocated to the team or workload that caused it, budgets and alerts in place, and a regular review so new waste is caught while it is small. The audit finds the savings; the operating model keeps them.

Frequently asked questions

What is FinOps?

FinOps is the operating model for managing cloud and AI spend as a shared, accountable discipline: visibility into what is being spent, ownership of each cost, and a regular review that turns one-time cleanups into durable savings. It is the difference between a cost-cutting project and a cost that stays under control.

Where does most recoverable cloud spend hide?

Rarely in the compute line teams watch. It hides in idle or over-provisioned warehouse capacity, redundant BI tools and licenses, unmetered AI and LLM API calls, and untuned pipelines that re-run more than they need to. The first two usually cover the cost of the engagement before the AI line is touched.

Does cost optimization mean using less AI?

No. It means AI cost scales with value rather than with usage. Token budgets, model selection, and caching let you keep shipping AI while the bill stays predictable. The goal is control, not a freeze.

How fast does a FinOps engagement pay back?

The warehouse right-sizing and BI tool rationalization are typically the fastest, and in many environments they recover more than the cost of the engagement within the first quarter. The AI metering and the operating model are what keep the savings from drifting back.

How much can a first pass recover?

A first-pass optimization sprint typically recovers 30 to 45 percent of cloud, warehouse, and AI compute spend. The largest savings come from idle or oversized capacity, inefficient queries, duplicate pipelines, and unused licenses, with environments that have never been optimized seeing the biggest first cut.

What keeps the savings from coming back?

The operating model: cost allocated to the team or workload that caused it, budgets and alerts in place, and a regular review so new waste is caught while it is small. The audit finds the savings; the operating model keeps them.

Related reading

Thinklytics

Data and AI consulting for Fortune 500s, health systems, and growth-stage companies. Clean data, governed metrics, analytics ready for AI.

Austin, TX · United States

[email protected]