FinOps for AI: The AI Cost Discipline [Part I of IV]
In this blog
- Introducing FinOps for AI: The new billing model every technology leader needs to understand
- Executive summary
- What enterprises got right about cloud
- Why AI is a different problem
- The market signals that demand attention
- Two fallacies driving most AI overspend
- Essential terms: The vocabulary of AI cost management
- Additional references
- Download
Introducing FinOps for AI: The new billing model every technology leader needs to understand
This is the first in a four-part series on FinOps for AI, written for organizations actively scaling AI workloads. The series argues that AI is an emerging cost challenge for enterprises, one that echoes the shift to cloud, and that the discipline built for cloud does not translate directly to AI.
This article defines what FinOps for AI is and why it stands apart from FinOps as it applies to cloud computing. The second names five major structural pain points in AI spend. The third lays out a four-phase playbook, Understand, Attribute, Optimize, Operate, for acting on those pain points. The fourth shows how WWT ran that playbook on itself, and it closes with recommendations for where to start.
Follow the series here:
- Part II: FinOps for AI - Five Reasons AI Spend Spirals
- Part III: FinOps for AI - A Four-Phase Playbook for Governing AI Spend
- Part IV: FinOps for AI - Lessons from WWT's Journey and Where to Start
Executive summary
AI is an emerging cost challenge for enterprises, one that echoes the shift to cloud, and the original playbook that made cloud spend governable does not fully transfer to AI spend. We see this at WWT, where in-year GenAI cost growth is measured in multiples rather than percentages. FinOps for AI is a distinct discipline, and technology leaders need to recognize this core distinction before building any more tooling, dashboards, or policies.
The evidence for this recommendation is found in three places. First, enterprises mastered the cloud model: they spent a decade learning to tag resources, attribute spend, identify idle infrastructure, and right-size what remained, and most large organizations built that muscle well. Second, the billing model for consumption-based AI spend behaves differently. Cost accrues by usage rather than by provisioned resource, and it appears in standard billing tools as a single line that neither finance nor engineering can explain. Third, most of the overspend we see traces to two fallacies: the belief that high token consumption is a proxy for ROI, and the belief that cost per token is the number that matters. Consumption tracks adoption, not return. Both fallacies mistake activity for outcomes.
We write from more than client experience or theory. We write from firsthand experience, because WWT ran this discipline on our own AI spend and identified 45% attainable cost savings before recommending it to anyone. The immediate actions are clear: organizations should adopt cost per completed task as the north-star metric and require every AI initiative to report against it. That definition is not one-size-fits-all. A completed task means something different in Workforce AI, AI-Native Engineering and Applied AI, and getting it wrong makes costs across categories look comparable when they aren't.
What enterprises got right about cloud
A decade of discipline made the cloud bill a governed document
A FinOps lead is hosting a monthly spend review. Every workload on the report is tagged to an owning team, the idle capacity flagged last quarter is gone, and the one anomaly on the bill already has an explanation attached. The business value of the largest cost items is clear from the project charters' business cases; some are critical infrastructure. The meeting ends early because the numbers are governed by a structured Cloud FinOps program.
That routine took a decade to earn, and the practices behind it now run as standing discipline in most large enterprises. The habit underneath them was right-sizing, which means buying the right amount or type of infrastructure. WWT built our own FinOps capability on that foundation, and we watched the same capability mature across the enterprises we work with.
The maturity here reflects the confidence executives now bring to AI. A leadership team that tamed cloud spend expects its tools and instincts to transfer, but that expectation is where the trouble starts.
Why AI is a different problem
The billing unit changed under the discipline
A technology finance lead circles a single line on a printed vendor invoice ahead of the quarterly review. The line has grown every month since the AI pilots went live, and no one in the room can say which teams sit within it or what share of it has produced any business results. The conversation about seat counts from last quarter has been made moot by the vendor's move to consumption-based billing. The tools that explain every other line on that invoice do not apply to this AI spend line.
This is a billing-model problem. Cloud spend is resource-based, paid per server and per hour, which is why tagging and right-sizing could govern it. AI spend is often consumption-based, and it accrues with every token a model reads and writes, every inference call, and every loop an autonomous agent decides to run. Standard cloud billing tools were built for the resource world, so they aggregate all that consumption into one line item. Finance receives a total with no drivers behind it. Engineering holds the usage detail but cannot translate that detail into dollar value.
The discipline's role changed with the unit being managed and the associated usage behaviors. "Right-sizing" governs what an enterprise can define and forecast, but it cannot govern consumption that resists both. The discipline that can is "right-using": directing each unit of consumption at a defined business outcome and pricing the work by the task it completes. It also directly challenges us to use GenAI much more effectively and efficiently. The rest of this article, and the series it opens, builds out what "right-using" requires.
Daniel Cholakov, WWT Sr. Principal
The market signals that demand attention
The pressure to address AI spend will not ease as these programs mature. From its 2026 CIO and Technology Executive Survey, Gartner® found that "84% of surveyed CIOs reported their enterprise would increase funding for Generative AI in 2026, while at the same time, 52% indicated reducing costs would become an important outcome of digital initiatives in 2026 and 2027" [2]. Those two mandates collide in the same budget cycle. The collision lands on whoever owns the AI line, because that team must both find the funds for the increased spend and explain the value that the investment is supposed to deliver.
Another survey suggests that 79% of finance leaders experienced AI cost overruns in 2025. More visible examples include Uber, which burned through its entire 2026 AI budget in four months as Claude Code adoption jumped to 84% of its engineering staff, and Amazon, which saw a single Claude Sonnet project run 860% over budget. This leads to rapid reactions, such as capping employee usage, which risks limiting innovation and productivity gains. WWT experienced an expected quintupling of GenAI subscription costs in Q2 2026 and an unexpected doubling eight weeks later. We needed to act immediately.
The spend itself is moving toward its most expensive phase. Gartner projects that "by 2030, over 80% of AI-optimized IaaS spending will support inference workloads providing agentic AI capabilities" [1]. While pilots validate value, the real expense begins when AI operates continuously at enterprise scale.
Unmanaged consumption doesn't just show up on the invoice. Left unchecked, it puts both budget and reputation at risk. And if AI spend isn't tied to a defined business outcome, it can't be justified.
Two fallacies driving most AI overspend
Overspend runs on two beliefs, both of which value activity over outcomes. Under that pressure, most of the AI overspend we see traces to two beliefs, and both sound reasonable at first.
Fallacy 1: High token consumption equals ROI
Rising consumption reads as evidence that the organization is adopting AI, but consumption proves only that money is being spent. When tokens have no defined business outcome, the volume signals inefficiency in forms such as runaway agents, bloated context windows and unoptimized prompts. 20+ user interviews surfaced genuine ideas and innovations worth keeping, but the same users often made tactical decisions that an experienced data scientist would have executed two to 10 times more efficiently. The ideas were sound, but the execution lagged, and that gap is where the spend goes.
Did you know: Token-maxxing is maximizing token consumption without a defined business outcome behind it. On the invoice, the pattern is indistinguishable from successful adoption until an outcome metric sits beside it to measure the value of the spend.
Fallacy 2: Cost per token is the metric that matters
Cost per token prices a unit of computation and says nothing about whether the task was done or done well. Cost per completed task is much closer to the value, because it measures the same task the same way, whether a cheap model needed three attempts or an expensive model needed one.
Both fallacies are symptoms of deeper structural gaps, and the next article in this series identifies the five we most often encounter.
One distinction changes everything: "Right-using," measured by the same completed tasks at lower costs, is the discipline.
Cloud FinOps asks whether the enterprise owns the right amount of infrastructure. FinOps for AI asks whether the work is being done efficiently, with the right model and for a defined business outcome.
Unit economics is what makes the answer measurable. Cost per completed task connects every AI dollar to the outcome it bought. Gartner calls unit economics "a discipline that connects costs with business outcomes". In our opinion, precisely the connection cost per completed task is built to make.
Cost per task can be hard to define. AI-Native Engineering provides cost per pull request or cost per merge as an established practice. Even then, a low cost per outcome may correlate with the erosion of value (e.g., if more merges are performed to correct errors from prior AI runs). That's where we attempt to understand the intent of a merge. AI inside pro-code apps may be easier to define outcomes for, for instance, a successful RFP response from WWT's RFP Assistant. Workforce AI, the largest component of GenAI for some enterprises, may, however, be the hardest to estimate its value for. How many of your Cowork sessions last week brought business outcomes?
We ran this discipline on our own AI spend before recommending it, and that build and its lessons will close this series. The directive is clear: adopt cost per completed task as the north-star metric and require every AI initiative to report against it before the next budget cycle.
Essential terms: The vocabulary of AI cost management
- Chargeback: Billing each team for its share of AI spend so the cost lands in that team's budget.
- Goodput: The share of AI output that actually advances the intended business task, as distinct from raw throughput.
- Model routing: Directing each task to the least expensive model that meets its quality requirement.
- Right-using: Directing AI consumption toward defined business outcomes and measuring efficiency by cost per completed task.
- Showback: Reporting each team's share of AI spend to that team without billing for it.
- Token: The unit of text an AI model reads and writes, and the unit AI usage is priced in.
- Token-maxxing: Maximizing token consumption without a defined business outcome behind it.
- Tokenomics: The economics of AI workloads that are priced and consumed at the token level.
- Unit economics: The discipline of connecting cost to a unit of business outcome, such as cost per completed task.
Additional references
- [1] Gartner. "AI Inference's Financial Reckoning: How Infrastructure and IT Operations Must Master Consumption-Based Economics," Dennis Smith and Jeff Vogel, 9 April 2026. GARTNER is a trademark of Gartner, Inc. and/or its affiliates.
- [2] Gartner. "Don't Let AI Agents Burn Your Budget," Yogesh Bhatt, Deven Tasgaonkar, Ben Yan and Mike Fang, 1 March 2026.
- [3] Gartner. "Emerging FinOps Trends Pave the Road to AI Value," Marco Meinardi, 23 April 2026.