Executive summary

Many enterprises assume that having a dashboard means they already have an AI cost management program, but that assumption doesn't hold up. A dashboard that shows spend without attribution, optimization or automated controls is just a reporting tool, not a governance system. WWT built a four-phase playbook and applied it to our own AI spend before we asked anyone else to trust it. The four phases are Understand, Attribute, Optimize and Operate, and the order matters because each phase produces what the next one needs.

The whole program moves through a sequence. Understanding and visibility of the numbers enable attribution, and attribution enables optimization decisions and targeted controls. Organizations that jump straight to optimization without completing attribution often end up optimizing the wrong things, and organizations that stop at attribution end up with accountability reports that no one enforces. We learned this by applying the framework to our own AI spend first, so every phase below includes our own lessons alongside general best practices.

We also want to flag where most maturity journeys stall. The most common failure point comes early, when dashboards and partial attribution convince an organization that the discipline is finished, even though governance remains manual and spend is still disconnected from business value.

Technology executives should honestly locate their organization on the maturity map and start with the phase that best matches their maturity stage. One metric, cost per completed task, guides every step along the way, and the real goal is steady movement through the sequence.

Follow the series here: 

Why a dashboard is not a governance system

The CIO opens the operating review with the new AI spend dashboard, and for the first time, the number has shape. Spend is broken out by provider, with a red flag on the workflow that spiked. A couple of anecdotes about individual contributors managing their own AI costs well make the room relax. Weeks later, the overrun happens anyway. Nobody set a spend limit on the flagged workflow, and nobody owned it.

The lesson? A dashboard is not a governance system. It can highlight the spiral but not stop it. Visibility is a reporting capability, and reporting alone is not the full FinOps for AI program you need. A dashboard states what was spent. A governance system does more. It assigns an owner to each cost, tests spend against return, sets limits and acts when a limit is breached.

The market pressure our first article documented does not pause while reporting improves, and a clearer view of the five gaps the second article named closes none of them. Closing them requires a sequenced, structured playbook appropriate to your maturity level.

The sequence that makes this work

The four phases run in order because each one manufactures what its successor consumes. Understanding produces the visibility data for attribution tags. Attribution adds the ownership needed to make optimization decisions defensible. Optimization then hands automated governance the measured results it needs to manage spend and value effectively.

The goal is movement through the sequence rather than perfection at each stage, and an organization can move fast through a phase whose inputs are ready.

Phase 1, Understand: Spend in motion becomes visible before anything else is attempted

Phase 1 consolidates AI spend into a single view, pulling cloud, SaaS and API costs into one picture instead of scattered invoices. This phase also baselines current maturity against a Foundational, Developing, Mature model and surfaces the priorities later phases will act on, starting with the workload archetypes that group costs into categories the business recognizes.

WWT has arrived at three main types of AI use cases: 

  1. Workforce AI spans the chatbots, personal assistants and Cowork solutions most knowledge workers use today.
  2. AI-Native Engineering (AINE) covers coding assistants and agentic software development platforms.
  3. Applied AI runs inside pro-code applications and is usually charged through API subscriptions or owned hardware.

Defining archetypes early matters because every later phase prices, owns and optimizes work at the archetype level rather than line-by-line on the invoice.

We built this the hard way in our own environment. In our opinion, Gartner guidance points in the same direction. Gartner recommends "policy mandating all created agents must be instrumented with telemetry prior to deployment" and "ensuring dashboard can capture granular spending patterns rather than aggregated cloud bills that mask AI agent behavior" [1].

Phase 1 ends when AI spend is consolidated, and a priority list is surfaced.

Phase 2, Attribute: Every visible dollar gets an owner at its source

Attribution turns the visible number into an owned number. The mechanics are allocation conventions that map spend to workflow, team and environment, applied at the point of creation rather than reconstructed at month-end. Those conventions pair with a dashboard that shows live cost attribution, so each owner sees their number as it accrues rather than after the fact. The convention only works when teams actually apply it, which makes Phase 2 as much an agreement as a technical process.

Showback starts the phase, and chargeback finishes it. Reporting a team's AI spend to a high-spending team gets attention, and billing it to their budget changes behavior, which is why we treat the chargeback trigger as the forcing function of the entire playbook.

 - Chuck Balog, WWT Cloud Strategy and FinOps Practice Director

Chargebacks carry a real cost in friction because teams challenge their first bills, and those challenges take time that finance did not anticipate. The challenges are the point. A number worth contesting is a number someone finally owns, and Phase 3 depends on that ownership.

Attribution itself becomes an optimization lever once it reaches individual contributors, who usually want to use enterprise AI resources well but don't always have the knowledge or tools to act on that instinct.

Phase 3, Optimize: With owners in place, optimization cuts waste against cost-per-completed-task

With costs owned, optimization has a foundation. Some of the available levers include model routing, token efficiency and workload rationalization. The largest of these is routing, directing each task to the least expensive model that meets its quality requirements rather than sending everything to the most capable model by default.

We have identified 36 levers across architecture, context, model and human behavior, inspired by the FinOps Foundation's proven cloud practices [2]. We enriched them with what we saw working for our clients and partners, then adjusted them again after implementing them internally. 

A key lesson was the need to differentiate which treatments to apply to each use case archetype. Task routing to the right models can significantly change Workforce AI costs; spend quotas affect business analysts differently than AINE developers; and caching is especially effective for Applied AI that reruns similar tasks.

The cost-per-completed-task metric keeps the work honest because an optimization counts only when the metric moves. Goodput optimization applies the same test to output, raising the share of AI work that advances the intended business task rather than raw throughput. A team can reduce token costs, switch models and rationalize workloads for an entire quarter, and if the cost of a completed task has not fallen, the optimization was cosmetic. At worst, the team spends time and effort on an optimization that could have been directed toward something more useful.

Phase 4, Operate: Optimization hardens into governance that runs at the speed of AI

Phase 4 operationalizes what the first three phases learned. Budget limits, circuit breakers and real-time anomaly detection replace manual review, because manual review runs on a meeting cadence and cannot keep pace with spend that accrues by the second. Our six prioritized levers, alongside 30 others, let us scale optimization effort up or down against spend trend versus budget. The approach is easy for our teams to understand and doesn't create undue concern around cost controls. The thresholds come from Phase 3's measured results, which is why this phase cannot run first.

Gartner also recommends that "Teams should configure anomaly detection rules to identify suspicious patterns, such as unexpected increases in task spawning, repeated retries, unusual memory usage or surges in token consumption" [1]. Automated governance is what lets an enterprise scale AI agent use without scaling the risk described earlier in this article.

The maturity trap most organizations fall into


Early-stage dashboards feel like good maturity until controls fail to contain costs

The journey has three maturity stages. Foundational means limited or no AI-specific spend visibility. Developing means dashboards and some attribution exist and optimization is underway. Mature means full governance, with unit economics tracked, good AI practices understood across employees, automated controls live and spend connected to business value.

The maturity trap is the Developing-stage belief that dashboards and partial attribution mean the discipline is sufficient. An organization in the trap has reporting, some tagging and a policy document, but it has automated nothing and its incentives are ineffective. Its governance still depends on someone noticing. Visibility without control is still Phase 1, just with better reporting.

Escaping the trap is as much a cultural as a technical maneuver, because each stage asks teams to accept a new definition of done, and moving finance and engineering onto shared metrics and enforced conventions is outcome-focused change work rather than tooling work. CIOs should place their organization on the model this month and fund the phase that placement calls for, accepting that the placement will probably read lower than the last status report implied. The trap is comfortable, and the organizations that escape it are the ones that treat the dashboard as the beginning of governance.

Glossary

  • Anomaly detection: Automated identification of spend or usage patterns that break from a workload's established baseline.
  • Automated governance: Cost and behavior controls that enforce themselves in real time rather than through manual review.
  • Goodput optimization: Raising the share of AI output that advances the intended business task rather than raw throughput.
  • Model routing: Directing each task to the least expensive model that meets its quality requirement.
  • Showback: Attributing AI costs to consuming teams and reporting those costs for visibility and accountability before implementing chargeback.
  • Workload archetype: A named category of AI workload with a shared cost and usage profile, defined during Phase 1.

Additional references

[1] Gartner, "FinOps Is Critical to Maximizing ROI of AI Agents," Deacon D.K Wan, Tigran Egiazarov, Aaron Harrison, Bill Blosen, 9 February 2026.

[2] FinOps Foundation, "Optimizing GenAI Usage: A FinOps Perspective on Cost, Performance, and Efficiency", 2026.