How Intelligent Model Routing Drastically Improves Tokenomics
In this article
- The AI explosion: Third-party analysis
- Two levers that do not solve the problem, and one that does
- The enterprise advantage of open-weight models on AMD
- How the off-ramp works
- What the numbers look like
- Beyond cost savings: Protecting IP and controlling spend visibility
- Is this the right fit for your organization?
- Where WWT fits in
- Begin with an assessment
- Download
Enterprise organizations everywhere are watching their bills grow faster than they can forecast. The biggest line item is tokenomics.
Premium tokens generated from inference of the largest and most expensive models are a significant, if not the primary, cost. Software-as-a-service (SaaS) inference providers have a revenue-maximization incentive, with pricing structures that accelerate at the top end of the inference-quality curve. The better the model, the more expensive the tokens it produces, and, due to reasoning traces, the more tokens it produces. The result is an accelerating cost curve.
Trying to control that spend with usage caps or access limits is rife with problems.
Capping tokens per developer means their work stalls until next week's quota.
Restricting access sidelines most of the organization and the innovation it would drive. Add to this the fact that the controls provided by inference providers and desktop harnesses are often too coarse.
A better option already exists: Using a routing layer that sends requests to the appropriate model based on a holistic assessment of the organization's factors on a per-request basis.
Total cost of ownership (TCO) analyses indicate that this approach yields significant reductions in annual AI inference costs without sacrificing the capabilities that make AI valuable.
Setups as small as a single AMD GPU server, combined with the right configuration, model selection and harnesses, can support up to 40 developers on an MI350x/355x.
The AI explosion: Third-party analysis
Recent market research from Menlo Ventures' 2025 State of Generative AI in the Enterprise confirms that of the $19 billion enterprises spent on generative AI applications in 2025, coding tools alone accounted for $4 billion. AI-assisted software development and the agents' citizen developers they build drive up costs.
Menlo Ventures separately reports that enterprise generative AI spend more than tripled between 2024 and 2025. Worldwide AI spending is projected to grow roughly 47% year over year in 2026, according to Gartner. This isn't a niche budget line anymore. It is becoming one of the fastest-growing costs on the P&L for any organization with in-house engineering teams.
Two levers that do not solve the problem, and one that does
Most leadership teams reach for one of two levers when cost control becomes an imperative:
- Usage caps: Limit productivity because the caps are frequently arbitrarily set and easy to exceed.
- Access restrictions: Limit the reach and potential of AI enablement.
Open-weights models, even small ones (when coupled with a capable harness), rival the capabilities of the most expensive SaaS models. Independent benchmarks bear this out. As of July 2026, a leading open-weights coding model scored 81 out of 100 on the "Terminal-Bench" test. This compares well with the top-tier models' score of 85. The scores come from Artificial Analysis, a trustworthy, vendor-neutral evaluator.
Keep in mind that routing to self-hosted models is not only a cost-driven pattern. Other factors, such as security, privacy, accuracy, speed, provenance, attestation, and qualitative and regulatory considerations, also apply.
Whatever the driver, inferencing on owned or leased infrastructure can reduce costs by an order of magnitude for those routed tokens.
The enterprise advantage of open-weight models on AMD
AMD's approach optimizes the deployment, tuning and operation of new generations of open-weight models such as GLM 5.2 and Kimi K3, making those and other leading models practical and economical to run at enterprise scale on-prem. Kimi, Meta, Cohere, DeepSeek, Mistral, MiniMax, OpenAI OSS, Qwen and z.AI (GLM) all run on a single MI355x server, illustrating the benefit of AMD's aggregate memory capacity and other key infrastructure features.
AMD's Blueprints and Inference Microservices (AIMs) facilitate both the evaluation and production phases of AI infrastructure rollouts. AMD's routing LLM Router Blueprint is an open and free resource for the industry, not a proprietary product.
AMD's AIMs and Blueprints allow further customizations, including fine-tuning of models that can train models on proprietary codebases and internal knowledge, adapting them to your company's specific patterns, libraries and conventions. This approach can, in fact, exceed the quality of closed SaaS inference models that usually cannot be tuned in this way.
How the off-ramp works
The heart of the solution is an LLM gateway that performs semantic routing. Before any request goes to a model, the gateway evaluates it.
- Appropriate requests route to an on-premises, open-weights model.
- Other requests can still go to a frontier provider when justified.
This is not a wholesale migration. It is a selective offloading of the requests based on a host of qualitative and quantitative measures.
What the numbers look like
The math is straightforward. A high-performance AMD GPU server was recently quoted for a WWT customer at roughly $380,000. That server comfortably supports 35 to 40 developers. For an organization operating on that scale, payback on the hardware investment can be realized within a few months.
Below, find a table comparing frontier inference token costs to on-premises, open-weight model costs.
Note the separate prices for input and output tokens, and the cost differential for access to the highest-capability tier. WWT's TCO models show that on-premises, open-weights inference brings all of those costs down to a small fraction (1/10th).
Finance teams will also note that building on-prem AI capacity doesn't have to be a large one-time capital expense. With consumption or leasing options, that infrastructure can be paid for quarterly or annually, keeping the cost shaped like existing AI bills, as a steady operating expense.
This also reduces enterprises' barrier to entry. WWT offers financing options during the procurement engagement.
Finally, the biggest objection WWT hears from its clients is the lead-time to production deployment. The good news is that a typical project only takes a few months.
The story doesn't end there.
Beyond cost savings: Protecting IP and controlling spend visibility
Cost may be a headline, but it is not the only factor. WWT and AMD experts operate in defense, financial services, healthcare, energy and other complex and regulated industries.
WWT experts have witnessed developers effectively barred from using any SaaS coding assistant at all, out of concern that proprietary code could leave the building through a third-party API. Those teams were left with tools so limited that developers described them as barely better than autocomplete, and their rate of innovation suffered as a result.
Recent, well-publicized incidents involving AI providers and their handling of submitted code have raised the stakes.
For a growing number of companies, the answer is to run inference on their own infrastructure. Visibility is another related issue. Running inference on infrastructure you control gives you back a clear view of what you are spending and what is driving it.
Finally, right-sizing the system so that a small, quick-to-inference model drives simple agentic tasks that don't require complex reasoning can yield significant cost and time savings.
Is this the right fit for your organization?
The benefits of this approach accelerate with organization size. As a rough guide, it starts to make sense once you have 35 developers actively using AI-assisted development tools. It's best if an existing data center footprint and an operations team comfortable managing AI infrastructure day-to-day already exist.
For organizations without those pieces in place, the cost of acquiring and maintaining the hardware can offset savings. The right starting point in that case is an assessment, not a purchase.
Where WWT fits in
WWT guides organizations through each stage of this transition. Its role spans assessment, design, testing and deployment.
The process begins with a comprehensive assessment of your current AI spend and usage patterns to project how much of your workload can move in-house and at what cost, before taking more complex implementation steps.
WWT also helps you to navigate the complex landscape of model selection and compliance. For industries with stringent data sovereignty or regulatory requirements, certain models may be off-limits. WWT helps you make informed decisions that balance performance and compliance.
WWT solutions are designed for flexibility.
Whether your infrastructure is built on AMD or other AI hardware, WWT can adapt the Blueprint reference applications and example implementations to your specific context.
WWT's capabilities extend deeply into the full stack, including orchestration, fine-tuning of open-weight models, customizations, routing policies, guardrail criteria and all other technical factors that optimize your workloads.
WWT also closes the skills gap. Its team brings the AI infrastructure expertise most organizations have not yet hired for, leaving your team free to focus on their core skills.
Crucially, before any changes are made to your systems, WWT's AI Proving Ground provides facilities to test proposed solutions. This hands-on lab environment allows you to test Blueprints against your own workloads and experience impacts firsthand.
Although AMD blueprints and reference architectures carry no licensing cost, infrastructure and deployment work do require investment.
Begin with an assessment
The first step is not a purchase. It is an assessment of what your organization is spending and what of that spend you can move. Contact your WWT account manager to scope that assessment for your environment.
Take the first step. Schedule your AI routing assessment:
About the authors
- Andrew Athan — Technical Solutions Architect III, GS&A Cloud & Infrastructure Solutions, WWT
- Eric Becker — Technical Solutions Architect III, GS&A Cloud & Infrastructure Solutions, WWT
- Nathan Chang — Senior Manager, Technical Business Development (AI), AMD
- Wes Henderson — Senior Manager, Enterprise AI Business Development, AMD