Written and provided by: AMD

Monthly token consumption has increased 158X in two years as AI moves from experimentation into products and services. Training compute has continued to grow, increasing 5X per year since 2020. Inference is becoming the largest AI workload as models serve billions of interactions.

Agentic AI increases that demand further. A single request can trigger multiple reasoning steps, sub-agents, retrieval operations and tool calls, along with CPU-driven routing, scheduling and memory management. The demand is therefore not only for more tokens; every useful result requires more inference, orchestration and data movement, placing additional pressure on compute, memory capacity, networking, latency and cost per token.

AI requires a new infrastructure blueprint. Leadership compute is essential, but compute alone is not enough. Customers also need open rack architecture, high-bandwidth scale-up and scale-out fabrics, efficient power and cooling, serviceability and turnkey solutions that can be deployed rapidly. AMD set out to design these capabilities together as one system.

AMD Instinct MI455X GPU is the engine. AMD Helios is the system.

At the heart of AMD Helios, the AMD Instinct MI455X GPU delivers a generational leap in AI compute, HBM4 capacity and memory bandwidth.

On DeepSeek-V4-Flash, one of the leading open-weight models in its class, AMD Instinct MI455X GPUs deliver up to 34X higher token throughput at high interactivity and up to 18X lower token cost compared with AMD Instinct MI355X GPUs. The results connect the architectural gains to the performance and economics required for production inference.

Inside AMD Helios: One rack, one co-designed system

AMD Helios is designed around the path an AI workload takes through the rack. Work enters through the host layer, moves into the GPU domain, accesses model and active data in HBM4 and communicates across the rack through the scale-up fabric.

The AMD Helios rackscale solution combines 6th Gen AMD EPYC "Venice" 9006 Series Server CPUs, 5th Gen AMD Instinct MI455X GPUs and AMD Pensando networking with UALoE fabric in a fully co-designed platform. Together, these components connect 72 GPUs in one scale-up domain with 260 TB/s of scale-up bandwidth for frontier AI inference and training. 

That rack-scale design sets a new competitive bar. More compute and HBM capacity give large models and active workloads additional room within the rack. Higher HBM bandwidth helps keep the GPU domain supplied with data, and greater scale-out bandwidth provides more capacity for expanding beyond a single rack. AMD Helios brings those capabilities together as one rack-scale platform rather than treating compute, memory and networking as separate infrastructure decisions.

Open at every layer, from rack to software

AMD ROCm software connects the rack-scale architecture of AMD Helios to production AI workloads. Native support for leading frameworks, including PyTorch, TensorFlow and JAX, lets developers use familiar tools for high-throughput inference and distributed training.

Optimized libraries and open standards help applications take advantage of AMD Instinct GPU compute, memory and fabric capabilities across the rack. AMD ROCm software also provides tools for deployment, observability and lifecycle management, giving infrastructure teams a consistent software foundation for operating AMD Helios at scale.

AMD Helios is being adopted across AI leaders, cloud partners and infrastructure partners. That breadth matters because rack-scale infrastructure reaches production through a connected supply chain, from model developers and cloud operators to the companies building and servicing the physical systems.

Final takeaway

AMD Instinct MI455X GPUs bring a major increase in compute, memory bandwidth and HBM capacity. AMD Helios scales that engine into a complete rack-scale platform, connecting 72 GPUs with AMD EPYC "Venice" server CPUs, AMD Pensando networking, AMD ROCm software and an open system architecture.

AMD Helios also marks a broader shift in how AMD advances AI infrastructure. Annual execution now connects successive AMD Instinct GPU generations with progress in memory, interconnect and GPU scale-up. The result is an open, multi-generation compute cadence for rack-scale AI.

With AMD Helios, the rack is no longer simply where the AI system is installed. The rack is the AI system, and AMD is delivering an open platform for the AI factory era.

Ready to advance your AI infrastructure? Connect with our experts

Technologies