The Rack Is the New Server: What the Cisco and Supermicro Partnership Means for AI Infrastructure
In this blog
- What Cisco announced
- Before the spec sheet: Why this is a business decision
- NVIDIA GB300 NVL72 and NVIDIA Vera Rubin NVL72: The rack-scale flagship
- NVIDIA HGX™ B300 and NVIDIA HGX™ NVL8: The dense eight-GPU building block
- MGX H200 NVL PCIe: The enterprise on-ramp
- How this differs from the UCS C880A and C845A
- The network is part of the product
- Security built into the AI factory
- Operations: One operating model instead of five consoles
- The facility decides the timeline
- What it means for the market
- Why this matters for leaders
- How WWT approaches this
- Start with the outcome, not the rack
- Prove it in the AI Proving Ground before purchase
- No AI without ARMOR
- What leaders should do next
- Why WWT
- The bottom line
- Download
For infrastructure leaders, the ground keeps shifting. Platforms that made sense a year ago are already being outpaced by model size and inference demand. Cisco's latest move with Supermicro and NVIDIA is the clearest evidence yet.
By adding Supermicro's liquid-cooled and air-cooled GPU platforms to the Secure AI Factory with NVIDIA, Cisco is extending its architecture from enterprise GPU servers into rack-scale AI systems. The flagship NVL72 platform is a 48U liquid-cooled system that integrates 72 GPUs, 36 CPUs and nine NVIDIA NVLink™ switch trays into a single compute domain, with power, cooling, networking, security and operations designed around the rack as a system.
That reframing didn't happen overnight. For much of the AI era, enterprise infrastructure conversations centered on servers. The question was usually some version of, "Which GPU server should we deploy?" Organizations then moved from individual GPU servers to large-scale AI infrastructure based on validated designs such as NVIDIA DGX BasePOD™, NVIDIA DGX SuperPOD™ and NVIDIA Enterprise Reference Architectures.
The bigger story is the unit of architecture. At the high end of AI infrastructure, customers are no longer designing around an isolated server and adding the rest later. They are increasingly evaluating a validated rack or scalable unit in which compute, fabric, cooling, security and lifecycle operations are part of the design from the start.
What Cisco announced
Three changes matter most in the expansion of the Secure AI Factory with NVIDIA into the rack-scale era.
- Rack-scale and dense NVIDIA accelerated computing enters Cisco's AI portfolio. Cisco will offer and support Supermicro rack-scale NVL72 systems, dense eight-GPU NVIDIA HGX™ systems and NVIDIA MGX PCIe systems as part of the broader Cisco AI infrastructure portfolio, sold through Cisco's authorized channel ecosystem.
- The architecture extends to NVIDIA Cloud Partner (NCP) reference designs. The NCP-compliant design uses Cisco Silicon One systems on the front end, Cisco N9100 switches based on NVIDIA Spectrum-X™ Ethernet silicon on the GPU back end, and Nexus One as the unified network architecture. Cisco states that it is the only NVIDIA technology partner using its own networking switches and network operating system in an NCP-compliant solution, even though it is not the only NCP partner overall.
- Validation and operations become part of the offer. Cisco Validated Infrastructure Services (CVIS), aligned with NVIDIA Infrastructure Services methodology, is intended to validate a deployed cluster against the reference architecture. Cisco Cloud Control is also positioned to correlate job health with compute, NIC, optics and network metrics across the stack.
Cisco says orderability begins in October 2026. WWT was one of the partners named in the announcement.
"Our deep partnership with Cisco is central to our growth, and we rely on their strategic foresight in bringing the right specialized technologies and partner ecosystem into their AI infrastructure portfolio. By combining our strengths with Cisco's secure AI architecture, we can deliver an even more comprehensive full-stack solution, which is exactly what our customers need now."
— Neil Anderson, VP, Cloud, Infrastructure and AI Solutions, WWT
The key takeaway is that Cisco is treating rack-scale AI as a full-stack architecture problem, not just adding another OEM to the lineup. Compute matters, but so do the network, power, cooling, security, validation and the operating model that keeps the cluster productive after it is installed.
Before the spec sheet: Why this is a business decision
The first wave of rack-scale AI data center spending came largely from hyperscalers and AI service providers with the engineering depth to integrate the stack. Most do not want a large integration project. They want capacity that can move into production predictably and produce measurable business value.
That changes the economics. Deloitte expects inference to account for roughly two-thirds of all AI compute in 2026, up from a third in 2023, and mixture-of-experts and reasoning models put sustained pressure on memory, scale-up interconnects and the Ethernet fabric between racks. The NVLink™ domain determines how efficiently GPUs can exchange traffic within the rack. The scale-out fabric determines how effectively the cluster can keep those GPUs busy across racks. At this scale, cost per token is not simply a GPU calculation. It is an architecture calculation.
WWT's view, shaped by several hundred customer engagements a year, is that the experiment phase is giving way to an execution phase. Executives and corporate boards are pressing for production returns rather than pilot metrics. Four issues stall AI programs as often as, or more often than, GPU availability: power and cooling, network congestion, an operating model built for a different generation of infrastructure, and security introduced too late. Cisco's announcement speaks to all four.
The products and where they fit
The Supermicro portfolio spans three practical tiers. They function like a ladder, where the right rung depends less on ambition than on model size, utilization and facility readiness.
NVIDIA GB300 NVL72 and NVIDIA Vera Rubin NVL72: The rack-scale flagship
The NVIDIA GB300 NVL72 is one liquid-cooled 48U rack. It includes 18 1U compute trays, each carrying two NVIDIA Grace Blackwell superchips, for a total of 36 NVIDIA Grace™ CPUs and 72 Blackwell Ultra B300 GPUs. The rack also includes nine NVLink™ switch trays. Each compute tray carries four integrated NVIDIA ConnectX®-8 SuperNIC™ for east-west traffic and one NVIDIA BlueField®-3 DPU for north-south traffic, which Cisco's Cloud RA describes as a 2-4-5-800 configuration: two CPUs, four GPUs, five SuperNICs and 800 Gb/s per GPU. Supermicro specifies eight shared 33 kW power shelves, with operating power in the 132 to 140 kW range.
Cooling is direct-to-chip liquid, with several heat-rejection options. Supermicro lists a 250 kW in-rack CDU, a 1.8 MW in-row CDU that can support up to eight racks and a 200 kW liquid-to-air sidecar for sites without facility water. The practical decision comes down to which cooling architecture matches the site, the scale of the deployment and the path to future racks, not which CDU has the largest number.
NVIDIA Vera Rubin NVL72 follows the same rack-scale direction one generation later: 72 NVIDIA Rubin GPUs with 288 GB of HBM4 each, 36 NVIDIA Vera CPUs, sixth-generation NVLink™ at up to 216 TB/s of rack-scale bandwidth, NVIDIA ConnectX®-9 and NVIDIA BlueField®-4. Cisco has included Vera Rubin NVL72 in the rack-scale architecture it plans to support. Availability and final system specification should still be treated as roadmap items until the products are orderable in the relevant Cisco and Supermicro configurations.
From an investment perspective, much of the surrounding architecture, including the rack-scale operating model, fabric design and facility planning, carries forward. Two cautions are worth noting. First, NVIDIA's headline Rubin performance and efficiency claims use specific comparison points and workloads, so customers should validate those claims against their own models and inference patterns. Second, Supermicro's current Vera Rubin planning guidance sizes liquid cooling at roughly 227 kW per rack. A customer deploying GB300 first should therefore consider whether the facility loop is being designed only for today's rack or for the next generation as well.
The 72-GPU NVLink™ domain is the reason this tier exists. Large-scale training, long-context reasoning and disaggregated prefill and decode place different demands on infrastructure than eight-GPU servers connected only through the scale-out fabric. This tier is most relevant to neoclouds, sovereign programs and enterprises with sustained utilization, very large models and facilities ready for liquid cooling. It is not where most enterprises should begin.
NVIDIA HGX™ B300 and NVIDIA HGX™ NVL8: The dense eight-GPU building block
The NVIDIA HGX™ B300 uses the familiar eight-GPU SXM form factor and comes in two configurations:
- A 4U liquid-cooled node uses direct-to-chip cold plates with quick-disconnects to a rack manifold. Eight nodes fit in a 48U rack, for a total of 64 GPUs.
- An 8U air-cooled node provides the same compute class for facilities where liquid cooling cannot be retrofitted; four fit within an 80 kW rack design.
Both configurations pair two Intel Xeon 6 processors with eight B300 GPUs, each with 288 GB of HBM3e, along with eight ConnectX®-8 SuperNICs, BlueField-3 DPUs, self-encrypting drives and TPM 2.0. NVIDIA HGX™ Rubin NVL8 follows in a 2U liquid-cooled chassis with BlueField®-4, with nine nodes per rack for a total of 72 GPUs. NVIDIA has published workload-specific claims showing materially fewer GPUs for some Rubin MoE pretraining scenarios compared with the HGX™ B200, but those results should be validated against the customer's actual workload.
Cisco's designs around these nodes use a dual-plane, rail-optimized back-end built on Nexus 9364E-SG2 switches. A 32-node design supports 256 GPUs across 12 switches, while a 64-node design supports 512 GPUs across 24 switches. Both are non-blocking and require no super-spine. This tier fits enterprises doing fine-tuning, RAG at scale and multi-model inference one node at a time; neoclouds offering eight-GPU tiers; and regulated sectors that want dense compute within a conventional operating model. Air-cooled is the bridge; liquid-cooled is the right call once the facility is ready.
MGX H200 NVL PCIe: The enterprise on-ramp
The 4U MGX system is a PCIe design with up to eight NVIDIA H200 NVL GPUs, 2 Intel Xeon 6 processors, up to 6 TB of system memory, ConnectX®-8 networking and BlueField®-3 DPUs, with air cooling and redundant power. It is well suited to private AI, RAG, copilots and agents; visualization and digital twin workloads; and departmental deployments in regulated or sovereign environments.
For many enterprises, this is the more practical starting point. It fits standard facility assumptions, aligns more closely with the operating model teams already know and provides a path upward when utilization, workload size and the business case justify denser infrastructure.
How this differs from the UCS C880A and C845A
Cisco already sells dense GPU servers, and the UCS C880A M8 and C845A M8 remain strategic parts of the portfolio. The Supermicro systems are best understood as an extension into rack-scale and liquid-cooled architectures rather than a replacement for UCS.
| Cisco UCS C845A M8 | Cisco UCS C880A M8 | Supermicro via Cisco | |
| Design and manufacture | Cisco UCS | Cisco UCS | Supermicro, validated, sold and supported by Cisco |
| GPU options | PCIe: NVIDIA H200 NVL, NVIDIA RTX PRO 6000, NVIDIA L40S (2, 4 or 8 per node) | 8x HGX™ B300 NVL8 | HGX B300 NVL8, HGX™ Rubin NVL8, GB300 NVL72, Vera Rubin NVL72, MGX H200 NVL |
| Host CPU | AMD EPYC | Intel Xeon 6 | Intel Xeon 6 (HGX™, MGX); NVIDIA Grace™ or Vera (NVL72) |
| Cooling | Air | Air | Air and direct-to-chip liquid |
| Unit of scale | Individual server | Individual server | Server, pre-integrated rack, or multi-rack cluster design |
| Cooling infrastructure | Customer-supplied | Customer-supplied | In-rack, in-row and liquid-to-air CDUs through Cisco |
| Management today | Intersight native, incl. policies and profiles | Intersight native | Nexus One (Nexus Dashboard or Nexus Hyperfabric) for the fabric; NVIDIA Mission Control™ for NVL72 compute; Supermicro Redfish BMC for HGX™ and MGX; Cloud Control and Intersight on Cisco's roadmap |
| Best for | Enterprise inference, RAG, visualization, mixed workloads | Premium eight-GPU training and inference, air-cooled | Liquid-cooled density, rack-scale NVLink™ domains, roadmap access to Rubin |
Four differences matter most:
- Liquid cooling: Cisco's current UCS GPU platforms are air-cooled, while the Supermicro portfolio brings direct liquid cooling with the CDUs, manifolds and rack integration required to deploy it.
- Rack scale: The C880A and C845A remain server platforms that can participate in validated reference architectures, while NVL72 is designed and delivered around the rack-scale NVLink™ domain.
- Roadmap access: The Supermicro portfolio extends Cisco's access to the latest NVIDIA platform generations.
- Industry reporting around the announcement has also framed the partnership as a way to broaden Cisco's compute supply options during a period of constrained AI component availability.
What does not change is the Cisco front door. UCS remains the core enterprise compute platform for general-purpose workloads, Unified Edge, PCIe GPU use cases and air-cooled HGX™. Supermicro extends the portfolio into liquid-cooled and rack scale designs.
The network is part of the product
Every OEM can build around the same NVIDIA accelerators. What differs is the architecture around them. In Cisco's case, that starts with the fabric.
Cisco has published two NCP-compliant Cloud Reference Architectures for the Supermicro GB300 NVL72: One built on Nexus 9000 switches managed by the Nexus Dashboard, and one built on Nexus Hyperfabric switches managed by the cloud-delivered Hyperfabric controller. Both sit under the Nexus One architecture, and both use the same design: Cisco N9164E-NS4-O switches on NVIDIA Spectrum-4 silicon for the NCP-compliant GPU back end, and Cisco Silicon One systems (N9364E-SG2X-O or HF6100-64ED) for the front-end, storage and management fabrics. Silicon One switches can also be used in the back end with the Spectrum-X™ fine-grain load-balancing license, but that configuration is outside NCP compliance. Customers therefore choose the management model, on-premises or cloud, without changing the architecture.
These environments also separate traffic by function. GPU scale-out traffic, front-end traffic, storage traffic and out-of-band management have different performance and operational requirements. Treating them as one undifferentiated network creates unnecessary risk. Cisco's Cloud RA defines three fabrics: an east-west compute fabric for GPU collectives; a converged north-south fabric that carries storage, host management and tenant traffic on separate VXLAN EVPN VRFs, sized at 44 Gb/s per GPU for north-south and 11 Gb/s per GPU to storage; and an out-of-band network that reaches every BMC in the rack, including the compute and NVSwitch trays, power shelves, BlueField® DPUs and the in-rack CDU. Each fabric has a defined role, while the operating model is brought together through Nexus One and, increasingly, Cisco Cloud Control.
The RA builds in scalable units of two NVL72 racks, 144 GPUs. Every ConnectX®-8 runs in 2×400G mode with one port to each of two identical fabric planes, and leaf switches are grouped by rail so that all GPUs in a scalable unit are one hop apart. An eight-rack, 576-GPU cluster uses 32 leaf and 12 spine switches in the compute fabric; the two-tier design scales to 64 racks and 4,608 GPUs, and a three-tier design continues to 512 racks and 36,864 GPUs, with a documented path to 73,728. Storage is Cisco EBox, the VAST Data AI OS on Cisco UCS C225-M8N servers, certified by NVIDIA at the NCP level, and the compute controller is NVIDIA Mission Control™ with NVIDIA Base Command™ and NMX-M. As the cluster grows, racks, leaf switches and spine switches are added; the underlying design remains the same.
From an infrastructure perspective, that is one of the larger shifts represented by the announcement. At rack-scale, the fabric is part of the performance envelope, the failure domain and the economics of the system, not a component selected after the GPU purchase.
Security built into the AI factory
86% of respondents in Cisco's 2025 Cybersecurity Readiness Index reported at least one AI-related security incident in the previous 12 months. That does not mean every incident was caused by a model. Many of the risks around production AI come from the systems surrounding the model: identity, data access, tool permissions, workload segmentation, software supply chains and operational visibility.
The Secure AI Factory is designed around the idea that those controls should be part of the architecture rather than added after the first production workload. In practical terms, the controls span three layers.
- Secure the model. Cisco AI Defense addresses validation and runtime risks such as prompt injection, jailbreaks and data leakage, with integration into the broader NVIDIA AI software stack.
- Secure the workload. Isovalent technologies provide identity-aware Kubernetes networking (Cilium) and runtime enforcement (Tetragon), while Cisco security controls and BlueField® DPUs can extend segmentation closer to the workload. In the Cloud RA, tenant isolation is enforced in the fabric itself through per-tenant VXLAN EVPN VRFs for the compute and front-end networks.
- Secure the user. Duo, Splunk, Talos and AI Defense contribute identity, security telemetry and SOC visibility across the environment.
For sovereign and regulated customers, there is a fourth layer that is really an architecture decision: Nothing has to leave the building. GPUs, data, models and the inference path can all remain on premises, with no SaaS hop. SONiC is available, and the nodes ship with TPM 2.0 and self-encrypting drives. But "secure by design" should not be taken on faith, which is why WWT assesses every architecture against the AI Readiness Model for Operational Resilience (ARMOR), discussed later in this article.
Operations: One operating model instead of five consoles
Rack-scale AI can become a multi-console problem very quickly. Server hardware may live in one console, cluster software in another, each network fabric in its own manager, and CDUs and power shelves in a facilities system. When a production job slows or fails, the operator needs to understand whether the problem is compute, networking, optics, power, cooling, software or the workload itself.
Cisco Cloud Control is the part of this announcement operations teams should examine closely. Cisco is positioning it as a common layer for inventory, topology and cross-domain observability, with Nexus One for network management and Intersight integration for compute management. AI Canvas adds an agentic troubleshooting layer that can correlate information across domains and guide an operator from alert to evidence-backed action.
The timing matters. Cisco states that management of the Supermicro compute systems through Intersight and Nexus One is planned for CY26Q4. In the published NVL72 Cloud RAs, the compute controller is NVIDIA Mission Control™ with Base Command™ Manager and NMX-M, which handles provisioning, firmware updates, liquid-cooling monitoring and the NVLink™ fabric, while the network controller is Nexus Dashboard or the Hyperfabric controller. Until the Cloud Control integrations are available, customers should be clear about what functions remain in Mission Control™, Nexus Dashboard or Hyperfabric, and the Supermicro management interfaces. Automation only creates value when the operational boundaries are explicit.
Day 0 and Day 1 also receive more attention in the new model. Stack Automation by Quali is intended to reduce the manual work required to deploy the full stack, and CVIS is designed to validate the resulting cluster against the reference architecture and produce evidence of the tested configuration. Cisco says the CVIS toolkit reduces deployment and validation timelines from months to weeks.
The metric that matters after that is time to first intelligence: how long it takes to move from hardware delivery to the first business-value workload in production. The software layer can include NVIDIA AI Enterprise on Red Hat OpenShift, Nutanix NKP or upstream Kubernetes, but the strategic point is broader. The value of the architecture depends on shortening the path from delivered infrastructure to useful work.
Cisco describes a coordinated support path in which Cisco handles initial triage and routes Supermicro product issues as needed; the Cloud RA states the same, with all customer cases front-ended by Cisco and routed to partner support after triage. For enterprise buyers, that matters because one historical concern with Supermicro was the support experience around the hardware, not the hardware itself. The partnership is designed to make that boundary more predictable.
The facility decides the timeline
At the flagship tier, liquid cooling is not simply a preference. GB300 NVL72 operates above 130 kW per rack, and Supermicro's current Vera Rubin planning guidance is approximately 227 kW per rack. At those densities, power delivery, heat rejection and facility water become architecture inputs.
A site without a liquid loop or a funded plan for one may be better served by the air-cooled tier while the facility roadmap is developed. An existing loop or a greenfield design opens the liquid-cooled option. The CDU model then depends on deployment scale, redundancy, serviceability and the facility's primary loop. Cisco's Cloud RA is neutral on in-rack versus in-row: It provisions out-of-band connectivity for an in-rack CDU in every rack and ties the OOB network to the building management system so leak-detection alarms can trigger flow-valve and power cut-off, but the physical cooling design remains a Supermicro and site decision.
The practical answer to "which tier can we deploy this year?" may be the tier below the one the team initially wanted. That is not a failure if the fabric, security model and operating approach create a deliberate path forward. Ordering rack-scale compute before the building can support it is the more expensive mistake.
What it means for the market
The OEM is handling more of the rack-level integration. NVIDIA defines the platform architecture, contract manufacturers build key components, Supermicro integrates the rack and liquid-cooling system, and Cisco brings the network, security, management architecture, validation and channel support. Customers increasingly evaluate the validated combination rather than a list of independent parts.
The unit of purchase moves up with it: server, then rack, then scalable unit, with procurement, facilities and network teams planning around the architecture as a whole. For Cisco, the partnership broadens the compute portfolio around AI networking wins and opens access to neocloud and sovereign businesses that were previously out of reach.
"With Cisco Secure AI Factory with NVIDIA, we no longer have to choose between performance, reliability or ease of management. NCP validation gives us the confidence that our infrastructure is optimized from day one, while the platform's rack-scale capability provides a seamless path to scale our AI operations as our business grows."
— James Manning, Co-founder and CEO, Sharon AI, quoted in Cisco's announcement
The logic is strong, but it also raises a natural question about where the market goes next. If rack-scale systems become the normal unit for the most demanding AI workloads, infrastructure teams will need procurement, facilities, networking, security and operations to converge earlier in the design process. The market may need fewer isolated product decisions and more repeatable architecture decisions.
That said, this is not an announcement from Cisco so much as an architectural implication of the direction the industry is taking.
Why this matters for leaders
For CIOs and CTOs, the announcement is a reminder that AI infrastructure is becoming a systems problem. The decision now turns on whether the organization can deploy, operate and upgrade the surrounding architecture without creating a new island of infrastructure, not only on which accelerator delivers the best benchmark.
For CISOs, the key question is whether AI security is embedded in the design early enough to govern data, models, agents, identities and tools before production scale exposes the gaps. Security that arrives after deployment is more expensive to retrofit and harder to operate consistently.
For infrastructure and operations leaders, facility readiness and operational consistency may determine time to value as much as GPU availability. A rack that arrives on time but cannot be powered, cooled, integrated or supported is stranded capital, not capacity.
For business leaders, the architecture matters because utilization matters. The business case improves when infrastructure reaches production quickly, stays busy and can grow without being redesigned at every step.
How WWT approaches this
Every tier above raises questions a data sheet cannot answer. Is the facility ready for direct liquid cooling? At what utilization does a rack-scale system justify itself compared with dense eight-GPU nodes? Does the fabric behave as the reference architecture predicts under the customer's workload and context lengths? Will the new systems fit the tools and processes the operations team already uses?
And beneath all of those questions is a fundamental one: "Do we have the use cases, the data and the governance to keep this thing busy once it is built?"
WWT's strategy is built around moving organizations toward AI Native operations and business models, but the infrastructure is a means to that end, not the destination. Production AI requires the right use cases, data, security, operating model and economics. A rack of GPUs may be necessary, but it isn't an AI strategy by itself.
Start with the outcome, not the rack
WWT organizes AI work around four broad outcomes: Workforce AI, Applied AI in mission-critical business processes, Physical AI through robotics, spatial computing and computer vision, and Agentic IT Operations. Those outcomes sit across five technology layers: infrastructure, data, intelligence, applications and security, with AI-native engineering as the build discipline.
Everything thus far is primarily the infrastructure layer. It is necessary, but it is not sufficient. The journey begins with use cases and business-value prioritization, then moves into architecture, build and scale. That is why rack-scale conversations often end with a smaller first purchase than expected and a deliberate plan to expand as utilization and business-value are proven.
Prove it in the AI Proving Ground before purchase
WWT's Advanced Technology Center (ATC) is a world-class testing and validation environment with 500-plus racks, 200-plus OEMs and more than 6,000 customer engagements. The AI Proving Ground (AIPG) is WWT's dedicated multi-OEM environment for building, testing and validating enterprise AI architectures at production-relevant scale.
It is designed to answer the questions customers bring into real projects: How should we size the AI environment? Can the existing storage fabric support the workload? What changes when we move from air cooling to liquid cooling? How should we secure the data, model and inference path? Which architecture delivers the best performance for our workload rather than a vendor benchmark?
The work under way in the AIPG is directly relevant to the Cisco and Supermicro conversation. It includes cluster sizing, InfiniBand-versus-Ethernet comparisons, thermal and power modeling, LLM evaluation, proofs of value, and AI security and risk mitigation. Power becomes a design input rather than a surprise.
As the Supermicro systems become orderable through Cisco, the goal is to evaluate them alongside the UCS, Intersight and Cloud Control environments already in the ATC and against competing OEM designs. The point is to give customers evidence they can use to choose among architectures, not to validate a predetermined answer.
For customers, that creates three practical advantages:
- Proof of concept before purchase. Customers can run the model, data pipeline and inference pattern on the candidate architecture and measure the characteristics that matter before ordering production capacity.
- Design validation. The Cisco cluster design under consideration is built and tested against the relevant reference architecture, surfacing integration issues before replication at the customer site.
- Facilities and operations readiness. Infrastructure and facilities teams gain hands-on experience with CDUs, manifolds, telemetry, leak detection and operating procedures before those components become production dependencies.
No AI without ARMOR
"Security and innovation can't sit on opposite sides of the table. WWT's path forward is clear: No AI without ARMOR."
— Chris Konrad, VP Global Cyber, WWT
The AI Readiness Model for Operational Resilience (ARMOR) is WWT's framework for safe, secure and responsible AI. It is vendor-agnostic and is designed to connect strategy-phase risk assessment with the controls and operating model required in production. WWT uses the framework to evaluate where an architecture is strong by design, where additional controls are required and where customer governance must carry the responsibility.
The seven domains are:
- Governance, risk and compliance: Responsible AI policy, regulatory mapping, auditability and policy-as-code approaches.
- Secure AI operations: Runtime detection, agent guardrails, tool controls and SOC integration.
- Model protection: Model scanning, red teaming, prompt-injection and jailbreak defenses, output filtering and supply chain integrity.
- Secure development lifecycle: Secure coding, CI/CD controls, software bills of materials and release governance.
- Infrastructure security: Network, host and DPU security, segmentation, workload protection and hardening.
- Data protection: Encryption, classification, DLP, sensitive-data handling, residency and sovereignty.
- Identity security: Identity and access management, zero trust, agentic and non-human identities, and key management.
Together, the seven domains are designed to produce a cyber-resilient AI architecture. ARMOR aligns these concerns to frameworks security leaders already use, including NIST AI RMF, ISO/IEC 42001, OWASP guidance, NIST CSF, the CSA AI Controls Matrix and the EU AI Act. Its value is practical: It gives the CISO, architect and operations team a common framework for evaluating the architecture instead of relying on the phrase "secure by design."
What leaders should do next
The announcement creates more choice, but the next step should not be a product-selection exercise. Leaders should use it to ask a set of architecture questions.
- Assess facility readiness before selecting the highest-density platform. Understand available power, cooling, floor loading, water, commissioning requirements and the timeline for liquid cooling.
- Match the unit of scale to the workload. Determine whether the use case truly needs rack-scale NVLink™, an eight-GPU HGX™ building block or an enterprise PCIe platform. Do not buy the largest architecture simply because it is available.
- Evaluate the operating model, not only the hardware. Identify which tools manage compute, networking, cooling and security on Day 1 and how those integrations change over the product roadmap.
- Design security and governance into the architecture. Define controls for models, data, agents, identities and tools before the first production workload.
- Validate before scaling. Test the real workload, fabric, automation, security controls and facility assumptions in a representative environment. The best time to discover a design gap is before the production racks arrive.
Why WWT
WWT brings together the AI Proving Ground, long-standing strategic partnerships and an Accelerate-Build-Scale methodology. The objective is to help customers choose the architecture that best fits the workload, facility, security requirements and operating model, not to steer every customer toward the same platform.
Sometimes the answer will be Cisco UCS. Sometimes it will be Supermicro through Cisco. Sometimes it will be a different OEM or a cloud service. The value of the ATC and AIPG is that those options can be compared using the customer's requirements and workloads rather than a slide deck.
A practical starting point usually includes
- AI Factory strategy and architecture workshop: use cases, business case, workload placement and first-purchase recommendation.
- Facility readiness assessment: power, cooling, floor loading and the liquid-cooling roadmap.
- AIPG proof of concept: the customer's model and data on the candidate design, with performance, power and cost measured under representative conditions.
- AI ARMOR assessment: where the architecture is strong across the seven security domains and where compensating controls are required.
The bottom line
The Cisco and Supermicro partnership looks like a compute announcement on the surface, but the larger story is architectural. It extends Cisco's Secure AI Factory into rack-scale and liquid-cooled systems, brings Supermicro into a Cisco enterprise support and channel model, and gives customers a more continuous path from air-cooled enterprise GPU infrastructure to large liquid-cooled AI factories.
The rack is becoming the new server because dependencies that determine success now live at rack and cluster scale. Power, cooling, NVLink™, Ethernet, security, validation and operations cannot be treated as separate decisions after the compute has been selected.
That also raises the bar for integrators. When the reference architecture is a multi-rack scalable unit with liquid cooling and a high-performance fabric, the difference between a successful deployment and an expensive science project is whether the architecture has been built, tested, secured and operated before it reaches production.
That is where WWT's ATC and the AIPG matter. Customers can validate architectures, test automation workflows, evaluate real use cases, assess security controls and understand the operational impact before making large-scale commitments.
Cisco's announcement provides the direction. The next step is to prove what works.