NetApp AFX 2K - What's New in AFX for Scalable AI Storage
This is a follow-up to our original deep dive, NetApp AFX: Disaggregated ONTAP for the AI Era. If you haven't read that piece, start there; it covers the disaggregated architecture, the Storage Availability Zone and the hardware fundamentals that everything below builds on.
A few months ago, we introduced NetApp AFX as the newest ONTAP personality. AFX is a disaggregated, all-flash platform built to bring three decades of enterprise ONTAP data management to AI-scale workloads. NetApp didn't rest on its laurels after that launch. With the July 2026 launch of new AFX pricing and the second-generation AFX 2K storage controller, paired with the ONTAP 9.19.1 release, the platform picked up a meaningful round of upgrades: a lot more write performance, bigger clusters, storage-availability-zone-wide deduplication and, for the first time, cold-data tiering.
AFX 2K is an evolution of the same architecture we walked through last time, with NetApp addressing some of the items that showed up in AFX's first few months in the field. Read on to learn what's new.
A quick refresher
If you're joining without the background: AFX splits storage into two independently scalable resource pools: controller nodes for compute/performance and NX224 NVMe shelves for capacity, all connected over a shared Ethernet backend fabric. Every controller node can see every disk in what NetApp calls a Storage Availability Zone (SAZ), so there's no per-node aggregate to manage, no data silos and volumes rebalance automatically as the cluster changes shape. None of that has changed with this update. What's changed is the hardware running it, and a batch of ONTAP-level enhancements shipping alongside it.
Meet the AFX 2K controller
The headline hardware change is the AFX 2K, a refreshed storage controller replacing the original AFX 1K for new deployments. NetApp still offers the 1K, but it's aimed at upgrades to existing clusters; AFX 2K is now the default, though it is possible to do an in-field upgrade from 1K to 2K.
Here's how the two nodes compare:
Spec | AFX 1K | AFX 2K |
|---|---|---|
| CPU | 2 × 52-core, 1.7GHz | 2 × 48-core, 2.3GHz |
| DRAM | 1TB | 1TB |
| NVRAM | 64GB (1x 64GB) | 128GB (2x 64GB) |
| Backend fabric (HA/cluster/storage) | 100GbE | 400GbE |
| Client-facing network | 100/200/400GbE | 100/200/400GbE |
| Boot media | 3.8TB NVMe | 3.8TB NVMe |
| Power | 2 × 2000W Titanium | 2 × 2000W Titanium |
You'll note there are two changes. First, NVRAM doubles: the 2K adds a second 64GB NVRAM module alongside the original, for 128GB of write cache per node. Second, the backend fabric jumps from 100GbE to 400GbE for HA, cluster and storage traffic. Those two changes are the reason for the performance improvement shown below.The CPU swap is a smaller detail but worth knowing: the 2K trades core count for clock speed consistent with a controller that's now more cache- and I/O-bound than compute-bound for its intended workloads.
Why it matters: The write performance story
NetApp's own benchmarking of the original AFX 1K against the AFF A90 told an interesting story: AFX's read advantage over the A90 was substantial, roughly 75% higher throughput, but the write advantage was far more modest, around 13%. Reads had headroom to spare but writes were bottlenecked by the size of the write cache. Doubling NVRAM in the AFX 2K goes directly after that limitation, and the numbers back it up.
Per-node throughput in a 4-node, single-shelf configuration:
Metric | AFX 1K (ONTAP 9.18.1) | AFX 2K (ONTAP 9.19.1) | Delta |
|---|---|---|---|
| Sequential read | ~35 GB/s | ~35 GB/s | — |
| Sequential write | ~10 GB/s | ~17 GB/s | +70% |
| Min. cluster total (4 nodes/1 shelf) | 140 GB/s read, 40 GB/s write | 140 GB/s read, 68 GB/s write | +70% write |
Reads were already great and didn't change. Writes, the side of the equation that matters most for checkpoint-heavy training runs, jumped 70%! That's a large number for a mid-cycle hardware refresh, and it lines up with the architecture: doubling the write cache without touching the read path is exactly where you'd expect the gain to show up.
The gains are largest at small node counts and shrink as the cluster grows. This tracks, since a bigger cluster was already spreading write load across more NVRAM modules even on the 1K. Where the 2K helps most is exactly where a lot of AFX deployments start: 4- to 8-node clusters sized around a single training or inferencing environment.
ONTAP 9.19.1: The platform gets bigger, not just faster
The AFX 2K ships alongside ONTAP 9.19.1, and several of the most consequential changes in this release aren't hardware-specific; they apply to any AFX cluster, 1K or 2K, running the new code version.
Cluster scale grows to 32 nodes. ONTAP 9.19.1 raises the supported AFX cluster size to 32 controller nodes, surpassing the 24-node ceiling that has long defined unified ONTAP clusters. NetApp is calling this limited support at general availability, with full support expected in a later patch release; 48-node clusters are already on the roadmap for later this year (subject to change, as roadmap items generally are). Clusters approaching the top of that range with 15 or more shelves step up to NetApp's larger Cisco 9808 backend switch rather than the 9332/9364 pair used in smaller configurations.
A single Storage Availability Zone can now scale to roughly 32PB. Support for up to 25 shelves, combined with the 61.4TB NVMe drive option, pushes a single SAZ to about 32PB raw (roughly 30.4PB usable), still presented as one pool of capacity to every controller node. Something worth flagging for anyone running an existing cluster: AFX systems initialized on ONTAP 9.17.1 or 9.18.1 can't grow past the prior 16PB ceiling without a full reinitialization, so this is a forward-looking number for new deployments rather than a free upgrade for systems already in the field.
Deduplication becomes SAZ-wide instead of per-node. This is arguably the most important software change in the release. At AFX's original launch, deduplication ran per node; data on Node 1's volumes was deduplicated against other data on Node 1, but not against identical data sitting on Node 2. With 9.19.1, AFX moves to a global hash table spanning the entire Storage Availability Zone, so duplicate data gets caught no matter which node it comes in on. NetApp's initial testing shows the efficiency gain scaling with how duplicate-heavy the dataset already is; this is logical. A dataset modeled with a 5:1 dedupe ratio and 2:1 compression profile improved from a 7.72:1 to a 9.57:1 overall storage-efficiency ratio, for example, with smaller, but still meaningful, gains on less duplicate-heavy data. For AI environments where multiple teams often keep near-identical copies of the same base datasets or model checkpoints spread across different volumes, moving the dedupe domain up to the whole SAZ is a genuine capacity win, not just a bigger number on a slide.
FlexGroup volumes get room to grow. To make full use of a 32-node cluster from a single namespace, NetApp raised the maximum FlexGroup member volume count from 200 to 512 which is enough for each node to host up to 16 member volumes, which is what it takes to saturate a node's performance potential inside one FlexGroup. That also lifts the theoretical addressable capacity of a single FlexGroup namespace to roughly 150PB (512 volumes × 300TB each), well beyond today's 32PB SAZ ceiling, but useful headroom as both tiering and raw capacity continue to grow.
FabricPool tiering arrives as a tech preview. This closes one of the gaps with unified ONTAP we flagged in our original piece: AFX launched without cold-data tiering. ONTAP 9.19.1 brings FabricPool to the platform, letting cold blocks move transparently off the performance tier to a lower-cost object target, policy-driven and invisible to applications, the same value FabricPool has delivered on unified ONTAP for close to a decade. At general availability it ships as a tech preview, with full support landing in a later patch release. Only ONTAP S3 and StorageGRID are supported as tiering targets for now, with other S3-compatible targets expected later.
With the cost of flash SSD seeing quarterly cost increases presently, being able to tier to one of the only platforms on the market with enterprise level spinning disk support, this is a good move on NetApp's part. Certainly, HDDs have increased in price, but they're not stratospheric (yet).
Paired with SAZ-wide deduplication, this means AFX can perform both major capacity-reducing techniques on the same system, at the same time.
The protocol layer keeps improving too
None of the gains above are really about NFS itself, but it's worth noting that NetApp continues to close the historical performance gap between NFSv3 and NFSv4 on AFX, release over release. Sequential throughput on NFSv4 improved 29% for reads and 10% for writes between ONTAP 9.17.1 and 9.18.1. On the metadata side, historically the harder problem for NFSv4, AI Image, Software Build and EDA benchmark results have climbed release over release to the point where NFSv4 metadata performance is now within 15% of NFSv3, while still carrying NFSv4's security and manageability advantages: a single firewall port, granular ACLs, integrated Kerberos and native support for parallel NFS and session trunking.
That closing gap matters for the "NFSv3 vs. NFSv4?" conversation that comes up, not infrequently: as of AFX 2K and ONTAP 9.19.1, there's less of a performance tax for choosing the more secure, modern protocol. Where client tooling supports it, NFSv4.1 with pNFS and session trunking remains the recommended default rather than a trade-off decision. For the unaware, it allows client-driven multipathing of NFS connections. This increases performance by opening up additional streams to the other controllers and it doesn't require a proprietary client driver.
Where AFX 2K fits and where it still doesn't
AFX isn't the only performance-tier option in NetApp's portfolio. For the very largest GPU environments, like those north of 16,000 GPUs, NetApp's EF-Series combined with Lustre makes a rather potent combination. AFX certainly sits nicely in the ONTAP ecosystem with all the data movement capabilities it's known for. For workloads below that threshold, our guidance is the same as the first post. AFX is purpose built for pretty big AI workloads and high-performance NAS needs.
Where this leaves the positioning
None of this changes the conclusion from our first article: AFX is the disaggregated ONTAP platform for AI-scale workloads, not a replacement for unified ONTAP's role as the backbone of the broader enterprise data estate. What the AFX 2K, its additional NVRAM and ONTAP 9.19.1 combination does is improve the platform's performance for things like write-oriented checkpointing during training runs. A 70% jump in per-node write throughput, a larger cluster ceiling, SAZ-wide deduplication and a real answer to the tiering gap all land squarely on the objections that came up during AFX's first few months in the field.
Putting it all together
The AFX 2K is the same disaggregated ONTAP foundation we covered in depth last time, tuned where the first-generation hardware was clearly limited and paired with an ONTAP release that pushes scale, efficiency and lifecycle management forward at the same time. If you're already comfortable with the SAZ concept, RAID-TEC and the migration paths from our original article, the 2K slots in as a straightforward hardware tweak with an outsized impact on write-heavy workloads. Finally, one of AFX's key strengths continues its tailwinds: AFX fits into the larger ONTAP ecosystem, making data movement in and out of AFX simple.
The logical next step for WWT's AI Proving Ground is to run the same testing methodology against AFX 2K hardware once it's in the lab, pre-filled capacity, sustained test runs and GPU-connected clients, to see how NetApp's numbers hold up under our own rigorous testing. When that happens, we'll follow up with results.
In the meantime, if you're evaluating AFX for an AI infrastructure buildout, new or existing, the best next step is still a conversation with your WWT account team or a session in the AI Proving Ground. Read the original architecture deep dive if you haven't already, and reach out to your WWT team to discuss storage for your AI workloads.