Future of Memory and Storage Conference 2026
It has been said, AI isn't driven by compute, but by how fast we can feed compute. The things that feed compute are some form of data storage, either memory or some form of disk device. The folks that care about keeping compute satiated got together recently in Santa Clara. August 4-6 saw the 2026 edition of the Future of Memory and Storage conference, formerly known as Flash Memory Summit or, in both cases, FMS. FMS describes its mission as 'the most comprehensive and technically credible memory and storage event on the planet. [link] It's a deeply technical event with excellent industry presenters, from SSD engineers, CEOs, analysts, and everyone in-between. The two authors of this post attended, listened, learned, and interacted with peers and exhibitors.
In this space currently, we all lament the same things: supply, demand, and the cost that falls out when they meet, as well as figuring out the best way to provide storage to a consumer. The former is of such importance that FMS this year featured an entire track dedicated to the flash and memory market of late, discussing the supply, demand, costs, trends, market inputs and outputs. On the other side of it were some of the solutions for addressing it, from CXL to HBF, both of which are discussed below and even tiering down to spinning media. Finally, it was interesting to see the developing ecosystem around the inbound PCIe 7.0 products, from signal analyzers to error injectors and protocol testers.
Storage isn't just one manufacturer or one singular product type. Across HBM*, DRAM, NAND, HDD, SAS, and PCIe, to name a few, there are more acronyms than you can shake a stick at. The ecosystem has myriad products, from fast to slow, with different attachment types and protocols. All are tools in the toolbox that marry to some form of compute like CPUs or accelerators, then joined by software to achieve a customer's needs. There's no one way to do this and this conference highlights the enormity of the solution space available to end users.
*HBM is a key component of the current crop of server accelerators. Pioneered by SK Hynix and AMD in the early-to-mid 2010s, it uses a multi-layer die stack and wide data paths to trade DDR and GDDR's low latency for a 10x bandwidth improvement, ensuring the accelerator remains fed.
The memory wall is no longer theoretical
Every keynote at FMS 2026 arrived at the same conclusion from a different angle: AI stopped being a compute problem and has become a memory and storage problem (fitting for a memory and storage conference). Micron's closing keynote put it plainly: the industry has entered an era where a single heavily used GPU can require tens of terabytes of KV cache, far beyond what high-bandwidth memory or even local DRAM can hold. SanDisk's numbers frame the scale: the four largest hyperscalers alone are projected to spend roughly $720 billion on AI infrastructure this year. That spending is increasingly going toward solving one problem: getting the right data close enough to compute, fast enough, without paying for capacity you don't need.
What follows are two views into how the industry is answering that problem: one from the compute side, one from storage.
Compute's answer: CXL grows up
CXL what and why
PCI Express has doubled in speed roughly every generation since its inception. Today's widely available PCIe 5.0 can push 63GB/s on a x16 card, and the working group has already finalized PCIe 7.0, though it isn't available in any general-purpose servers yet. That kind of bandwidth is now in the same neighborhood as DDR5 RAM's per-channel speed, which is exactly what's being exploited to extend a system's memory over the PCIe bus itself. That's what Compute Express Link® (CXL) does: it rides the PCIe transport with a cache-coherent protocol layered on top, so memory inside a CXL device stays synchronized with the rest of the system automatically. In practice, that lets servers otherwise capped at eight terabytes of DIMM-based memory reach well beyond it, at a latency cost of roughly 320 nanoseconds versus about 80 nanoseconds for local DRAM, still a massive improvement over the 400,000 nanoseconds NVMe storage offers.
A fair question is, how do we know this won't go the way of Optane? We believe Optane struggled to gain broader adoption, not because the technology was bad, but because application-level rework was needed to take advantage of it in its most useful form. Decreasing DRAM prices at the time of its termination didn't help either.
For the last several years, Compute Express Link® (CXL) has been the industry's most-discussed answer to the memory wall, promising to let memory be expanded, pooled, and shared across servers instead of being stranded on a single machine. This year at FMS, the technology showed real signs of maturing past the hype cycle, though not without some honest disagreement about how far it has actually traveled.
Industry perspective
Intel's own CXL Consortium lead was candid on stage that shipments have lagged three-year-old projections, largely because industry attention and budget shifted toward GPUs and HBM during the AI buildout rather than CPU-side memory expansion. Keynote speaker and independent market analyst Jim Handy's own research pegs the CXL market at a modest but steady climb toward roughly $15 billion by 2031. Perhaps most notably, one SK Hynix presenter mentioned that Google has published research arguing against CXL memory pooling, a reminder that this remains a live technical debate, not settled science, even after several years of development.
Yet the deployment evidence collected across the week was the strongest we've seen. Micron, working with the Department of Energy and Pacific Northwest National Laboratory, has a production disaggregated-memory system (internally called 'Abaco') that isn't a lab demo. It reduced a real government graph-analytics workload's runtime from seven hours to twelve minutes and delivered an 8x improvement on a financial database benchmark. Astera Labs' CXL memory controller, branded Leo, showed similarly concrete numbers in production deployments: 75% better GPU utilization, 3x more AI tokens generated per second, and twice as fast time-to-first-token, versus not using CXL-attached memory at all. SK Hynix and Marvell's joint CMM-Ax platform, built on Marvell's Structera controller, posted a 5.5x throughput improvement for AI memory caching specifically (more on this platform ahead).
Perhaps the clearest signal of all came from a panel hosted by a major accelerator designer where representatives from KIOXIA, Micron, and Everpure (formerly Pure) each independently confirmed they are actively selling or demonstrating CXL-attached memory pooling products today, not on a future roadmap. When three competitors on the same stage all say the same thing without prompting, it's worth taking seriously.
What's still catching up, by near-universal agreement among the vendors presenting, is not the silicon. It's the software. Samsung, SK Hynix, and Astera Labs were each asked directly what the biggest obstacle to broader CXL adoption is, and each gave a version of the same answer: standard programming models across vendors, consistent error-reporting and diagnostics, and enterprise-grade reliability tooling are still maturing. Meta's own engineers, describing their hyperscale deployment experience, pointed to the lack of standardized error reporting across CXL vendors as a real operational gap: an uncorrectable memory error can still crash a host in ways that are harder to diagnose than with decades-proven local DRAM.
The practical takeaway for anyone evaluating AI infrastructure today: CXL memory pooling is no longer a science project, but it's also not yet a drop-in, vendor-agnostic commodity. The organizations getting real value from it right now (Micron's DOE partnership, and the vendors who confirmed live deployments in the aforementioned panel) are doing careful, workload-specific engineering, not buying it off a shelf. That's likely to change over the next 12-24 months as the software ecosystem matures, but it's the honest state of things today.
Storage's answer: doing more with less
Constraint breeds innovation. Escalating capacity and performance demands, combined with a limited supply of silicon wafers and fab capacity, have forced the ecosystem to re-evaluate its offerings and consider better architectures. High Bandwidth Memory (HBM) and DRAM or GDDR provide excellent low-latency, high-bandwidth performance, but at a high cost per bit. Those characteristics are still needed in today's hardware, but additional capacity is required to accommodate model growth.
Enter High Bandwidth Flash (HBF). HBF restructures the internal architecture of NAND chips, moving from fewer, larger arrays to many smaller arrays that can be accessed in parallel, improving overall flash performance. Rather than being bundled into a drive behind a separate controller, HBF dies are integrated directly onto the accelerator package, similar to HBM, with the controller sitting as the IO die at the base of the stack. Compared to a 64GB stack of HBM, a stack of HBF provides roughly 8x more capacity, around 512GB. Cell wear remains a concern with any flash, so the emerging thinking is that HBF is best suited to read-intensive stages of the inference pipeline, like decode, while write-heavy stages like prefill stay on HBM.
Hard drives still have a role in the overall strategy. HDDs deliver a lot of storage at low cost, and demand is climbing as customers globally work around limited flash supply. Flash drives are already reaching 120 to 240TB while spinning drives top out around 30TB, so density favors flash, but cost per gigabyte still favors HDDs. HDD manufacturers have reportedly been slow to add production capacity, wary of ending up oversupplied when the AI hype cycle cools. These drives aren't considered performant enough to feed AI pipelines directly, but they remain well suited to cold or bulk storage, or as a data-staging area ahead of faster tiers.
QLC flash is the other capacity lever getting real attention, considered good enough for AI workloads, particularly read-intensive ones like inference, where it will outperform spinning drives on random reads by a wide margin. Because QLC draws from the same pool of raw silicon wafer capacity as everything else, a broader industry shift from TLC to QLC would put meaningfully more bits on the market globally, a logical move given that QLC is likely adequate for most general-purpose workloads.
Solidigm's QLC engineering work puts real numbers behind that TLC-to-QLC shift: a technique they call 4-16 programming is squeezing 33% more usable bits out of every silicon wafer and delivering roughly 4x better endurance, directly attacking the flash shortage from the supply side. VAST Data's contribution is architectural rather than media-level: a locally-decodable erasure code with just 2.75% overhead (versus the 20-30% typical in the industry), combined with compression, deduplication, and a similarity-based reduction technique, compounding to roughly a 3x increase in usable capacity on top of the erasure-coding gains.
IBM's own session on scaling enterprise vector search offered a genuinely new capability worth noting: a 100-billion-vector search database running on a single server. It was built on a hierarchical multi-index architecture that can be constructed with GPUs in 4 days instead of the 20-120 days a CPU-only build would take, all while sustaining over 90% search accuracy under 700 milliseconds. Notably, the presenter's own testing showed the system is fundamentally limited by SSD read speed, not by compute or memory, a useful data point for anyone designing large-scale retrieval systems for enterprise AI.
System efficiency
Tiering memory and storage saves money on the individual components, since colder data can sit on cheaper media until it's needed, but that savings comes with a cost of its own: more data movement to bring information to where it's actually needed. Energy analysis of full AI infrastructure stacks is now happening at places like the U.S. national labs, and the finding is consistent across venues: it's still the same problem discussed two years ago, and the solutions on offer are converging. One presenter's framing, citing MIT research, was blunt about the scale of it: the industry spends 10 to 100 times more energy moving data to and from compute than it spends performing the calculation itself. Because AI and analytics workloads require far more data than can fit in accelerator cache, a large volume of bits has to keep moving, sometimes starting as historical data pulled from cold archive or tape, staged on spinning disk for cleanup, and then bounced progressively between DRAM, local storage, and accelerator memory before it's actually used.
Computational storage is one answer the industry is investigating: instead of hauling data to a processor, crunching it, and sending it back, you could instead push the compute task into the storage itself, whether that's the SSD or the DRAM. In practice that could mean an ARM chip living on each SSD's own circuit board, with the server pushing work instructions down to it rather than pulling data up. One presenter grounded the idea in biomimicry, pointing out that the human brain computes within its own storage and does it on roughly twenty watts, a useful benchmark for how far conventional compute architectures still have to go. The CMM-Ax platform from Marvell and SK Hynix, mentioned above, is an example of computational storage or near-memory processing (NMP). It's a PCIe card with four DDR-5 memory slots/channels and 16 ARM cores for processing at 200GB/s without involving the core CPU to do it, and all for 45 watts.
AI data platforms (AIDP) are the other area seeing heavy development. Historically, a data mart would pull specific data from a larger warehouse, often enriching it with additional sources and creating yet another copy in the process. Multiple data marts would exist across an enterprise environment, each for different reasons. Newer platforms take one of two approaches instead: some use connectors to operate on data in place and drop results into an open table format, while others handle data movement automatically within the platform itself, rather than requiring an administrator or a manually orchestrated workflow to move it. f
What this means for WWT customers
The throughline across both compute and storage this year was the same: AI's appetite for memory and capacity has outpaced what any single tier, vendor, or architecture can solve alone. The organizations making real progress, whether it's Micron's DOE partnership, Astera Labs' production CXL deployments, VAST and Solidigm's flash economics, or IBM's storage-bound vector database, are doing careful, workload-aware engineering rather than treating memory or storage as an afterthought behind the GPU. For customers evaluating AI infrastructure investments today, the practical question isn't whether to adopt these technologies, but which workloads justify the engineering effort right now versus which can wait for the ecosystem to mature further. That's exactly the kind of assessment WWT's teams consult with our customers about. With our experts' time in industry coupled with ATC product testing, we achieve our customers' goals faster and with less risk. Reach out to your WWT team or to us to start a conversation about what's next.