The Diligence Stack - By Creative Strategies

The Diligence Stack - By Creative Strategies

Breaking the Memory Wall With CXL

Why disaggregated memory is moving closer to commercial deployment and who benefits

Ben Bajarin's avatar
Ben Bajarin
Aug 06, 2026
∙ Paid

We spent the last few days at FMS (Future of Memory and Storage) and had meetings with all the key players in the memory, storage, and now also interconnect/networking, ecosystem. A clear takeaway was better line of sight to the CXL standard to start to solidify. If you are not familiar with CXL it is a standard interface that gives the industry a common way to attach and share memory over a PCIe-based physical link. Up to this point the deployment model around that interface is still being developed, which is why the ecosystem can support several architectures, from a memory box inside a rack to pooled or optical memory systems.

From our conversations we believe deployments are more likely to start next year and build into 2028, but the ecosystem is maturing enough for CXL to move from a standard into something that can be deployed with enough customers to meaningfully start to deploy in their AI compute infrastructure. The first and most logical customer base is the hyperscalers, which can put memory and inference workloads into custom compute clusters and use the surrounding ecosystem to qualify the architecture.

It is our conviction that solving the memory wall will take many different shapes by many different players but the same problem statement remains. We need more memory, and designers are up against how much memory can go on CPU/XPU/GPU package or near the package. Having ways to expand the available memory pool while keeping latency low enough for selected near-memory workloads is the promise of CXL.

Memory is becoming a fleet problem

The old server model ties memory capacity to one processor and its local channels. AI workloads make that boundary more expensive because context, KV cache, and orchestration state can grow faster than the memory attached to one compute device. A fleet can have plenty of memory in total while individual CPUs or accelerators still run short.

The practical distinction is important: “hot” describes the role memory plays in the workload, while local, off-die, off-board, and rack-level describe where that memory sits. HBM can remain the hot tier in an attached appliance when the fabric preserves the bandwidth and latency the workload requires. DDR can be split by role in the same way. Local DDR can serve the CPU’s most active data, while DDR4 or DDR5 behind a CXL controller can provide a larger attached tier and eventually a shared pool. That remote memory will not have the same latency as on-package HBM, but it can still be the highest-performance tier available outside the package.

CXL gives the system those placement choices through a coherent interface. The first commercial use is likely to be a card, module, or box that adds memory to one host. Later designs can attach several hosts to a pool, with software deciding which data belongs close to compute and which data can move into the attached memory. The value comes from expanding the working set and using capacity more fully without buying a full server for every increment of memory.

The graphic below frames the opportunity by workload rather than by device. HBM remains closest to compute for the hottest accesses, while local DDR/SOCAMM and CXL-attached memory support larger working sets, overflow KV cache, retrieval buffers, and shared state. CXL is not a fixed warm tier; its role depends on the memory type, topology, and workload.

Exhibit 1. Inference moves the memory problem from model weights to active state. CXL expands the set of places where that state can live.

Deployment flexibility

It is noteworthy how early we are in the rack scale compute era. By our estimates, using our accelerator installed base model, we believe rack scale accelerators in the range of 9-11% of the total AI accelerator installed base. The shift to rack scale solutions are the necessary catalyst that will enable CXL, and other solutions even if custom, to be designed by the customer. Knowing the data center customer is scaling their rack scale infrastructure helps as a catalyst for CXL whether that deployment is north south or east to west in its location. Enough of the ecosystem was present, through announcements and demonstrations at FMS, to make the deployment path more visible. Memory suppliers, interconnect companies, switch vendors, custom silicon providers, and networking ASIC companies were all discussing a path toward viability with much different language than earlier in the year. As of now, we expect the larger deployments are more likely to arrive in 2027 and build into 2028, yet the ecosystem is now moving from evaluation toward qualification, driven by a key set of customers.

We believe hyperscalers are the first catalysts for several reasons. They control rack specifications, work with ODMs, and can make system-level TCO decisions that are difficult in standardized OEM deployments. They also have access to a large legacy memory resource that is not always useful in its original server configuration, but could become valuable again in a CXL-based system. We explain the size of that resource and the qualification path in the full report. That flexibility gives hyperscalers room to qualify a CXL memory box, attach memory beside a CPU or custom ASIC, or build a dedicated memory system. The economic benefit is that they can add controller and system content around capacity they already own.

Our key take here is: We have increased confidence that the CXL standard will emerge as a preferred additional approach to infrastructure build out. Its economic and TCO benefits, detailed in the full report, along with its flexibility in implementation give it many advantages over other solutions. While we understand the tradeoffs, we believe the biggest customers in the world are positioned to drive the adoption of CXL and have distinct advantages over others in the market. The competitive advantage available to those customers through CXL could become evident.

Inside the Full Report

  • Why hyperscalers are likely to become the first large CXL customers, and how their ODM model helps them qualify custom systems faster.

  • How much existing memory could become reusable, what that changes for deployment timing, and the potential TCO benefit.

  • Where CXL fits alongside HBM, local DDR and MRDIMM, HBF, proprietary memory attach, and networked-memory approaches.

  • Our CXL market model through 2030, including the path from single-host expansion to rack-level pooling.

  • The beneficiary map across controllers, memory suppliers, interconnect, systems, and software, plus the production signals needed to validate each opportunity.

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Creative Strategies, Inc. · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture