by

Micron ships “World’s First” 256GB SOCAMM2 Modules

Micron is pushing a new SOCAMM2 module that targets the memory choke points showing up in long-context AI workloads.

The outfit claims it has set a “new benchmark” by raising per-module capacity to 256GB from the previous 192GB, pitching it as a way to ease constraints in modern AI infrastructure gear. Micron says SOCAMM2 is built to cut latency by shifting pressure away from slower paths, particularly when KV-cache gets pushed out to LPDRAM.

Nvidia, head of product, data centre CPUs Ian Finder said: “Micron’s achievements in delivering massive memory capacity and bandwidth using less power than traditional server memory with 256GB SOCAMM2 is enabling the next generation of AI CPUs.”

Micron says the latest SOCAMM2 lifts a single LPDRAM monolithic die to 32GB, and the 256GB module configuration can deliver 2TB of LPDRAM per eight-channel CPU. That is aimed at servers chewing through long context windows, where keeping more of the working set close to the CPU can prevent inference from becoming a waiting game.

The company claims the time-to-first-token is up 2.3x for long-context inference when KV-cache is offloaded to LPDRAM, which it reckons helps agentic-style workloads where CPU-side work matters. Micron is developing SOCAMM2 with Nvidia, and it has already been linked to upcoming AI infrastructure designs such as Vera Rubin.

Micron says 256GB SOCAMM2 samples have shipped to customers, and the modules will be shown at GTC 2026, with the unspoken sting being that chunky AI-focused DRAM can chew into supply that might have gone to more general-purpose parts.

 

 

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews