Micron says AI’s hunger for bandwidth is turning HBM into both the hero and the headache.
At Hot Chips, Micron HBM design architecture fellow Raghu Sreeramaneni laid out why memory is becoming a nasty bottleneck for AI systems. Compute is racing ahead at about three times every two years, while HBM bandwidth is advancing at less than two times over the same stretch.
That gap keeps the memory wall standing, and Sreeramaneni reckons it may be getting worse. AI processors spend too much time waiting for data, then only hit their stride once memory can feed them fast enough.
HBM remains the best DRAM option for current AI accelerators because it lifts the bandwidth ceiling. It stacks DRAM dies, links them with TSVs, and hooks the cube to an interposer via short, dense connections.
Micron said a typical GPU system-in-package with four 12-high HBM stacks can have memory silicon accounting for about 90 per cent of the total silicon. That is roughly eight times the GPU silicon, which is a big clue that memory is no longer the side dish.
The reliability problem is ugly too. Meta’s Llama 3 training paper blamed GPU HBM3 memory for 17.2 per cent of unexpected interruptions during a 54-day run, which is not ideal when training costs already resemble defence budgets with better slides.
Micron’s current HBM4 offers up to 2,800GB/s of bandwidth, twice the I/O and double the channels of HBM3E. Each generation is pushing data rates, pseudo-channels, density and stack height higher.
The snag is that taller stacks bring thermal grief. Micron says the issue is less about raw power and more about power density, especially around the base die where much of the advanced functionality lives.
The company is looking at liquid cooling, hybrid bonding and fusion bonding to cut thermal resistance and tighten pitches to single-digit microns. It is looking beyond 16-Hi stacks too, although Sreeramaneni made clear there is still plenty of work before that stops being a packaging migraine.







