One of the nastier bottlenecks in agentic AI is the KV cache, a huge temporary log used to build context while queries run. A lot of that data sits in HBM, but capacity demands in AI clusters are scaling too fast for HBM to keep everything on board.
At CES 2026, Nvidia announced that BlueField-4 DPUs will hook into a new storage layer called Inference Memory Context Storage, or ICMS. It is pitched as a way to push more context data through the system, but it also risks moving the crunch from memory to NAND.
Citi estimates that a Vera Rubin system could pack roughly 16TB of NAND per GPU in a rack. That works out at 1,152TB in a single NVL72 configuration, which is an absurd amount of flash to bolt to one box, even by data centre standards.
The numbers get nastier when you scale them. Citi projects Vera Rubin shipments could hit 100,000 units in 2027, implying Nvidia-driven demand of 115.2 million TB of NAND, which it puts at 9.3 per cent of total projected global NAND demand.
If that is even vaguely right, ICMS-equipped Vera Rubin systems could create a supply shock the flash industry has not priced in. The AI supply chain has a habit of turning “nice to have” components into “no stock anywhere” overnight.
Nvidia has been flagging agentic AI as a major application focus, which means a bigger KV cache pool is non-negotiable for future racks. ICMS is effectively a plan to externalise that cache onto SSD-class storage, and that will put a hard strain on NAND availability.
This is landing while NAND is already under pressure from data centre buildouts and the broader inference frenzy. If Nvidia starts taking a meaningful slice of global output, everyone else gets to fight over what is left.
The uncomfortable parallel is DRAM. AI vendors are not slowing down, and consumer markets rarely win a bidding war against hyperscalers and GPU rack budgets.
If the NAND market tightens the way DRAM has, the knock-on is simple. General-purpose SSDs and everyday storage could get pricier and harder to find, and consumers will be told it is somehow “demand strength” rather than a supply chain being strip-mined by AI racks.







