DDR prices have dipped in recent days, and some tried to pin it all on the release of Google’s TurboQuant, but now it is seems to have been a mirage.
For those not in the know, TurboQuant’s basic trick is to squeeze memory usage so LLMs can run on accelerators while using less memory.
TurboQuant’s trick is squeezing the key value cache, the short-term memory that lets models like ChatGPT and Claude keep conversational context, then rebuilding it when needed with little obvious accuracy loss. As interactions get longer and user counts climb, the KV cache bill grows fast, so anything that cuts “cost per token” gets attention.
Sungkyunkwan University professor Kwon Seok-joon said TurboQuant “potentially slashes the cost of running large language models by a factor of four to eight. At first glance, this appears to threaten demand for high-bandwidth memory chips.”
“Dramatically cheaper inference unlocks workloads previously too expensive to run”, Kwon said, pointing to real-time coding assistants and multiple AI agents running at once. The cheaper it gets, the more people try to do, and the memory bill finds a way.
Google’s researchers claim TurboQuant could cut memory usage by as much as sixfold, which is the sort of line that makes short sellers reach for the champagne. Analysts and researchers are now leaning the other way, arguing it could expand demand by making more workloads economically viable.
However, the supply side is now behaving as if it expects pain to last, with chipmakers leaning into longer contracts to gain visibility and lock customers in. Samsung, co-chief executive Jun Young-hyun recently said the company was pursuing “contracts of three or five years with major clients, shifting from the existing quarterly and annual terms”.
SemiAnalysis analyst Ray Wang said, “The market has largely misread TurboQuant,” and he is framing it as an acceleration story, not a collapse. We continue to believe that increasing memory demand will be required for both training and inference as AI models evolve and innovation advances.”
Wang said longer contracts should cushion any shock to the South Korean suppliers as AI service providers try to secure supply. “Memory is becoming a bit less cyclical, driven by accelerating and sustainable AI demand. Contract pricing now matters more than spot pricing.”
There is a practical limit to how much TurboQuant can change the near-term picture, because shortages ease when new capacity comes online, not when Twitter decides it has found a magic algorithm. The more credible forecast is that demand stays high for several quarters, with relief tied to how quickly suppliers can add production lines.
TurboQuant researcher Han In-su told the FT the algorithm “can serve as a foundation for realising previously impossible high-difficulty tasks, such as processing much longer contexts within limited memory resources without sacrificing accuracy, or implementing high-performance AI on smaller devices”.







