by

Google’s TurboQuant squeezes LLMs

RAM prices are enough to make you choke on your toast, so Google Research has turned up with TurboQuant to cram LLMs into less memory.

TurboQuant is pitched as a compression trick for the key-value cache, which Google calls a “digital cheat sheet” that stops the model from recomputing the same junk every token.

That matters because LLMs do not know anything; they just juggle vectors that map tokenised text into semantic space and fake competence when the numbers line up.

Those vectors can be brutally high-dimensional, packing hundreds or thousands of embeddings to describe complex data, and the cache balloons until performance hits a wall.

Developers normally fight bloat with quantisation, running lower precision maths and accepting uglier outputs when token estimates go wonky.

Google claims TurboQuant dodges the usual trade-off, showing an 8x performance jump and a 6x memory cut in some tests with no quality drop.

The method starts with PolarQuant, which no longer treats vectors as plain XYZ coordinates and instead converts them to polar coordinates in a Cartesian system. On that circular grid, a vector gets boiled down to radius for core strength and direction for meaning, keeping the important geometry while trimming the fat.

Google’s analogy is the sort you can explain to a project manager without crying. “Go 3 blocks East, 4 blocks North.” becomes “Go 5 blocks at 37-degrees.”

PolarQuant handles the heavy compression, but it can leave residual errors, so Google follows it with Quantised Johnson-Lindenstrauss (QJL).

QJL is a one-bit error-correction layer that reduces each vector to a single bit, +1 or -1, while preserving the relationships that drive attention scoring.

Google says it ran TurboQuant on long-context benchmarks using Gemma and Mistral open models and got perfect downstream results while cutting key-value cache memory by 6x.

It claims the cache can be quantised to just 3 bits with no extra training, so you can slap it onto existing models without a fresh optimisation circus.

For speed, Google says computing attention with 4-bit TurboQuant is 8x faster than 32-bit unquantised keys on Nvidia H100 accelerators, which is the sort of number marketing teams frame.

If it lands in real deployments, the savings could cut running costs or just tempt firms to stuff bigger models into the same box, and nobody has ever met spare memory they did not want to waste.

 

 

TOPICS:
google research  ·  key value cache  ·  llm memory  ·  mobile ai  ·  nvidia h100  ·  polarquant  ·  qjl  ·  quantization  ·  turboquant

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews