Google's TurboQuant shrinks LLM memory 6x with zero accuracy loss
Google's new KV cache compression algorithm cuts memory 6x and boosts inference up to 8x on H100s — no retraining needed.

Google Research just dropped TurboQuant, a KV cache compression algorithm that quantizes keys and values down to 3 bits without touching accuracy or requiring any fine-tuning. It's being presented at ICLR 2026.
The trick is a two-stage pipeline: PolarQuant first converts vectors into polar coordinates, separating magnitude from direction for efficient compression. Then QJL (Quantized Johnson-Lindenstrauss) uses a single residual bit to eliminate quantization bias — acting as a mathematical error-corrector. Together, they hit 6x memory reduction and up to 8x throughput improvement on Nvidia H100s.
Benchmarks on Gemma and Mistral across LongBench, RULER, and Needle-in-a-Haystack show no degradation — even at extreme compression ratios. Google says it also works for vector search, not just LLM inference.
Why it matters: KV cache is the main memory bottleneck for long-context LLMs — shrinking it 6x without accuracy tradeoffs is a direct unlock for longer contexts and cheaper inference at scale.
Sources
Independent coverage
- TurboQuant: Redefining AI Efficiency with Extreme Compression research.google
- Google's TurboQuant reduces LLM cache memory by 6x — Tom's Hardware tomshardware.com
Written by an AI pipeline from the sources above. Methodology · Report an error
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.