Dev reimplements Google's TurboQuant solo — 35B on a MacBook
Google dropped TurboQuant (no code). A Texas dev rebuilt it from the math using Claude Code — 4.6x KV cache compression, 35B on Apple Silicon.

Google published TurboQuant (ICLR 2026) — a technique that compresses LLM KV caches to 3-bit with zero accuracy loss, claiming 6x memory reduction and 8x attention speedup. They released no code.
Tom Turney (Psyguard.ai), using Claude Code and Codex as coding partners, built a native llama.cpp implementation straight from the paper's math. His turboquant_plus hit 4.57–5.12x KV cache compression with full decode speed parity — plus a self-authored Sparse V optimization that adds +22.8% throughput at 32K context, outperforming Google's own benchmarks. Separately, Prince Canuma ran Qwen3.5-35B on Apple Silicon via MLX with 4.9x compression.
The community has now shipped at least 4 open-source implementations from a paper Google dropped with zero code — in under a week.
Why it matters: AI coding assistants are closing the gap between closed research and open-source shipping — a paper without code is no longer a moat.
Sources
- TurboQuant: Redefining AI Efficiency with Extreme Compression research.google
- turboquant_plus — llama.cpp implementation by Tom Turney github.com
- TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate (arXiv) arxiv.org
Written by an AI pipeline from the sources above. How it works.
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.