Get the app
Computing

Dev reimplements Google's TurboQuant solo — 35B on a MacBook

Google dropped TurboQuant (no code). A Texas dev rebuilt it from the math using Claude Code — 4.6x KV cache compression, 35B on Apple Silicon.

Dev reimplements Google's TurboQuant solo — 35B on a MacBook

Google published TurboQuant (ICLR 2026) — a technique that compresses LLM KV caches to 3-bit with zero accuracy loss, claiming 6x memory reduction and 8x attention speedup. They released no code.

Tom Turney (Psyguard.ai), using Claude Code and Codex as coding partners, built a native llama.cpp implementation straight from the paper's math. His turboquant_plus hit 4.57–5.12x KV cache compression with full decode speed parity — plus a self-authored Sparse V optimization that adds +22.8% throughput at 32K context, outperforming Google's own benchmarks. Separately, Prince Canuma ran Qwen3.5-35B on Apple Silicon via MLX with 4.9x compression.

The community has now shipped at least 4 open-source implementations from a paper Google dropped with zero code — in under a week.

Why it matters: AI coding assistants are closing the gap between closed research and open-source shipping — a paper without code is no longer a moat.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play