Get the app
Computing

Nvidia Is Now Shipping a Groq Chip — and It's 4x Faster

Nvidia's Groq 3 LPX is in full production: 3,400 tokens/sec on 100K-context runs, 4x the nearest rival. Nebius takes it first.

Nvidia Is Now Shipping a Groq Chip — and It's 4x Faster

The strangest line in Nvidia's Vera Rubin launch: one of the seven chips now in full production is called the Groq 3 LPX. Nvidia's flagship platform ships with its former inference rival's LPU branding baked in, sitting alongside the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9 SuperNIC, BlueField-4 DPU and Spectrum-6 Ethernet.

It exists for exactly one job: token generation. Agents reason one token at a time, so decode latency — not raw training FLOPs — decides whether a tool-calling loop feels instant or unusable. Nvidia's number: 3,400 output tokens/sec running Gemma 4 31B at a 100,000-token context, which it claims is 4x the nearest alternative platform. Nebius is the first AI cloud to adopt it, with Vera Rubin NVL72 landing in US and European data centers from H2 2026. AWS, Google Cloud, Microsoft, OCI, CoreWeave, Lambda and Nscale queue up behind it.

The framing is the real signal. Nvidia spent a decade selling training capacity; the headline metric here is cost per token — the exact number inference startups spent years using against it.

Why it matters: long-context agents live or die on decode speed, and Nvidia just moved to own that number too.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play