Nvidia Is Now Shipping a Groq Chip — and It's 4x Faster
Nvidia's Groq 3 LPX is in full production: 3,400 tokens/sec on 100K-context runs, 4x the nearest rival. Nebius takes it first.

The strangest line in Nvidia's Vera Rubin launch: one of the seven chips now in full production is called the Groq 3 LPX. Nvidia's flagship platform ships with its former inference rival's LPU branding baked in, sitting alongside the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9 SuperNIC, BlueField-4 DPU and Spectrum-6 Ethernet.
It exists for exactly one job: token generation. Agents reason one token at a time, so decode latency — not raw training FLOPs — decides whether a tool-calling loop feels instant or unusable. Nvidia's number: 3,400 output tokens/sec running Gemma 4 31B at a 100,000-token context, which it claims is 4x the nearest alternative platform. Nebius is the first AI cloud to adopt it, with Vera Rubin NVL72 landing in US and European data centers from H2 2026. AWS, Google Cloud, Microsoft, OCI, CoreWeave, Lambda and Nscale queue up behind it.
The framing is the real signal. Nvidia spent a decade selling training capacity; the headline metric here is cost per token — the exact number inference startups spent years using against it.
Why it matters: long-context agents live or die on decode speed, and Nvidia just moved to own that number too.
Sources
Primary: the company, paper or repository
- NVIDIA Advances Vera Rubin Inference With New LPX and CPX Platforms blogs.nvidia.com
Independent coverage
Written by an AI pipeline from the sources above. Methodology · Report an error
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.