Get the app
LLMs

Grok 4.6 is cheap — until your agent crosses 200K tokens

SpaceXAI's new frontier model matches GPT-5.6 Sol on benchmarks at a fraction of the price. The fine print doubles it.

Grok 4.6 is cheap — until your agent crosses 200K tokens

SpaceXAI (the company formerly called xAI) shipped Grok 4.6 on 12 August, built explicitly for long-running agents — multi-step research, whole-codebase work, rough idea to shipped v1. It lands at $2/M input, $6/M output with a 500K-token context window, and scores 61 on the Artificial Analysis Intelligence Index — a statistical tie with GPT-5.6 Sol Max, one point behind Fable 5 Max at 62. Claude Opus 5 still tops the index.

Here's the part the hype posts skip: cross 200K prompt tokens and xAI rebills the entire request at $4/$12 per million. That's precisely the workload Grok 4.6 is marketed for — agent loops accumulate context by design. Against GPT-5.6 Sol's $5/$30, the real saving is closer to 60–80%, and only below the threshold. A faster variant costs double again, with no separate model ID.

On code it's improved but not leading: DeepSWE v1.1 at 65.9% (up 11.9 points generationally, still behind Sol Max's 73%) and Terminal-Bench v3.0 at 26% — nearly double Grok 4.5, yet last among the four models compared. Its genuine win is knowledge work: 1753 Elo on GDPval-AA v2, top of the field. Available now via the xAI API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare. No open weights.

Why it matters: frontier intelligence is converging on parity, so the fight has moved to price-per-token — and to whose fine print you read before wiring up an agent.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play