Grok 4.6 is cheap — until your agent crosses 200K tokens
SpaceXAI's new frontier model matches GPT-5.6 Sol on benchmarks at a fraction of the price. The fine print doubles it.

SpaceXAI (the company formerly called xAI) shipped Grok 4.6 on 12 August, built explicitly for long-running agents — multi-step research, whole-codebase work, rough idea to shipped v1. It lands at $2/M input, $6/M output with a 500K-token context window, and scores 61 on the Artificial Analysis Intelligence Index — a statistical tie with GPT-5.6 Sol Max, one point behind Fable 5 Max at 62. Claude Opus 5 still tops the index.
Here's the part the hype posts skip: cross 200K prompt tokens and xAI rebills the entire request at $4/$12 per million. That's precisely the workload Grok 4.6 is marketed for — agent loops accumulate context by design. Against GPT-5.6 Sol's $5/$30, the real saving is closer to 60–80%, and only below the threshold. A faster variant costs double again, with no separate model ID.
On code it's improved but not leading: DeepSWE v1.1 at 65.9% (up 11.9 points generationally, still behind Sol Max's 73%) and Terminal-Bench v3.0 at 26% — nearly double Grok 4.5, yet last among the four models compared. Its genuine win is knowledge work: 1753 Elo on GDPval-AA v2, top of the field. Available now via the xAI API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare. No open weights.
Why it matters: frontier intelligence is converging on parity, so the fight has moved to price-per-token — and to whose fine print you read before wiring up an agent.
Sources
Independent coverage
- SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work marktechpost.com
- SpaceXAI debuts Grok 4.6, matching GPT-5.6 Sol on Artificial Analysis venturebeat.com
Written by an AI pipeline from the sources above. Methodology · Report an error
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.