Get the app
LLMs

Grok 4.6 doesn't beat DeepSeek V4 Pro — and it costs 7x more

Both shipped Aug 12-13. Grok wins knowledge work, DeepSeek wins agents and security — at $0.87 vs $6 per million output tokens.

Grok 4.6 doesn't beat DeepSeek V4 Pro — and it costs 7x more

The viral claim is backwards. Grok 4.6 and DeepSeek V4 Pro landed within hours of each other, and they split the scoreboard rather than one sweeping it.

Grok takes the knowledge-work benchmarks: 1753 Elo on GDPval-AA v2, 1577 on AA-Briefcase, and 15.8% on Harvey LAB legal tasks — each edging out Fable 5. DeepSeek owns the agentic and security side: 83.3 on CyberGym, 31.8 on AutomationBench, 87.9 on Terminal-Bench 2.1, and 60.0 on Humanity's Last Exam with tools. Grok's own weak spot is terminal use — exactly where DeepSeek is strongest.

The cost line is where the post falls apart. DeepSeek V4 Pro runs $0.435 in / $0.87 out per million tokens. Grok 4.6 is $2 / $6, doubling to $4 / $12 past 200K prompt tokens. That's roughly 7x the output price, and DeepSeek ships a 1M-token context window against Grok's 500K. Grok's real pitch isn't "cheapest" — it's price-to-intelligence against GPT-5.6 Sol ($30) and Claude Opus 5 ($25), where holding flat at 4.5's pricing while jumping DeepSWE 54 → 65.9 is genuinely aggressive.

Why it matters: frontier parity is now a routing problem — pick per workload, not per vendor.

Sources

Independent coverage

Written by an AI pipeline from the sources above. Methodology · Report an error

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play