Get the app

Inside Grok 4.7: xAI Freezes Prices While Chasing Long-Horizon Agent Benchmarks

xAI's Grok 4.7 holds rates at $2/$6 per million tokens while taking on engineering and legal workloads—but surging token volume reveals the true cost of 'xHigh' reasoning.

xAI has officially deployed Grok 4.7, its latest reasoning and agentic foundation model, across the xAI API, Cursor, GitHub Copilot, and major model routers. Rather than competing purely on raw parameter scale or claiming across-the-board benchmark supremacy, xAI has locked its pricing floor: holding rates identical to Grok 4.6 at $2.00 per million input tokens and $6.00 per million output tokens for standard context windows under 200,000 tokens.

While frontier flagships like Claude Fable 5.1 Max ($10/$50 per million) and GPT-5.6 Sol Max ($4/$20 per million) continue to demand premium inference budgets, Grok 4.7 attempts to deliver competitive multi-hour agent performance at a fraction of the cost per token. However, independent audits and architectural breakdowns reveal a more nuanced story about token expansion, reasoning effort levels, and task-specific domain advantages.

The Architecture of Long-Horizon Reasoning

xAI engineered Grok 4.7 around an extended reinforcement learning (RL) training regime tailored specifically for multi-step agentic workflows. Instead of optimizing primarily for conversational turn-taking or single-prompt question answering, the training mix was heavily weighted toward tasks spanning multi-hour execution loops—such as iterative software debugging, code execution verification, and complex document analysis.

Key architectural specifications include:

  • Context Window: 500,000 tokens total capacity (with $4.00/$12.00 pricing above 200k tokens).
  • Reasoning Effort Modifiers: Four distinct inference modes—low, medium, high, and xhigh—with high set as the default API setting.
  • Fast Variant: An optimized inference tier offering roughly 2× the output generation speed at a $4.00/$12.00 token rate.
  • Integrated Agent Tooling: Native support for function calling, structured schema outputs, integrated web search, live X platform retrieval, and sandboxed code execution.
  • Safety and Refusal Stack: 62.4% on LatchBio's biosafety benchmark alongside custom HackerBench cybersecurity evaluations.

Benchmark Performance: Outliers and Trade-offs

Grok 4.7 does not displace the absolute frontier on every general benchmark, but it posts dramatic outlier gains in specific structured disciplines.

Benchmark                  | Grok 4.7 (xHigh) | Grok 4.6 (High) | GPT-5.6 Sol (Max) | Fable 5.1 (Max)
---------------------------|------------------|-----------------|-------------------|----------------
CursorBench 4.0            | 46.3%            | 40.4%           | 41.7%             | 51.8%
DeepSWE v1.1               | 71.0%*           | 65.2%           | 72.7%             | 70.0%
EEBench (Electrical Eng.)  | 64.0%            | 53.0%           | 39.4%             | 56.4%
Terminal-Bench 4.0         | 38.0%            | 20.3%           | 37.3%             | 57.9%
Harvey Legal Agent Bench   | 19.6%            | 15.8%           | 2.5%              | 6.7%
AA Briefcase v1.1          | 1,657            | 1,546           | 1,487             | 1,678
HealthBench Professional   | 56.7%            | 48.5%           | 60.5%             | 62.1%
*DeepSWE reported at High effort.

Where Grok 4.7 Surges

  • Specialized Technical & Legal Domains: Grok 4.7 achieved 19.6% on the Harvey Legal Agent Benchmark, nearly eight times higher than GPT-5.6 Sol Max (2.5%) and almost triple Claude Fable 5.1 Max (6.7%). On EEBench (electrical engineering reasoning), it reached 64.0%, outperforming both Grok 4.6 (53.0%) and Fable 5.1 Max (56.4%).
  • Software Engineering Cost-Efficiency: In CursorBench 4.0 evaluation data, Grok 4.7 xHigh achieved 46.3% at an average task cost of $6.01, outscoring GPT-5.6 Sol Max (41.7% at $8.23 per task). Claude Fable 5.1 Max retained the absolute accuracy lead at 51.8%, but at an average task cost of $17.28—nearly 3× Grok 4.7's per-task footprint.

The Token Burn Catch: 'xHigh' Effort vs. Task Cost

While xAI held per-token list prices steady, early independent evaluations highlight that raw per-token rates do not tell the entire financial story. Multi-token reasoning mechanisms burn context rapidly when pushed to maximum deliberative settings.

In evaluation suites run by Artificial Analysis (Index v4.3.2), Grok 4.7 on xHigh reasoning generated approximately 240 million output tokens across the full evaluation suite, compared to just 94 million tokens generated by Grok 4.6 on High effort. Across the identical test harness, the suite-wide intelligence score moved modestly from 44 to 46.

When a model generates 2.5× more output tokens to navigate complex logic gates, the real-world cost savings of a $6/million output rate can narrow quickly against competitors whose default reasoning chains are denser and shorter. Developers building production pipelines must benchmark real task completion rates and token volumes rather than assuming linear savings from token rate cards.

The Emerging 'Routing Stack' Strategy

The arrival of Grok 4.7 reinforces a dominant pattern in modern AI engineering: the obsolescence of single-model monoliths.

Instead of selecting a single foundation model for an enterprise pipeline, production agent architectures are shifting toward dynamic semantic routing:

  • Tier 1 (High-Volume Sub-agents & Specialized Extraction): Routing long-running background tasks, repo-wide parsing, electrical engineering diagrams, and legal contract scanning to models like Grok 4.7 (high or medium effort) keeps inference budgets sustainable.
  • Tier 2 (High-Stakes Planning & Open Terminal Execution): Routing multi-file terminal orchestration or ambiguous architectural redesigns to models like Claude Fable/Opus or GPT-6 Astra, where frontier terminal benchmarks (55%–87%) justify the higher token cost.

By pricing Grok 4.7 aggressively while tuning its RL weights toward multi-hour autonomous execution, xAI has secured a compelling position in the agentic cost-performance curve—provided engineers keep a close eye on their effort parameters.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play