Claude Sonnet 5.5 Arrives: Terminal-Bench at 70.6% and a Massive Free-Tier Upgrade
Anthropic drops Claude Sonnet 5.5, delivering 30% faster execution, near-Opus benchmark parity, and unprecedented agentic coding power directly to its free tier.
The battle for frontier AI is no longer being waged solely at the theoretical apex—it has moved decisively into the workhorse tier. Anthropic has released Claude Sonnet 5.5, delivering a dramatic generational leap that brings everyday API economics into near parity with flagship frontier models. Running at the existing $2.00 input / $10.00 output per million tokens pricing floor ($0.20 per million cached tokens), Sonnet 5.5 generates tokens more than 30% faster than Sonnet 5 while reducing total token consumption per task by up to 30%.
Crucially, Anthropic has swapped Sonnet 5.5 into the free tier of claude.ai—a calculated move that severely pressures OpenAI's free-tier GPT-6 Luna offering by giving non-paying users access to state-of-the-art coding and agentic performance.
The Agentic Coding Surge: Terminal-Bench and CursorBench
While Sonnet 5 was widely adopted as a reliable intermediate model, it lagged behind Opus on extended terminal interaction and multi-step tool execution. Sonnet 5.5 completely upends that dynamic:
- Terminal-Bench 4.0: Sonnet 5.5 achieved 70.6%, up from Sonnet 5's 10.3%. Notably, this edges past Claude Opus 5.5's 66.4% (at Xhigh effort), marking the first time a mid-tier Sonnet release has outpaced an Opus-tier model on an agentic terminal benchmark.
- CursorBench 4.0: Recreating real developer sessions inside IDE workflows, Sonnet 5.5 posted 55.5%, compared to 34.1% for Sonnet 5 and just two points behind Opus 5.5’s 57.8%.
- SWE-bench Pro & DeepSWE: Sonnet 5.5 scored 81.3% on SWE-bench Pro and 71.0% on DeepSWE v1.1, demonstrating robust repository-level autonomous debugging.
- Tool Call Batching: Early telemetry highlights that Sonnet 5.5 aggressively batches dependent tool invocations, shrinking multi-turn latency in automated developer loops.
Benchmark Comparison: Claude 5.5 Family & Competitors
───────────────────────────────────────────────────────────────────
Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
───────────────────────────────────────────────────────────────────
Terminal-Bench 4.0 70.6% 10.3% 66.4%¹ —
CursorBench 4.0 55.5% 34.1% 57.8% —
FrontierCode 1.1 (Xhigh) 52.1% 42.4% 54.4% 49.3%
GDPval-AA v2.1 (Elo) 1,844 1,449 1,846 1,487
AA Briefcase v1.1 (Elo) 1,811 1,359 1,822 1,483
Humanity's Last Exam (HLE) 64.5% 54.9% 67.7% —
OSWorld 2.1 (Partial) 80.1% 57.0% 81.8% —
Chartography (No tools) 61.6% 15.6% 64.4% 53.6%
───────────────────────────────────────────────────────────────────
¹ Opus 5.5 reported at Xhigh setting. Scores reflect self-reported vendor evaluations.
The Over-Decomposition Bug: When 'Max' Effort Hurts
One of the most technically intriguing findings in the Sonnet 5.5 release is a non-linear scaling anomaly on FrontierCode 1.1 Main. While Sonnet 5.5 scored 52.1% at the Xhigh thinking budget, its score degraded to 46.2% when configured to Max effort.
Anthropic’s post-mortem reveals that at maximum reasoning depth, the model exhibits hyper-active meta-deliberation: it spontaneously invokes internal code-review sub-agents, splitting tasks across sub-delegations. In complex benchmarks with strict timeouts and rigid modification boundaries, these sub-agents frequently exceeded wall-clock limits or introduced out-of-scope edits.
Developer testing by Simon Willison demonstrated a similar phenomenon in unbounded creative tasks: when prompted at Max thinking effort for complex WebGL rendering, Sonnet 5.5 consumed its entire 128,000-token output thinking window before terminating, whereas at Xhigh it produced fully functional, optimized code in 41 seconds.
Closing the Knowledge Gap: GDPval-AA and Visual Grounding
Beyond software engineering, Sonnet 5.5 virtually eliminates the performance tax historically associated with running mid-tier models on complex knowledge workflows:
- Professional Knowledge Work: On OpenAI’s GDPval-AA v2.1 benchmark (evaluating tasks across 44 professions), Sonnet 5.5 hit 1,844, just two points under Opus 5.5 (1,846) and nearly 400 points above GPT-6 Sol (1,487).
- Visual Diagramming: On Chartography, a pure-vision diagrammatic reasoning test without tool assistance, Sonnet 5.5 leapt from 15.6% to 61.6%, proving substantial improvements in multimodal token encoding.
- Autonomous UI Execution: In OSWorld 2.1, Sonnet 5.5 notched 80.1% on partial evaluations, closing within two points of Opus 5.5 (81.8%).
- Vision-Only Interactive Tasks: Anthropic demonstrated Sonnet 5.5 playing through Pokémon Red purely via visual screen state processing and keyboard API commands, a domain that previously required dedicated reinforcement learning pipelines.
Ecosystem Impact and Deployment Architecture
Sonnet 5.5 ships with a 1,000,000-token context window and a maximum generation ceiling of 128,000 output tokens. The model is live across the Anthropic API (claude-sonnet-5-5), Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Azure.
Key architectural and deployment details include:
- Adaptive Thinking: Enabled by default, with
higheffort set across developer endpoints andmediumacross client applications. - Thinking Block Interoperability: Sonnet 5.5 supports cross-model thinking block comprehension, ingesting reasoning traces from Sonnet 5, Opus 4.8, and Haiku 4.5.
- Opus-Grade Safeguards: Sonnet 5.5 is the first Sonnet variant equipped with high-tier cybersecurity guardrails and anti-distillation defense mechanisms previously reserved for Opus-class weights.
With Claude Haiku 5.5 slated for rollout in the coming weeks, Anthropic’s rapid tier refresh establishes a formidable barrier to entry. For enterprises balancing inference budgets against agentic autonomy, Sonnet 5.5 transforms the mid-tier from an operational compromise into the primary production default.
Sources
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.