Get the app

Anthropic's Claude Sonnet 5.5 Delivers a 7x Leap in Agentic Coding

Clocking 70.6% on Terminal-Bench 4.0 with 30% faster generation and a 1M token window, Sonnet 5.5 redefines the performance-per-dollar frontier.

Anthropic has rolled out Claude Sonnet 5.5, delivering an aggressive overhaul to the engine room of its model lineup. While the flagship Opus 5.5 was designed to hold the line on open-ended architectural reasoning, Sonnet 5.5 is optimized for execution speed, developer workflows, and long-horizon autonomy. The standout benchmark: Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, surging past its predecessor Sonnet 5’s 10.3% and outperforming even Opus 5.5 (66.4%) in deterministic command-line execution.

Combined with a 30%+ boost in generation throughput and token-efficiency gains that trim effective per-task compute costs by up to 30%, Anthropic’s mid-tier workhorse is setting a new efficiency benchmark for agentic coding and multi-step enterprise workflows.


The Numbers: Where Sonnet 5.5 Outpaces Its Weight Class

The benchmark disclosures in the Sonnet 5.5 release highlight a decisive shift in model behavior: Anthropic is prioritizing token economy and tight execution loops over verbose internal monologues.

Key performance highlights across coding, reasoning, and visual grounding include:

  • Agentic Coding (Terminal-Bench 4.0): Jumps from 10.3% to 70.6%, proving that Sonnet 5.5 avoids the terminal loop derailments and syntax errors that crippled earlier mid-tier models.
  • IDE Autonomy (CursorBench 4.0): Reaches 55.5%, compared to 34.1% for Sonnet 5 and 57.8% for Opus 5.5.
  • Code Synthesis (FrontierCode 1.1): Hits 46.2% at Max effort and 52.1% at Xhigh effort, directly challenging frontier flagships.
  • General Knowledge Work (GDPval-AA v2.1): Clocks an Elo score of 1844 across 44 professions, landing within two points of Opus 5.5 (1846) and leaving Sonnet 5 (1449) far behind.
  • Computer Use (OSWorld 2.1): Reaches 80.1% partial completion, up from 57.0% on Sonnet 5.
  • Multimodal Perception (Chartography): Scores 61.6% without external tools, a fourfold improvement over Sonnet 5’s 15.6%.
  • Autonomous Visual Agency: In a telling long-horizon benchmark, Sonnet 5.5 became the first Sonnet-tier model capable of completing Pokémon Red entirely end-to-end via raw visual screen captures.
Benchmark Comparison: Sonnet 5 vs. Sonnet 5.5 vs. Opus 5.5
┌───────────────────────┬────────────┬──────────────┬────────────┐
│ Benchmark             │  Sonnet 5  │  Sonnet 5.5  │  Opus 5.5  │
├───────────────────────┼────────────┼──────────────┼────────────┤
│ Terminal-Bench 4.0    │   10.3%    │    70.6%     │   66.4%    │
│ CursorBench 4.0       │   34.1%    │    55.5%     │   57.8%    │
│ GDPval-AA v2.1        │    1449    │     1844     │    1846    │
│ OSWorld 2.1 (Partial) │   57.0%    │    80.1%     │   81.8%    │
│ Chartography (Vision) │   15.6%    │    61.6%     │   64.4%    │
└───────────────────────┴────────────┴──────────────┴────────────┘

Adaptive Thinking, Token Economics, and Latency

Nominally, Anthropic has preserved Sonnet's pricing tier at $2.00 per million input tokens and $10.00 per million output tokens, with prompt cache writes priced at $2.50/MTok (5-minute TTL) or $4.00/MTok (1-hour TTL), and cache reads sitting at $0.20/MTok (a 90% read discount). On the Message Batches API, workloads receive a flat 50% discount across input and output tokens.

However, the effective bill for production workloads drops sharply due to Adaptive Thinking and token brevity:

  • Fewer Tokens per Task: Sonnet 5.5 reaches accurate stopping criteria significantly earlier than Sonnet 5, eliminating unnecessary rambling and cutting realized task costs by up to 30%.
  • Effort Sliders as Cost Levers: Rather than treating reasoning as a binary on/off switch, developers adjust an effort parameter (low, medium, high, xhigh, max). At low and medium effort, Sonnet 5.5 routinely matches or exceeds Sonnet 5’s peak benchmark scores at roughly one-tenth the total inference cost.
  • Throughput Gains: Generating tokens 30%+ faster than Sonnet 5, Sonnet 5.5 is explicitly tailored for synchronous tool-calling loops where agent latency directly degrades developer experience.

Critical API Changes and Developer Breaking Shifts

Migrating agentic scaffolding to claude-sonnet-5-5 requires noting several strict architectural constraints deployed in this generation:

  1. Sampling Parameter Restrictions: If reasoning or adaptive thinking is enabled, setting custom sampling parameters such as temperature, top_p, or top_k returns an immediate 400 Bad Request. Inference distribution is managed entirely via the thinking budget and effort level.
  2. Thinking Preserved Across Multi-Turn Tool Calls: Inter-tool deliberation tokens now stream back inside structured thinking blocks rather than raw text streams. Scaffolds parsing intermediate reasoning must inspect the thinking blocks or utilize between_tools to disable up-front preamble thinking.
  3. Deprecation of Legacy Computer Use Schemas: The older computer_20251124 tool namespace is blocked; integrations on AWS Bedrock, Google Cloud Vertex, and the Claude API must migrate to the standardized OSWorld-compatible computer use primitives.
  4. Context and Buffer Specs: Sonnet 5.5 ships with a default 1,000,000-token context window and supports up to 128,000 max output tokens in standard calls, stretching to 300,000 output tokens via the Batch API with the output-300k-2026-03-24 beta header.

Enterprise Cloud Rollout and Ecosystem Impact

Alongside immediate availability on the Claude API, Sonnet 5.5 has launched across major hyperscaler environments, including Amazon Bedrock, Google Cloud, Microsoft Foundry, and Snowflake Cortex AI.

Within enterprise stacks like Snowflake, Sonnet 5.5 is being deployed as the runtime engine for data-aware code generation (Snowflake CoCo) and autonomous analytic agents (Snowflake CoWork), sustaining multi-step data pipelines and SQL synthesis inside customer virtual private clouds without exposing raw telemetry.

Anthropic also confirmed that Claude Haiku 5.5 will round out the 5.5 family in the coming weeks, targeting high-volume micro-tasks and low-latency edge deployment.

For engineers building autonomous developer agents, the tactical calculus has shifted: Opus 5.5 remains reserved for high-stakes system architecture and open-ended research, but for interactive CLI manipulation, rapid PR resolution, and high-frequency tooling loops, Sonnet 5.5 is currently the strongest price-to-performance workhorse available.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play