Get the app

GPT-6.1 Sol Lands: OpenAI Brings Near-Astra Agent Performance to the $2 Tier

OpenAI’s rapid GPT-6.1 Sol release delivers 75.2% on DeepSWE and 71.4% on OSWorld at $2/$10 pricing, resetting the economics of multi-step agentic workflows.

OpenAI has officially launched GPT-6.1 Sol (gpt-6.1-sol), fundamentally upending the pricing curve for frontier reasoning and autonomous software engineering. Arriving just weeks after the flagship GPT-6 Astra debuted at $10/$50 per million tokens, the new 6.1 Sol model matches or nears Astra-class benchmarks across coding, computer use, and multi-step tool execution at one-fifth the list price: $2.00 per million input tokens and $10.00 per million output tokens.

The release signals a rapid compression cycle between flagship frontier capabilities and high-throughput developer tiers. For engineering teams running autonomous coding agents and multi-turn computer-use harnesses that consume millions of tokens per task, GPT-6.1 Sol reshapes the unit economics of production AI.


Benchmark Breakdown: Where 6.1 Sol Closes the Gap

Self-reported evals from OpenAI illustrate how aggressively 6.1 Sol encroaches on top-tier territory, particularly on benchmarks requiring tool orchestration, terminal execution, and environment feedback:

  • DeepSWE v1.1 (Agentic Coding): GPT-6.1 Sol scores 75.22% at high reasoning effort, slightly edging out GPT-6 Astra's baseline 74.10% and significantly outpacing previous generation Sol models.
  • OSWorld 2.0 (Computer Use): Reaches 71.42% on offline partial evaluation at max reasoning effort, trailing Astra (72.60%) by barely a single percentage point.
  • Terminal-Bench Science 0.1: Lands at 57.02% at max effort. Here, Astra’s raw scale retains a clear lead at 68.10%, showing that frontier compute still matters on complex scientific synthesis and mathematical proofs.
  • AutomationBench 1.0.6: Achieves 36.10% at max effort, demonstrating reliable long-horizon script generation and API automation.
  • GDP.pdf (Professional Document Work): Clocks 32.00% at high effort on dense enterprise PDF ingestion and analytical synthesis.

While Astra remains OpenAI's flagship for raw academic, biological, and mathematical discovery, GPT-6.1 Sol captures roughly 90–95% of Astra’s practical agentic capabilities at an 80% discount.


The Reasoning Ladder: Why More Thinking Isn't Always Better

With GPT-6.1 Sol, OpenAI has introduced a formal five-tier reasoning dial inside the API via the reasoning.effort parameter: low, medium (API default), high, xhigh, and max. Notably, OpenAI has removed the legacy none and minimal rungs entirely—confirming that the model's core architecture is permanently routed through an adaptive chain-of-thought engine.

Crucially, OpenAI’s benchmark disclosures reveal non-monotonic scaling across reasoning effort tiers:

  • On DeepSWE v1.1, the model peaks at 75.22% on high, dropping to 71.90% on xhigh and max (medium scored 73.01%, and low reached 64.38%). Over-allocated reasoning tokens in recursive agent loops can introduce overthinking, redundant patching attempts, and tool timeouts.
  • On OSWorld 2.0 and Terminal-Bench Science, capability scaled monotonically, requiring max effort (at ~$5.47 per task on TB-Science) to reach peak performance.

For systems engineers, setting explicit reasoning tiers per task type—rather than defaulting to maximum allocation—will be critical to optimizing both latency and task success rates.


Context Architecture and Cache Economics

GPT-6.1 Sol matches Astra’s massive 1,050,000-token input context window and provides up to 128,000 maximum output tokens. However, OpenAI has introduced specific architectural tiers to manage memory pressure across long-context workloads:

  • Standard Pricing: $2.00 / 1M input tokens, $10.00 / 1M output tokens.
  • Prompt Caching: Cached inputs drop to $0.10 / 1M tokens (a 95% reduction), with cache write costs fixed at $2.50 / 1M tokens.
  • Long-Prompt Surcharge: Prompts exceeding 272,000 input tokens trigger a dynamic scaling tier, applying a 2× multiplier on input/cache and a 1.5× multiplier on output across the full payload.
  • Batch API: 50% discount on standard list rates for asynchronous queues.

This makes GPT-6.1 Sol particularly well-suited for repository-scale codebases and document indices, provided prompts are carefully structured to leverage caching beneath the 272K-token threshold.


Native Tooling and the Responses API

GPT-6.1 Sol is optimized around OpenAI's unified Responses API, deprecating complex client-side tool loops in favor of platform-managed sandboxes. Supported capabilities include:

  • Hosted Shell & Apply Patch: Direct bash execution and unified diff parsing, cutting subagent roundtrips by more than 40%.
  • Native Model Context Protocol (MCP) Support: Built-in discovery and tool calling across local and remote MCP servers.
  • Computer Use & Screen Parsing: Direct coordinate targeting and OS navigation.
  • Skills & Code Interpreter: Dynamic multi-step Python execution with sandboxed memory persistence.

OpenAI has also previewed an upcoming Ultrafast variant of GPT-6.1 Sol, which claims up to 8× faster token generation for latency-critical agent execution loops.


What This Means for the AI Stack

The arrival of GPT-6.1 Sol accelerates a critical industry shift: frontier capability is no longer locked behind flagship pricing. By delivering Astra-grade software engineering and tool execution at $2/$10 per million tokens, OpenAI is directly pressuring the economics of autonomous developer tooling.

For production engineering teams, the default playbook for building autonomous coding assistants, document extraction pipelines, and automated ops agents has changed. The choice is no longer between expensive frontier intelligence and fast, shallow sub-tier models. With GPT-6.1 Sol, near-frontier autonomous execution is now a standard commodity tier.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play