OpenAI Drops GPT-5.6: Sol, Terra, Luna, and the Multi-Agent 'Ultra' Mode
OpenAI’s new flagship model family crushes Claude Fable 5 on agentic benchmarks, introducing parallel subagents and massive cost-efficiency gains.
OpenAI didn't just release a smarter model; they fundamentally changed how models scale compute during inference. The release of the GPT-5.6 family—comprising Sol, Terra, and Luna—shifts the battleground from raw parameter counts to agentic coordination, token efficiency, and parallel processing.
The New Lineup: Sol, Terra, and Luna
OpenAI has officially moved to a three-tier architecture, abandoning the old "Turbo" nomenclature for a more distinct stratification of capabilities tailored to specific enterprise and developer needs:
- GPT-5.6 Sol: The flagship model. Built for complex reasoning, long-horizon agentic workflows, and deep scientific research. It is OpenAI's most capable model to date, yet it operates at half the cost of GPT-5.5. Sol is the model you reach for when the task is genuinely hard—intricate code generation, vulnerability research, or nuanced creative analysis.
- GPT-5.6 Terra: The mid-tier workhorse. Designed to balance capability and cost, Terra is positioned for everyday coding and reasoning tasks. Remarkably, Terra outperforms Anthropic's Claude Fable 5 at roughly one-sixteenth the estimated cost, making it a highly disruptive offering for enterprise deployments.
- GPT-5.6 Luna: The lightweight, high-volume model. Optimized for latency-sensitive tasks like chat routing and classification, Luna outperforms Claude Opus 4.8 while being drastically cheaper and faster. It is designed for high-frequency workflows where speed and cost-per-query are the primary metrics.
Scaling Inference: 'Max' Reasoning and 'Ultra' Mode
The most significant technical leap in GPT-5.6 isn't just the base model's intelligence—it's how OpenAI is exposing inference-time compute to developers. We are seeing the realization of test-time compute scaling in a production environment.
GPT-5.6 introduces a new max reasoning effort, giving the model extended time to explore alternatives, run internal checks, and revise its approach before outputting a final token. But the real game-changer is ultra mode.
Instead of relying on a single monolithic agent, ultra mode coordinates four subagents in parallel by default. This multi-agent architecture trades higher token consumption for significantly stronger results and faster time-to-resolution on demanding tasks. Across all evaluations, adding parallel agents shifts the score-latency frontier upward and to the left. In the API, developers can build these ultra-like experiences using the new multi-agent beta in the Responses API, effectively democratizing complex agentic orchestration.
Crushing the Agentic Benchmarks
OpenAI is aggressively targeting Anthropic's recent dominance in agentic workflows. The benchmark results for GPT-5.6 Sol are staggering, particularly in environments that require planning, tool coordination, and iteration over long time horizons.
- Agents’ Last Exam: On this rigorous evaluation of long-running professional workflows across 55 fields, GPT-5.6 Sol sets a new high score of 53.6, eclipsing Claude Fable 5 (with adaptive reasoning) by a massive 13.1 points. Even at medium reasoning, Sol beats Fable 5 by 11.4 points at a quarter of the cost.
- Terminal-Bench 2.1: Testing complex command-line workflows, Sol Ultra hits 91.9%, and base Sol hits 88.8%, comfortably beating Claude Mythos 5 (84.3%) and Claude Fable 5 (82.5%).
- Artificial Analysis Coding Agent Index: Sol with
maxreasoning sets a new state-of-the-art at 80 (2.8 points above Fable 5), while using less than half the output tokens and costing about one-third less. - OSWorld 2.0: Sol hits 62.6%, surpassing Opus 4.8 while using 85% fewer output tokens.
In the cybersecurity domain, Sol is shifting the performance-efficiency frontier. On ExploitBench and UC Berkeley's ExploitGym, Sol demonstrates capabilities competitive with Mythos Preview but uses only a third of the output tokens. OpenAI notes that Sol is heavily safeguarded—it excels at finding and patching vulnerabilities (defensive blue-teaming) but is constrained from executing end-to-end offensive attacks. It deliberately does not cross the Cyber Critical threshold under OpenAI's Preparedness Framework.
Programmatic Tool Calling: The Unsung Hero
For developers building agentic systems, the most impactful feature might be Programmatic Tool Calling.
Historically, tool-heavy tasks required developers to script every step or pass massive amounts of intermediate tool responses back through the LLM context window, bloating costs and latency. GPT-5.6 can now write and run lightweight programs in-memory that coordinate tools, process intermediate results, monitor progress, and choose the next action autonomously as work unfolds.
This filters out the noise, retaining only the data that matters. Because this happens in-memory, it is Zero Data Retention (ZDR) compatible, making it highly attractive for enterprise, financial, and legal workflows. OpenAI also introduced more predictable prompt caching, including support for explicit cache breakpoints and a guaranteed 30-minute minimum cache life.
Accelerating AI Research and Scientific Discovery
OpenAI isn't just selling GPT-5.6; they are using it to build the next generation of models. Inside OpenAI, researchers are deploying Sol across the development loop: diagnosing failures, optimizing training systems, running machine-learning experiments, and interpreting results.
On an internal bundle of evaluations measuring progress towards recursive self-improvement (the RSI Index), GPT-5.6 Sol demonstrated a 16.2-point improvement over GPT-5.5. This means the model is actively accelerating internal research across the board, hinting at a compounding loop of AI development.
In external scientific domains, Sol shows broad improvements in biology workflows. On GeneBench v1, which evaluates long-horizon genomics and quantitative-biology analyses, it achieves stronger results than GPT-5.5 while using fewer tokens. In financial and legal sectors, early testers report that Sol improves rubric quality and answer accuracy while cutting prompt tokens by up to 38% for multi-step document analysis.
The Vibe Check: How It Actually Feels
Benchmarks are one thing, but how does GPT-5.6 Sol feel in the trenches? Early reports from power users and developers highlight its persistence and steerability.
Unlike earlier models that would hallucinate or give up when a tool failed, Sol is remarkably tenacious. It searches available files, reads standing project instructions, and debugs its own errors without constantly prompting the user for intervention. It feels less like a chat assistant and more like an end-to-end technical operator.
However, its default writing style has some quirks. Sol tends to write in tight, compact paragraphs but reaches for longer, more abstract vocabulary. In early tests, it scored the highest Flesch-Kincaid reading difficulty score among frontier models. It isn't the best model for zero-shot creative writing, but it is an exceptional collaborator. When provided with a style guide, context files, and a clear system prompt, Sol adapts its tone flawlessly, making it a highly steerable engine for knowledge work.
The Bottom Line
With GPT-5.6, OpenAI has effectively commoditized the mid-tier agentic market. By offering Terra and Luna at fractions of the cost of competing models while matching or exceeding their performance, OpenAI is putting immense pressure on Anthropic, Google, and DeepSeek.
The introduction of ultra mode proves that the future of LLMs isn't just larger parameter counts—it's native, parallel multi-agent orchestration and test-time compute scaling. The frontier has moved again, and right now, it firmly belongs to Sol.
Sources
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.