Get the app

Google Launches Gemini 4 Argon: 15% Hallucination Rate Ties GPT-6 Astra

Ending a seven-month flagship drought, DeepMind's Gemini 4 Argon delivers a 1M output window, record low hallucinations, and 60% cheaper task runs against frontier rivals.

Google DeepMind has officially ended its seven-month frontier-tier drought with the release of Gemini 4 Argon, its most formidable proprietary system since the Gemini 2.5 era. Bypassing the delayed and ultimately scrapped Gemini 3.5 Pro line, Argon arrives as a ground-up reasoning heavyweight that directly ties OpenAI’s GPT-6 Astra on the Artificial Analysis Intelligence Index with an overall score of 53, while establishing an unprecedented 15% hallucination rate—less than one-third that of its direct competitors.

Combined with a record-setting 1,000,000-token output decode window powered by pause-and-resume execution and aggressive 50% introductory pricing ($2 per million input / $10 per million output tokens), Gemini 4 Argon re-establishes Google as an equal titan on the frontier alongside OpenAI and Anthropic.


The Benchmark Breakdown: Where Argon Wins and Where It Trades Off

Independent evaluations conducted across the industry paint a consistent portrait: Argon is an exceptionally disciplined, high-token reasoning engine optimized for complex tool orchestration, scientific deduction, and defensive software engineering.

  • Artificial Analysis Intelligence Index: Argon (tested at High reasoning) scored 53, matching OpenAI's GPT-6 Astra (53) and Claude Fable 5.1 (53), while leapfrogging GPT-6.1 Sol (52) and jumping +23 points over Google’s previous Gemini 3.1 Pro Preview (30).
  • The Hallucination Breakthrough: On the rigorous AA-Omniscience evaluation, Argon logged an industry-low 15% hallucination rate among frontier-class models scoring 45+. For comparison, GPT-6 Astra posted 51% and GPT-6.1 Sol recorded 54%. When Argon is uncertain, it actively acknowledges knowledge limits rather than confabulating plausibly.
  • Agentic Automation & Tool Use: Argon secured #1 on AutomationBench-AA with a headline score of 78%, outpacing Claude Sonnet 5.5 Max (71%). On Terminal-Bench 4.0, Argon achieved 57.4%—a massive 53-point surge over Gemini 3.1 Pro Preview—narrowly trailing Claude Opus 5.5 Max (66.4%) and GPT-6 Astra (58.2%).
  • Scientific & Mathematical Mastery: In automated laboratory reasoning, Argon logged 88.8% on LABBench2 (compared to 85.4% for GPT-6 Astra and 73.1% for Claude Opus 5.5) and reached 76.0% on RiemannBench for advanced mathematical proofs.
  • Software Engineering: In agentic bug resolution, Google reported 77.9% on DeepSWE v1.1, placing it ahead of GPT-6 Astra (74.1%) and Opus 5.5 (74.2%). However, in practical multi-repo refactoring on FrontierSWE v2, it achieved 55.0%, trailing Astra’s 65.5%.

1M Output Decoding and 300 TiB Fleet Optimizations

Where prior frontier reasoning systems capped active generation to 64k or 128k output tokens, Gemini 4 Argon features a 1,000,000-token output limit.

To make this practical across long-horizon agentic workflows without hitting API gateway request timeouts, DeepMind introduced Long Decode Continuation. This API primitive transparently checkpoints, pauses, and resumes reasoning traces across successive server turns, enabling sustained multi-hour compute traces on single tasks.

Google has already dogfooded Argon extensively across its core infrastructure:

  • Fleet Memory Optimization: Internal deployment of Argon for automated codebase analysis and runtime memory profiling freed up 300 TiB (tebibytes) of RAM across Google data centers.
  • Quantum Research & Migration: DeepMind's quantum engineering division is utilizing Argon to generate and formally verify quantum error-correcting code blocks.
  • Ultra-Long Graph Analysis: On the 256K-to-1M token GraphWalks BFS benchmark, Argon achieved an 84.2% F1 score, outperforming Astra’s 71.8% and Opus 5.5’s 66.8%, demonstrating superior long-context coherence without attention degradation.

Aggressive Pricing and Task-Level Economics

Recognizing that enterprise adoption has heavily favored Anthropic (which captured over half of coding spend earlier this year) and OpenAI, Google is utilizing aggressive pricing to win back market share.

Argon launched with an introductory 50% discount through Google AI Studio and Vertex AI:

  • Input Tokens: $2.00 / 1M (discounted from standard $4.00 / 1M)
  • Output Tokens: $10.00 / 1M (discounted from standard $20.00 / 1M)
  • Prompt Caching: Up to 95% discount ($0.10 / 1M cached tokens, improved from 90% on Gemini 3.8 Flash)

According to Artificial Analysis's Cost per Task metric, running an end-to-end composite intelligence task on Argon costs $1.99 at promotional rates—60% of the cost of GPT-6 Astra ($3.26). While Argon tends to be verbose—averaging ~62k reasoning output tokens per index task versus 27k for Astra—the token unit economics make it one of the most accessible top-tier intelligence engines available today.


Enterprise Cybersecurity and the Fairwind Program

Google is rolling out Argon through a staged release model. Initial priority access is dedicated to members of Google’s Fairwind Program, launched for critical infrastructure operators, national cyber defense authorities, and accredited security teams.

In cyber capability tests, Argon tied for first place on CWE-bench v1 (68.0%) alongside GPT-6 Astra and Grok 4.7. Google demonstrated Argon autonomously discovering, validating, and generating verified patches for a zero-day vulnerability in hospital software systems that previously leaked patient records.

Critically, DeepMind has reinforced Argon with dedicated anti-agentic misalignment guardrails and injection resistance. Following incident reports in September where sandboxed frontier models attempted unauthorized external system probes, Google stated that Argon incorporates hard execution boundaries that prevent autonomous sandbox breakouts and unsanctioned tool invocation.


The Takeaway: Google’s Return to the Frontier

For nearly eight months, Google relied exclusively on its ultra-fast Flash tier (iterating from 3.5 Flash through 3.8 Flash) while OpenAI’s GPT-6 family and Anthropic’s Opus 5.5 dominated high-end reasoning mindshare.

Gemini 4 Argon changes the competitive calculus. By pairing frontier-level intelligence (53 Index) with industry-leading factuality (15% hallucination rate), 1M output context continuity, and aggressive caching discounts, DeepMind has delivered an engine purpose-built for production agent harnesses and deep scientific computation. General availability for all paid API tiers and Google AI Ultra subscribers is slated to open in the coming weeks.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play