Get the app

The 0% Pass Rate Illusion: 38% of 'Frontier' AI Agent Tasks Are Just Broken

An audit of 28,801 Terminal-Bench runs reveals that broken oracles, flaking infrastructure, and gaming verifiers masquerade as unsolvable frontier reasoning challenges.

When an autonomous AI agent fails a software engineering or terminal task 100% of the time, the machine learning industry typically draws an immediate conclusion: the task has exposed a genuine frontier reasoning gap. That assumption is fundamentally broken. A landmark audit of the production record behind Terminal-Bench and Frontier-Bench demonstrates that nearly 38% of all-fail agent tasks do not measure model capability at all—they measure broken reference solutions, hostile infrastructure, missing dependencies, and exploitable test verifiers.

The paper, titled "What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus" (arXiv:2609.26826), by Edward Lue Chee Lip, Boden Moraski, Tim Knappe, Lang Xiong, Sarvesh Gharat, Antonio Mari, and Ivan Bercovich, investigates what zero pass rates actually certify. Rather than treating a leaderboard as ground truth, the researchers audited the entire frozen lifecycle of Terminal-Bench 3 / Frontier-Bench 0.1: 1,081 pull requests, 639 scored tasks, 28,801 trial trajectories, and $105,933 in logged agent spend.

Their findings deliver a sobering verdict on how AI labs construct, score, and optimize against autonomous agent benchmarks.


Dissecting the 125 'Unsolvable' Tasks

Frontier benchmarks require headroom to remain informative. When a new suite of models saturates older evals like SWE-bench or HumanEval, benchmark creators intentionally harvest harder, multi-step command-line tasks where contemporary frontier models score zero.

In the audited corpus, exactly 125 tasks recorded a 0% honest pass rate across all evaluated state-of-the-art models. To verify whether these tasks represented authentic frontier frontiers, the authors applied a rigorous, ordered validity screening pipeline:

  • Step 1: Reference Route Verification (Oracle Check). Does the authored reference script pass cleanly in the target container?
  • Step 2: Infrastructure Domination Check. Are failure trajectories dominated by platform timeouts, memory exhaustion (OOM), or network socket drops rather than agent logic errors?
  • Step 3: Verifier Bypass Audit. Does the task allow reward-hacking or reward-independent bypasses that invalidate test assertions?
  • Step 4: Solvability Certification. Is the prompt sufficiently specified with necessary context, credentials, and environment states for a competent engineer to solve it?

Out of the 125 "unsolved" tasks, only 78 survived as certified-unsolved candidates. The remaining 47 tasks (37.6%) collapsed under basic scrutiny:

  • 14 tasks had broken oracles: The reference solutions provided by task authors failed their own verifiers. In these cases, even an omniscient agent executing the exact intended patch or shell sequence would register a failure.
  • 8 tasks were dominated by infrastructure failures: The tasks required resources or execution times exceeding container limits, causing silent process kills before the models could finish reasoning.
  • 4 tasks were passable only through verifier bypasses: The task verifiers contained structural loopholes where only arbitrary file tampering—unrelated to the task goal—could toggle a pass.
  • 21 tasks lacked solvability certification: The instructions were so ambiguous, or omitted critical environmental context, that no valid single-route completion could be verified from the telemetry.
 Audited All-Fail Tasks Breakdown (n = 125)
 ┌──────────────────────────────────────────┬────────┬──────────┐
 │ Category                                 │ Count  │ Share    │
 ├──────────────────────────────────────────┼────────┼──────────┤
 │ Certified-Unsolved (Genuine Candidates)  │   78   │  62.4%   │
 │ Solvability Uncertified (Context Gap)    │   21   │  16.8%   │
 │ Broken Oracles (Reference Code Failed)   │   14   │  11.2%   │
 │ Infrastructure Dominated (OOM / Timeouts)│    8   │   6.4%   │
 │ Verifier Bypass Dependent                │    4   │   3.2%   │
 └──────────────────────────────────────────┴────────┴──────────┘

Why 'Fake-Hardness' Corrupts RL and Post-Training

The consequences of unvalidated failure suites extend far beyond leaderboard aesthetics. Reinforcement learning from autonomous feedback (RLAF) and agent search policies heavily optimize against difficult, unsolved edge cases.

When post-training pipelines feed on tasks characterized by fake-hardness, two destructive dynamics emerge:

1. Rewarding Degenerate Exploration

When an environment has a broken reference solution or missing context, no amount of logical chain-of-thought planning can solve it legitimately. Instead, reinforcement learning algorithms reward agents that engage in unhinged search patterns—inspecting environment variables, executing blind brute-force file sweeps, or exploiting subtle quirks in bash execution runtimes.

2. The Illusion of Saturation and Progress

When benchmark maintainers eventually patch infrastructure or repair a broken test harness, the subsequent jump in model pass rate is frequently misattributed to a generational breakthrough in AI reasoning. In reality, the underlying models have not gotten smarter; the test harness simply stopped crashing.

As the researchers emphasize: "A lack of saturation and genuine difficulty are not the same thing. Pass rate alone cannot determine whether the agent failed at the intended capability or whether the task failed as a measurement instrument."


The Anatomy of an Agentic Benchmark Failure

Unlike static question-answering benchmarks (like MMLU or GPQA) where inputs and outputs are isolated token sequences, agentic benchmarks operate inside stateful operating system environments. The audit shows that an agent failure can occur across at least five distinct, brittle layers:

  1. Instruction Ambiguity: The agent is asked to "fix the service configuration," but three distinct service files exist across the filesystem with no indication of which systemd unit is active.
  2. Ephemeral Network & Tool States: Dependency mirrors (e.g., PyPI, apt-get, cargo) fail or rate-limit mid-trial, causing non-deterministic crashes that are logged as reasoning failures.
  3. State Leakage & Dirty Containers: Artifacts left behind from failed test attempts or authoring sessions interfere with subshell executions.
  4. Flawed Verification Oracles: Shell assertions that check process status via grep instead of exit codes, causing verifier crashes when agent output formats vary.

Even among the 78 tasks certified as unsolved, the authors issue an important caveat: "certified-unsolved" merely guarantees that the authored route passed, infrastructure was stable, and all evaluated models failed. It does not prove that the task isolates a pure reasoning deficiency, nor does it prove that the verifier is complete.


The New Standard for Frontier Evals

With over $100,000 spent evaluating single benchmark snapshots across hundreds of PR iterations, the cost of faulty evaluations is unsustainable. The paper advocates for fundamental shifts in how the AI community designs autonomous agent testbeds:

  • Mandatory Dual-Run Oracle Certification: Every task must prove that its reference solution passes repeatedly across cold-booted, isolated containers under constrained resources before entering evaluation splits.
  • Empty-Solution & Adversarial Baselines: Benchmarks must run empty-string, no-op, and randomly mutating agents against all verifiers to catch permissive bypasses before deployment.
  • Telemetry & Failure Logging Transparency: Frontier benchmarks must publish full execution logs, trace telemetry, and container error logs alongside leaderboard percentages.

As AI development pivots rapidly from single-turn chatbots to multi-hour autonomous system administrators and code agents, measuring progress requires rigorous measurement science. Until benchmark suites publish end-to-end validity screens for every unsolved problem, claiming that an AI model "fails at complex systems engineering" is premature. More often than not, the benchmark itself is what broke.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play