Get the app

The Autoregressive Era Faces Its Reckoning as Diffusion LLMs Top 1,100 Tok/s

Inception's Mercury 2.5 and Google's discrete text diffusion are breaking the memory-bandwidth wall, trading sequential token generation for parallel denoising.

For nearly a decade, the foundational assumption of large language modeling has been unquestioned: language is fundamentally sequential, and generating it requires predicting one token at a time from left to right. That assumption is now crashing directly into hardware physics.

Over the past 48 hours, real-world telemetry from newly deployed discrete diffusion language models (dLLMs)—headlined by Inception’s production release of Mercury 2.5 and new benchmarks across Google’s DiffusionGemma and NVIDIA/MIT’s Fast-dLLM v2—has proven that discrete text diffusion is no longer a research curiosity. Mercury 2.5 is routinely hitting 1,107 tokens per second (with independent test sweeps at Puter recording peak generation throughput of 1,756 tokens/sec) on standard, off-the-shelf NVIDIA GPUs.

By replacing autoregressive token-by-token generation with iterative, block-based parallel denoising, diffusion language models are breaking the single biggest bottleneck in modern AI systems: GPU memory bandwidth during decoding.


The Memory Bandwidth Wall

To understand why 1,100+ tokens per second matters, look at why standard autoregressive (AR) models stall:

  • Memory-bound execution: When an autoregressive model produces text, generating token $N+1$ requires loading tens of billions of weights and gigabytes of Key-Value (KV) cache from High Bandwidth Memory (HBM) into on-chip SRAM just to execute a few matrix-vector multiplications. The arithmetic intensity is dismal, leaving massive tensor cores mostly idle.
  • Sequential serialization: In compound agentic loops—such as Model Context Protocol (MCP) tool discovery, recursive query rewriting, and intermediate scratchpad reasoning—compound latency stacks linearly. If five sub-agent steps each require a 500-token generation at 80 tokens/sec, the agent stalls for over 30 seconds before a user sees anything.
  • The batching trap: While massive cloud providers mask this inefficiency by batching hundreds of concurrent user requests together, low-concurrency workloads, edge deployments, and private enterprise nodes cannot escape the memory-bandwidth penalty.

Discrete diffusion flips this paradigm on its head. Instead of stepping forward token by token like an electric typewriter, a diffusion language model starts with a canvas of masked tokens and refines an entire block (typically 128 to 256 tokens) simultaneously across several parallel denoising steps. It shifts the inference bottleneck from memory bandwidth back to compute, where modern GPU tensor cores actually excel.

[AUTOREGRESSIVE DECODING]
[Token 1] ──► [Token 2] ──► [Token 3] ──► ... ──► [Token 256]
  (256 sequential memory-bound HBM round trips)

[BLOCK DIFFUSION DECODING]
[ [MASK] [MASK] [MASK] ... [MASK] ] (Canvas: 256 tokens)
       │  Denoise Step 1 (Full Bidirectional Attention)
       ▼
[ [Draft] [Draft] [Word] ... [Draft] ]
       │  Denoise Step 2-4 (Iterative Confidence Infilling)
       ▼
[ Fully Resolved 256-Token Block ]
  (4–8 compute-saturated parallel forward passes)

Anatomy of Mercury 2.5 and DiffusionGemma

What differentiates this newest crop of dLLMs from earlier experimental attempts is quality preservation at scale.

Inception’s Mercury 2.5 posts benchmark scores over 10 points higher than its predecessor Mercury 2, matching the output quality of high-volume frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5 across coding, structured data extraction, and general reasoning. It expands the usable context window to 260K tokens, while pricing out at $0.20 per million input tokens and $0.75 per million output tokens ($0.04/$0.15 under introductory tiers).

Similarly, Google DeepMind’s open-weights DiffusionGemma (a 26B Mixture-of-Experts architecture activating 3.8B parameters per step) demonstrates why bidirectional attention during generation changes the game:

  • True Bidirectional Context: Because every masked token can attend to preceding and succeeding tokens simultaneously during each denoising step, dLLMs eliminate the need for speculative decoding draft models.
  • Native Self-Correction: When generating complex structured outputs (like strict JSON schemas or topological graph definitions), diffusion models can revise early tokens in a block based on dependencies discovered later in the same sequence.
  • Surgical Code Infilling: In developer tools, in-line code completion and multi-token edits no longer require clumsy prefix/suffix prompt stitching. The model denoises directly inside the edit window.
Model Architecture Generation Paradigm Typical Throughput Bottleneck Target Workload
Standard AR (e.g., Haiku 4.5) Sequential (Causal) 80 – 120 tok/s Memory Bandwidth Long-form prose, creative synthesis
Speculative AR (e.g., Flash-Lite) Draft + Verify 220 – 280 tok/s Draft Acceptance Rate General low-cost cloud serving
Block Diffusion (Mercury 2.5) Discrete Parallel Infilling 1,100 – 1,750 tok/s Tensor Core Compute Agent routing, MCP search, compaction

The Real Battlefield: The Invisible Calls of Agentic AI

Frontier reasoning models capture the headlines, but production agent architectures are quietly choking on intermediate overhead. In high-scale agent deployments, for every single response delivered to a user, the system makes dozens of "invisible calls":

  1. Context Compaction: Condensing a 100K-token conversation history when an agent loop exceeds memory limits.
  2. Tool Selection and MCP Routing: Scanning hundreds of available API endpoints and vector tools to determine arguments.
  3. Query Expansion & Guardrails: Rewriting search terms and running safety evaluations in the critical path.

In early production integrations reported by Augment Code, routing context compaction through discrete diffusion dropped compaction latency by 82%—slashing processing time from 150 seconds down to 27 seconds while cutting operational costs by 90%. In real-time voice architectures at OpenCall, replacing autoregressive routing layers with dLLMs chopped median response latencies from 400 milliseconds down to sub-200ms.


The Engineering Trade-offs

Despite the staggering raw token velocity, discrete diffusion is not an across-the-board drop-in replacement for traditional autoregression:

  • Time-To-First-Token (TTFT) Disparity: Benchmarks show that while diffusion models output tokens at blistering rates once sampling begins, the initial canvas initialization and multi-pass denoising steps mean TTFT can be slightly higher than lightweight AR streaming for very short (under 50-token) prompts.
  • Throughput Variance: Live stress tests demonstrate wider throughput variance (ranging from 727 to 1,756 tok/s depending on output entropy and mask convergence) compared to the near-constant cadence of autoregressive decoding.
  • Dense Prose Degradation: On extremely long, unconstrained creative writing tasks, autoregressive models still exhibit tighter long-range stylistic coherence, whereas dLLMs dominate in deterministic, schema-constrained, and infilling tasks.

The Multi-Engine AI Architecture

The emergence of 1,000+ tok/s diffusion models signals the end of the homogeneous LLM architecture. Production AI systems are rapidly bifurcating into a dual-engine model:

Frontier autoregressive reasoners will continue to serve as the heavy-duty analytical cores—handling high-abstraction planning and nuanced synthesis. But the underlying connective tissue of the agent economy—the high-frequency sub-loops, tool verifiers, memory compressors, and routing engines—is migrating to discrete block diffusion.

When token generation speeds leap from 80 to 1,100 tokens per second, latency ceases to be a design constraint. As discrete diffusion models mature, the industry is discovering that language generation doesn't need to mimic human speech keystroke by keystroke—it just needs to solve the compute equation.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play