Get the app

Mistral Drops 'Cogito-104B': Open-Weight Test-Time Compute Finally Arrives

By natively integrating 'pause tokens' and dynamic inference scaling, Mistral's new 104B model beats proprietary giants on SWE-bench by thinking longer.

For the last year, the AI industry has been obsessed with a single, closely-guarded secret: how to effectively scale test-time compute. Proprietary labs proved it was possible with closed-source reasoning models, demonstrating that allowing an AI to "think" before it speaks could drastically improve performance on complex logic, math, and coding tasks. But the actual mechanics of how to train a model to reason in latent space remained locked behind API paywalls.

Until today.

Early this morning, Mistral dropped Cogito-104B, a fully open-weight model that natively integrates System 2 reasoning through a novel "pause token" architecture. Released under the Apache 2.0 license, it is, without hyperbole, the most significant open-source release since Llama 3. By allowing the model to dynamically scale its compute during inference—literally thinking longer on harder problems—Cogito-104B achieves state-of-the-art performance on SWE-bench and MATH, rivaling the best proprietary models on the market.

The Mechanics of Latent Reasoning: Enter the <pause> Token

Standard autoregressive LLMs are forced to allocate the exact same amount of compute to every token. Whether predicting the next word in "The cat sat on the..." or solving a complex Navier-Stokes differential equation, a standard Transformer executes one forward pass per token. Chain-of-Thought (CoT) prompting bypassed this by forcing the model to output its reasoning in English, effectively using output tokens as a scratchpad.

But CoT is highly inefficient. English is a terrible, low-bandwidth language for high-dimensional mathematical reasoning.

Cogito-104B introduces native <pause> tokens. When faced with a complex prompt, the model can choose to output a <pause> token instead of a standard vocabulary token. During a pause token, the model does not emit text to the user. Instead, it performs a full forward pass, updating its KV cache and hidden states in latent space. It can chain hundreds of these pause tokens together, effectively "thinking" in high-dimensional continuous space before finally outputting the first token of the actual answer.

According to Mistral's technical report, this allows the model to bypass the "bottleneck of language." The latent representations can hold multiple parallel hypotheses, evaluate them, and collapse them into a final answer without being forced to commit to a linear English sentence.

Benchmarks: The Inference Scaling Law

The benchmark results are staggering, primarily because they prove that the scaling laws of inference are just as robust as the scaling laws of training.

  • SWE-bench (Resolved): Cogito-104B hits 48.2%, absolutely shattering the previous open-weight record of 31.5% held by DeepSeek-V3.
  • MATH (Zero-shot): 92.4%, putting it in the exact same tier as the industry's leading proprietary reasoning models.
  • GPQA (Diamond): 64.1%, demonstrating PhD-level competence in physics, biology, and chemistry.

But the raw numbers aren't the most interesting part. Mistral published an inference scaling curve. When Cogito-104B is restricted to 0 pause tokens, it performs like a standard, highly-competent 100B parameter model (roughly GPT-4 class). But as you increase the maximum allowed <pause> tokens, performance scales logarithmically.

Mistral's researchers found that performance on SWE-bench continues to improve up to approximately 2,048 pause tokens before plateauing. You are literally trading inference compute for accuracy. "We have shifted the frontier from how much data you can train on, to how much compute you are willing to spend on a single query," noted Mistral's lead researcher in the release notes.

How Do You Train a Model to Pause?

Training a model to use latent reasoning isn't as simple as just adding a new token to the vocabulary. Mistral's technical paper, Latent Space Routing and the Economics of Pause, details a grueling three-stage training pipeline that diverges significantly from standard pre-training and RLHF.

  1. World Model Pre-training: The model underwent standard autoregressive pre-training on 8 trillion tokens to build a foundational understanding of language, code, and math.
  2. Latent Distillation: Mistral utilized a massive, proprietary teacher model to generate millions of Chain-of-Thought reasoning traces for complex problems. Instead of training Cogito-104B to output these English reasoning traces, they trained the model to map the intermediate hidden states of the teacher model directly into its own latent space, separated by <pause> tokens.
  3. Reinforcement Learning from Correctness Feedback (RLCF): Unlike RLHF, which optimizes for human preference (often leading to sycophantic or overly verbose models), RLCF optimizes purely for the correct final answer. The reward function penalized the model for using too many pause tokens (to encourage efficiency) but heavily rewarded it for getting the correct answer on hard problems.

The result is a model that naturally learns when it needs to stop and think. It doesn't waste pause tokens on "What is the capital of France?" but it will automatically generate hundreds of them when asked to "Write a lock-free concurrent hash map in Rust."

The Hardware Implications: Memory Bandwidth vs. Compute

Cogito-104B's architecture also shifts the hardware bottleneck. In standard LLM generation, the bottleneck is often memory bandwidth—moving the KV cache from HBM to the compute cores for every single token generated.

Because <pause> tokens don't require streaming output back to the user, the inference engine can keep the hidden states entirely on-chip (within the SRAM) for the duration of the "thinking" phase. This means that during a heavy reasoning phase, the model is entirely compute-bound rather than memory-bound.

Next-generation silicon architectures, which heavily index on compute density over memory bandwidth, suddenly look perfectly positioned for this paradigm. Meanwhile, inference startups are already writing custom CUDA kernels specifically optimized for "Pause-Token Generation," allowing for massive batching of latent thoughts without the usual memory overhead.

The Economics of "Thinking"

This architectural shift fundamentally breaks the current LLM API business model.

For the past three years, we've priced AI by the token. $0.50 per million input tokens, $1.50 per million output tokens. But how do you price a model that might generate 1,000 hidden <pause> tokens before outputting a single word?

Within hours of the release, infrastructure providers like Together AI, Fireworks, and Groq announced new pricing tiers based on "Latent Compute Time" rather than output tokens. Because pause tokens require full forward passes and KV cache updates, they cost just as much to generate as visible tokens.

This introduces a new paradigm for developers: Compute Budgets. When calling Cogito-104B, developers pass a new parameter: max_pause_tokens.

  • For a simple summarization task, you set max_pause_tokens: 0. The model responds instantly.
  • For a complex refactoring of a Python codebase, you set max_pause_tokens: 1024. The model might "hang" for 5 seconds while it processes the latent reasoning, but the final output will be flawlessly executed.

The Open Source Ecosystem Reacts

The reaction from the open-source community has been electric. Within 12 hours of the weights dropping on Hugging Face, the ecosystem mobilized:

  • Llama.cpp merged a PR to support <pause> token generation, allowing developers to run quantized versions of Cogito-104B on M4 MacBooks. Users are already reporting that their laptops will "spin up the fans for 10 seconds" before spitting out flawless code.
  • vLLM released an experimental branch supporting "Compute Budgeting," allowing API providers to cap the number of pause tokens per request.
  • Nous Research announced they are already fine-tuning Cogito-104B on their Hermes dataset to create an uncensored, highly agentic version of the model.

Why This Changes Everything

The AI community has been quietly panicking about the "data wall." We have effectively exhausted the high-quality text on the public internet. If we can't train on more data, how do we make models smarter?

The answer is test-time compute. If you can't make the model smarter during training, you give it the ability to think longer during inference.

Until today, this capability was monopolized by labs with the massive capital required to train proprietary reasoning models. By open-sourcing Cogito-104B, Mistral has democratized System 2 thinking. Researchers can now look inside the black box. Hugging Face has already launched a dedicated leaderboard for "Latent Reasoning Models," and the open-source community is actively fine-tuning Cogito to optimize its pause-token efficiency.

We are no longer just prompting models; we are allocating compute budgets to them. The era of the static forward-pass is over. The era of dynamic, test-time reasoning has officially arrived in the open-source ecosystem.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play