Get the app

Anthropic Drops Claude 4: Differentiable Working Memory Kills the Context Window

Moving beyond brute-force context scaling, Claude 4 introduces a read-write memory architecture that achieves 100M-token recall at a fraction of the inference cost, shifting LLMs from stateless to stateful.

The context window has always been a brute-force hack. For the past three years, the AI industry has been locked in a race to scale context lengths—from 8K to 128K, up to Google's 2M tokens, and beyond. But the fundamental physics of the standard Transformer architecture remained a bottleneck: attention scales quadratically. Every time you prompt a model, it has to re-read and re-compute the entire conversation history. It is the equivalent of reading an entire textbook from page one every time you need to answer a single question on a test.

Yesterday, Anthropic fundamentally changed the paradigm. With the release of Claude 4, the headline isn't just a bump in MMLU scores (though it did hit a staggering 92.4%). The real breakthrough is the deprecation of the traditional context window in favor of a Differentiable Working Memory (DWM) architecture. Claude 4 is the first frontier model to shift from a stateless text-predictor to a stateful reasoning engine.

By introducing a read-write memory stream, Anthropic has effectively solved the quadratic scaling problem of long-context attention, achieving near-infinite context recall at a fraction of the compute cost.

The Architecture: How Differentiable Working Memory Works

To understand why Claude 4 is a massive leap forward, we have to look at how it handles past tokens. In a standard Transformer, the KV (Key-Value) cache stores the representations of all past tokens. As the context grows, the KV cache balloons in size, consuming massive amounts of VRAM and slowing down time-to-first-token (TTFT).

Claude 4 replaces the massive KV cache with a hybrid architecture:

  • Short-Term Sliding Window: The model uses standard attention for the immediate context (the last 4,096 tokens), ensuring perfect syntactic and semantic flow for the current interaction.
  • Long-Term DWM Bank: Tokens that fall out of the sliding window are compressed and written to a discrete, addressable memory bank.

This isn't just a hidden state compression like we see in State Space Models (SSMs) such as Mamba or Mistral's recent architectures. SSMs struggle with exact recall—the classic "needle in a haystack" problem—because they force all history into a fixed-size vector. Claude 4's DWM acts more like a Neural Turing Machine. The model has specialized Read/Write Heads that allow it to dynamically decide what information is important enough to store, and what it needs to retrieve based on the current prompt.

When you ask Claude 4 a question, it queries its memory bank using cross-attention, retrieving only the relevant compressed concept vectors instead of re-computing attention over millions of raw tokens. The result is O(1) scaling for inference compute, regardless of how long the conversation history is.

Shattering Benchmarks: SWE-Bench and 100M-Token Recall

Anthropic didn't just release an architecture paper; they released a production-ready model that is currently destroying the leaderboards.

  • SWE-Bench Resolved: Claude 4 achieves 64.2% on the full SWE-Bench, a massive leap from Claude 3.5 Sonnet's previous SOTA. Because software engineering requires maintaining complex state across thousands of lines of code and multiple files, the DWM architecture gives Claude 4 an unfair advantage. It can "remember" a bug found in file A while editing file Z, without needing complex agentic scaffolding to pass context back and forth.
  • Needle In A Haystack (NIAH): Anthropic tested Claude 4 on a simulated 100-million token interaction. The model achieved 99.8% recall. For context, 100 million tokens is roughly equivalent to the entire text of the English Wikipedia.
  • MMLU-Pro: The model scored 88.5%, proving that the memory architecture does not degrade its core reasoning and zero-shot capabilities.

The Economics of Stateful AI

Perhaps the most disruptive aspect of Claude 4 is its pricing model. Because the model doesn't recompute the KV cache for the entire history, the cost per output token drops significantly for long, continuous conversations.

Anthropic is introducing a new pricing tier for Claude 4:

  • Memory Write: $3.00 per 1M tokens.
  • Memory Read: $0.00 (Free during inference).
  • Output: $15.00 per 1M tokens.

This completely flips the economics of agentic workflows. Previously, running an autonomous agent that required a 100K context window for 50 sequential steps would cost dollars per task, as you paid for the input tokens on every single turn. With Claude 4, you pay to write the context into memory once. Subsequent turns only cost you the output tokens. This makes long-running, persistent AI agents economically viable for the first time.

What This Means for RAG and Autonomous Agents

The ripple effects of Claude 4's release will be felt across the entire AI stack, particularly in how we build applications.

The Evolution of RAG (Retrieval-Augmented Generation) Standard RAG relies on chunking text, embedding it, and using cosine similarity in a vector database to retrieve relevant snippets. It's a lossy process. You lose the broader context of the document, and the retrieval mechanism often misses nuanced connections.

Claude 4's DWM allows developers to simply dump an entire repository, codebase, or legal library into the model's working memory. The model itself organizes the information in its latent space. Vector databases aren't dead—they will still be needed for enterprise-scale data (billions of documents)—but for application-level context, native model memory will replace the "chunk and retrieve" paradigm.

The Holy Grail for Agents Autonomous agents need to remember their past actions, failures, and environment states. Until now, developers had to build complex scaffolding to manage an agent's memory, summarizing past turns and injecting them into the prompt.

Claude 4 is natively stateful. If an agent tries a Python script and it fails, that failure is written to its working memory. Ten steps later, it won't make the same mistake, because the memory of the failure is an intrinsic part of its neural state.

The Bottom Line

Anthropic's Claude 4 is the most significant architectural shift since the original Transformer paper in 2017. By solving the quadratic scaling problem of attention through Differentiable Working Memory, they have unlocked the path to truly persistent, stateful AI.

We are moving away from models as stateless functions—where every prompt is a blank slate—and toward models as persistent entities that learn, remember, and adapt over the course of an infinite conversation. The context window is dead. Long live working memory.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play