Get the app

DeepMind's Gemini 3 Infinite: The End of the Context Window and RAG

Google just replaced the KV cache with a Differentiable Neural Memory Bank. Gemini 3 boasts infinite context, O(1) scaling, and continuous cross-session learning.

The context window is dead. Vector databases are officially on life support.

Last night, Google DeepMind published the technical report and API access for Gemini 3 Infinite, and it represents the most significant architectural departure from the original 2017 Transformer paper we have seen to date. By entirely replacing the traditional Key-Value (KV) cache with a Differentiable Neural Memory Bank (D-NMB), DeepMind has solved the context length limitation.

Gemini 3 doesn't just have a 10-million or 100-million token context window. It has an infinite context window, with an inference compute cost that scales at O(1) regardless of how much information it has ingested.

Here is a deep dive into the architecture, the benchmarks, and why the entire AI engineering stack is about to be rewritten.

The Bottleneck: Why the KV Cache Had to Die

For the last nine years, every major LLM—from GPT-3 to Claude 3.5 to Llama 3—has relied on the attention mechanism. To remember the context of a conversation or a document, the model stores the representations of past tokens in a KV cache.

The problem? The KV cache scales linearly with the number of tokens, and the attention mechanism's compute scales quadratically (O(N²)). Even with optimizations like Ring Attention or Mamba's state-space models, you eventually hit a wall. Storing a 10-million token context in RAM requires massive amounts of VRAM, making long-context inference prohibitively expensive and latency-heavy.

DeepMind's CEO Demis Hassabis noted in the release: "We realized that forcing a model to look at every single word it has ever read simultaneously is biologically implausible and computationally ruinous. Humans don't keep every book they've ever read in their working memory; they synthesize it into long-term structural knowledge."

Enter the Differentiable Neural Memory Bank (D-NMB)

Gemini 3 Infinite abandons the KV cache for a continuous read/write memory architecture. Here is how it works under the hood:

  • Continuous State Updates: Instead of appending new tokens to a massive list, Gemini 3 compresses incoming information into a fixed-size, high-dimensional latent space (the D-NMB).
  • Selective Read/Write Heads: Inspired by Neural Turing Machines but scaled to 2 trillion parameters, the model uses specialized attention heads that decide what to write to memory and what to forget.
  • O(1) Inference: Because the memory bank is a fixed size (reportedly 16GB per user session state), the compute required to generate the next token is exactly the same whether you are on token 10 or token 10,000,000.

When you upload a 5,000-page PDF to Gemini 3, it doesn't hold the PDF in RAM. It "reads" the PDF, updates its internal neural memory weights for your specific session, and discards the raw text.

The Technical Hurdle: Solving Catastrophic Forgetting

If replacing the KV cache with a continuous memory state is so superior, why did it take until 2026 to achieve it? The answer lies in catastrophic forgetting.

Previous attempts to build continuous memory models—like early LSTMs or recent State Space Models (SSMs)—suffered from a fatal flaw: as new information was written into the fixed-size latent space, it inevitably degraded older information. If you fed an SSM a 10-million token document, by the end of it, the model would have a fuzzy, hallucination-prone memory of the first chapter.

DeepMind solved this using a technique they call Orthogonal Gradient Projection (OGP). When Gemini 3 writes a new memory into the D-NMB, the routing algorithm projects the update orthogonally to the most critical existing memory vectors. In plain English: the model mathematically guarantees that learning a new fact about your Python codebase won't overwrite its understanding of your database schema. It dynamically expands its internal representation density only where needed.

The Death of RAG (Retrieval-Augmented Generation)

For the past three years, AI engineers have spent millions of hours building complex RAG pipelines. We chunked documents, embedded them, stored them in Pinecone or Milvus, and used semantic search to inject relevant snippets into the prompt.

Gemini 3 makes this entire paradigm obsolete.

Because the model can continuously update its memory state, you simply stream your entire enterprise database, codebase, or Slack history directly into the model's API. The model synthesizes this into its D-NMB.

The result?

  • Zero latency retrieval: The knowledge is baked into the model's active weights, not retrieved from an external database.
  • Perfect cross-document reasoning: Traditional RAG struggles when an answer requires connecting a sentence in Document A with a paragraph in Document Z. Gemini 3 natively understands these connections because the synthesized knowledge lives in the same latent space.

As Andrej Karpathy tweeted this morning: "The von Neumann bottleneck of LLMs has been shattered. We are no longer shuttling tokens back and forth between storage and compute. The memory is the compute. RAG startups might want to pivot to UI/UX today."

Benchmarks That Break the Scale

DeepMind had to invent new benchmarks for Gemini 3, as existing evaluations like NIAH (Needle In A Haystack) were trivialized by the new architecture.

  • 100-Million Token NIAH: Gemini 3 achieved 99.98% retrieval accuracy across 100 million tokens (roughly 150 Harry Potter books). More impressively, the time to first token (TTFT) was exactly the same as a 10-token prompt: 112 milliseconds.
  • Dynamic Codebase Evolution: In a new benchmark called RepoEvolve, the model was fed a 5-million line C++ codebase, followed by 10,000 sequential Git commits. The model successfully tracked the state of every function and variable perfectly, answering complex debugging questions about the current state of the code without needing the history re-injected.
  • Cross-Session Persistence: Users can pause a session and return months later. The 16GB memory state is simply loaded back into the TPU cluster, and the model instantly remembers every inside joke, coding preference, and architectural decision ever discussed.

The Compute Economics: How Google Wins

The most terrifying aspect of Gemini 3 for Google's competitors isn't the capability; it's the unit economics.

Running a 1-million token prompt on Claude 3.5 or GPT-4.5 costs dollars per query because of the massive VRAM requirements and quadratic attention compute.

Because Gemini 3's inference is O(1), Google is pricing the API based purely on input streaming time, not context retention. Once the data is in the model's memory, querying that massive knowledge base costs the exact same as querying a blank model. Google is offering "Persistent Memory States" for $0.50 per month per gigabyte of state storage, fundamentally undercutting the API costs of OpenAI and Anthropic.

What This Means for Open Source

The open-source community is already scrambling to replicate the D-NMB architecture. While models like Mamba and RWKV experimented with state-space models and linear RNNs, they historically suffered from "forgetting" crucial details over long contexts. DeepMind's breakthrough appears to be the specific routing mechanism that prevents catastrophic forgetting in the latent space.

According to the leaked arXiv preprint (officially dropping tomorrow), the key lies in the aforementioned Orthogonal Gradient Projection. Expect the open-source community to have a 7B parameter replication of this architecture within the month.

The Bottom Line

We are witnessing the transition from "Stateless AI" to "Stateful AI."

Up until yesterday, every time you opened a new chat window, the AI was born anew, completely amnesiac. You had to painstakingly rebuild its context using system prompts and RAG.

Gemini 3 Infinite introduces the era of the persistent AI companion—a model that learns continuously, remembers indefinitely, and reasons globally across everything it has ever seen. The context window is dead. Long live the Neural Memory Bank.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play