Get the app

Google DeepMind and KAIST Teach LLMs to Control Their Own Attention

Declarative Attention turns sparse KV caching into an intrinsic reasoning tool, cutting memory-bound attention tokens by up to 52% zero-shot without fine-tuning.

Language models spend vast amounts of memory bandwidth reading millions of Key-Value (KV) cache tokens at every generation step just to locate a handful of relevant words. In long-context regimes, loading 15 GB of KV cache per sequence at every single decoding token creates a debilitating memory wall. Rather than relying on external heuristic scorers or hardware hacks to guess what to keep, a joint team from KAIST AI and Google DeepMind has introduced a fundamentally different paradigm: Declarative Attention (DA), a protocol that enables models to explicitly choose and declare their own attention scope directly inside their chain-of-thought.

Published under arXiv:2609.02737, the paper—authored by Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, and Cicero Nogueira dos Santos—demonstrates that off-the-shelf models can autonomously control their attention windows zero-shot, slashing attended decoding tokens by 52.0% on Gemma-4-31B and 31.1% on Qwen-3.6-27B across 15 long-context benchmarks while suffering virtually negligible accuracy trade-offs.


The Memory Wall of Million-Token Inference

As context windows have expanded into the multi-million-token territory, inference latency has become almost entirely bound by memory bandwidth rather than compute.

  • The KV Cache Bottleneck: For a 1M-token sequence on large models, reading the full KV cache at each auto-regressive step requires transferring tens of gigabytes per token. In models like Qwen-3.5-397B-A17B, streaming the KV cache exceeds the memory footprint of the active 17B model parameters themselves.
  • The Flaw of Extrinsic Sparsification: Prior dynamic pruning methods (such as SnapKV, Quest, or MagicPIG) deploy auxiliary scoring modules or static heuristics (e.g., recency windows, heavy-hitter accumulation). Yet, extrinsic proxy scorers still incur an $O(N)$ scan overhead across the context before deciding what to prune, or they catastrophically drop needle-in-a-haystack tokens because static rules cannot anticipate future reasoning steps.

DA bypasses external guesswork by asking a simple question: Since the model already knows what information it plans to query next, why not let it steer the KV cache read directly?


How Declarative Attention Works: Three Attention Modes

Declarative Attention treats attention masking as a natural language tool call embedded within the model's scratchpad. Context documents are divided into tagged segments ("magic chunks"). During generation, the inference runtime intercepts special tokens generated by the model and configures the hardware attention mask accordingly:

  1. <global> Mode (Broad Context Navigation): The model attends to the entire context window. This mode is invoked during initial problem setup, exploratory searches, or when cross-referencing multi-document relationships across the prompt.
  2. <focus magic_chunks="K"> Mode (Targeted Segment Deep Dive): The model restricts its context attention strictly to chunk $K$. The KV cache manager skips reading all other segments, letting the model perform deep reasoning over specific passages without paying the full attention penalty or drowning in unrelated context noise.
  3. <local> Mode (Context-Free Pure Synthesis): The engine completely unlinks the long context KV cache. The model attends only to the original prompt scaffold (system prompt and instructions) and its own preceding generation tokens. This mode is used when performing mathematical derivations, planning subsequent steps, or formatting final answers.

Because the scaffold remains persistently attended, the model never forgets the original task or its ongoing chain-of-thought, but the massive bulk of the background document remains idle in memory.

[System & Prompt Scaffold]  --> Always Active
[Segment 1 ... Segment N]  --> Loaded dynamically on <global> or <focus K>
[Chain-of-Thought Output]  --> Self-attention on prior generated steps

Zero-Shot Results: Massive Savings, Minimal Accuracy Loss

What makes Declarative Attention particularly striking is that it requires zero architecture modifications and zero fine-tuning. Modern reasoning models already possess sufficient meta-cognitive awareness to switch modes when instructed via prompt engineering.

Across a suite of 15 long-context benchmarks spanning multi-document question answering, code comprehension, and long-horizon summarization:

  • Gemma-4-31B: Cut total attended tokens during decoding by 52.0%, with an overall accuracy delta of only 1.27 percentage points compared to full dense attention.
  • Qwen-3.6-27B: Reduced decoding attention volume by 31.1%, while preserving accuracy within 2.75 percentage points.
  • Scale Invariance: The researchers observed that larger, more capable reasoning models make significantly better attention routing decisions. The accuracy gap between dense attention and DA narrows steadily as model size and base reasoning quality increase.

Crucially, unlike speculative or proxy-based sparse attention frameworks, DA completely avoids the $O(N)$ calculation per decoding step. Once the model transitions into <focus> or <local> mode, memory reads drop to $O(K)$ or $O(1)$ relative to the input document length.


The Shift: From Low-Level Sparsity to Cognitive Routing

For years, hardware and ML systems researchers treated sparse attention purely as an indexing and tensor-slicing problem—trying to design faster PageAttention kernels, FlashAttention variants, or learned sparse masks.

Declarative Attention re-frames KV cache management as an interpretability and agentic reasoning capability:

  • Human-like Cognitive Gaze: When a human analyzes a 500-page financial report, they skim the index (global), flip to page 142 to read an EBITDA table (focus), and then draft an executive memo on a blank notebook without continuously staring at the other 499 pages (local). DA aligns transformer memory access with this exact cognitive workflow.
  • Built-in Interpretability: Because the model explicitly logs which chunk it is focusing on during each step of its thought process, developers gain granular visibility into attribution and retrieval paths without needing post-hoc attention map explainers.
  • Co-design with Reinforcement Learning: While the current paper demonstrates zero-shot viability, the authors emphasize that reinforcement learning with verifiable reward signals (RLVR) can optimize attention transitions. By adding a small computational penalty to <global> queries during RL exploration, models can be explicitly trained to become hyper-efficient attention planners.

What This Means for Long-Context Serving

As reasoning-heavy models (such as o-series and Claude Fable iterations) generate thousands of reasoning tokens before emitting an answer, long-context serving costs have exploded. Running full KV cache reads over 1M tokens across 8,000 chain-of-thought tokens is economically unsustainable for enterprise deployment.

Declarative Attention provides an immediate, deployable path forward for inference serving stacks like vLLM, TensorRT-LLM, and SGLang. By allowing the LLM's own decoder to emit dynamic cache slice instructions, production engines can reclaim more than half of their memory bandwidth overhead without modifying underlying model weights.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play