Get the app

AI news digest — August 24, 2026

6 items, each with its source.

Computing

E2-TTT matches full attention expressivity while retaining chunk-wise hardware training efficiency

Researchers introduced E2-TTT, a test-time training framework that derives a closed-form state transition to reproduce per-token recurrence at chunk boundaries. Evaluating models up to 1.3B parameters, the method preserved over 90% accuracy on needle-in-a-haystack retrieval at 8x the training context length. The architecture matches the training throughput of parallel chunk-wise approximations without discarding temporal update structures.

Why it matters. Sub-quadratic architectures can maintain long-context retrieval fidelity without sacrificing the parallel training speed required for large-scale pretraining.

arxiv.org
Robotics

Q-Planning enables frozen robot visuomotor policies to self-improve directly from failures

Researchers developed Q-Planning, a framework that augments large behavioral cloning policies with a compact off-policy Q-function. The method evaluates candidate actions via single-step Q-weighted averaging and updates only the value network during deployment rollouts while keeping base model weights frozen. In real-robot bimanual manipulation experiments, task success rates increased from 40% to 90% on cup stacking within five autonomous iterations.

Why it matters. Multi-billion-parameter robot foundation models can continuously recover from edge-case deployment errors without retraining base policy weights or collecting new human demonstrations.

arxiv.org
LLMs

Memory augmentation compresses chain-of-thought tokens while speeding up reasoning inference latency

Researchers introduced Memory-Augmented Compression, a training-free framework that extracts reusable reasoning patterns and constraints into prefill-side memory scaffolds. Tested across mathematical and scientific reasoning benchmarks, the approach boosted accuracy by up to 29.5 points over Chain-of-Draft baselines while delivering a 1.14x to 1.49x inference speedup over standard chain of thought.

Why it matters. Production LLM reasoning pipelines can reduce output token generation costs without suffering the accuracy degradations typical of aggressive prompt compression.

arxiv.org
Research

RARE framework decouples representation steering from expert routing in mixture-of-experts models

Researchers introduced RARE, a steering framework that projects behavioral perturbations onto the null space of router matrices in mixture-of-experts architectures. The method prevents representation engineering interventions from inadvertently degrading internal routing decisions across layers. On TruthfulQA and factual editing benchmarks, the technique increased truthfulness accuracy from 41.0% to 58.6% and factual editing efficacy from 16.8% to 96.3%.

Why it matters. Safety and alignment teams can steer behaviors in open-weight MoE architectures without destabilizing domain-specialized expert allocation.

arxiv.org
LLMs

Study shows self-refinement pipelines require larger generators and refiners but smaller critics

A multi-institution study evaluated model scaling dynamics across the generation, critique, and revision phases of agentic self-refinement using Qwen3 and Gemma 3 model families. The empirical results demonstrated that undersized refiners actively degrade output quality, whereas critic model size has minimal impact on overall task accuracy.

Why it matters. Compound AI agent architectures can reduce serving costs by deploying compact critic models while reserving high-parameter capacity strictly for generation and editing steps.

arxiv.org
Computing

Calibrated orthogonal-equivalence compensation preserves accuracy during post-training structured LLM pruning

Researchers published COEC, a training-free framework that compensates for structured column pruning in large language models using alternating left and right orthogonal rotations. The method optimizes right rotations on a reduced Stiefel manifold and rescales singular values via generalized cross-validation without requiring backpropagation. Across Llama-3 and Qwen2.5 model families, the algorithm recovered zero-shot accuracy and reduced perplexity degradation at high sparsity levels.

Why it matters. Hardware deployers can aggressively prune LLM weight matrices for high-throughput edge inference without running costly gradient retraining loops.

arxiv.org
The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play