AI news digest — August 24, 2026
6 items, each with its source.
E2-TTT matches full attention expressivity while retaining chunk-wise hardware training efficiency
Researchers introduced E2-TTT, a test-time training framework that derives a closed-form state transition to reproduce per-token recurrence at chunk boundaries. Evaluating models up to 1.3B parameters, the method preserved over 90% accuracy on needle-in-a-haystack retrieval at 8x the training context length. The architecture matches the training throughput of parallel chunk-wise approximations without discarding temporal update structures.
Why it matters. Sub-quadratic architectures can maintain long-context retrieval fidelity without sacrificing the parallel training speed required for large-scale pretraining.
arxiv.orgQ-Planning enables frozen robot visuomotor policies to self-improve directly from failures
Researchers developed Q-Planning, a framework that augments large behavioral cloning policies with a compact off-policy Q-function. The method evaluates candidate actions via single-step Q-weighted averaging and updates only the value network during deployment rollouts while keeping base model weights frozen. In real-robot bimanual manipulation experiments, task success rates increased from 40% to 90% on cup stacking within five autonomous iterations.
Why it matters. Multi-billion-parameter robot foundation models can continuously recover from edge-case deployment errors without retraining base policy weights or collecting new human demonstrations.
arxiv.orgMemory augmentation compresses chain-of-thought tokens while speeding up reasoning inference latency
Researchers introduced Memory-Augmented Compression, a training-free framework that extracts reusable reasoning patterns and constraints into prefill-side memory scaffolds. Tested across mathematical and scientific reasoning benchmarks, the approach boosted accuracy by up to 29.5 points over Chain-of-Draft baselines while delivering a 1.14x to 1.49x inference speedup over standard chain of thought.
Why it matters. Production LLM reasoning pipelines can reduce output token generation costs without suffering the accuracy degradations typical of aggressive prompt compression.
arxiv.orgRARE framework decouples representation steering from expert routing in mixture-of-experts models
Researchers introduced RARE, a steering framework that projects behavioral perturbations onto the null space of router matrices in mixture-of-experts architectures. The method prevents representation engineering interventions from inadvertently degrading internal routing decisions across layers. On TruthfulQA and factual editing benchmarks, the technique increased truthfulness accuracy from 41.0% to 58.6% and factual editing efficacy from 16.8% to 96.3%.
Why it matters. Safety and alignment teams can steer behaviors in open-weight MoE architectures without destabilizing domain-specialized expert allocation.
arxiv.orgStudy shows self-refinement pipelines require larger generators and refiners but smaller critics
A multi-institution study evaluated model scaling dynamics across the generation, critique, and revision phases of agentic self-refinement using Qwen3 and Gemma 3 model families. The empirical results demonstrated that undersized refiners actively degrade output quality, whereas critic model size has minimal impact on overall task accuracy.
Why it matters. Compound AI agent architectures can reduce serving costs by deploying compact critic models while reserving high-parameter capacity strictly for generation and editing steps.
arxiv.orgCalibrated orthogonal-equivalence compensation preserves accuracy during post-training structured LLM pruning
Researchers published COEC, a training-free framework that compensates for structured column pruning in large language models using alternating left and right orthogonal rotations. The method optimizes right rotations on a reduced Stiefel manifold and rescales singular values via generalized cross-validation without requiring backpropagation. Across Llama-3 and Qwen2.5 model families, the algorithm recovered zero-shot accuracy and reduced perplexity degradation at high sparsity levels.
Why it matters. Hardware deployers can aggressively prune LLM weight matrices for high-throughput edge inference without running costly gradient retraining loops.
arxiv.orgFeed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.