AI news digest — September 25, 2026
6 items, each with its source.
Stanford and Technion researchers pretrain language models via self-play without training data
Researchers introduced a framework where a program generator and an autoregressive learner train against each other via a universal Turing machine without any human text. The generator uses reinforcement learning to propose computable programs at the frontier of the learner's predictive capabilities, creating an adaptive curriculum. The resulting models demonstrate predictable compute scaling, emerging in-context learning, and zero-shot transfer to natural language datasets.
Why it matters. Pretraining data curation bottlenecks disappear if unsupervised algorithmic self-play reliably scales to real-world language tasks.
arxiv.orgMIT researchers automate robotic code synthesis from single visual demonstrations using RAPID
MIT CSAIL researchers developed Robot Agentic Programming from Demonstrations (RAPID), a framework that synthesizes and refines relational robot programs from a single video demonstration. RAPID automatically extracts task specifications, operational primitives, and a simulated validation environment to iteratively verify executable code. Across both simulated benchmarks and physical Franka arm experiments, the framework generalized manipulation skills across novel object geometries, poses, and materials.
Why it matters. Roboticists can deploy generalizable manipulation behaviors from one demonstration without manually writing task-specific code or safety wrappers.
arxiv.orgNubank deploys customer experience agent simulation framework across 140 million customer accounts
Nubank engineers introduced Snowglobe, a hypothesis-driven simulation workflow for evaluating autonomous support agents before live deployment. The platform simulates synthetic customer personas and backend tool executions across tens of thousands of conversations to identify failure modes without customer exposure. In live production A/B tests, simulation-screened agent configurations increased self-service rates by 8.82 percentage points and transactional net promoter scores by 36.69 points.
Why it matters. Regulated financial institutions can safely iterate and swap core agent models without exposing end customers to live testing failures.
arxiv.orgSAGE framework mitigates long-horizon reasoning failures in large language models using topological guidance
Researchers introduced Structural Admissibility-Guided Exploration (SAGE) to counter exploration bias and compounding errors in multi-step LLM reasoning. The framework combines algebraic subspace projections to prune spurious reasoning paths with hyperbolic geometry embeddings that provide dense depth signals during search. On complex mathematical and logical benchmarks, SAGE delivered up to an eightfold performance gain on open long-horizon challenges such as the Andrews-Curtis conjecture.
Why it matters. Reasoning models can solve deep symbolic tasks without degenerating into combinatorial dead ends or requiring dense intermediate human annotations.
arxiv.orgRolling-WAM speeds robotic world model replanning by fourfold via staggered denoising
Researchers developed Rolling-WAM, an architecture that distributes joint video-action denoising across successive robot replanning cycles rather than recalculating the entire prediction horizon from scratch. The system maintains a sliding window of trajectory chunks across staggered noise levels, fully resolving immediate actions while refining downstream plans. Physical tests on a Unitree G1 humanoid demonstrated closed-loop manipulation with a 4.5x replanning speedup over conventional world action models.
Why it matters. Latency bottlenecks that previously made diffusion-based world models impractical for high-speed closed-loop robot control are substantially resolved.
arxiv.orgPrivDrift benchmark reveals persistent secret leakage in long-context LLMs during multi-turn conversations
A new empirical audit examined whether user-disclosed secrets remain recoverable by adversary probes after extensive conversational topic drift. The benchmark revealed that major commercial models leak confidential information in 38.7% to 54.6% of multi-turn sessions under persuasive extraction prompts. The findings demonstrate that topic drift does not attenuate privacy risks, identifying persistent context retention as a critical behavioral failure mode in agentic sessions.
Why it matters. Conversational agents retain user confidential data in active contexts even after lengthy context shifts, making standard context-isolation defenses insufficient.
arxiv.orgFeed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.