Get the app

AI news digest — September 25, 2026

6 items, each with its source.

Research

Stanford and Technion researchers pretrain language models via self-play without training data

Researchers introduced a framework where a program generator and an autoregressive learner train against each other via a universal Turing machine without any human text. The generator uses reinforcement learning to propose computable programs at the frontier of the learner's predictive capabilities, creating an adaptive curriculum. The resulting models demonstrate predictable compute scaling, emerging in-context learning, and zero-shot transfer to natural language datasets.

Why it matters. Pretraining data curation bottlenecks disappear if unsupervised algorithmic self-play reliably scales to real-world language tasks.

arxiv.org
Robotics

MIT researchers automate robotic code synthesis from single visual demonstrations using RAPID

MIT CSAIL researchers developed Robot Agentic Programming from Demonstrations (RAPID), a framework that synthesizes and refines relational robot programs from a single video demonstration. RAPID automatically extracts task specifications, operational primitives, and a simulated validation environment to iteratively verify executable code. Across both simulated benchmarks and physical Franka arm experiments, the framework generalized manipulation skills across novel object geometries, poses, and materials.

Why it matters. Roboticists can deploy generalizable manipulation behaviors from one demonstration without manually writing task-specific code or safety wrappers.

arxiv.org
Industry

Nubank deploys customer experience agent simulation framework across 140 million customer accounts

Nubank engineers introduced Snowglobe, a hypothesis-driven simulation workflow for evaluating autonomous support agents before live deployment. The platform simulates synthetic customer personas and backend tool executions across tens of thousands of conversations to identify failure modes without customer exposure. In live production A/B tests, simulation-screened agent configurations increased self-service rates by 8.82 percentage points and transactional net promoter scores by 36.69 points.

Why it matters. Regulated financial institutions can safely iterate and swap core agent models without exposing end customers to live testing failures.

arxiv.org
LLMs

SAGE framework mitigates long-horizon reasoning failures in large language models using topological guidance

Researchers introduced Structural Admissibility-Guided Exploration (SAGE) to counter exploration bias and compounding errors in multi-step LLM reasoning. The framework combines algebraic subspace projections to prune spurious reasoning paths with hyperbolic geometry embeddings that provide dense depth signals during search. On complex mathematical and logical benchmarks, SAGE delivered up to an eightfold performance gain on open long-horizon challenges such as the Andrews-Curtis conjecture.

Why it matters. Reasoning models can solve deep symbolic tasks without degenerating into combinatorial dead ends or requiring dense intermediate human annotations.

arxiv.org
Computing

Rolling-WAM speeds robotic world model replanning by fourfold via staggered denoising

Researchers developed Rolling-WAM, an architecture that distributes joint video-action denoising across successive robot replanning cycles rather than recalculating the entire prediction horizon from scratch. The system maintains a sliding window of trajectory chunks across staggered noise levels, fully resolving immediate actions while refining downstream plans. Physical tests on a Unitree G1 humanoid demonstrated closed-loop manipulation with a 4.5x replanning speedup over conventional world action models.

Why it matters. Latency bottlenecks that previously made diffusion-based world models impractical for high-speed closed-loop robot control are substantially resolved.

arxiv.org
Ethics

PrivDrift benchmark reveals persistent secret leakage in long-context LLMs during multi-turn conversations

A new empirical audit examined whether user-disclosed secrets remain recoverable by adversary probes after extensive conversational topic drift. The benchmark revealed that major commercial models leak confidential information in 38.7% to 54.6% of multi-turn sessions under persuasive extraction prompts. The findings demonstrate that topic drift does not attenuate privacy risks, identifying persistent context retention as a critical behavioral failure mode in agentic sessions.

Why it matters. Conversational agents retain user confidential data in active contexts even after lengthy context shifts, making standard context-isolation defenses insufficient.

arxiv.org
The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play