AI news digest — August 26, 2026
6 items, each with its source.
Effective learning rate dictates language model loss dynamics across training setups
Researchers identified effective learning rate (ELR)—the ratio between learning rate and parameter norm—as the primary factor governing loss trajectories during language model pretraining. When ELR is matched, loss trajectories collapse across different optimizers, architectures, datasets, and scales within small error margins. Replacing raw learning rates with ELR enables functional scaling laws to transfer consistently across various norm-control techniques.
Why it matters. Hyperparameter tuning and scaling law projections become transferable across different optimizers and norm-control schemes without requiring full re-sweeps.
arxiv.orgBrowserForge scales open web agent trajectories across two hundred thousand distinct websites
Researchers released BrowserForge, an automated framework that drives parallel sandboxed browsers across reachable websites to generate web interaction data. The system uses a Proposer-Solver agent loop to formulate executable tasks, verify trajectories, and standardize reasoning into visual chain-of-thought demonstrations. Fine-tuning a compact vision-language model on the resulting 203,238-trajectory corpus raised Online-Mind2Web task success rates from 25.66% to 33.33%.
Why it matters. Vision-based web agents can train on diverse real-world site interactions without relying on synthetic site lists or costly human demonstrations.
arxiv.orgWorldSync aligns robotic world models to follow off-expert actions during policy simulation
Researchers demonstrated that existing action-conditioned robotic world models frequently fail when generating rollouts from off-expert actions, either ignoring commands or yielding invalid visuals. They introduced WorldSync, a framework that expands action consequence coverage, grounds video representations in robot dynamics, and aligns intervention effects. In benchmarks on RoboTwin and real hardware, WorldSync improved trajectory fidelity and boosted downstream policy learning success rates.
Why it matters. Robotic policies can be trained and evaluated against out-of-distribution actions inside learned simulators without physical hardware rollouts failing from compounding simulator drift.
arxiv.orgOPDVR unifies on-policy distillation and verifiable rewards without adding hyperparameter overhead
Researchers developed On-Policy Distillation with Verifiable Reward (OPDVR) to combine token-level distillation guidance with trajectory-level correctness feedback in LLMs. The method reformulates the distillation loss using a ReLU gating mechanism that enforces non-negative rewards for correct paths and non-positive rewards for incorrect ones. Across six mathematical and logical reasoning benchmarks, OPDVR consistently outperformed standard on-policy distillation.
Why it matters. Post-training pipelines can combine teacher distillation with reinforcement learning verification without tuning arbitrary loss-weighting hyperparameters.
arxiv.orgStepGuard introduces step-level pre-execution guardrails to stop dangerous autonomous agent actions
Researchers introduced StepGuard, a safety guardrail model designed to evaluate tool invocations by LLM agents before execution rather than after completion. The model was trained using an automated data engine that pairs safe and unsafe variations of risky steps with a dynamically balanced policy optimization algorithm. Evaluated on AgentDojo and AgentDyn benchmarks, StepGuard reduced attack success rates by 77.3% with only a 2.8 percentage point drop in task utility.
Why it matters. Autonomous tool-using agents can block malicious or destructive operations in real time without severe degradation in legitimate task execution.
arxiv.orgCAFE framework improves agentic search by co-evolving corrective critics with search policies
Researchers introduced CAFE, an agentic framework where a single shared-parameter model alternates between performing search tasks and generating corrective critic feedback. The framework utilizes prompt-level success estimation during online reinforcement learning alongside offline preference optimization on paired failure and recovery trajectories. Tested across seven agentic search benchmarks, CAFE consistently outperformed conventional RL-based search agents and reduced factual hallucinations.
Why it matters. Search agents avoid hitting performance plateaus caused by static critics that fail to adapt as base agent behavior evolves.
arxiv.orgFeed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.