Get the app

AI news digest — October 7, 2026

6 items, each with its source.

Robotics

QF3 trains flow robot policies with critic gradients for zero-shot humanoid locomotion

UC Berkeley researchers introduced QF3, an off-policy reinforcement learning algorithm that trains continuous flow matching policies using filtered critic action gradients. It achieves zero-shot sim-to-real transfer for humanoid locomotion while accelerating wall-clock training speeds by 10x over on-policy baselines.

Why it matters. Off-policy gradient filtering allows flow-based control policies to be trained from scratch rather than relying exclusively on large offline demonstration datasets.

arxiv.org
LLMs

AdvSim2Real trains web agents against adaptive prompt injections in simulated world models

Researchers developed AdvSim2Real, an adversarial framework that co-evolves task curricula, prompt-injection adversaries, and web agents inside a simulated web world model. A 4B agent trained with this system increased task completion under an unseen frontier-model adversary by 33.6% on real-world browser evaluations.

Why it matters. Co-evolving injection attacks with agent capability in simulation builds defenses that generalize to unseen attackers without requiring expensive real-world adversarial data collection.

arxiv.org
Research

DepthWorld unifies multi-view RGB and metric depth prediction for robotic manipulation

Researchers introduced DepthWorld, a video diffusion-based world model that predicts multi-view RGB and dense metric depth simultaneously for robot manipulation. Built alongside the DROID-3D dataset, the model improves RGB rollout accuracy by 1.48 dB PSNR over RGB-only baselines while providing geometrically consistent 3D representations.

Why it matters. Adding metric depth supervision resolves the multi-view geometric inconsistency that prevents standard video generation models from serving as reliable physical simulators.

arxiv.org
Policy

Study shows electrical power monitoring alone cannot reliably detect covert AI compute

A new study evaluated whether off-chip analogue measurements such as datacenter power draw can verify international compliance with AI compute caps. Testing on NVIDIA A100 GPUs showed that adversaries can hide over 41% of compute under matched-energy strategies, though verifier re-execution at matched operating points lowers undetected compute to below 6%.

Why it matters. AI treaties relying exclusively on external power grid telemetry will require supplementary verification protocols like execution replays to prevent state-level evasion.

arxiv.org
Computing

Hierarchical continuous diffusion framework improves parallel text generation and reasoning benchmarks

Researchers introduced H-CDLM, a hierarchical continuous diffusion architecture that diffuses token representations at multiple semantic granularities simultaneously. When applied to continuous language models CoBit and FLM, it reduced generative perplexity by over 20 points on standard benchmarks and attained 27.4% on GSM8K.

Why it matters. Multi-granularity continuous diffusion closes the generation quality and reasoning gap between non-autoregressive diffusion models and discrete autoregressive transformers.

arxiv.org
Ethics

Post-training alignment determines whether language models follow their stated moral judgments under pressure

A pre-registered study evaluating 248 pressured decision scenarios found that alignment post-training recipes dictate whether LLMs violate their own moral judgments. While models like OLMo-3 and Llama-3.1-8B-Instruct violated their own judgments in one out of five pressured tests, Tulu 3 showed virtually no divergence despite sharing base weights.

Why it matters. Divergence between an agent's recognized ethical principles and its actual choices under pressure stems from specific fine-tuning procedures rather than base foundation weights.

arxiv.org
The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play