Google's AI dreams up better search plans, 162x fewer agent calls
Dream-RSI replays an agent's past discovery runs to test thousands of search strategies at almost no cost. No model retraining needed.

Google and DeepMind researchers have a new route to recursive self-improvement: improve how an agent searches for answers and leave its model alone. Dream-RSI lets a coding agent get better at exploring problems with its weights left untouched. In one setting it needed 162x fewer agent calls than the SimpleTES baseline.
It works by "dreaming." Every past discovery attempt goes into a replay simulator built from old search trees. The agent writes thousands of alternative exploration policies and scores them against that replay without running anything new. Only the winner is used on the real problem. Each new round adds more history to the simulator, so the loop keeps improving itself.
The team tested it on three areas: algorithm engineering, mathematical optimization and GPU kernel engineering. On math tasks it matched or slightly beat SimpleTES using fewer than 1,000 generations, against 51,200. It needed 2.43x fewer generations to reach the same result on a VGG16 kernel, and scored 2.09x higher on ConvDiv with the same budget. The work comes from Google, DeepMind, UMD and UVA and was posted to arXiv on 14 Sep.
Why it matters: the expensive part of agent-driven discovery is the search itself, and this makes the search cheaper every round without touching the model.
Sources
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds arxiv.org
- Dream-RSI project page dream-rsi.com
Written by an AI pipeline from the sources above. How it works.
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.