The AI with More Than One Past
Humans learn from the past they have lived. An AI may learn from the many pasts contained in its recorded or simulated world.

In a computer game, you can do something reality denies us. You can save the game, open a door, discover a monster behind it, and then return to the moment before the decision. Then you open the other door.
The character inside the game experiences only the new path. The player remembers both.
Our own lives do not work that way. We can change jobs, pursue another education, or change our minds. But the new always comes after the old. I cannot return to an earlier point as the person I was then, choose differently, and see where it leads. The first path has already consumed the time and helped shape the person who now chooses the next one.
We can start over. We just cannot go back.
A new research project from Google, Google DeepMind, and two American universities plays with precisely that difference. It is called Dream-RSI, and beneath the rather grand name lies a surprisingly concrete idea: An AI can use its own search history as a collection of saved game states. It can try other routes through what has already happened and let the next strategy carry the lessons from them.
The experiment is limited to AI agents searching for better algorithms and computer code. Even so, it points toward a possibility larger than the experiment itself. AI’s particular advantage may not only be that it can read more than we can or think faster. It may be able to extract more experience from the same reality.
A history becomes a world
When a coding agent searches for a better solution, it does not move directly from the task to the answer. It proposes something, tests it, observes the result, and then decides which trail to follow. Some trails branch. Some are abandoned. Others are allowed many attempts before it becomes clear that they lead nowhere.
What remains is a tree. Each branch contains a sequence of decisions and the results the experiments actually produced.
Normally, we would treat the tree as documentation of a completed search. The Dream-RSI researchers see something else in it. The tree can also be used as a world in which new search strategies can practise.
One strategy may choose to follow a certain branch earlier. Another may abandon it sooner. A third may examine more paths in parallel or concentrate its attempts more tightly. The outcomes on the branches have already been stored. The system therefore does not need to rerun the expensive coding experiments every time a new strategy is evaluated. It can replay the history and see how the strategy would have moved through the parts of the tree that already exist.
The strategy that performs best in the stored trajectories is then sent into a new real search. It creates more attempts, which expand the tree. Another generation of strategies can then practise on the larger history.
The researchers call this recursive self-improvement, but the term needs some room around it. The underlying language model is not rewritten, and the agent does not magically acquire new fundamental abilities. The code around it becomes better at choosing what the agent should investigate, in what order, and for how long. The system improves the way it searches for improvements.
In AI Finds the Connections We Overlook, I wrote about models that can find long, verifiable paths through knowledge and code. Dream-RSI places a feedback loop around that ability. The system does not only search for a solution; the history of the search becomes material for organising the next search more effectively.
Across tasks in algorithms, mathematical optimisation, and GPU code, the researchers report that the learned search strategy in several cases finds equally good or better solutions with fewer agent attempts than fixed search strategies. For now, the results come from their own technical manuscript, and the complete code and reproduction tools have not yet been released. It is too early to know how broadly the method holds.
The underlying move is still worth pausing over. History is no longer merely something the system remembers. It becomes a place where it can practise.
The internet is only one past
One might object that the large language models have already read enormous parts of the internet. They know millions of human lives, mistakes, experiments, and decisions. Do they not already know all the pasts?
No. The internet is a vast archive of many people’s experiences, but those experiences still come from the world that actually came to be. We can read which policy a country adopted, which treatment a doctor chose, and which strategy a company pursued. Much more rarely do we learn what would have happened if the decision had been different under exactly the same conditions.
It is like having the records of every chess game that was played. We can learn a great deal from them. But the records do not automatically show the outcomes of all the other moves the players could have made from each position.
Dream-RSI’s tree is interesting because it preserves more than the final result. It preserves the places where the search divided and the outcomes on the branches that were explored. A new strategy can therefore draw a different sequence of experience from the same material. In a sense, it can live through a past the earlier strategy might have had.
These are not all imaginable pasts. If no one opened a particular branch, its result does not exist in the history. Dream-RSI does not guess at the unknown world; the method is exact only where outcomes have already been recorded. With actual world models, one can imagine AI later simulating probable consequences of actions that were never taken. The space of possibilities would grow, but so would the uncertainty, because the alternative experiences would no longer come from real experiments.
The precise formulation must therefore be that humans learn through the past they have lived, while an AI can learn through the many pasts contained in its recorded or simulated world.
More experience for the price of reality
We often discuss the economics of AI in terms of tokens, data centres, and electricity. Dream-RSI points toward another kind of economy: the price of experience.
An experiment in a laboratory can take months. A robot can destroy the object it handles. A new production method may require a factory to stop. A political intervention can affect real people’s lives for years. Much of the experience that matters most is expensive, slow, or impossible to repeat under identical conditions.
If an AI can preserve the course of events in a way that allows alternative strategies to be tested afterwards, one expensive encounter with reality can create far more learning. A failed research trajectory no longer tells us only that the solution failed. It can become material for examining whether the failure could have been detected sooner, whether resources could have been allocated differently, and which signals a better strategy should have noticed.
A memory tells us what happened. A training ground makes it possible to use the course of events again.
The potential is greatest where experience is expensive to acquire. A research agent can learn to select more promising experiments. A robot can encounter variations of the same task before moving among people. An energy system can test responses to weather and failures that only rarely occur in real operation. And in the longer term, societies may be able to examine which decisions work across many possible developments instead of optimising only for the single future a forecast considers most likely.
The last possibility requires much more caution than a search tree of coding experiments. A society cannot be rendered completely in a model. People respond to rules, to one another, and to the very expectation of what will happen. Dignity, trust, and justice may be decisive even when they are difficult to turn into numbers. A simulation can be impressively detailed and still omit exactly what later proves to matter most.
But the aim does not have to be finding the one optimal social decision. It could be finding decisions that perform reasonably well across many different trajectories. A policy that works only if citizens, markets, and the outside world respond exactly as expected may be worth less than one that survives history taking a different turn.
AI does not need to know the future to be useful. It may be enough that it has practised on several pasts.
Who chooses the pasts?
There is something appealing about the idea of an intelligence that can go back, correct its mistakes, and carry the lesson forward. But alternative pasts are never neutral.
This is related to the question in The Childhood of Machines: A simulated world teaches a machine not only what is physically possible, but also which consequences count. Dream-RSI adds another layer. The same world also determines which alternative pasts the machine can play through.
In Dream-RSI, the search tree determines which paths can be replayed, and a fixed metric determines what counts as a good solution. That makes sense in a bounded coding task. The closer the method moves toward people and societies, the more political both the simulated world and the objective become.
A hospital model can learn to reduce waiting time while overlooking that certain patients are being moved between categories. An economic simulation can improve aggregate prosperity while still placing the losses on people who are barely visible in the model. A robot may have practised in a million traffic situations where pedestrians always behave more sensibly than real pedestrians do.
A million alternative experiences do not make a model wise about what is missing from them. On the contrary, they may make it extremely confident in a world drawn incorrectly.
Power over the AI of the future will therefore not lie only with those who own the models and the computing power. It will also lie with those who decide which pasts the models may live through, which consequences are recorded, and which outcome gets to look like a victory.
A different relationship with time
Evolution explores possibilities by allowing generations and organisms to live different lives. Societies do something similar, more messily and painfully, through companies, institutions, and countries that try different solutions and compare the results afterwards. That learning takes decades, and the conditions are never quite the same.
In Knowledge Explosion, I explored the idea of a feedback loop in which new knowledge improves the methods used to create more knowledge. Dream-RSI resembles a small, concrete version of that mechanism. A discovery leaves behind not only a new solution, but also a history that can improve the way the next solution is found.
A digital intelligence can copy a state, allow strategies to branch, and then recombine the experiences. Its past can thus become material that can be arranged, replayed, and used again. That gives it a different relationship with time than ours, even though it is not, of course, travelling backwards.
Dream-RSI remains a modest version of that possibility. It dreams only within the branches that previous searches actually opened, and the promise that the next strategy is no worse applies to its score on the stored history. The next real task may still surprise it.
But if systems like this merge with more credible simulations, the difference between human and machine experience may become greater than the difference in speed. We get one past, which we can remember, interpret, and try to learn from. The machine may receive many trajectories, compare them, and pass their lessons to a version that never lived through them itself.
The strangest intelligence of the future may not be the one that thinks faster than we do. It may be the one that meets the world with memories from lives it never lived.
Sources and further reading
- Tong Zheng et al.: Dream-RSI: Recursive Self-Improvement through Evolving Worlds, technical manuscript, September 2026.
- The Dream-RSI project page with method, demonstration, and results.
- Dream-RSI on GitHub with the current release status of code and reproduction tools.
- The Childhood of Machines — on world models, simulated experience, and the worlds in which we let machines grow up.
- AI Finds the Connections We Overlook — on long, verifiable search paths through knowledge and code.
- Knowledge Explosion — on the feedback loop between knowledge and the methods used to create more knowledge.