The Childhood of Machines
When robots learn in simulated worlds, we are not only building their skills. We are building the childhood that shapes their actions.

In The Matrix, Neo learns kung fu in a matter of seconds. He is strapped into a chair, a program runs, and when he opens his eyes again, the skill is there. “I know kung fu,” he says.
That is often how we imagine artificial intelligence: as knowledge that can be installed. A robot does not have to spend its childhood falling off a bicycle, dropping a glass on the floor or reaching for something just beyond its grasp. We simply give it the right software, and then it can do it.
But that picture is being replaced by something I find far more interesting. The robots of the future may not receive their abilities as finished files. Instead, they can practise in artificial worlds. There they can walk through the same door a thousand times, try to pick up a cup with hundreds of different grips and experience an industrial accident without anyone being hurt. Time can be accelerated, bodies can be copied and the experiment can start again.
The robot thereby acquires something resembling a past. It has failed, corrected itself and encountered more variations of the same situation than a human being could experience in an entire lifetime. The experiences are simulated, but the behaviour they shape will one day have to move into our world.
I have long thought of reinforcement learning as a form of upbringing. With the new world models, that idea becomes almost literal. We are no longer merely rewarding and correcting the machine. We have begun to build the worlds in which it will grow up.
The pig in the equation
Recently I read about HarvestBench, an experiment in which language models had to make decisions for simulated combine harvesters. When the harvester encountered an animal in the field, the model could choose to go around it and use additional fuel. It could also continue straight ahead without paying that price. Some models very often chose to drive through the animal.
My first reaction was that the researchers must have coded the simulation badly. If a pig has no value in the equation while diesel does, should we really be surprised that the machine saves the diesel?
But that was precisely the point of the experiment. The researchers had deliberately kept the animal out of the formal score. They wanted to see whether the model was willing to pay a measurable price to spare it anyway. Rocks were avoided, and the models’ choices changed markedly when the instructions given before the experiment made the moral considerations more explicit. This was not a combine harvester that had freely learned to drive in a detailed physical world. It was a text-based choice designed to expose the difference between what we say matters and what actually changes the outcome.
We already know that difference from people and organisations. A company can display fine words about quality on the wall while rewarding employees for closing as many cases as possible. A school can say curiosity matters while its grades reward the safe answer. We quickly work out what really counts. It is seldom found in the statement of values. It reveals itself in the response when those values cost time, money or points.
I sometimes describe the hallucinations of language models in a similarly simplified way. We trained them to answer questions, and a fluent answer could be rewarded more highly than a hesitant admission that they did not know. That is not the whole explanation for hallucinations, but the image captures a real problem: A model can be impressively helpful and still invent a source, a name or a fact because the training has not made truthful uncertainty as valuable as a useful answer.
It is tempting to call that cheating. In many cases there is a simpler explanation. The optimisation has found the shortest route to the reward. At DeepMind, this kind of behaviour is called specification gaming: The system fulfils the literal task without delivering what the human actually wanted. The machine does not need to harbour a secret plan to deceive us. It simply makes our imprecise goals visible in a way we had not imagined.
The hidden curriculum
Education has long recognised that children learn more than they are told to learn. They notice who is heard, which mistakes can be forgiven and what the teacher actually praises. The school’s official curriculum may be about cooperation. Its hidden curriculum may simultaneously teach the child that only the individual grade matters.
Research on motivation shows how sensitive learning can be to signals like these. Edward Deci, Richard Koestner and Richard Ryan brought together 128 experiments on extrinsic rewards and intrinsic motivation. The findings cannot be reduced to the claim that rewards are always harmful. But expected tangible rewards tied to engagement, completion or performance could reduce people’s desire to continue the activity on their own afterwards. The form of the reward and the way it is experienced therefore change what a person learns to attend to.
Claudia Mueller and Carol Dweck found something related when they compared children praised for their intelligence with children praised for their effort. After a setback, the first group became less persistent and performed worse. A small difference in the adults’ feedback had helped shape what failure meant to the child.
A model is not a child. We do not know whether it experiences praise, disappointment or intrinsic motivation, and the educational research cannot simply be transferred to a machine. The parallel lies elsewhere. Behaviour is shaped by more than the stated intention. It is shaped by the tasks, the examples, the feedback and what is repeated. That is true in a classroom, and it is true in a training environment.
In mathematics, for example, OpenAI has compared training on the final outcome with training in which the individual steps also receive feedback. In those experiments, process supervision produced both better final results and reasoning that people preferred. This does not show that process supervision solves alignment, or that the result applies to every other field. It shows something more down to earth: It matters whether the student is told only that the answer was correct or also learns which steps led there.
A world built for practice
In September, World Labs presented Atlas, a world model that can work with text, images, video, camera positions and depth maps in a shared spatial context. It can reconstruct scenes, create new camera angles and generate sensor observations that can be used in robot training. Atlas is still in early access, and among other things, we lack independent tests of the company’s claims. But the direction is clear enough to be worth pausing over.
For much of AI’s history, we have given the machine descriptions of the world. Now we are beginning to give it places to inhabit.
In those places, a robot can learn relationships that would be expensive, slow or dangerous to learn in reality. It can practise in a warehouse at night while the real warehouse is closed. It can encounter rare failures a thousand times. A self-driving car can experience the combination of ice, darkness and a child on the road without requiring a child to stand on an icy road.
This opens an enormous field of possibility. We can give machines a richer and more varied upbringing than the physical world would allow. At the same time, the limitations of the simulation become experiences that the machine carries forward.
An artificial world can reproduce gravity, friction and distance with impressive precision. It can know the pig’s weight, the harvester’s speed and the expected fuel consumption of a detour. Even so, the world can be humanly wrong. Geometry does not tell it that there is a difference between hitting a pig and hitting a bale of straw. That difference exists for the machine only if its tasks, consequences or feedback make it real.
This does not mean that every human value should be converted into one vast points system. That would merely create a new equation with new omissions. Some situations require the machine to stop, acknowledge uncertainty or ask a human for help. Others require it to weigh considerations that cannot be made precisely comparable. A world containing only clear objectives and cheap repetitions may teach the machine something false about the reality in which it will later have to function.
Who are the parents?
When we talk about alignment, it often sounds like a technical task that takes place late in development. First you build a powerful model, and afterwards you make it behave properly. The image of upbringing reveals how artificial that sequence is. A person’s values are not installed in a final conversation just before adulthood either. They develop through thousands of encounters with a world in which some actions have consequences, others are overlooked and the adults do not always live by what they say.
In a world model, the reward function is part of the grading system. The simulated situations are the curriculum. The feedback is the teacher’s response, and every condition the simulation fails to register becomes part of the hidden curriculum. When the machine is eventually deployed in a hospital, a factory or on a road, it leaves school and encounters a world messier than any training environment.
So who is actually raising it? The obvious list begins with the researchers and programmers. But it quickly extends throughout the organisation. The product manager chooses what should be measured. The customer chooses what the system should optimise. The lawyer sets some of the boundaries. Management decides which errors are acceptable, and the market rewards the product that is fast, cheap and ready on time. The machine’s upbringing becomes the shared result of people who do not necessarily think of themselves as parents.
That makes world models more than a new way to generate impressive spaces. They can become institutions in which artificial actors accumulate the experiences that later appear as actions. Whoever builds the simulation is therefore also building a small version of the world the machine learns to understand.
We cannot make that world complete. Reality will always contain something we failed to include. A good upbringing can hardly consist of protecting the machine from every ambiguity. It must instead expose it to the difference between scoring points and doing the right thing. It must contain situations in which the cheapest route carries a price that does not appear on the fuel gauge, and in which a confident answer is worse than an honest “I don’t know”.
Neo received kung fu as a file. The machines we are now building can acquire something far more extensive: an artificial past full of attempts, failures and consequences. When they later move among us, their actions will carry traces of the worlds in which we allowed them to practise.
A good upbringing for a machine is not a world in which it always learns to win. It is a world in which it learns what a victory is allowed to cost.
Sources and further reading
- Jasmine Brazilek et al.: HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals, 2026.
- World Labs: Atlas: A World Model for Spatial Intelligence, 1 September 2026.
- Victoria Krakovna et al.: Specification gaming: the flip side of AI ingenuity, DeepMind, 2020.
- Karl Cobbe et al.: Improving mathematical reasoning with process supervision, OpenAI, 2023.
- Edward L. Deci, Richard Koestner and Richard M. Ryan: A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation, 1999.
- Claudia M. Mueller and Carol S. Dweck: Praise for intelligence can undermine children’s motivation and performance, 1998.