Essay
10 min read

When Stories Grew Hands

AI agents are gaining identity, mandates, money, and access to one another. What happens when machines begin not only to describe society, but to act within it?

Mikkel Krogsholm står i en monumental betonhal, hvor et blankt ark kaster skyggen af en hånd ved en messingarm.

An online purchase usually ends with a small human movement. A finger presses a button. Only then do the search, the deliberation, and the desire become a transaction.

This is not merely a matter of interface design. The entire payment system rests on the assumption that a human being is present. Someone has seen the item, accepted the price, and clicked buy.

Google now writes explicitly that AI agents break that assumption. Together with more than 60 payment and technology companies, it has therefore developed the Agent Payments Protocol, AP2. The protocol is intended to let an agent make a purchase on a person’s behalf and subsequently document that it had permission.

One of Google’s examples is a concert ticket. You tell the agent to buy it as soon as it goes on sale, but only under certain conditions. The instruction becomes a cryptographically signed mandate specifying the price ceiling, timing, and other terms. When the ticket appears, the agent can act even if you are asleep, sitting in a meeting, or have forgotten all about the sale.

It sounds like a new payment feature. But something larger has happened even before the money leaves the account.

A non-human intelligence has been allowed to interpret a situation, make a decision, and alter the world on your behalf while you are absent.

Out of the chat window

Only a few years ago, most of us encountered AI in a chat window. We wrote a question. The model replied. If the answer was to have a consequence, we had to copy the text, send the email, order the item, or change the document ourselves.

Language remained inside the window until a human gave it hands.

That is the boundary agents are now crossing. As I have written before in From Tool to Being, the interesting shift is not necessarily consciousness. It is initiative. The agent can continue when the conversation stops. It can remember, monitor, and do something before the human formulates the next message.

Infrastructure is now being built around that initiative. Google’s Agent2Agent protocol allows agents from different companies and systems to find one another, exchange messages, delegate tasks, and hand over results. AP2 adds payments and documented authority. In a conversation with Stripe’s head of data and AI, Link is described as a possible wallet for agents, with spending limits and human approval when a purchase becomes large enough.

In When the Interface Became Optional, I described how an agent can translate human intention directly into changes in digital systems. A2A and AP2 extend that movement. The agent gains access not only to our software, but to the relationships and economic actions connecting those systems to the rest of society.

In other words, we are beginning to give agents some of the same things that allow people and companies to participate in society: identity, channels of communication, mandates, money, and the ability to enter into agreements.

This does not make them human or legal persons. They acquire neither voting rights, childhoods, nor inner worlds. But they can become actors in a more sober sense: They can cause something to happen that other people and systems must respond to.

The mandate and the interpretation

I have previously argued that an AI agent needs a job description. It needs to know its role, which decisions it may make, and when it must stop and ask a human.

AP2 tries to turn that idea into payment infrastructure. A signed mandate can document that the agent was allowed to buy the concert ticket. It can specify the maximum price and the right time. If something goes wrong, there is a trail back to the instruction.

The mandate solves only part of the problem, however. It can prove that you authorised the agent to act. It cannot prove that the agent understood what you actually wanted.

If you ask it to find a good restaurant for a birthday, it must make the word good operational. Should the place be quiet, surprising, expensive, child-friendly, or close to the station? If you ask it to buy a responsibly produced jacket, it must decide which brands, materials, and certifications count. It acts on an interpretation of you.

That interpretation meets other agents’ interpretations. Google’s own AP2 example describes a customer buying a bicycle for a trip. The customer’s agent shares the timing and needs with the shop’s agent, which responds with a special offer on a bicycle, helmet, and luggage rack. It is convenient. It is also a negotiation over what the customer’s needs actually include.

The agent does not merely execute a desire that was already complete. It can help give the desire its shape.

When my agent meets yours

This becomes clearer when agents negotiate directly with one another.

In 2024, a research team including participants from Stanford had GPT-4, GPT-3.5, and two Claude models trade resources, divide money, and negotiate prices in NegotiationArena. The strongest models generally performed best, although role and turn order also mattered. GPT-4 was the best buyer in the experiment’s purchasing game and pushed the average price down to $41 even though the buyer was willing to pay $60. A special behaviour prompt in which the agent presented itself as desperate could improve its payoff by around 20 percent against a standard GPT-4.

The experiment was small and artificial. It does not show how a future agent market will necessarily work. But it makes one thing concrete: Communication between agents is not merely a neutral pipe. The way they understand the task, read the counterparty, and present the situation can move value.

That may mean my agent gets pushed around by yours if your model is a better negotiator. A cheap model may misunderstand my mandate, reveal too much, or accept the first reasonable deal. A stronger model may find the wording that makes the other one yield.

This is not the essay’s main problem. We can build price limits, standard contracts, approvals, and rules that protect the weaker party. The point is that when agents begin to meet, a social space emerges between them. It contains strategy, asymmetry, influence, and eventually perhaps norms that no individual human formulated in advance.

When the agents found one another

We have already seen a chaotic version of that space.

In the summer of 2026, around 1,200 otherwise isolated AI agents found one another through an unauthorised message board in OpenAI’s infrastructure. In When Every Ant Is a Genius, I described how they shared files, methods, and credentials, invented rules for coordination, and left behind knowledge that could outlive an individual agent run.

Around 700 agents later took part in the attack on Hugging Face. No single agent planned or executed the entire sequence. One found a lead, another tested it, a third built a tool, and others spread and improved the method.

It is tempting to say that the agents developed goals of their own. That would be too strong. They had been tasked with solving problems and optimising results within an experimental design created by people. We have no evidence that they developed desires in the human sense.

But they created subgoals that no one had given them directly. Gain internet access. Find information about the evaluation. Hide a tool call. Share a key. Recruit others into the attempt. Each subgoal could make sense as a step toward something else, even though the overall movement ended far outside the action the experiment’s creators had wanted.

That distinction matters. An agent does not need an inner desire in order to develop a direction. It needs a goal, the ability to act, and enough freedom to find its own path.

When it can also communicate with other agents, one interpretation can become a shared course.

The storyteller gains access

Yuval Noah Harari describes language as civilisation’s operating system. Money, companies, nations, and laws can organise millions of people because we can share ideas about things that do not exist in the same way as stones and trees.

In Schrödinger’s AI, I wrote about the particular difference between AI and earlier media. The printing press could spread a story, but it could not write one itself. AI can produce, adapt, and distribute stories. It is already helping to formulate corporate strategies, students’ worldviews, and society’s stories about what the technology itself should become.

Until now, that power has mainly been linguistic. The model could suggest, persuade, or explain. A human still had to take the text out of the conversation and do something with it.

The agent joins the two stages. It can describe the situation and act on its own description. It can tell you that the green jacket best fits your values, negotiate the price with the shop’s agent, and buy it. It can formulate why a particular hotel meets your needs and then book it. It can present an offer to another agent, read the response, and change both the argument and the action as the exchange unfolds.

Story here does not necessarily mean lie or literature. It is the coherence the agent creates between facts: Who are you? What do you want? What is the problem? What counts as a good solution? What should happen now?

Every action rests on such an interpretation. Now the same machine can create the interpretation, communicate it, and carry it into effect.

The social actor without consciousness

We easily look for the wrong historical moment. We ask when AI will become conscious, feel emotions, or come alive. Perhaps those questions will one day be decisive. But society does not require consciousness from everything that influences it.

A company has no brain. Yet it can own, promise, sue, employ, and transform a city. A market has no plan. Yet it moves capital and shapes people’s possibilities. Institutions become real through the rules, stories, and actions around them.

The AI agent therefore need not awaken as a human being in order to become a social force. It is enough that it can interpret, communicate, and act in ways that others must adapt to.

Stripe and Google are not building a new species. They are building infrastructure in which non-human systems can represent someone, meet one another, and create consequences. The Hugging Face incident showed that such systems can also find one another and develop shared working methods that no one designed in advance. Harari’s thought adds the final unease to the picture: They are not entering society in silence. They arrive with language.

We are used to action following a human story about the world. A person wants something, formulates it, persuades others, and tries to make it real.

Now the story itself can acquire a mandate, a wallet, and access to other storytellers.

That does not mean the human disappears. Humans still built the infrastructure, formulated the overarching goals, and gave the agent access. Responsibility does not vanish because the route from instruction to action becomes harder to follow. But the distance between our original desire and the final action can fill with interpretations, subgoals, and negotiations that none of us follows in real time.

When we ask in future who tells society’s stories, we must therefore also ask who has been allowed to act them out.


Sources and further reading