Essay
14 min read

When Code Is No Longer a Human Language

AI may soon write more code than humans can review. We therefore need a new language that makes the machine’s work understandable and controllable.

Mikkel Krogsholm betragter fire gennemsigtige lag, hvor et tæt netværk bliver til én sporbar messinglinje.

I recently submitted a proposed change to Graphify, a software project whose code is open for others to read and contribute to.

In software, such a proposal is often submitted as a pull request. It is a package containing the new code, an explanation and an overview of what has changed. Other people can read the package, ask questions and approve it before the change becomes part of the program.

Usually, you are met by the code itself. Lines that have been removed appear in red. Lines that have been added appear in green. If the change is large, there may be thousands of them.

My proposal for Graphify was about giving the reader another way in.

Instead of beginning with all the red and green lines, you should be able to begin with a few ordinary questions: What is different? Which parts of the program are affected? What might the change have consequences for? And where can I see the evidence?

The proposal therefore creates a report in several layers. First, a short story about the change. Then, a picture of the parts of the program that are connected. Next, a simple step-by-step description of the most important logic. At the bottom is the actual code.

Programmers sometimes call that simple step-by-step description pseudocode. It is not code that the computer can run. It is more like a recipe:

  1. Receive an order.
  2. Check whether the item is in stock.
  3. Remove the item from the inventory.
  4. Send a confirmation.

The real code contains all the precise rules, formats and exceptions the computer requires. The pseudocode tries to show the train of thought without requiring the reader to know the programming language.

I did not build the experiment because the real code has become irrelevant. Quite the opposite. I built it because code has become too important for human control to depend on how quickly a person can read thousands of lines in a language that has never been our mother tongue.

That problem grows as AI writes an increasing share of the code.

The machine is becoming faster at producing. The human is still expected to read afterwards.

That may become the next great bottleneck in software development.

From Zeros and Ones to Python

We have tried something similar before.

The first programmers worked very close to the computer’s own language. They had to describe the work as numbers and very concrete, small operations. It was cumbersome, slow and reserved for very few people.

Later, the small operations were given names. Then came languages such as Fortran, in which a human could write a mathematical formula in a recognisable form. Another program then translated the formula into the many small instructions the computer could execute.

Such a translation program is called a compiler. The name is not important. The division of labour is: the human describes the task at a level that makes sense to a human. The tool translates it to a level that makes sense to the machine.

IBM describes how a task that had previously required a thousand machine instructions could be expressed in 47 lines of Fortran. Many doubted that automatically translated code could become as efficient as code written directly for the machine. It almost could.

We have repeated that movement ever since.

New programming languages moved the developer a little further away from the physical workings of the computer. Python made it possible to express a great deal in relatively few and readable lines. Ready-made building blocks made it unnecessary to invent everything from scratch.

Python is often called a higher-level programming language. Higher does not mean finer or better. It means further away from the computer’s individual operations and closer to the problem the human is trying to solve.

Each time we have moved up a level, the human has lost a little direct contact with the layer below. In return, more people have been able to build more complex things.

We have moved the place where the human expresses their intention.

AI continues that movement. I can now describe a desired change in ordinary words, and a digital colleague can find the relevant files, propose the new code and check whether the familiar parts of the program still work.

In When the Interface Became Optional, I wrote about precisely that translation. AI can stand between a human wish and the technical material of the system. We no longer need to know every button, file and command ourselves to make something happen.

But when the translation down to the machine becomes that good, a new problem appears in the opposite direction.

How do we translate the machine’s work back into human understanding?

Code Has Had Two Readers

Source code is the written instructions that determine what a program does.

It has long had two different readers.

The computer must be able to turn the instructions into action. But a human must also be able to read them to understand the program, find errors and decide whether a change is safe.

That is why programmers spend time giving things good names and dividing the work into understandable parts. The computer does not care whether the name of a part of the program makes sense. The next programmer does.

The next programmer may simply have changed character.

If code is increasingly written, changed and maintained by AI programs that can solve tasks on their own, source code may become an intermediate product between machines. It remains essential. The precise instructions still live there. But it does not have to be the place where most humans first encounter the change.

Few Python programmers today read the most basic machine instructions into which their program is eventually translated. They work at a higher level because the connection between the levels is reliable enough.

The question is whether something similar can happen when we review new code.

Not by removing the code, but by giving the human a more comprehensible reading layer above it.

For a small change, a recipe-like explanation can stay close to the code. For a very large change, even that explanation can become long. Then we can begin with a short overview and open more detail as needed.

The human starts with the meaning and moves down towards the technical instructions when something is unclear, risky or important.

We have spent decades building layers that translate human wishes down into machine instructions. Perhaps the time has come for a layer that translates the machine’s work up into human meaning.

More Code Is Not the Same as More Progress

It is tempting to turn this into a simple story about speed.

AI writes the code faster. Therefore we get better software faster.

Reality does not look that simple.

DORA is a large research programme that examines how software is developed in real organisations. In its 2025 report, more than 80 per cent of participants said AI had increased their productivity. At the same time, 30 per cent had little or no confidence in code written by AI.

Organisations using more AI moved more changes through their systems. Yet the report still found an association with lower stability — more difficulty keeping the software working properly while it was being changed.

Other studies make the picture even less simple.

In a controlled experiment, GitHub found that developers with AI assistance completed a particular task faster. The research organisation METR, however, found that experienced developers in early 2025 spent 19 per cent longer on real tasks when they were allowed to use the AI tools of the time. The developers themselves believed they had become faster.

The tools have already changed since then. The tasks were different, and METR has itself explained how difficult it is to measure time when several AI programs work at once. The numbers therefore cannot tell us whether AI always makes developers faster.

They show something more interesting:

Producing more code is not necessarily the same as creating more progress.

When proposing another change becomes cheap, the slow part moves. It becomes understanding, testing, coordination and responsibility.

Google studied nine million proposed code changes. The study showed a workflow built around small changes and rapid feedback. The review was not only about finding errors. It was also about ensuring that more people understood the program and shared knowledge about it.

That is worth noting. Even before today’s AI programs, human understanding was a scarce resource.

If AI can now produce more and larger changes at the same time, the need does not disappear. It grows.

In Agent Teams Are Changing My Software Process, I described how several AI programs can investigate, build and check in parallel. That is a real gain. But ten of them can also write faster than one human can understand their combined work.

At some point, telling the human to read faster does not help.

We must change what the human reads.

A Summary Can Hide Things Too

The obvious solution is to have another AI write a short summary of the code.

It is also the dangerous solution.

An AI can write a calm and convincing explanation of a change it has misunderstood. It can emphasise what the developer intended to build and overlook what the program actually ended up doing. It can make five thousand changed lines feel like three simple points.

The easier the summary is to read, the easier it can be to forget how much has been left out.

GitHub itself recommends that AI-based code review supplement rather than replace human review. The tool can miss problems, especially in large and complicated changes. It can also suggest criticism that sounds right without being right.

But that advice contains a paradox.

If AI increases the amount of code, the answer cannot always be that a human should simply read all of it with the same thoroughness as before. Then we have automated production while keeping control as manual labour.

The understandable layer must therefore be more than a summary. It must be a path down to the evidence.

The short explanation must be expandable to show the more detailed recipe. The recipe must be able to point to the parts of the program it describes. If the report says a change may affect payments, it must show why. If it says everything works, it must show the automated checks that were run and the exact version of the program they were run on.

My Graphify experiment therefore distinguishes between three things:

  • What the program can determine from the structure with certainty.
  • What an AI has inferred.
  • What remains unclear.

That is not a minor technical detail. It is the difference between showing what we know and what we are guessing.

A human reading layer must not hide uncertainty. It must make it visible.

A Map You Can Zoom Into

I imagine the future review of software as a map with several zoom levels.

At the top is the intention: What was the change meant to achieve?

The next layer shows the behaviour: What does the program do differently before and after?

Beneath that are the connections: Which other parts might be affected?

Then comes the recipe: How does the most important logic work, explained without all the details of the programming language?

At the bottom are the actual code, the automated checks and the history of the change.

Not everyone has to read every layer every time. Correcting a spelling mistake does not require the same attention as changing who can access a bank account or medical record.

For a dangerous change, a specialist may need to go all the way down into the code. A familiar and limited change might be approved on the basis of visible behaviour, automated checks and a clear path back to the details.

The important thing is not that the top layer contains everything. Then it would become as heavy as the code. The important thing is that every significant claim can be checked in the layer below.

It must be a shortened version with a route back.

That can do more than relieve the programmer. It can bring other professions into the review.

A lawyer does not need to understand every character in the code to judge whether a rule has been interpreted correctly. A doctor can more easily see whether a digital patient pathway matches the real workflow. An editor can assess the rules for publication without first learning the tool on which the website is built.

Today, many such assessments are translated through a developer. The professional describes the intention. The developer reads the code. Together they try to determine whether they are talking about the same thing.

A good reading layer can make the program’s actual behaviour the subject of a more direct conversation.

That is the positive possibility. Higher-level programming languages enabled more people to build software. A higher-level language for review may enable more people to take responsibility for it.

At Digital Medarbejder, I have written that human control does not mean a human must approve every single step. Control is about roles, permissions, documentation, clear stopping points and knowing when a human must take over.

The same applies here.

Human control of AI-produced software does not have to mean that a person mechanically reads every line. It can mean that the human owns the intention, the assessment of risk and the decision about how far down into the details it is necessary to go.

That may be more honest control than a quick tick from someone who has formally seen the change but has not truly understood it.

What Must Humans Still Be Able to Do?

There is a strong objection to the whole idea.

If humans stop reading code, do we not lose the ability to notice when the translation is wrong? Do we not become pilots who can only read the dashboard and no longer understand the engine?

Yes, that risk is real.

A profession can be hollowed out if no one learns the lower layer any longer. We will still need people who can read and write code, understand the computer’s inner workings and discover errors in the translations the rest of us use.

Higher levels have never made the lower levels worthless. They have made them more specialised.

Most Python programmers do not write the computer’s most basic instructions in their daily work. That does not mean those instructions have ceased to exist, or that no one needs to understand them. It means that understanding has been distributed differently.

The same may happen with software review.

Some people will go all the way down into the code. More people will control the program through its behaviour, automated checks, connections to other parts and explanations that can be verified.

The important task will be deciding when the top layer is enough and when the risk requires us to go further down.

That is not a question technology can answer on its own. It is a question of responsibility.

An AI can suggest that a change is small. It cannot assume the consequences if that assessment is wrong. An automated check can show that the known situations work. It cannot decide whether we tested the right thing. A simple recipe can explain the logic. It cannot decide on its own whether that logic ought to exist.

Perhaps the most important human language will therefore not be code, but purpose.

What are we trying to achieve? What behaviour do we accept? Who may be affected? What uncertainty can we live with? When must the machine stop?

Those questions have always sat behind good software. AI merely makes them harder to hide between the lines.

The Next Layer

We have spent roughly 70 years lifting programming away from the machine’s own language.

Each new layer allowed humans to express more without describing every detail below. It gave us more software, more complex systems and more people who could help build them.

Now the machine itself is beginning to write in the languages we created for humans.

That does not mean the development is over. It means that we need another layer — this time in the opposite direction.

We need a language in which the machine can show its work in a form humans can understand. A language of intention, behaviour, connections, risk and evidence. Recipe-like pseudocode can be part of it. Overviews, automated checks and worked examples can be other parts.

The source code will not disappear. But it may no longer need to be the main door every human must pass through in order to exercise control.

When we invented higher-level programming languages, we accepted that humans should not have to think like machines in order to make machines work.

Now we should make the opposite demand.

Humans should not have to learn to read at machine speed in order to control the machine’s work.

The machine must learn to explain itself in human language — and show the receipts.


Sources and Further Reading