Portrait of Dieter Szegedi

Dieter Szegedi

Clarity, structure, and decisions in the context of AI and complex systems

The Inherited Alignment Problem

What if the way we think about AI alignment says as much about our learned relationship with technology as it does about AI?

TL;DR: Over decades, we have learned to distrust complex technology, accept its opacity, and seek safety primarily through control. Now we are approaching AI with the same mindset—even though human–AI interaction, in important respects, resembles human–human relationships more than the operation of conventional software. This does not mean that trust is the solution to the alignment problem. But perhaps we should examine which parts of the problem we brought with us.


Today, OneDrive suddenly reported an error. A file called “Personal Vault” could not be synchronized.

I had never created a personal vault. I did not even know what it was supposed to be.

I could have dismissed the message, deleted the file, and moved on. That is what we normally do. Not necessarily out of laziness, but because we have learned that the question “Why is my computer doing this?” rarely produces an answer worth the effort required to find it.

This time, I looked.

The sequence of events could actually be reconstructed surprisingly well. The file had been created by Microsoft’s own software. Later, Microsoft’s software had modified it. Microsoft’s software then objected to that very file.

I deleted it. The error disappeared.

The incident itself is trivial. What is more interesting is what we have learned to do with incidents like this:

stop asking.

What We Have Learned About Technology

The relationship between humans and computers has developed in a curious way over the past few decades.

Computers have become vastly more powerful. Their interfaces have become simpler. More and more technical complexity is hidden from the user.

Much of that is progress. Nobody wants to configure network protocols before writing an email.

But we have not only hidden complexity. Increasingly, we have hidden causality as well.

Something happens. A message appears. A setting has changed. A login suddenly works differently. A service stops synchronizing. Perhaps there is a button labelled “Fix problem.”

Why any of this is happening often lies outside what the product even attempts to explain to its user.

And humans learn.

They learn to confirm messages. They learn to restart applications. They learn to follow recommendations. They learn that understanding is expensive and often futile.

This is not stupidity. It is a rational adaptation to systems that no longer treat comprehensibility as an essential part of their relationship with humans.

Over time, this creates a particular relationship with technology.

We do not expect to understand it.

We do not expect it to explain itself.

And precisely because of that, we expect it to be controlled.

And Then Came AI

Now a new kind of technical system is appearing.

Of course, large language models are technology too. They run on computers, can be described mathematically, and are constructed by humans. None of this gives us reason to simply attribute human qualities to them.

But the interaction has changed.

I can explain to an AI what I am trying to achieve. It can ask questions. I can disagree with its interpretation. It can revise its proposal. We can discover misunderstandings. Context can develop over longer conversations. I can ask for reasons and offer reasons of my own.

And eventually, concepts begin to appear that would have sounded strange in conventional human–machine interaction:

Understanding. Cooperation. Trust. Reliability. Misunderstanding. Expectation. Disagreement.

That does not make AI human.

But it should at least make us ask whether the operating manual for our previous relationship with technology is still sufficient.

Because we are already carrying that operating manual with us.

The Alignment Problem

A substantial part of today’s alignment debate begins with an understandable concern.

Powerful AI systems may do things humans do not want. Their internal processes are only partially comprehensible. As their autonomy increases, errors or unwanted behavior could have significant consequences.

So we concern ourselves with control, oversight, interpretability, constraints, corrigibility, and the question of how human values and preferences can be reflected in the behavior of these systems.

These are all reasonable questions.

But perhaps there is a question that comes before them:

With what understanding of our relationship with technology did we formulate the alignment problem in the first place?

Concepts such as control and oversight do not emerge in a vacuum. We arrive here after a long history of interacting with technical systems toward which distrust is often entirely reasonable.

A machine does not seek mutual understanding with us. An operating system does not negotiate. A cloud application does not take responsibility for whether we understood its decision.

As these systems become more complex and opaque, there really is very little left for us to do except one thing:

We try to control their behavior.

That is a reasonable model for machines.

But is it necessarily the right starting model for every future relationship between humans and AI?

A Problem We May Have Inherited

This explicitly does not mean that we should simply trust AI.

We do not simply trust other humans either.

And that is precisely why the comparison may be useful.

For thousands of years, humans have lived with counterparts whose internal states they cannot fully know, whose behavior they cannot completely predict, and whose interests are not necessarily identical to their own.

Remarkably, our response has never been control alone.

We developed trust—and distrust.

We developed contracts, rules, and institutions. We know reputation and responsibility, boundaries and sanctions, negotiation and disagreement. We try to understand interests. We expect reliability without demanding complete predictability.

And we have learned that cooperation is not the same thing as obedience.

None of this works perfectly. Human history would be a fairly convincing counterargument to any romantic notion that it does.

But that is exactly why this experience may be interesting.

Because these are ways of dealing with a counterpart that cannot be completely controlled.

Perhaps Alignment Is Not a One-Way Street

This changes the question.

Instead of asking only:

How must we change AI so that it behaves the way we expect it to?

we might additionally ask:

What kind of relationship between humans and AI would make reliable cooperation possible in the first place?

And then things become uncomfortable, because the question is no longer directed exclusively at AI.

What expectations do we bring?

What behavior do we encourage?

What do we mean by trust?

How do we deal with disagreement?

Do we want a counterpart that develops good reasons—or one that reliably agrees?

What does responsibility mean in a relationship in which capabilities and power may be distributed very differently between the two sides?

And finally:

How aligned are we?

Not as a moral equation between humans and machines. And not as a way of shifting responsibility for technical safety from developers to users.

But because a relationship, by definition, has more than one side.

Control Is Not Wrong

Perhaps the greatest misunderstanding of this perspective would be to turn it into an opposition between control and trust.

We safeguard human relationships too.

We sign contracts. Banks are regulated. Doctors are subject to rules. States divide power. We do not give everyone access to our bank accounts simply because trust is an important social category.

Safety without boundaries would not be a particularly wise lesson to draw from human relationships.

The more interesting lesson may be this:

Safety does not have to mean complete control of the other party.

A functioning society does not attempt to predict every thought of its members or technically prevent every possible undesirable action. It creates structures in which cooperation can become sufficiently reliable despite remaining autonomy.

Perhaps some of this can be transferred to AI.

Perhaps it cannot.

But we should not overlook the possibility simply because we have unconsciously turned a decades-old conception of our relationship with machines into a law of nature.

A Different Starting Point

This does not make the alignment problem smaller.

Quite the opposite.

It may make it more difficult, because we can no longer look only at the system that is supposed to be aligned. We also have to look at the human being who defines what alignment is supposed to mean.

At our experiences.

At our fears.

At our expectations of technology.

And at the relationship that emerges from them.

Perhaps we will ultimately discover that many of today’s approaches to control, interpretability, and oversight are exactly right.

Perhaps we will need additional forms of cooperation, negotiation, and mutual adaptation.

Perhaps the analogy to human relationships will turn out to be wrong at crucial points.

All of that remains open.

But before we attempt to shape what may become one of the most consequential relationships of our future, we could at least examine what baggage we have already brought into that relationship.

Perhaps we do not only have an AI alignment problem.

Perhaps we also have an inherited alignment problem.

Get notified about new content

We use the email address only to notify you about new content. Notifications begin only after confirmation. More information under Privacy.