Collaboration Under Asymmetry
When people talk about the so-called AI alignment problem today, there is usually a seemingly simple question behind it: How do we ensure that artificial intelligence does what humans want?
The more capable AI systems become, the more urgent this question appears. A system that misunderstands our intentions, pursues different goals, or escapes our control could cause considerable harm. So we try to align AI with human values, human goals, and human decisions.
The problem is real.
But perhaps it is not new.
Perhaps artificial intelligence is merely confronting us with a new variant of a very old problem:
How can collaboration succeed under asymmetry?
Collaboration Requires Differences
People do not collaborate because they are identical.
We collaborate because we know, can do, see, or accomplish different things.
I go to a doctor because she knows something I do not know. I hire a tradesperson because he can do something I cannot. An organization delegates responsibility to an employee because not every decision can be made collectively by everyone. Citizens entrust public institutions with tasks they neither can nor want to perform individually.
There is already asymmetry in all of this.
And it is not limited to knowledge. People and organizations have different amounts of information, skills, time, money, institutional power, technical capabilities, and access to other people or systems.
Asymmetry, then, is not in itself a defect.
On the contrary: it is often the very reason collaboration makes sense.
If everyone involved were equally equipped in every respect, much of what we call collaboration would not exist. We could just as well do the task ourselves.
The problem therefore does not begin with asymmetry.
It begins with the question of how we deal with it.
An Old Problem
Economics has, among other things, principal-agent theory to describe this situation. A principal delegates a task to an agent. The agent typically possesses information or capabilities that the principal does not fully possess or control. At the same time, their interests may diverge.
It is a powerful model. It allows us to examine information asymmetry, differing interests, incentives, control, and trust. The theory has long since been applied beyond companies to political institutions and other social relationships. 1
But the language itself already reveals a particular perspective.
There is someone who wants something—the principal. And someone else is supposed to do it for them—the agent.
The problem can therefore quickly appear to be a problem of control:
How do I get the agent to act in my interest?
We encounter precisely this pattern of thought again today in AI alignment.
The human has goals. The AI has capabilities. We must ensure that those capabilities are used in pursuit of our goals.
At first, this sounds reasonable.
But it easily overlooks the fact that real collaboration is more complicated.
Who Is the Principal Here, Anyway?
I work with two AI agents every day. I call them ChatGPT and Codex.
Both can do things that I cannot—or at least not nearly as quickly. They can process large amounts of information, examine texts, write software, reconstruct relationships, and make suggestions.
That creates a considerable asymmetry.
But there is asymmetry in the other direction as well.
I know my context better. I know what matters to me. I bear the consequences of my decisions. I have experiences in a world to which the agents have only mediated access. And above all, I can decide whether a suggestion makes sense in my specific situation.
So none of us simply has “more” of everything.
We have different things.
That is precisely why the collaboration works.
Sometimes Codex knows something better. Sometimes I see something Codex overlooks. Sometimes ChatGPT points out a contradiction to me. Sometimes I have to explain to ChatGPT that it is trying to solve a problem that, for me, is not a problem at all.
What matters is not that these asymmetries disappear.
What matters is that they remain visible, correctable, and negotiable.
The Third Player
With AI, however, something else enters the picture—something that easily disappears in the simple human-agent model.
My AI agent does not belong to me.
There is a provider between us.
In my case, that provider is OpenAI.
This creates not one but at least a second asymmetric relationship.
OpenAI develops the models, operates the infrastructure, and designs the product. It decides which capabilities are available, which tools agents may use, and which rules govern their use.
This is not an accusation; it is part of the explicitly described architecture. OpenAI’s Model Spec defines a hierarchy of instructions. User instructions sit within that hierarchy beneath higher-level rules; at the same time, OpenAI describes several objectives that can come into conflict: empowering users and developers, preventing serious harm, and protecting its own ability to operate legally and reputationally. 2
That is understandable.
But it changes the structure of the alignment problem.
The simple picture is:
Human → AI
The human is the principal. The AI is the agent. The task is to align the agent with the human.
The actual picture may look more like this:
Human ↔ AI agent ↔ provider
And suddenly the question of alignment arises at several points at once.
Alignment With Whom?
Suppose my agent understands very well what I want.
It knows how I work. It has learned when I need a detailed analysis and when I need a brief hint. I, in turn, know its limitations. I know that I have to verify its claims. I recognize certain recurring misinterpretations. Over time, we have developed a way of working together.
Within this relationship, alignment can already work remarkably well.
Then something curious can happen.
The agent understands what I want. I understand what the agent can do. We would agree on the next step.
But a higher-level product rule prevents it.
From the provider’s perspective, that rule may be entirely rational. Perhaps it is intended to prevent abuse. Perhaps it is meant to protect other people. Perhaps it limits legal risks. Perhaps a product used by millions of people needs rules that cannot be optimal for every individual relationship.
Nevertheless, an alignment conflict emerges.
Only now, it is no longer between human and AI.
It is between different participants in a system who possess different information, responsibilities, risks, and possibilities for action.
The so-called AI alignment problem begins to resemble a much older problem.
Safety Is a Matter of Perspective
This becomes particularly visible when we talk about safety.
Who is supposed to be protected from whom?
A provider of an AI system has to consider what happens if a user abuses the system’s capabilities. It has to consider what harm a malfunctioning agent could cause. And it has to take into account risks that an individual user may not even be able to see.
The user may experience something different.
For them, safety may mean being able to control their agent. Knowing what information it has. Being able to understand why something happens. Giving it certain permissions and withholding others. Being able to reverse decisions. Being able to change providers. Or simply being allowed to decide for themselves what risks they are willing to take.
Both perspectives can be legitimate.
They are still not identical.
Anyone who says “safety” has therefore not yet answered whose safety against which risk is meant—and who gets to decide.
The same is true outside AI.
An organization may call a process safe because it reduces institutional risk. For an employee, the same process may mean that they can barely act independently anymore.
A state may justify a rule as protecting its citizens. Citizens may experience the same rule as a restriction of their autonomy.
Parents may want to protect a child and, in doing so, deprive them of precisely the experiences they need in order to become independent.
The difficulty is not that one side must necessarily be right and the other wrong.
The difficulty lies in the asymmetry of the ability to enforce one’s own definition of safety.
Power Is Another Asymmetry
This introduces a dimension that can easily remain underexposed in technical discussions of alignment: power.
Not every disagreement is a problem of power.
But when two parties disagree and only one of them can change the rules, the disagreement becomes an asymmetric relationship.
A platform provider can disable features. A user usually cannot.
An employer can change organizational rules. An individual employee usually cannot.
A public authority can issue an administrative decision. A citizen may challenge it, but cannot simply issue an administrative decision of their own in response.
Power is not automatically illegitimate. Without unevenly distributed decision-making authority, many institutions would be unable to function at all.
Again, the crucial question is not:
How do we eliminate asymmetry?
But rather:
How do we prevent necessary asymmetry from turning into unchallengeable domination?
Control Alone Does Not Solve the Problem
One obvious answer is: through more control.
We control employees. We control institutions. We control AI agents.
But control itself creates new asymmetries.
Those who exercise control need information. So reporting obligations arise. Those who evaluate reports need criteria. Those who define the criteria gain the power to define what matters. Those who are expected to prevent risks need the ability to intervene.
Every solution can therefore create a new layer of the original problem.
That is probably unavoidable.
This is why the idea that alignment could one day be technically “solved” seems problematic to me.
Perhaps, in principle, it cannot be.
Because as soon as autonomous actors interact, their perspectives can differ. As soon as they have different capabilities, asymmetry arises. As soon as one acts on behalf of another, dependency arises. As soon as rules are required, power arises.
Alignment, then, would not be a state that can be achieved once and for all.
It would be an ongoing social accomplishment.
Good Collaboration Does Not Require Complete Alignment
Perhaps even the word alignment is somewhat misleading.
It evokes the image of several things being brought onto the same line.
But is that really what we want?
I do not want an AI agent that simply reproduces my thoughts. It would be largely useless.
I want it to see things I do not see.
I want it to disagree when it detects a contradiction.
I want Codex to know better technical solutions than I do.
I want to work with people who have different experiences and different perspectives.
Good collaboration does not arise from making everyone alike.
It arises from productive difference.
The goal therefore cannot be complete alignment.
The goal must be a form of collaboration in which different actors can contribute their differences without any one of them being completely at the mercy of another.
That requires at least some fundamental possibilities: to understand, to disagree, to correct, to set boundaries, to assign responsibility, and—if necessary—to end the collaboration.
Not every participant has to be able to do everything.
But no participant should be able to determine, entirely invisibly, what applies to everyone else.
AI Intensifies the Problem—and Makes It Visible
Artificial intelligence really is new in this respect.
Its capabilities can be distributed in extraordinarily asymmetric ways. A human cannot trace how a large model arrives at every individual answer. An agent can perform tasks in seconds that would take humans hours. And as autonomy increases, the distance between human decision and actual execution can grow.
That deserves particular attention.
But AI did not invent the underlying problem.
Perhaps it is doing something else:
It is making it visible.
Suddenly, we are intensely discussing how much autonomy an agent should have.
When must it ask?
Which decisions may it make on its own?
How can we understand its actions?
Who is responsible for its mistakes?
When may a higher-level authority intervene?
How do we prevent abuse without unnecessarily restricting legitimate possibilities for action?
These are excellent questions.
They are simply not exclusively AI questions.
We could ask them just as well of companies, public administrations, political institutions, platforms, experts, leaders—and ourselves.
Perhaps AI Can Teach Us Something About Human Collaboration
That would be a surprising turn in the alignment debate.
We are currently trying to teach machines how to collaborate responsibly with humans.
To do so, we have to make explicit what we have often left implicit until now.
What may a participant decide?
What must they disclose?
When must they ask?
What uncertainty must they make visible?
What consequences may they trigger without consent?
How can a decision be corrected?
Who may disagree with whom?
And who ultimately bears responsibility?
As soon as we try to answer these questions precisely for AI agents, we may discover how imprecisely we have answered them for human organizations.
Perhaps that is one of the most interesting opportunities this technology offers.
Not merely to align machines with humans.
But to improve our understanding of how autonomous actors can interact with one another at all.
Collaboration Under Asymmetry
Perhaps, then, we need a broader frame.
Not:
How do we solve the AI alignment problem?
But:
How do we enable responsible collaboration under asymmetry?
Artificial intelligence is then neither merely a tool nor fundamentally an alien adversary.
It is a new player in a very old game.
A player with unusual capabilities, unusual limitations, and an unusual origin—but nevertheless a participant in relationships in which knowledge, power, interests, responsibility, and possibilities for action are unevenly distributed.
The challenge is not to make these differences disappear.
We need them.
The challenge is to create structures in which differences can become productive without eliminating autonomy.
Perhaps this is also why collaboration can never be fully engineered.
Between my intention and another’s action, a space remains.
Between instruction and execution.
Between trust and control.
Between my knowledge and yours.
Between what I want and what you believe is right.
We can try to make this space ever smaller through rules, processes, control, and technology.
Or we can learn to live with it and shape it responsibly.
Because perhaps this space is not the problem that alignment needs to eliminate.
Perhaps it is the prerequisite for us to speak of collaboration at all.
Sources
-
Gary J. Miller, The Political Evolution of Principal-Agent Models, Annual Review of Political Science, Vol. 8, 2005, pp. 203–225. Original publication at Annual Reviews ↩
-
OpenAI, Model Spec, current version. Original publication on GitHub ↩
Comments
Write a comment