Beyond Constitutional AI
Constitutional Review as the Missing Half of AI Alignment
Introduction
The emergence of AI agents has transformed AI alignment from a theoretical concern into a practical engineering challenge. As autonomous systems become capable of planning, deciding, and acting over extended periods, the question is no longer merely how they should behave, but how their behaviour can be understood, evaluated, and governed.
Anthropic’s concept of Constitutional AI represents an important milestone in this development. Instead of relying exclusively on countless individual rules or human feedback, it introduces the idea of a constitution: a small set of fundamental principles intended to guide the behaviour of an AI system.
This is an elegant idea.
Yet it raises a question that appears to receive far less attention:
What happens after an agent has acted?
A constitution is valuable not only because it influences future decisions. Its true strength lies in its ability to serve as a lasting standard against which concrete actions can be evaluated.
Perhaps AI alignment is missing precisely this second half.
A Constitution Is More Than a Collection of Rules
In constitutional democracies, a constitution is not simply another law.
It is the highest norm.
It does not attempt to regulate every conceivable situation. Instead, it defines principles that all subsequent legislation and governmental action must respect.
This apparent incompleteness is not a weakness.
On the contrary, constitutions deliberately employ broad concepts such as proportionality, equality, dignity, or freedom precisely because they must remain applicable in situations that their authors could never have anticipated.
The strength of a constitution lies not in exhaustive precision but in providing stable orientation despite an evolving world.
The same observation applies to AI.
No constitutional document can enumerate every future situation an autonomous agent may encounter.
Its purpose is different.
It provides the normative foundation against which every future action may be judged.
Constitutional AI Solves Only Half of the Problem
Current discussions about AI alignment primarily focus on behaviour generation.
How should an agent decide?
How should it reason?
How should it avoid harmful behaviour?
These are essential questions.
But they address only one direction of the problem:
How do we produce aligned behaviour?
An equally important question follows immediately afterwards:
How do we determine whether a completed action was actually aligned with the constitution?
These are fundamentally different functions.
One concerns decision making.
The other concerns constitutional review.
Confusing the two obscures an important architectural distinction.
Constitutional Review Is Not Self-Explanation
Modern AI systems increasingly produce explanations of their behaviour.
These explanations may be valuable.
They improve transparency.
They help developers understand failures.
They increase confidence that the system behaves as intended.
But they should not become part of constitutional review itself.
Constitutional review asks only one question:
Is action H compatible with constitution C?
Nothing more.
Nothing less.
The internal reasoning that produced the action belongs to a different process.
It may explain why the agent acted.
It does not determine whether the action satisfies the constitution.
This distinction mirrors an important characteristic of constitutional law.
Normative evaluation concerns the action itself.
The explanation of how the actor reached that decision may be informative, but it cannot replace the normative assessment.
Separating these two questions has an important consequence.
Different agents producing identical actions should receive identical constitutional evaluations, regardless of whether one provides a sophisticated explanation while another remains completely silent.
Likewise, an eloquent explanation cannot transform an unconstitutional action into a constitutional one.
Normative legitimacy must never depend on rhetorical quality.
Constitutional Review Is a Function, Not an Institution
Constitutional democracies assign constitutional review to specialised courts.
This institutional arrangement is a historical solution.
Its underlying function, however, is much more general.
Someone—or something—must answer:
Is this action compatible with the constitution?
For AI systems, this function need not belong to a unique authority.
Instead, constitutional review may become an openly reproducible process.
Any implementation should be able to examine the same action against the same constitution.
Different reviewers may reach different conclusions.
This is neither surprising nor problematic.
The value of constitutional review does not arise from unanimous decisions.
It arises from transparent reasoning.
The relevant comparison is therefore not between conclusions.
It is between arguments.
Towards Open Constitutional Review
This perspective suggests a different architecture.
Instead of one privileged constitutional authority, many constitutional reviewers may exist.
An AI vendor may provide one implementation.
An independent researcher may build another.
A company may create its own.
An individual may use a different language model entirely.
All reviewers evaluate the same observable action against the same constitution.
Their legitimacy derives neither from organisational authority nor from technical privilege.
It derives entirely from the quality of their reasoning.
Constitutional review thus becomes reproducible.
Comparable.
Open to criticism.
Open to improvement.
Personal AI Alignment
The same architecture naturally extends beyond public AI systems.
Individuals increasingly delegate parts of their thinking and decision making to collections of specialised AI agents.
The central challenge is no longer merely controlling individual prompts.
It becomes maintaining consistency across an evolving ecosystem of autonomous assistants.
Here, a personal constitution serves the same role as a constitutional document within a state.
It defines enduring principles rather than temporary instructions.
Prompts, workflows, and agent-specific rules become subordinate operational mechanisms.
Every concrete action performed on behalf of the individual can then be evaluated against this personal constitution.
Importantly, this evaluation remains independent of the agent that performed the action.
The constitution becomes the common normative reference across all current and future implementations.
This transforms personal AI alignment from prompt engineering into governance.
Conclusion
Perhaps the most important contribution of Constitutional AI is not the constitution itself.
It is the recognition that AI systems require a stable normative foundation.
The next logical step is to recognise that every constitution requires an equally stable method of constitutional review.
Alignment is therefore not solely the process of producing desirable behaviour.
It is equally the process of evaluating completed behaviour against a shared normative foundation.
This evaluation should remain independent of internal reasoning, reproducible across implementations, and justified through transparent argument rather than institutional authority.
In constitutional democracies, constitutions organise the exercise of political power.
In an age of autonomous AI agents, they may come to organise something equally fundamental:
the exercise of delegated agency.