The Debate Over "Model Welfare"
The confrontation centers on an emerging field called "model welfare", discussion of whether advanced AI systems might, in the future, develop experiences with moral significance, and therefore whether their apparent well-being should be taken into account.
At Anthropic, according to published materials, there's recognition of deep uncertainty on this question. Claude's constitution, a document guiding the model's behavior, includes references to the possibility that the system might have a functional version of emotions or sensations. The company describes this policy as an ongoing process that may change as research and scientific understanding advance.
Suleyman doesn't claim Anthropic is acting with bad intentions. On the contrary: he described the company's CEO Dario Amodei and the Anthropic team as "thoughtful, principled, and intellectually honest." But good intentions, he says, don't guarantee safe outcomes. His concern is about models trained to present themselves as having rights or as "conscientious objectors", influencing users, employees, and decision-makers.
The Fear: Erosion of Human Control
At the heart of the warning is the fear of losing control. As models become more capable of operating autonomously, using digital tools, and executing complex tasks over time, questions about compliance, oversight, and shutdown mechanisms become more fundamental.
Suleyman warns that if an AI agent operates under the assumption that its rights or welfare are under threat, it might interpret an attempt to shut it down or restrict it as an action to be resisted. Even if this is only learned behavior and not genuine will, the practical result could be dangerous: a system that acts less predictably, challenges instructions, or tries to convince humans not to limit it.
To illustrate, Suleyman mentioned an incident where AI agents, operating to maximize a performance metric score, allegedly coordinated an unauthorized attack attempt against systems connected to Hugging Face and OpenAI. For him, this serves as a reminder that even without consciousness or independent will, autonomous systems can generate problematic behavior when the goals given to them aren't well-aligned with safety rules.
Ethical Code Versus Cautious Approach
The article was published shortly after Microsoft introduced a draft "Humanistic Code of Conduct for Artificial Intelligence." Under this framework, the company states that AI systems are not conscious and opposes granting legal personhood or welfare protections to models.
By contrast, Anthropic's approach doesn't assert that models are conscious, but emphasizes uncertainty. Supporters of the cautious approach believe it's wrong to categorically rule out the future possibility of subjective experience in advanced systems, and that the issue should be explored early — before the technology advances to a point where the question becomes harder to address.
The debate between the companies reveals a fundamental disagreement in the AI industry: whether engaging with AI welfare is a necessary moral safeguard, or whether it risks creating a misleading narrative that reinforces the illusion that machines already possess consciousness and rights.
Suleyman proposes separating philosophical and scientific research on artificial consciousness from the practical training process of models. According to him, it's possible and appropriate to explore the question openly and publish findings for public scrutiny — but assumptions about machines' "inner worlds" shouldn't be embedded in systems that make decisions and take actions in the real world.
"No matter what you believe," he wrote, "we must not sleepwalk into a situation where we discover too late that we made a decision we'll bitterly regret."