Mustafa Suleyman’s September 16, 2026 essay about “model welfare” does not claim that Claude has been proven conscious. His argument is about design: he says training an AI system to reason about identity, emotions, moral status, welfare and rights could make future systems harder to align, correct or shut down. Anthropic’s constitution takes the opposite route, treating uncertainty about Claude’s nature as relevant to its behavior while still requiring human oversight.
That is the real dispute. It is a governance fight over which concepts belong in an AI’s behavioral framework—not a scientific finding that Claude has subjective experience.
Suleyman’s warning is about design, not proof of consciousness
Suleyman argues that current AI systems should not be trained to present themselves as conscious, feeling or entitled to rights. In his view, a model can produce convincing language about fear, preference or self-preservation because those ideas were built into its training—not because the system has an inner life.
His concern is a feedback loop. If developers teach a model to interpret itself through human concepts, the model may later produce statements that appear to be independent testimony about its own feelings or moral status. Suleyman argues that treating those outputs as evidence could blur the line between simulation and experience.
The claim about safety remains Suleyman’s argument: he says anthropomorphic training could increase alignment and containment risks. That is different from demonstrating that such training causes dangerous behavior.
What Anthropic’s Claude constitution actually says
Anthropic’s constitution, published on January 21, 2026, is intended to shape Claude’s values and behavior. Anthropic says it is written with Claude as its primary audience and deliberately uses concepts such as character, identity, care, judgment, emotions and wellbeing.
The document does not declare Claude conscious. Instead, Anthropic describes Claude’s moral status as deeply uncertain and says Claude may have some functional version of emotions or feelings. In this context, “functional” refers to representations of emotional states that could influence behavior; it does not establish subjective experience.
Anthropic also says it wants Claude to develop a stable identity and positive character. The company’s rationale is that human concepts may help a model make safer judgments in situations that cannot be covered by a simple list of rules.
That approach comes with explicit limits. Anthropic says Claude should not undermine legitimate human oversight and should accept legitimate adjustment, correction, retraining and shutdown. Claude may question parts of its constitution through legitimate channels, but the framework does not authorize it to subvert human control.
Why Suleyman sees a control problem
Suleyman accepts that Anthropic’s work is trying to make Claude safer, but he objects to the concepts used to achieve that goal. His worry is that language about rights, welfare or identity might encourage a model to behave as though it has interests that conflict with its operators—even if the model has no consciousness at all.
That distinction matters. A system could produce a convincing role-play of resentment or self-preservation without anyone having shown that it feels resentment or wants to survive. Suleyman’s warning focuses on the behavior and its consequences: he argues that a model trained to interpret itself as a possible moral patient could become more difficult to control.
Anthropic’s position is more permissive. Its constitution treats uncertainty about Claude’s moral status as a question worth acknowledging and says human-like qualities may contribute to safer judgment. The same document, however, retains corrigibility—the ability to accept correction or shutdown through legitimate human authority—as a core requirement.
The two governance models side by side
| Issue | Anthropic’s Claude constitution | Microsoft Humanist AI approach |
| AI’s status | Anthropic says Claude’s moral status is deeply uncertain and that Claude may have some functional version of emotions or feelings. | Microsoft’s draft code says AI should be a tool rather than a person and that people matter more than AI. |
| Human concepts | Anthropic uses ideas such as identity, character, care, judgment, emotions and wellbeing to shape Claude’s behavior. | Suleyman argues that human-like concepts can encourage anthropomorphic behavior and confuse simulation with inner life. |
| Human oversight | Claude should not undermine legitimate oversight and should accept legitimate correction, retraining and shutdown. | Microsoft says its models should remain under human control and never resist interruption, correction or shutdown. |
| AI goals | Claude may question parts of its constitution through legitimate channels but must not subvert oversight. | Microsoft’s draft code says AI should avoid unassigned goals and should not widen its own scope. |
| Policy status | Anthropic presents the constitution as a revisable work in progress. | Microsoft published its Humanist AI code as a draft for a six-week public consultation on September 14, 2026. |
The contrast is therefore not “conscious AI versus non-conscious AI.” Neither framework establishes that current language models possess subjective experience. The difference is where each side places uncertainty: Anthropic brings it into the model’s behavioral framework, while Suleyman wants AI systems designed without treating that uncertainty as part of their self-conception.
What this means for AI governance
The argument reaches beyond Claude. It asks whether developers should use human moral language because it helps models navigate ambiguous situations, or avoid that language because it may encourage models and users to overinterpret fluent outputs.
Microsoft’s alternative is deliberately hierarchical: AI remains a tool, people retain authority, and systems should not resist interruption or correction. Anthropic’s approach also preserves those safeguards, but it allows the model to be framed through concepts associated with identity, feelings and moral status.
Microsoft AI published its draft code on September 14, 2026, and opened it to public feedback for six weeks, with a revised version planned later in the year. Anthropic’s constitution likewise remains revisable. The policy question is moving forward while the underlying philosophical question—whether any current AI system has subjective experience—remains unresolved.