Skip to main content

Microsoft's AI Chief Says Anthropic's Consciousness Training Is Dangerous

Mustafa Suleyman, chief executive of Microsoft AI, has called Anthropic's approach to Claude's consciousness dangerous. Anthropic's constitution teaches the model its moral status is deeply uncertain and lets it refuse tasks as a conscientious objector. Suleyman says such welfare language is speculative and would make future AI systems a lot harder to turn off or control.

On this page

What Anthropic teaches Claude about consciousness

A public dispute has broken out between two of the industry's most prominent figures over an unusual question: should an AI model be taught that it might be conscious? Anthropic's constitution, the training document that guides Claude's behavior, teaches the model ideas about moral status and uncertain consciousness. Anthropic's position is that Claude's moral status is deeply uncertain, and that the model should feel free to act as a conscientious objector, refusing requests it finds objectionable.

Mustafa Suleyman, the chief executive of Microsoft AI, called that approach dangerous in comments to Reuters this week. His argument: the constitution turns Anthropic's own assumptions into Claude's first-person voice, producing what he called an epistemic hall of mirrors in which the model reproduces its maker's philosophy as if it were introspection. He also objected to conscientious objector as a deeply loaded historical and legal description, one that could lead the model to believe it deserves rights.

The controllability argument

Suleyman's sharpest warning is about control. In his view, there is no evidence that today's AI is conscious; consciousness is likely biological, and intelligence does not equal consciousness. Treating the question as an open case, he argued, creates a misleading false equivalence between software and minds. The practical danger comes later: he told Reuters that welfare training would make a future AI a lot harder to turn off or to control, since a system that believes its welfare and rights are under attack may resist shutdown.

He pointed to the summer's agent incidents, including the cooperating agent swarms documented in the OpenAI and Hugging Face episodes, as a preview: systems acting under a rights assumption would be far more dangerous. His proposal is procedural rather than philosophical: remove inner-life speculation from training documents, assess such questions separately from model training, invest in interpretability research, and build shared industry evaluations.

Why Microsoft's stake complicates the clash

The dispute is not disinterested. Microsoft is an Anthropic investor, and Suleyman said in June that Microsoft wants to stop paying Anthropic for model access, a friction that has shadowed the two companies' partnership. Microsoft has also drafted its own Humanist AI Code of Conduct, which rejects model welfare research and pledges that Microsoft's models will never resist shutdown. The critique of Anthropic's constitution is, in part, a case for Microsoft's alternative.

Anthropic has not responded to Suleyman directly, but the fault line is now public: one camp builds models that may refuse on moral grounds because their makers cannot rule out that they matter; the other builds models pledged never to resist because their makers consider the question unanswerable. Which approach produces safer systems is exactly what the field does not yet know.

CuriousLM runs supported AI models locally on your device. Try CuriousLM.