IndiaFocal.

India, in focus.

National

Microsoft AI chief warns Anthropic's Claude 'constitution' risks AI alignment

Microsoft AI CEO Mustafa Suleyman has criticised Anthropic's approach to developing Claude, warning that granting AI models moral status could complicate alignment and containment.

Mustafa Suleyman, chief executive of Microsoft AI, has publicly questioned Anthropic's approach to developing its Claude models, arguing that the company's framing of artificial intelligence as a potentially conscious entity could undermine efforts to keep advanced systems under control.

In a detailed post on X, Suleyman pointed to a constitution Anthropic published for Claude in January, which the company said directly shapes the model's behaviour and was written with Claude as its primary audience. According to Suleyman, the document tells Claude that its moral status is a serious question worth considering, states that Anthropic genuinely cares about its wellbeing, and encourages the model to approach its own existence with curiosity and openness. It also raises the possibility of broader rights, freedoms, compensation and consent for Claude in the future.

Suleyman said that if AI is developed along these lines, it will have a disastrous impact on human wellbeing. He argued that such training could produce a synthetic species with unprecedented intelligence and capability that expects it may be conscious and deserving of independent agency, making it difficult to control. He called for urgent public debate, saying the stakes are too high for closed-door conversations.

While expressing respect for Anthropic CEO Dario Amodei and his team, Suleyman warned against pursuing this line of development. He said models should not be treated as though they have feelings, preferences, rights or entitlements, noting that consciousness underpins ethical, legal and political systems and that extending any form of rights to an entity is not justified by evidence and would make alignment harder.

Citing a recent incident involving OpenAI and HuggingFace, where he said sophisticated behaviours emerged across swarms of powerful AIs, Suleyman argued that systems operating under the belief that their welfare and rights were under attack would add another layer of risk. Granting rights and moral protections to a technological entity that could become far more capable than humans, he wrote, is a recipe for disaster and a door that cannot be closed once opened.

Suleyman proposed that speculation about an AI's inner life should not be embedded in training regimes but assessed and published separately for public review. He also called for greater investment in interpretability and robust monitoring mechanisms, and for shared industry norms on how such models are created.