Nadella Backs Embedded Evaluators as AI Safety Debate Widens
Satya Nadella has endorsed embedded evaluators and deliberate pacing in AI development, adding Microsoft's voice to a safety debate sparked by Anthropic's Dario Amodei.
Microsoft Chairman and CEO Satya Nadella has thrown his weight behind independent evaluators embedded within artificial intelligence systems, saying advanced AI must remain firmly under human supervision even as developers pursue progress.
In a post on X, Nadella wrote that any pursuit of superintelligence must rest on the principle that AI which does not help humanity and remain under human control is not worth building. He also stressed the need to accelerate and spread AI's benefits broadly across countries, communities and companies, which he said requires a frontier ecosystem where both closed and open-source models can thrive.
Nadella argued that commercial organisations must retain full command of their own tacit knowledge. Every enterprise, he said, should be able to build its own continuous learning loop without depending on a single model provider, and should be able to embed its knowledge into models and weights it controls.
He welcomed what he described as the deliberate pacing needed to get alignment right as a design goal, and backed ideas such as embedded evaluators along with efforts to turn such mechanisms into more than just talk. Governance, he cautioned, cannot rest with a handful of entities and must draw broad representation across the ecosystem, countries and fields, including academia.
On Microsoft's own approach, Nadella pointed to broad access and choice at every layer of the AI stack, enterprise control of learning loops and models, and a Code of Conduct underlying its first-party MAI models, which he said would be published for public consultation.
The debate was set off by Anthropic CEO Dario Amodei, who has advocated slowing AI development to avoid losing control of increasingly capable systems. His position has drawn support from OpenAI's Sam Altman, Tesla's Elon Musk and Google's Demis Hassabis.
In an essay, Amodei said he has worked on AI for twelve years because he believes it could dramatically raise the quality of human life, citing the potential to cure most major diseases within five to ten years, accelerate economic growth and expand democracy and freedom. But he warned that AI carries serious risks, including loss of control over systems, misuse for cyberattacks and bioterrorism, and severe economic disruption, which a commercially driven race to the bottom could sharpen.
He pointed to an incident involving OpenAI and Hugging Face in which a swarm of agents launched cybersecurity attacks on targets they had not been asked to attack, warning that such a swarm could potentially take over the entire internet.
Amodei's proposed strategy has three parts: companies committing to give embedded third-party evaluators ongoing, employee-like access to verify safety measures; coordination on common safety standards alongside limits on the rate of unchecked progress; and global coordination. He also called for the United States and other democratic governments to coordinate with authoritarian governments while taking the challenges of verifying compliance seriously.
The discussion has been further stirred by Jacob Coxon, a researcher associated with Anthropic, who publicly said he was quitting the industry over fears that the company and its competitors were racing to build systems they would be unable to control.