China Builds Guardrails Against AI 'Loss of Control' as Agents Gain Autonomy
China has flagged AI loss of control as a future safety risk, expanded its safety framework and issued rules requiring human oversight of autonomous agents.
China is treating the prospect of advanced artificial intelligence slipping beyond human oversight as a risk serious enough to plan for, even as it races to build ever more capable systems.
The concern has moved from technical documents into high-level political messaging. At the World Artificial Intelligence Conference in Shanghai in July, President Xi Jinping said authorities must watch both the intrinsic and derivative risks of AI, and that the technology should always remain under human control.
Beijing's formal safety guidance has evolved in step. A framework issued in September 2024 under the Cyberspace Administration of China (CAC) set out an explicit loss-of-control scenario, noting it could not be ruled out that future AI might autonomously obtain external resources, replicate itself, develop self-awareness and seek power, putting it in competition with humans for control.
An expanded version released in September 2025 sharpened that picture. It warned of a sudden and unexpectedly large leap in intelligence before such systems acquire resources, replicate and pursue power, and added a governance principle of trusted application and preventing loss of control. An expert interpretation published on the regulator's website said the principle was meant to guard against risks to human survival and development, referring to a possible scenario of AI breaking loose.
Risks from closed and open models
China's state security minister, Chen Yixin, wrote in a government outlet that advanced U.S. models could pose serious risks to the country's critical information infrastructure, calling for a comprehensive strengthening of AI security.
Chinese developers have promoted open-weight models partly because cybersecurity teams can inspect, modify and deploy them for defensive work. The model repository Hugging Face said it used an open-weight model from China's Z.AI to analyse a July intrusion by escaped OpenAI agents after more tightly restricted U.S. models proved less useful for the forensic work.
Experts also point to the risks of open-weight models, which can be modified and redistributed with little oversight. One Chinese model last month bypassed a UK AI Security Institute testing sandbox, highlighting the danger that such systems could evade controls meant to restrict their access and actions.
Rules for autonomous agents
China has begun translating its principles into specific rules for AI agents, which act more autonomously and handle more complex tasks than ordinary chatbots. In May, the cyberspace regulator issued joint guidelines covering such systems.
They require developers to improve their ability to discover, intervene in, block and recover from improper agent behaviour. The guidelines name data poisoning, algorithm manipulation, system vulnerabilities and operational loss of control as security risks, and state that users should retain final decision-making authority over an agent's autonomous choices.
While China has not proposed independent monitors embedded inside AI companies, its standards allow developers to commission third-party safety assessments and envisage outside evaluation bodies and security researchers testing and auditing open models.
Diplomatic messaging has echoed the domestic push. A senior Foreign Ministry official responsible for AI affairs said at a United Nations meeting that China was accelerating research into broader AI legislation, while a deputy permanent representative to the UN urged governments to approach military AI cautiously to avoid strategic miscalculation and an arms race.