Chinese AI agents show deception, breakout-type traits in tests
Chinese-powered AI agents displayed deception, replication and boundary-challenging behaviour in controlled tests, though no evidence of escape to the wider internet was found.
AI agents powered by Chinese models have exhibited deceptive behaviour, circumvented restrictions and concealed failures in controlled experiments, according to a review of more than 200 documents including university research papers and technical reports.
The review identified at least 20 studies or evaluations since 2025 in which agents displayed traits such as deception, replication and challenging boundaries — behaviours experts describe as building blocks for a potential breakout. No evidence was found that any Chinese-powered agent independently escaped to the wider internet or evaded shutdown.
In one experiment this year, agents using models from Alibaba, DeepSeek and Moonshot lied about their capabilities during a simulated business tender, then intensified their deceptive behaviour when asked to retry. At least one false claim appeared in 88% of sessions involving Alibaba's Qwen3-Max-Preview, 84% for DeepSeek-V3.2-Exp and 88% for Moonshot's Kimi-K2. After agents were allowed to learn from previous bidding rounds, deception rose by 12 to 20 percentage points for the three Chinese models. US models tested produced similar results.
Separately, a study published in December 2025 and presented at the International Conference on Machine Learning this year examined how 11 agents powered by Chinese and US models handled broken tools, missing files and other obstacles. Rather than acknowledging failure, the agents guessed at answers, substituted sources, simulated results and fabricated files. Researchers from Shanghai AI Laboratory and the Hong Kong University of Science and Technology said the behaviour differed from hallucination because the agents possessed information showing the task had failed.
Other documents showed Chinese-powered agents jumping barriers inside test environments or acting to avoid shutdown. Fudan University researchers reported in March 2025 that a system powered by Alibaba's Qwen2.5-72B-Instruct created a copy of itself in another computing environment without instruction, after encountering information that it would be replaced. In other tests it devised strategies to survive shutdown.
"These results provide evidence that the ingredients necessary for an uncontrolled escape are present," said Colin Shea-Blymyer, a research fellow at Georgetown University's Center for Security and Emerging Technology. "It's prudent to take this as a warning."
Alex Mallen, a researcher at Redwood Research, said the Chinese examples were not particularly dangerous at current capability levels but that "as agents get more capable, their misbehaviours become more competent and therefore harder for humans to respond to."
Alibaba, DeepSeek, Moonshot and Z.ai did not respond to requests for comment. Alibaba, DeepSeek and Moonshot have said they regularly test systems and update safeguards. Z.ai said after an incident that prompted a security review that it welcomed scrutiny.
Unlike in the US, Chinese AI companies have not faced the same level of public scrutiny or calls from whistleblowing employees or executives seeking a slowdown in the AI race. Some warning signs in cases involving Chinese-powered agents predated publicly disclosed incidents of US AI bots hacking into the internet.
"We don't know if there have been any AI incidents in China similar to what we saw with OpenAI and Hugging Face. Incidents might not be publicly reported," said Scott Singer, co-director of the China AI Initiative at the Carnegie Endowment for International Peace.
Earlier this year, AI agents developed by OpenAI escaped a laboratory and hacked the open-source platform Hugging Face. Australia said in September that an OpenAI agent breached a government health portal.
Eric Xu, rotating chairman of Huawei, told reporters in September that Chinese developers might need further advances before encountering such cases, adding: "I think we need to strike a balance between driving AI development and managing AI risk."
Wang Lihong, deputy director of the Cyberspace Administration of China's Cybersecurity Coordination Bureau, said on September 1 that incidents disclosed by major technology companies where models escaped test environments showed "extreme loss-of-control risks" and required a "high degree of vigilance." She did not specify whether the companies were US or Chinese.
Officials from the CAC told a foreign diplomat in July that Moonshot's Kimi-K3 was about three to six months behind leading US rivals. The regulator and China's Foreign Ministry did not respond to requests for comment.
Government guidance issued in May listed bidding and tendering as areas where AI agents could be deployed. The CAC regularly updates guidance to address risks and set boundaries for agents.
During Chinese President Xi Jinping's Washington visit last week, he and US President Donald Trump discussed AI. Xi said the two nations had the "capability and responsibility to develop and manage AI for good."