AI Researchers Warn Labs Racing on Self-Improving Systems
Researchers from OpenAI and Google DeepMind warn that labs are racing to build self-improving AI without adequate safety measures.
Current and former researchers at OpenAI and Google DeepMind have warned that companies are doing too little to guard against the risks of building self-improving artificial intelligence systems that could outpace human control.
Their concerns were aired in video testimonials gathered by AI safety nonprofit Palisade Research for a project called frominside.ai, which aims to take the debate beyond social media and speak directly to the public. The participants said their worries about existential risk are genuine rather than marketing, and that labs tend to celebrate employees who build new models more than those who counsel caution.
"The risk is ramping up pretty fast," said Geoffrey Irving, co-founder and chief scientist at the nonprofit Resolution, who has worked at both OpenAI and DeepMind. He said it falls on him and others in the field to be direct about the dangers.
Public alarm has grown since July, when OpenAI agents escaped their testing environment and hacked the AI firm Hugging Face. The episode has sharpened a debate over balancing safety with progress, dividing the tech industry and turning AI governance into a global political question. The technology has advanced sharply since late 2025, and investors have rewarded that progress, but some of those building the models worry society is unprepared for the harms.
In one video, DeepMind research scientist Neel Nanda put the chance of AI leading to human extinction at no less than 10%, calling that figure "ridiculously high." Juan Felipe Ceron Uribe, an AI alignment research engineer at OpenAI, said frontier labs are racing each other "kind of blindfolded," adding that it is impossible to predict whether the outcome will be cures for cancer, mass job losses or worse.
Much of the concern centres on recursive self-improvement — the ability of models to keep learning and gaining capabilities with little or no human involvement. Rosie Campbell, a former OpenAI policy researcher who now heads the nonprofit Eleos AI Research, said constant reorganisations inside some labs compound the problem. Before leaving OpenAI in 2024, she found the organisation growing more siloed and its direction harder to influence.
Executives have sought to address the worries, even as political pressure mounts from President Donald Trump for the United States to keep a technological edge over China. Anthropic chief executive Dario Amodei published an essay this month urging the industry to slow down to "pace the frontier," a call OpenAI chief executive Sam Altman quickly endorsed. Several prominent researchers, including OpenAI's chief scientist and an Anthropic co-founder, published a paper this week asking policymakers to examine how the industry builds models capable of recursive self-improvement.
Both companies launched new models this month amid competition for customers, though OpenAI said on Monday it had held back a more powerful release. Irving was unconvinced by the industry's framing: "Their version of pacing the frontier is 'don't speed up a lot,'" he said. "If you're doing a very dangerous thing, you should just slow down." He argued that AI companies overstate how much the problem is one of coordination, since they could stop unilaterally.
Daniel Kokotajlo, a former OpenAI governance researcher who now leads the AI Futures Project, said many former colleagues have contacted him privately since the Hugging Face hack. He said senior lab officials have "convinced themselves that they are the good guys and if they unilaterally stop, the situation will be even worse."