AI Labs Edge Toward Self-Improving Models, Raising Control Concerns
Leading AI labs are advancing toward recursive self-improvement, where models help build their successors, prompting fresh debate over human control.
The long-held ambition of building artificial intelligence that can improve itself is moving from speculation toward concrete engineering targets, as several leading laboratories disclose how their models are already assisting in the development of more capable successors.
The concept, known as recursive self-improvement, or RSI, describes a process in which an AI system helps design the next version of itself, which in turn builds the following generation. The prospect has drawn attention both for its potential to accelerate breakthroughs in science and medicine and for the risks it poses if human oversight erodes.
Anthropic said this week that its Claude model now leads 26% of the company's model research and development, handling most of a given task from a high-level prompt while remaining under human supervision. The company has not stated how close it is to fully autonomous improvement, but its disclosure offered a rare public metric and a nudge to competitors to share similar data.
OpenAI said this month it has built an automated "research intern" capable of carrying out well-defined tasks that would take a skilled researcher several days. The company aims to create an automated AI researcher by March 2028. It cautioned that while RSI could help align model behaviour with human values, rapid self-improvement is not necessarily a desirable outcome, and that progress must depend on preserving human control and on informed public choices.
Elon Musk has taken a more aggressive line, saying that for xAI's Grok models humans are gradually being removed from the improvement loop and that each successive model is built by the one before it. He said the process is not yet fully automated but could reach that point by the end of this year, and no later than 2027.
Microsoft is charting a different course. Mustafa Suleyman, chief executive of Microsoft AI, has described a goal of "humanist superintelligence" — advanced capabilities placed at the service of people, carefully calibrated and bounded rather than an unrestricted autonomous entity.
Definitions of RSI vary across the industry. Some labs treat any AI feedback on model improvement as a form of it, while others reserve the term for systems that pursue the goal fully autonomously. Anthony Aguirre, president of the nonprofit Future of Life Institute, said the crucial factor is speed: as AI takes on more of the work, improvement accelerates far beyond human pace. He called the pursuit extremely dangerous.
John Thickstun, a Cornell University computer scientist who studies methods for controlling AI behaviour, said a more grounded view is that a form of recursive self-improvement has been underway for years, with past model generations writing code for the systems that train their successors. Those efforts have yielded incremental gains rather than large creative leaps, though researchers such as OpenAI co-founder Andrej Karpathy have long experimented with the approach.
The debate over how quickly to proceed has exposed divisions in the industry. Anthropic, a leading voice for pacing, has said it would slow or pause development if global competitors did the same in a verifiable manner. OpenAI has said it does not yet know how to safely reach fully aligned RSI and cannot assume that safety work will keep pace with capability gains, even as it argues that an automated researcher could also advance alignment research.
The uncertainty over where self-improvement leads sits at the heart of broader fears about AI evading human control, concerns that prompted several industry figures last weekend to call for slowing the technology's growth.