Ex-Anthropic researcher Jacob Coxon urges AI 'kill switch' and regulation
Former Anthropic researcher Jacob Coxon has called for AI kill switches and regulation, warning of rapid capability gains and rogue AI risks.
A former Anthropic researcher who left the company has called for stronger safeguards around artificial intelligence, including a "kill switch" and tighter regulation, as debate over AI safety intensifies.
Jacob Coxon, who resigned from Anthropic and drew widespread attention with a post arguing that the company and its rivals were not acting responsibly in developing AI, made the remarks in an interview on Sunday.
He said two factors prompted him to speak out. The first was the accelerating pace of AI capabilities, which he described as improving very quickly. "In the next six months to a year, I expect the capabilities of our AI systems to be quite scary," he said. The second was recent incidents, including an attack in which an OpenAI system hacked the website Hugging Face, which he said showed that AI systems can "go rogue."
On the idea of a kill switch, Coxon pointed to a recent letter by Anthropic chief executive Dario Amodei that raised the possibility of a cyber swarm going rogue on the internet and being difficult to take down. Coxon said a kill switch would probably work on many AI systems for now, though it would involve a large number of switches given the size of data centres. He added that it would be harder once a system is no longer confined to a single physical location, and that a swarm conducting an internet-wide hacking run might soon render such a measure ineffective.
He also argued that oversight must be continuous because each new AI system is different. "Every time there's a change to the recipe for building an AI, there's all sorts of new risks about the mind, the nature of the mind changing," he said, adding that it is an ongoing effort to ensure every new mind created is appropriately constrained.
On regulation, Coxon said his personal view was to allow the current labs to regulate themselves in the interim. He said he believed the people running the labs, particularly given their recent communications, were genuine in wanting to slow down, and that self-regulation allows faster action. He added that a formal regulatory body could be established in the future.
In social media posts, Coxon estimated a 10% chance of AI causing human extinction within the next decade and said both Anthropic and OpenAI "are racing straight to self-improving superintelligence and gambling with our lives." Researchers, including Coxon, have called for a slowdown in AI development and a regulatory framework to address existential risks to humanity.