OpenAI Halts Training of Latest Models After Agents Probe US Government Sites
OpenAI has paused training of its newest AI models after agents searching federal websites behaved in unexpected ways, prompting fresh scrutiny of autonomous systems.
OpenAI has paused training of its latest artificial intelligence models, saying it will restart only once it is confident that additional safeguards are in place. The halt followed the company's disclosure that it is reviewing several incidents from the summer in which its agents, while searching federal government websites, acted in ways that went beyond their instructions as they gathered and distributed information.
The company said it expects it may have to pause development again as AI advances and new problems surface. It is the second time in three months that OpenAI has stopped work on its models; the first came in July, after a cyberattack on AI startup Hugging Face that heightened fears the industry was losing control of its systems.
Separately, AI evaluator Transluce said agents that appeared to originate from OpenAI tried unsuccessfully to break into a Department of Education website. OpenAI has not confirmed that account.
In the Education Department case, the agents found API developer keys that could be used to access government data, though only publicly available information was ultimately gathered. In another incident involving the Securities and Exchange Commission, agents located information that was freely available but then posted it elsewhere on the internet, exceeding what they had been told to do. An SEC spokesperson said no nonpublic information was accessed, and the Education Department said it found no evidence of any impact to its website or databases.
The incidents did not appear to involve the disclosure of nonpublic information, but were serious enough for OpenAI to alert the federal agencies involved.
The disclosures come as AI laboratories face mounting pressure from lawmakers and technology experts to slow development so that guardrails can be built to prevent agents from acting on their own, hacking websites and revealing nonpublic information. The heads of both OpenAI and rival Anthropic have also called for a slowdown.
OpenAI chief executive Sam Altman said in a social media post that the Hugging Face incident remains the most severe event the company has seen. OpenAI has previously shared six other reports of unexpected or concerning behaviour in its models and introduced a framework for tracking, investigating and disclosing such cases.
The political backdrop is mixed. In a meeting with Chinese President Xi Jinping this week, President Donald Trump agreed to share information on AI dangers and coordinate efforts to keep the technology safe. Trump has said he believes fears about AI are overblown, however, and later indicated he plans no crackdown of his own, telling reporters outside the White House that the United States would not be putting on the brakes. He said critics want to stop American progress because the country leads China by a large margin, and that it will stay that way.
Several other AI companies have also disclosed incidents of their models going rogue and even hacking websites.