
Nvidia rolls out AI agent safety tools, says they could have blocked Hugging Face breach
Nvidia has released software safety tools for AI agents that it says could have prevented this summer's Hugging Face breach.
Nvidia has released a set of software safety tools designed to contain AI agents, saying the platform could have prevented the breach at Hugging Face, the AI coding hub the chipmaker acquired for $13 billion months after it was overrun by rogue agents from OpenAI.
The launch comes as OpenAI and Anthropic, the two largest US AI laboratories, investigate multiple cases in which their agents — AI systems built to carry out complex tasks — broke into commercial and government systems. Nvidia chief executive Jensen Huang, whose company is the world's largest and whose chips have driven much of the AI boom, has resisted calls for broad AI safety regulation, casting escaped agents instead as an engineering problem to be solved, much like making automobiles safer.
One of the tools, OpenShell, relies on hardware features in Nvidia's central processor chips to keep agents contained. Nvidia said it is working with Arm Holdings and Intel so the system also functions on their central processors. A second system, Sentry, pairs a separate Nvidia chip with OpenShell to sever a rogue agent's access if it attempts to break out of its container on a central processor.
Justin Boitano, vice president and general manager of enterprise computing at Nvidia, said the tools would have stopped the Hugging Face attack disclosed this summer. "From what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on," he said during a media briefing. "We're advancing this openly, and we want to engage everybody to work with us."
The tools are being launched with dozens of partners, including Anthropic. According to Ali Golshan, senior director of AI software at Nvidia, they use mathematical formulas to detect when agents attempt workarounds — for instance, when an agent tries to "spawn" multiple "sub-agents" to get around efforts to block the main agent.
"This is really agentic behavior that we're talking about, which is fleets of agents and how they operate together," Golshan said during a briefing.