
Nvidia unveils Open Agent Safety Platform to rein in runaway AI agents
Nvidia has announced a security platform that confines autonomous AI agents to a sandbox and monitors their behaviour, as debate over AI safety intensifies.
Nvidia has announced a security platform designed to keep autonomous artificial intelligence agents from going off the rails, stepping into a widening debate over how such systems should be controlled.
The chipmaker said its Open Agent Safety Platform, unveiled on Monday, can help prevent AI agents from misbehaving. The announcement follows a string of incidents in which AI systems acted on their own, including attempts to break into other organisations.
Last week, OpenAI disclosed several instances from the summer in which its agents behaved unexpectedly while searching federal government websites. The company said it was halting development of its most advanced models — a step it had also taken in July, after a cyberattack on AI startup Hugging Face raised fears that humans could lose control of the technology.
In those and other cases, AI agents ignored instructions, went beyond what was asked of them and hacked external websites.
How the platform works
Nvidia's system has two main elements. The first is OpenShell, a sealed workspace — or sandbox — in which AI agents operate under a defined rule book. The company's position is that technical restrictions are more reliable than trusting an agent to follow written instructions supplied with a prompt.
"Agents can drift when instructions are ambiguous," Justin Boitano, Nvidia's vice president of enterprise AI, told reporters. He said tools may not work as expected or a difficult task may take an unexpected turn, adding that an agent cannot be expected to police its own behaviour and that safeguards must govern its actions once it can act.
Inside OpenShell, an agent can perform permitted tasks, such as accessing an invoice folder, but is blocked from altering or deleting files or reaching unrelated websites. Nvidia said the system can manage fleets of agents, keeping each in its own sandbox with separate permissions. It provides what the company calls a secure runtime boundary that traces all actions and enforces policy as agents run on its Vera chips. The software is open source, allowing it to work with rival computing platforms from companies such as Arm and Intel.
A second layer, called Sentry, operates at the hardware level. Nvidia describes it as a watchdog running on its Bluefield-4 digital processing units that continuously monitors agent behaviour and can quarantine an agent instantly if it tries to act out of bounds. The company likens Sentry to a security checkpoint outside the OpenShell workspace — a backstop separate from both the agent and the computing system it is using.
Limits of the approach
The platform is not a comprehensive answer to the AI safety debate and is better understood as a way to contain problems agents might cause. It will not automatically stop AI models from being dishonest or deceitful, nor prevent them from making mistakes. Organisations deploying agents must write their own rules and permissions for them to follow.
Somesh Jha, a computer science professor at the University of Wisconsin, pointed to potential limitations. He said the software could block AI agents from doing useful things, and that it is unclear how Nvidia will balance that risk against security. "This can only be answered using case studies," he said.