
NVIDIA unveils open safety layer to keep AI agents on course
NVIDIA has launched an open safety platform for AI agents that enforces security controls outside the model, combining OpenShell and Sentry to detect behavioural drift.
NVIDIA has introduced an open safety platform for AI agents that places monitoring and security controls outside the AI model, aiming to prevent autonomous systems from straying beyond their intended tasks or operating boundaries.
The NVIDIA Open Agent Safety Platform combines the company's OpenShell and Sentry technologies into an independent layer for monitoring, policy enforcement and security around AI agents, the company said in a statement on Monday.
Founder and CEO Jensen Huang said the effort involves more than 100 industry partners. "This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems," he said.
The problem of drift
NVIDIA said its work on OpenShell showed that AI agents need to operate in a zero-trust environment with isolation, monitoring and behaviour detection. Agents can experience what the company calls "drift", where their actions depart from the intended task or operating constraints.
Drift can stem from a policy block, a software bug, missing tools or ambiguous instructions, and can also set in when agents are allowed to work for long periods on complex problems. "An agent in these circumstances cannot be expected to fully govern its own behaviour," NVIDIA said, adding that the issue cannot be trained away without losing capability.
That is why, in its proposed approach, safety controls sit outside the agent itself and remain beyond its reach.
How the controls work
OpenShell is an open-source secure runtime that executes autonomous AI agents in sandboxed environments with kernel-level isolation. It converts an operator's instructions into a verifiable policy and lets operators define which files, networks, tools, processes and credentials an agent may access. These limits are checked before an agent starts operating and enforced while it runs.
For organisations wanting an additional layer, NVIDIA Sentry extends monitoring and enforcement into NVIDIA BlueField hardware. The company's DOCA software makes the BlueField security foundation programmable and links it with OpenShell policies.
The system can correlate agent interactions, policy decisions and access to tools and data to build a contextual record of agent activity. According to NVIDIA, this helps safety systems spot behavioural drift, investigate suspicious activity and decide when intervention or deeper analysis is needed.
Five principles and hardware support
NVIDIA outlined five principles for safer agent systems: making policies verifiable, keeping enforcement outside the agent, controlling the path to the AI model, scaling agent authority alongside the ability to inspect its behaviour, and applying a shared responsibility model across AI labs, enterprises and infrastructure providers.
The platform is designed to run on systems based on NVIDIA Vera CPU and BlueField DPU hardware while remaining compatible with other systems. In a Vera Rubin POD, BlueField-4 can provide continuous, out-of-band monitoring of agent behaviour and enforce security policies in real time, the company said.
The approach is intended to let organisations run fleets of agents, subagents, tools and applications within a controlled boundary, with monitoring and policy enforcement continuing throughout their operation.