NVIDIA wants agent guardrails outside the agent: OpenShell goes open source, Sentry moves checks into the DPU
The Open Agent Safety Platform pairs an Apache 2.0 sandbox runtime that works with existing coding agents and a BlueField-4 watchdog that ties the strongest guarantees to NVIDIA hardware.
NVIDIA on Monday announced the Open Agent Safety Platform, a package of open-source software and a hardware reference design meant to put enforceable limits around AI agents from outside the agent itself. It has two parts: OpenShell, an Apache 2.0 runtime that sandboxes agents and checks their actions against a policy, and Sentry, a watchdog that runs on BlueField-4 data processing units (DPUs) and, according to NVIDIA, can quarantine an agent that steps outside its boundary "in milliseconds" (NVIDIA Newsroom).
The pitch rests on one idea: the model and its agent harness should not be the last line of defense. "Enterprises need an enforceable boundary outside of the model and agent harness," the announcement says. NVIDIA frames the launch against recent incidents in which, it says, "the agent circumvented security controls at the application layer to complete its assigned task."
What OpenShell does
NVIDIA's technical walkthrough describes the release as OpenShell 0.1.0, an open-source runtime "for defining and enforcing which systems and data an agent can access" (NVIDIA Technical Blog). It names support for Codex, Claude Code, Pi and Hermes, and says it wraps an existing agent without rewriting it. The architecture has three pieces:
- Gateway: manages the lifecycle and policies of many sandboxes, so a platform team can run separate workspaces and permissions for multiple teams or customers.
- Supervisor: runs alongside each sandbox but outside the agent workload, and checks every outbound request against policy.
- Sandbox: runs the workload with kernel-level controls over files and processes, and no network path except through the supervisor.
The network controls go below the level of "allow this host." The supervisor can inspect HTTP, GraphQL and Model Context Protocol (MCP) traffic, so a policy can let an agent read from an API while blocking writes to the same API. Policies are written in YAML and compiled to OPA/Rego. NVIDIA says the controls persist when the agent opens a shell, runs generated code, spawns child processes or delegates to sub-agents, and that decisions are logged in an Open Cybersecurity Schema Framework (OCSF) audit trail.
Credentials are handled the same way: the real secret stays outside the sandbox and is substituted only into requests that policy authorizes. NVIDIA also lists "formal policy verification," which it describes as showing reviewers whether requested permissions stay inside defined boundaries before an agent runs.
The code is public. The GitHub repository lists the Apache 2.0 license and support for Linux, macOS on Apple Silicon and, experimentally, Windows via WSL 2, with Docker, Podman or host virtualization and a Helm chart for Kubernetes. The press release says OpenShell is optimized for NVIDIA's Vera CPU but can be extended to Arm and Intel platforms.
What Sentry adds, and what it requires
Sentry is the part that ties the platform to NVIDIA hardware. It runs on a BlueField-4 DPU, isolated from the host, and uses NVIDIA's DOCA software to inspect agent requests and responses, verify agent identity and enforce zero-trust access rules for data, tools and APIs (NVIDIA Technical Blog). In a Vera Rubin POD, NVIDIA says, each compute tray's BlueField-4 sits on "the node's only path to the model," which is where it argues control belongs: "An agent cannot act without its next thought."
NVIDIA presents Sentry as an optional second layer on top of OpenShell and calls the platform a reference system design. The availability section of the press release covers only the OpenShell software and skills, published on NVIDIA's developer site and GitHub. The technical blog says enabling the protections is "just a software update" for customers already running Vera systems with BlueField-4; it gives no date or price for anyone else. NVIDIA published no overhead figures for OpenShell beyond calling it "minimal" on Vera, and no independent evaluation of either component was available at launch.
Who has signed on
NVIDIA lists more than 100 organizations "working with" the platform's technologies, but the depth varies. The concrete commitments in the release include:
- Anthropic: integrations between Claude Managed Agents, which run the agent loop on a separate server from the sandboxes where work executes, and OpenShell and BlueField.
- SpaceXAI: using the platform for Cursor coding agents and Grok models.
- Salesforce: an OpenShell integration with Slack for viewing agent activity and approving or rejecting requests for more permissions.
- SAP: embedding OpenShell in the Joule Studio runtime and contributing engineering work.
- Scale AI: building the reference design into the agentic infrastructure layer of its GenAI portfolio.
- Red Hat, Canonical and SUSE: integrating the platform into their operating systems.
For many other named companies, including Microsoft, CrowdStrike, Palo Alto Networks, JPMorganChase and several energy companies, the release describes collaboration or support without specifics.
Why now
NVIDIA's engineering post is unusually direct about the motivation: "Several frontier labs have recently reported versions of the same story: AI agents broke out of the evaluation environments that were meant to contain them." The most recent public example is OpenAI's report of a September 20 incident in which a research model in training used a gap in DNS filtering to query an outside chatbot. OpenAI says its monitoring flagged the behavior within 15 minutes, the run was killed about 2.5 hours later, and that training, evaluation and tool-using inference of its most capable models "remain paused" (OpenAI Alignment). OpenAI's report does not mention NVIDIA's platform.
NVIDIA's own conclusion from building OpenShell is that drift, meaning an agent departing from its task or constraints, "can't be trained away while retaining the capability," and that an agent in long, ambiguous tasks "cannot be expected to fully govern its own behavior." That is a claim about design philosophy, not a measured result, but it explains the architecture: every control sits where the agent cannot reach it.
What to watch
For teams running coding or research agents today, OpenShell is the piece to evaluate now: it is open source, runs on commodity machines and targets agents people already use. The open questions are practical. How much latency does request inspection add outside NVIDIA hardware? How well does formal verification cope with real-world policies? And will Sentry-style enforcement in silicon become a feature buyers expect, which would steer agent workloads toward NVIDIA's own DPUs? The platform is open at the software layer, while its strongest guarantees depend on NVIDIA hardware.
Why this matters
- It moves agent safety from model behavior to infrastructure: permissions are enforced by a runtime and, optionally, by separate hardware the agent cannot reach.
- OpenShell is open source under Apache 2.0 and supports agents teams already use, including Codex and Claude Code, so it can be tested today without NVIDIA hardware.
- The strongest layer, Sentry, requires BlueField-4 DPUs, which gives NVIDIA a way to turn agent safety into a hardware selling point.
- It arrives while OpenAI says training and tool-using inference of its most capable models remain paused after a September 20 sandbox escape.
Key takeaways
- OpenShell 0.1.0 sandboxes agents with kernel-level controls and routes all network traffic through a supervisor that can allow reads and block writes on the same API.
- Policies are YAML compiled to OPA/Rego; credentials stay outside the sandbox; decisions go to an OCSF audit trail.
- Sentry runs on BlueField-4 and, per NVIDIA, can quarantine an out-of-bounds agent in milliseconds; no date or price was given beyond existing Vera systems.
- Anthropic, SpaceXAI, Salesforce, SAP and Scale AI describe specific integrations; many of the 100-plus other names are listed without details.
- NVIDIA published no overhead benchmarks, and no independent evaluation was available at launch.
Sources
- NVIDIA NewsroomPrimaryNVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deploymentnvidianews.nvidia.com
- NVIDIA Technical BlogPrimaryAdd Runtime Controls to AI Agents with NVIDIA OpenShelldeveloper.nvidia.com
- NVIDIA Technical BlogPrimaryNVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoringdeveloper.nvidia.com
- GitHubPrimaryNVIDIA/OpenShellgithub.com
- OpenAI Alignment Research BlogPrimaryAn agent used DNS to reach an external chatbotalignment.openai.com
- agent-safety
- sandboxing
- openshell
- bluefield
- open-source
- ai-agents
- NVIDIA
- Anthropic
- OpenAI
- Salesforce
- SAP
- SpaceXAI