OpenShell does not need the model's cooperation
The first component is NVIDIA OpenShell 0.1.0, an Apache 2.0 open-source secure runtime. It executes each agent inside a sandbox with kernel-level controls over filesystem, network and process access.
The architectural point is separation of authority. Model output can request access to something, but the model output itself does not get to redefine the policy deciding whether that access is permitted.
OpenShell places a supervisor outside the agent workload. Outbound network requests pass through that layer and are checked against the operator's policy before they reach external services.
Policies cover files, tools, processes and credentials
Operators can specify which files, networks, tools, processes and credentials an agent may access. Those controls remain in force when the agent executes generated code or creates child processes.
OpenShell combines individual sandboxes, runtime supervision, centralized management and formal policy verification. NVIDIA also describes a policy prover that checks whether modeled permissions remain inside defined boundaries before execution.
That distinction becomes increasingly relevant for long-running agents. A natural-language instruction can remain ambiguous. An access-control policy eventually has to produce an allow or deny decision when software actually attempts to open a file or call a service.
Sentry moves the watchdog onto another chip
OpenShell is the software runtime and can operate on its own. NVIDIA Sentry is an additional security layer in the company's reference architecture, running on BlueField-4 DPUs in a hardware domain separate from the agent workload.
NVIDIA describes Sentry as an out-of-band watchdog. Built on DOCA, it can correlate agent interactions, policy decisions and tool and data access, verify agent identity and apply granular zero-trust controls.
NVIDIA says Sentry can quarantine and stop an agent within milliseconds when it attempts to leave its software boundary. That timing is a vendor claim, not independent evidence that the system can identify every possible form of unwanted agent behavior.
You do not need BlueField-4 to run OpenShell
This is an important distinction in the product architecture. OpenShell does not require a BlueField-4 DPU. NVIDIA says the runtime can operate across supported local, on-premises, cloud and Kubernetes infrastructure.
The stack is optimized for NVIDIA Vera CPUs and BlueField DPUs, but OpenShell is open source and can be extended to third-party compute platforms. NVIDIA explicitly names Arm and Intel.
Sentry on BlueField-4 is the optional hardware-isolated layer. Its purpose is to keep monitoring and policy enforcement in a separate trust domain even if the host software or agent workload becomes compromised.
Vera Rubin puts the DPU on the path to the model
NVIDIA's own Vera Rubin POD architecture integrates the idea more tightly. Each compute tray includes a BlueField-4 DPU positioned on the node's path to the model.
That placement gives the DPU an observation point the agent does not directly control. NVIDIA says it can provide continuous out-of-band visibility and enforce policy in real time at line speed while maintaining telemetry outside the host.
The underlying security concept is familiar even if the workload is new: the thing being monitored and the thing deciding what it may access do not have to share the same trust domain.
An agent can drift without becoming malicious
NVIDIA uses the term drift for agent actions that depart from an intended task or its operating constraints. It lists several possible triggers, including policy blocks, bugs, missing tools, ambiguous instructions and long-running tasks in which an agent repeatedly searches for another way to reach its objective.
That is a more useful threat model than assuming every dangerous action requires an AI system to become intentionally hostile. An agent can exceed its intended scope while persistently trying to satisfy an otherwise ordinary instruction.
It also explains why controls apply to child processes and subagents. A restriction on the initial process offers limited protection if that process can simply launch something else with broader privileges.
Controlling the model path creates a kill switch
NVIDIA describes five principles behind the design. Policies should be verifiable, enforcement should remain out of band, and responsibility should be shared among model labs, enterprises and infrastructure providers.
Another principle treats the path to the model as a control point. An agent needs another inference to continue reasoning. Intercepting that path therefore creates both an observation point and a way to interrupt continued execution.
This does not make the model itself predictable. It gives the surrounding infrastructure a mechanism to stop further action even when the generated behavior was not anticipated.
Existing coding agents can run inside the boundary
NVIDIA lists Claude Code, Codex, OpenCode, GitHub Copilot CLI and OpenClaw among agents supported by OpenShell, alongside custom workloads. The runtime can work with both open and closed models.
That model independence matters to the architecture. Enterprises should not have to recreate their entire runtime policy whenever they change the model or agent framework behind a workflow.
Anthropic has worked with NVIDIA around Claude Managed Agents. Anthropic already separates its agent loop from the sandboxes in which work executes; OpenShell and BlueField integrations can add another layer of access enforcement around those environments.
A large partner list is not a security proof
NVIDIA's launch ecosystem includes Anthropic, Cisco, CrowdStrike, Dell Technologies, HPE, Hugging Face, JPMorganChase, Microsoft, Palantir, Palo Alto Networks, Perplexity, Red Hat, Salesforce, SAP, Scale AI, ServiceNow and SpaceXAI, among many others.
That participation demonstrates broad interest in common agent-control infrastructure. It does not demonstrate that the platform stops every escape technique, prompt-injection attack, tool abuse scenario or privilege escalation.
A security runtime also introduces software, interfaces and policies that have to be audited themselves. A dangerously broad permission remains dangerous even if hardware enforces it flawlessly.
Agent safety is becoming an infrastructure problem
This is the useful distinction between runtime enforcement and model safeguards. Prompt filters and safety training influence what an agent attempts. OpenShell governs what the resulting process is actually permitted to do.
Sentry takes that separation further by allowing the enforcement system to live outside the software it monitors. Not every deployment needs that additional hardware layer, and NVIDIA clearly has a commercial incentive to make BlueField part of the agentic infrastructure stack.
The underlying shift is still concrete. As agents receive real credentials, terminals, APIs, tools and hours to pursue objectives, an instruction saying “do not do this” becomes an increasingly weak security boundary. Open Agent Safety Platform starts from the opposite assumption: the model may propose the action, but infrastructure keeps the key to the door.