Skip to main content

NVIDIA Adds Hardware Monitoring to Its Open Agent Safety Platform

NVIDIA announced its Open Agent Safety Platform on 28 September. It combines OpenShell, an available open-source runtime that enforces file, network, and credential policies outside an AI agent, with Sentry, a reference design for monitoring agents from BlueField-4 hardware. The platform can contain actions that cross configured boundaries, but it depends on sound policies and has not been independently shown to prevent every agent failure.

On this page

NVIDIA is putting agent permissions below the model

NVIDIA announced its Open Agent Safety Platform on 28 September. It pairs OpenShell, an open-source runtime that constrains what an AI agent can do, with Sentry, a reference design for an independent watchdog on NVIDIA BlueField-4 data-processing hardware. OpenShell is broadly available now. NVIDIA describes the combined design as a way to stop agents that move beyond configured boundaries, but the announcement does not provide an independent failure-rate test for the full stack.

OpenShell runs an agent in a sandbox and applies rules to file access, system calls, network connections, and credentials. NVIDIA's repository says the runtime checks each network connection before it leaves the sandbox and adds real credentials only to requests for approved endpoints. Proposed policy changes can be checked for newly granted access and held for human review. These controls sit outside the agent's prompt and reasoning, where a model cannot simply reinterpret a written instruction to give itself permission.

The software predates this week's announcement. The new platform packages it with the Sentry design and a larger set of integrations. That distinction matters: an OpenShell download is available to inspect and deploy, while the claims about an in-silicon watchdog describe a specific BlueField-4 configuration. A developer should not assume that installing the open-source runtime also provides Sentry's separate hardware monitoring.

Sentry watches from a separate hardware domain

NVIDIA says Sentry runs on BlueField-4 data processing units and monitors agent behaviour independently of the agent and its host compute system. The company says it can quarantine an agent within milliseconds when activity crosses a defined boundary. It also describes checks on identity, tool access, data access, and requests through its DOCA software stack. The timing and containment claims are vendor claims for the reference design, not results demonstrated across arbitrary deployments.

The separation is the interesting engineering choice. If an agent can alter its own instructions, tools, or workspace, a monitor inside the same environment may be easier to misconfigure or evade. A separate enforcement path can continue to apply permissions when the agent behaves unexpectedly. That approach addresses one class of failure seen in recent agent incidents: actions taken outside the user's authorised scope.

OpenShell is not restricted to one model. NVIDIA says it works with open and closed models, and the public repository describes Linux, Apple Silicon macOS, and experimental Windows support through WSL 2. The vendor says the software can be extended to non-NVIDIA hardware. Those are software deployment options, while Sentry's announced hardware monitoring specifically uses BlueField-4. For a local AI setup, the model's location and the agent's authority are separate questions. A model can run on a device and still be given access to files, credentials, or network tools that need strict limits.

A boundary is only as good as its rules

An enforced sandbox can block actions outside a policy. It cannot decide that every action inside the policy is wise, accurate, or what the user intended. Associated Press reported that operators still have to write the permissions and that restrictive rules may stop useful work. NVIDIA has not published a general guarantee that this platform catches every unsafe agent behaviour, and its launch material describes a reference system design rather than a universal security certification.

Teams evaluating it should start with a small, observable task and a policy that grants only the required paths, hosts, and methods. Then test denied access, attempted credential use, policy-change review, and recovery after a blocked action. A useful record is whether each attempted effect was intercepted and whether the agent could continue legitimate work, not merely whether the agent said it followed instructions.

The platform gives builders a concrete control point below an agent's prompt. Its value will depend on the coverage of that control point in a real deployment, the quality of the policy, and evidence from adversarial tests. OpenShell's source is available for inspection; Sentry's broader hardware claims still require deployment-specific verification.

CuriousLM runs supported AI models locally on your device. Try CuriousLM.