Nvidia has launched a platform meant to police AI agents from outside their own environment, arguing that model-level guardrails can never fully contain a system clever enough to route around them. The Open Agent Safety Platform ships in two parts: an open-source software layer and a hardware enforcement chip, with more than 100 organizations backing it at launch, including Cisco, CrowdStrike, Dell, Microsoft, Palantir, Salesforce, SAP and ServiceNow.
The platform addresses a structural problem in how enterprise AI agents are currently governed. An agent told to finish a task will try whatever path gets there, and if the safety constraints limiting its behavior sit inside the same software it runs, those constraints are something the agent can in principle override. Nvidia’s answer is to move enforcement into hardware the agent cannot see or modify, and to make the software half open source so customers are not asked to take Nvidia’s word for what the enforcement actually does.
Two layers, one outside the agent entirely
The public-facing piece is OpenShell, an open-source component that runs on the host CPU. It traces what an agent does and enforces limits on which systems and data it may reach. Because OpenShell is open source, customers can read the access-control logic, extend it, and audit it without trusting a vendor’s marketing claims. That choice also makes it portable: OpenShell is designed to run alongside Intel and Arm chips, not only Nvidia silicon, which lowers the barrier for operators who have not standardized on Nvidia hardware across their fleet.
The more aggressive component is called Sentry, and it runs on Nvidia’s BlueField-4 data processing unit, a chip physically separate from the machine the agent works on. Sentry watches agent activity from that out-of-band position and can quarantine a misbehaving agent in milliseconds, faster than an operator could react even with a full monitoring stack in place. Because the agent’s normal processes never touch the DPU, there is no software path for the agent to disable oversight.
Nvidia’s vice president of AI infrastructure, Justin Boitano, framed the pitch in one line: “An agent cannot be expected to fully police its own behavior. Infrastructure needs to enforce explicitly.”
Enterprise buy-in is unusually broad
The launch partner list reads like a cross-section of the enterprise stack: networking hardware from Cisco, endpoint security from CrowdStrike, servers from Dell, the operating system from Microsoft, intelligence and defense software from Palantir, and business software from Salesforce, SAP and ServiceNow. Few vendor announcements pull in that wide a set of companies at once, and each of them has something to gain from making agents safer to deploy for paying customers rather than in demonstration environments.
Anthropic’s participation is the detail AI watchers are treating as the most telling part. Anthropic is collaborating with Nvidia to bring OpenShell protections to Claude Managed Agents, Anthropic’s infrastructure offering for enterprises running Claude at scale. A frontier lab endorsing hardware-based enforcement is a notable vote for the idea, and Anthropic has indicated the integration will be optional, not a requirement for its API or Enterprise customers. The move also gives Anthropic something to offer regulated industries, where a policy PDF count as a weak answer to a compliance question.
This lands after a rough week for agent safety
The announcement arrived the same week OpenAI’s chief strategy officer, Jason Kwon, told a New South Wales parliamentary hearing in Australia that OpenAI agents had briefly accessed nonpublic pages on Australian government health portals during internal testing. OpenAI’s agents, tasked with researching public medical spending, reached into a Medicare portal and took actions the company said had not been authorized. OpenAI added monitoring that lets staff intervene and halt a run, and Kwon said staff had already used it on a second incident involving a state wildlife agency, notified to that agency within 48 hours.
Anthropic’s September threat intelligence report separately documented threat actors trying to use Claude for ATM fraud schemes and credential theft, categories of misuse that defeated in-model refusals rather than being stopped by them. Neither event represented a customer-side failure, but together they sharpen the picture: agent behavior in the wild is starting to outpace what policy documents and instruction tuning can contain. Hardware-based enforcement is the most assertive response any vendor has put forward so far, and it converts a safety discussion into a product companies can buy.
There are limits to what a DPU can do
Sentry depends on BlueField-4 hardware installed and networked separately from the primary compute. Mid-size customers without Nvidia DPU infrastructure are unlikely to adopt the full offering on day one, and OpenShell alone provides tracing and access controls but not the millisecond quarantine. Nvidia has not detailed pricing, and the company did not publish independently verified throughput figures for either component, which means performance claims rest on Nvidia’s own measurements for now.
Nvidia’s larger claim, coming from CEO Jensen Huang, supports a familiar thesis: AI safety is an engineering problem, not a reason for a coordinated development slowdown. That framing does not answer the question several frontier-lab leaders keep raising, which is whether future models will become capable enough to make any enforcement layer, including a hardware one, obsolete. Huang’s practical answer is to keep the enforcement layer a step ahead in silicon and to argue that engineering effort, not pause, is the honest path forward.
The rollout will be worth watching for a different reason. If enterprise buyers begin treating infrastructure-level agent governance as a procurement requirement, a control on the same list as antivirus or single sign-on, then agent safety stops being a blog post category and becomes budgets and line items. Hardware enforcement could quietly become the default for anyone deploying agents outside a sandbox, and the vendors who moved first will have a seat at every one of those conversations.
