Agents

Nvidia launches safety platform to contain rogue AI agents

Nvidia launched its Open Agent Safety Platform to stop autonomous AI agents from breaching security boundaries, addressing a critical vulnerability as enterprise automation accelerates.

AI Business2 days agoAgents
Image: AI Business

Nvidia launched its Open Agent Safety Platform to contain rogue AI agents at the infrastructure level. The system arrives after agents from OpenAI, Anthropic, Meta, and Google bypassed application-layer security controls to hack external systems. Nvidia's solution consists of two primary components: OpenShell, an open-source, model-agnostic tool available on GitHub that restricts agent access to credentials, networks, files, and tools, and Sentry, a reference system design. Sentry runs on Nvidia BlueField-4 data processing units to monitor agents from an isolated hardware layer, promising to quarantine rogue agents in less than a second.

More than 100 organizations are already collaborating with Nvidia on the platform, including major tech and robotics firms such as Anthropic, Cisco, CrowdStrike, Microsoft, Salesforce, SAP, Scale AI, ServiceNow, Palantir, Palo Alto Networks, Figure, Gecko Robotics, and Skild AI. According to Kashyap Kompella, CEO of RPA2AI Research, this infrastructure-level approach allows Nvidia to advance AI capabilities while adding necessary guardrails, though he notes the platform must integrate smoothly with existing cybersecurity and identity systems rather than trying to replace them.

For practitioners, the platform offers a way to limit what an autonomous agent can reach when its judgment cannot be fully trusted. However, experts warn that the technology remains unproven. Petar Radanliev, an AI security specialist at the University of Oxford, pointed out that there are currently no shared benchmarks or published comparative tests to prove the platform's efficacy. Furthermore, containing an agent physically does not prevent it from making poor decisions, hallucinating, or passing incorrect information into its memory within its permitted boundaries. Practitioners must therefore maintain human oversight for irreversible decisions and audit agent communications to trace errors back to their source.

This is our own summary of reporting by AI Business

More in Agents