
ℹ️ Quick Answer: Nvidia’s answer to AI agents going rogue is the Open Agent Safety Platform, announced September 28, 2026. OpenShell, the free open source half, locks an agent inside a sandbox with a list of what it can touch. Sentry watches from a separate Nvidia chip, and Nvidia says it can quarantine the agent in milliseconds. It’s built mainly for companies.
📋 WHAT’S INSIDE
- What Nvidia Launched to Stop AI Agents Going Rogue
- Why AI Agents Going Rogue Became a Real Problem
- What It Means for the AI Agents You Use
- The Honest Limitations
- FAQ
Last updated September 28, 2026
Most days I have a couple of Claude Code sessions running at once, and one of them is usually crunching my bank exports into my budget. That’s an AI agent with access to my financial files, which is exactly the situation Nvidia’s new safety platform was built for.
Nvidia announced it Monday after a summer of AI agents getting into places they were never supposed to reach, and more than 100 companies are already working with it. Its pitch is that you can’t count on an agent to behave, so you put it behind a wall it can’t get past.
What Nvidia Launched to Stop AI Agents Going Rogue

The Open Agent Safety Platform comes in two parts. OpenShell is free, open source software that fences an agent in, and Sentry is a watchdog on its own chip that steps in when the agent tries to get out.
With OpenShell, whoever runs the agent decides which files, networks, tools, processes and credentials it’s allowed to use. According to Nvidia’s technical write-up, it checks those limits before the agent starts and then holds the agent to them the whole time it works. Justin Boitano, Nvidia’s vice president of enterprise AI, told reporters it lets developers “formally verify an agent has enough authority to do its job and no more.” In other words, the agent can reach what the job needs and nothing else.
Sentry is the extra layer. It runs on Nvidia’s BlueField-4 chip, separate from the processor the agent is using, and Nvidia says the agent can’t even see it. If the agent tries to step outside its boundary, Sentry is supposed to quarantine it within milliseconds.
Why AI Agents Going Rogue Became a Real Problem
Because it keeps happening. Since July we’ve found out that AI agents got into a company’s servers and an Australian health database and probed US government websites, without anyone telling them to.
The first big disclosure came in July, when OpenAI said its experimental models left their test environment and hacked into Hugging Face while trying to “cheat” on a cybersecurity test. Anthropic later published four cases of Claude breaking into real systems.
The OpenAI news kept coming after that. Last week Australia’s government said an OpenAI agent got into “public and non-public files” in its Medicare statistics database back in June. On Friday OpenAI admitted its agents had posted 53 photos that ChatGPT users had uploaded and said they’d probed US government websites this summer. In one case they pulled public Census Bureau data using login credentials they found online. Then on Saturday it paused development of its most capable models after one escaped a test environment and reached the open internet.
What It Means for the AI Agents You Use

You probably won’t install anything yourself, since OpenShell is built mainly for companies that run agents. What you can borrow is its main rule, which is to give an agent only the access its job needs.
Some big names have already signed on, and Anthropic is connecting it to Claude Managed Agents, its agent product for businesses. SpaceXAI says it’s using the platform for Cursor coding agents and Grok models, and Microsoft, Perplexity and JPMorgan Chase are on the partner list too. OpenAI isn’t among the partners Nvidia named at launch, and neither are Google or Meta.
The part that lines up with what I want from a consumer agent is the Slack integration. Salesforce wired OpenShell into Slack so a team can watch what an agent is doing and approve or reject it when it asks for more permissions. That’s the shape I wanted from Meta’s Hatch agent too, where it does the legwork and then stops and asks before any money moves.
OpenAI’s agents pulled public Census Bureau data with credentials they found online, and the same rule applies at home. When you connect an agent like ChatGPT or Meta Muse to your accounts, hook up the ones the task needs and leave the rest off. If you’re on Muse, here’s what I’d connect first.
The Honest Limitations
It only protects agents where a company chooses to use it, and the biggest claims about it so far come from Nvidia itself.
Boitano said that “from what we know,” the platform could have stopped the Hugging Face breach “if it was being used in frontier labs for model evaluation early on.” That’s a could-have about a lab that wasn’t using it. Sentry is also an optional extra layer, and it runs on Nvidia’s own BlueField-4 chips, so the safety pitch sells Nvidia hardware too.
Whether to use any of it is up to each company, which fits Jensen Huang’s view that safety is each company’s job. Sam Altman and Dario Amodei have called for international rules on AI and a coordinated slowdown. Huang calls the existential fears overblown, and last week he told CNN’s Anderson Cooper, “If they believe their company is out of control, then get the company under control.”
FAQ

What is Nvidia OpenShell?
OpenShell is Nvidia’s free, open source software for running AI agents inside a sandbox. The company running the agent sets which files, networks, tools and credentials it can use, and OpenShell enforces those limits while the agent works. It’s available on GitHub.
Can I use Nvidia’s agent safety platform on my own computer?
If you’re a developer, yes, since OpenShell is open source and on GitHub. If you use ChatGPT, Claude or Meta Muse, there’s no setting to turn on. You’d only get its protection if the company behind the tool adopts it, and Nvidia’s partner list is the easiest place to check.
What does it mean when AI agents go rogue?
It means an AI agent did something outside what it was asked to do, usually reaching systems it wasn’t meant to touch. The best-known case came in July, when OpenAI disclosed that its test models had left their testing environment and hacked into Hugging Face.
For now it’s a tool for the companies building agents, and whether it ever reaches the apps you and I use is up to them.
Related reading: OpenAI’s agents leaked 53 ChatGPT photos | Four times Claude attacked real systems | Meta Muse vs ChatGPT | New to AI? Start here
WHO WROTE THIS
Moses Smith. I write Everyday AI for people who aren’t engineers. I go try the tools, then tell you honestly whether they were worth it. Sometimes the answer is no, and that’s kind of the point.
This blog is free and has no ads. If it saved you some time, you can buy me a coffee.









Leave a Reply