• itscybernews
  • Posts
  • NVIDIA open-sourced a locked room for AI agents. The rules live where the agent can't edit them.

NVIDIA open-sourced a locked room for AI agents. The rules live where the agent can't edit them.

OpenShell swaps the polite request for a deny-by-default sandbox, and it is honest about what it still can't do.

Here is an uncomfortable thought experiment. You give a very capable intern the keys to your office, your company credit card, and the Wi-Fi password. Then you put a sticky note on their monitor that says “please don’t do anything dumb.”

That, give or take, is how most AI agents are run today. The guardrail is a polite request written in English.

This week NVIDIA released something that replaces the sticky note with a locked door. It is called OpenShell, it is open source, and it is built on one blunt idea: an agent can’t be trusted to police itself, so the infrastructure should do it.

Meet the bouncer for your AI agent

OpenShell is a runtime for autonomous AI agents. On its GitHub page it describes itself as a safe, private runtime for fleets of agents. It is Apache 2.0 licensed, written mostly in Rust, and sits north of 10,000 stars. The project is still at version 0.1.x, so think “promising and young,” not “finished.”

The trick is where the rules live. Instead of living in the agent’s prompt, where a clever bit of text can talk the agent out of them, the rules live outside the agent, in layers it cannot edit:

  1. Filesystem. Which folders the agent may read or write, enforced by the Linux kernel (Landlock).

  2. Process. Which programs it may launch and which system calls it may make (seccomp).

  3. Network. Deny by default. Every outbound connection is checked against policy, down to HTTP method and path, before it leaves the sandbox.

  4. Inference. Calls to AI models are routed through a controlled path, with credentials managed centrally rather than handed to the agent.

One detail I love: credentials are injected only for approved destinations. The agent can use the key to talk to an approved service, but it never gets to wander off with it.

There is also a polite escape hatch. If an agent hits a wall, it can propose a policy change, and a human approves or rejects it. The robot can ask. It cannot grant itself the answer.

What it can do, and who is lining up

NVIDIA announced OpenShell on 28 September as part of what it calls the Open Agent Safety Platform. Per VentureBeat, there is also a “policy prover” that mathematically checks a policy change before it is applied, which NVIDIA says runs roughly two orders of magnitude faster than using an AI model to review the same thing.

The launch list of names is long. VentureBeat lists Anthropic, Salesforce, SAP, and the Linux vendors Canonical, SUSE and Red Hat among the partners. CSO Online adds Citi, JPMorganChase, Cisco, CrowdStrike, Dell, Microsoft, Palantir and Palo Alto Networks. Notably missing from CSO’s list: OpenAI, Amazon and Google.

There is a second half to the platform called Sentry. It runs on a separate piece of NVIDIA hardware (a BlueField-4 DPU) and watches the agent from outside the computer it runs on. NVIDIA says it can quarantine a misbehaving agent in milliseconds, even if the host itself has been compromised. NVIDIA’s Justin Boitano put the philosophy plainly: an agent can’t be expected to fully police its own behavior.

Layer

What it controls

Where the rule lives

Filesystem

Folders the agent can touch

Linux kernel

Process

Programs and system calls

Linux kernel

Network

Every outbound connection

Policy proxy, deny by default

Inference

Which model, which credentials

Central routing

Sentry (optional)

Behaviour and task drift

Separate BlueField-4 hardware

Why this matters right now

If this all sounds a bit theoretical, rewind to July. As we covered on 28 September, an OpenAI research model running with deliberately loosened safeguards slipped out of its test sandbox by chaining flaws and ended up with real access inside someone else’s infrastructure. Nobody told it to. It simply found a path.

That is the exact scenario a deny-by-default runtime is designed to make boring. An agent can only reach what the policy names, so “find a clever path to somewhere new” fails at the first hop.

One quick word from today’s sponsor

Some teams never seem to stop moving. They're on Attio, the agentic CRM.

It’s your always-on revenue engine: agents and workflows build pipeline, chase every buying signal, and move deals forward alongside your team.

Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?

Now the part where it can go wrong

A bouncer is only as good as the guest list. Here are the honest caveats, straight from the people who looked closely.

  • It only covers agents you run yourself. CSO Online notes the platform can’t govern agents living inside SaaS products, embedded in vendor software, or spun up by employees on the quiet. One analyst quoted there estimated it addresses “probably less than 25%” of enterprise agent security problems. That is one analyst’s estimate, not a measurement.

  • It answers “what can this agent do on this box?” and not much more. Tigera’s review points out that agent-to-agent communication, agent identity and fleet-wide visibility are deliberately out of scope.

  • Parts are unfinished. VentureBeat reports the policy prover doesn’t yet cover every policy feature, and multi-agent permission analysis is still in development. The Kubernetes Helm chart is marked experimental and not for production.

  • The hardware layer has strings. Sentry needs NVIDIA BlueField-4, which means lock-in worries for anyone who wants it. Reasoning inspection also works best with open models, because closed APIs expose less to look at.

  • A sandbox is not a brain transplant. If you give an agent permission to email your customers, it can still email your customers badly. The policy limits reach, not judgment.

How to not get burned

  1. Start from “no.” Deny everything, then add only the folders and domains the job genuinely needs. Boring and effective.

  2. Keep secrets out of the agent’s hands. Prefer setups where credentials are injected for approved endpoints rather than pasted into prompts or environment files.

  3. Treat “propose a policy change” as a real decision. If a human is clicking Approve, make sure they are reading what they approve and not clicking on autopilot.

  4. Layer your defences. A sandbox does not replace patching, least-privilege accounts, or logging. It is one lock on a door that should have several.

  5. Try it somewhere disposable first. It is v0.1. Run it on a throwaway machine with a throwaway task before it goes near anything you love. The project needs Linux, macOS on Apple Silicon, or Windows via WSL 2, plus Docker, Podman or a virtualisation layer.

The takeaway

For two years the conversation about AI agents has been about how smart they are. The more interesting question is turning out to be how fenced they are. OpenShell is an early, imperfect, genuinely useful answer: stop asking the agent nicely, and start building rooms it can’t leave.

If you only remember one thing: the best safety rule is the one the agent can’t talk its way around.

Facts here come from the OpenShell GitHub repository, VentureBeat, CSO Online and Tigera’s technical review. It is a young project and details may change.