

AI agent safety gets two gates in NVIDIA’s new platform
NVIDIA combines policy-bound software with a BlueField watchdog design. OpenShell is available now; Sentry’s millisecond claim still needs a test readers can inspect.
By Katie Delaney / 2026-09-28 / 16 min read

How AI agent safety moves below the prompt#
OpenShell runs an agent inside a sandbox and applies declarative policies to its access. The NVIDIA OpenShell repository describes kernel-level checks on file access, system calls and network connections. It also says that the agent does not receive real provider credentials directly. The runtime resolves them only for requests allowed to reach approved endpoints. These practical controls narrow the agent’s path through a system and reduce the scope of a mistake. For AI agent safety, this is the key distinction: the runtime can constrain an attempted action even when the model’s plan points elsewhere.
That distinction matters because an agent is a bundle of moving parts. A model interprets instructions; a harness gives it tools and memory; a runtime governs what the running process can reach. NVIDIA’s own agent-stack guidance argues that the model and harness can guide behaviour, while infrastructure controls determine what an agent is able to do. This is a sound security principle, though the guidance comes from the company selling a runtime that embodies it.
The public documentation adds detail an executive launch paragraph cannot. OpenShell’s policy model covers filesystem, network, process and provider access. Some rules are locked at sandbox creation, while network rules and provider attachments can change at runtime. Its support matrix lists Linux on x86 and Arm, and macOS on Apple Silicon, as supported host platforms; Windows under WSL 2 remains experimental. The same documentation names Landlock and seccomp as required Linux kernel facilities. A security buyer should read those deployment requirements before putting the tool near sensitive work. For teams comparing options, AI agent safety also depends on where those controls can run and how each host handles them.
OpenShell is also published under the Apache 2.0 licence, according to its repository. That makes the runtime inspectable and reusable, and the code can be adapted to more than NVIDIA hardware. The launch release says the software can be extended for Arm and Intel compute platforms. That portability is the cleanest meaning of “open” in the announcement. Each organisation still needs the people to configure the software well.
The difference between a prompt promise and a policy perimeter is easier to see in a marketing team than in a keynote. An agent might be asked to draft a campaign, check last month’s results and prepare a change to an advertising account. If its job is analysis, it may need read access to approved reports. It does not automatically need the ability to change bids, publish creative, export customer records or send a live email. Careful controls turn that distinction into a rule the agent cannot quietly rewrite for itself. An AI agent safety review should start with what the process can reach, not what the model claims it intends.
For teams taking an agent from pilot to production, the question is bigger than a model choice. folkfox’s AI consultancy work starts with the system around the model: access, ownership, tests and the people who respond when a boundary holds or fails.
Sentry watches from outside#
Sentry is the second gate, a separate system of observation. The launch release describes it as an out-of-band watchdog running on BlueField-4 DPUs, hardware that sits apart from the host workload. The idea is to inspect agent activity from a separate trust domain and quarantine an agent that crosses its software boundary. Arm’s description of the NVIDIA and Arm architecture says BlueField-4 provides an independent infrastructure environment beyond the host. That gives the design a physical point of separation, rather than asking the agent or its own software loop to judge its conduct. That makes AI agent safety a systems question: what can the independent layer observe, and what can it actually stop?
The distinction between components is essential. NVIDIA calls OpenShell open-source runtime software. It calls Sentry a reference system design. The launch release says the OpenShell software and skills are available through NVIDIA’s developer resources and GitHub, but presents Sentry as a design built for BlueField-4. Buyers should assess availability separately for each component. Developers can inspect and run OpenShell on supported systems; Sentry is a hardware-dependent blueprint whose adoption will depend on the infrastructure around it.
The launch release says Sentry can stop and quarantine a boundary-crossing agent in milliseconds. That is the sharpest speed claim in the launch, and also the one readers have least detail to assess. The release gives no benchmark, test conditions, false-positive rate or account of the kinds of behaviour that trigger quarantine alongside it. Until those appear, treat “milliseconds” as a vendor statement without an independently measured result.
The partner pool deserves a patient reading. NVIDIA says more than 100 organisations are working with its platform technologies. The release names companies from software, security, infrastructure, banking, energy and robotics, but their roles vary: some are integrating OpenShell into products; others are working on platform technologies or offering infrastructure that supports them. That signals interest; the release does not establish that a hundred organisations have deployed the complete OpenShell-and-Sentry design in production.
Some integrations make a clearer case for coordination than a logo wall. The launch release says Salesforce and NVIDIA have integrated OpenShell with Slack, so teams can view agent activity and approve or reject requests for extra permissions. It says SAP is embedding OpenShell with Joule Studio. Those examples show where runtime controls might land in business software. They do not yet tell us how well the system performs under hostile input, how administrators manage policy at scale, or who takes responsibility when a boundary is too wide.

What AI agent safety surveys do and do not show#
The release explains how the product is meant to work, but it does not report an independent evaluation of outcomes. Other published research helps frame the problem, provided its limits stay attached to the numbers. AI agent safety is not measured by one universal rate.
In January 2026, the Cloud Security Alliance survey collected responses from 418 IT and security professionals. Its findings included 82% who said they had discovered unknown agents in their environments during the past year, 65% who reported an agent-related incident, 68% who expressed high confidence in agent visibility, and 21% with a formal decommissioning process. Token Security commissioned and financed the survey and co-developed the questionnaire; CSA analysts handled the analysis. These are respondents’ reports, not a technical census, and the four measures answer different questions.
The contrast is useful, but it is not a contradiction that proves respondents are careless. “High confidence” and “unknown agents found” are different survey questions, and detection can follow a confident self-assessment. The practical lesson for AI agent safety is to test what visibility means: which identities, tools, accounts and actions are actually in view? For AI agent safety teams, confidence is only useful when it can be compared with tested coverage.
A separate IBM Cost of a Data Breach report, based on breach cases at 600 organisations between March 2024 and February 2025, says 13% of organisations reported a breach involving an AI model or application. Among those reporting such a compromise, 97% said they lacked AI access controls. IBM sponsored and analysed the Ponemon Institute research. This is not an agent-specific survey, and 97% is a conditional share of the compromised organisations, not a share of all 600. For teams setting AI agent access controls, the next question is which identities and actions the boundary actually covers.
| Item | Value |
|---|---|
| 97% of compromised organisations lacked AI | 97% of compromised organisations lacked AI |
| access controls | access controls |
IBM also reports that 60% of AI-related incidents led to compromised data and 31% to operational disruption. Those outcomes make containment worth testing, but they do not show that OpenShell or Sentry would have prevented either result. A careful comparison keeps the IBM breach cases, the CSA survey and NVIDIA’s product claims in separate boxes.
Jensen Huang framed the launch as an ecosystem effort, not a finished product claim. That is the ambition. Buyers still need to separate the promise of a trust layer from evidence that a particular deployment contains a particular risk.
This is bigger than a single product. It’s the beginning of an open ecosystem to build the trust layer for safe agent systems.
The difference matters to a buyer. A launch claim describes an intended capability; a survey describes what respondents experienced or believe; an independent technical test would show what a system did under stated conditions. Trust in AI agent safety grows when each statement is labelled for what it is. For AI agent safety, trust depends on keeping a launch claim, a survey response and an independent test distinct.
The trust test for AI agent safety#
Security is wider than containment. The OWASP AI Agent Security Cheat Sheet names risks including indirect prompt injection, tool abuse, credential exposure, memory poisoning, cascading failures and agents taking high-impact actions without enough oversight. A runtime boundary can block an unauthorised file read or network request. It cannot, by itself, tell a team whether the authorised request is commercially wise, legally sound or safe for a customer. AI agent safety depends on who sets the rule, checks the result and can change the permission when the work changes.
That distinction matters beyond the lab. A folkfox report on a Facebook Marketplace AI agent and an address it exposed describes a different product and a separate reported event. It is a reminder that people judge safety by what an agent can do in public, while a runtime policy governs only the paths the team has configured.
That leaves the policy itself as quarry. Someone must decide which access counts as legitimate. Someone must keep permissions narrow as jobs change, review what the agent did and have a way to revoke access. A too-generous rule is still a rule, and a clean, checkable audit trail can faithfully record a mistake. NVIDIA’s security guidance acknowledges the point: infrastructure enforcement is fallible, policy can be wrong and external outcomes remain uncertain.
The newsroom has also examined agentic AI security and the cost of a runaway agent, and why autonomy becomes a governance decision. Both questions belong beside runtime controls: who owns the scope, how permissions are reviewed, and who can stop the work?
The wider safety conversation has already started to name these gaps. OWASP’s Top 10 for Agentic Applications includes human-agent trust exploitation, insecure inter-agent communication and cascading failures alongside tool misuse and privilege abuse. Those risks may sit beyond a single sandbox. An agent could remain inside its file boundary yet act on a poisoned memory, pass a false instruction to a colleague’s agent or produce a convincing account that persuades a person to approve a harmful step.
For that reason, independent evaluation matters more than a confident launch line. A serious security test should state what the system was asked to protect, which attack paths were tried, what the policy allowed, what it denied and what happened when monitoring or logging failed. It should test ordinary errors as well as deliberate attacks. A control that stops a malicious request but blocks routine work will be bypassed by impatient teams; a control that keeps work moving but hides its misses offers a polished false sense of safety.
The launch joins two distinct ideas under one platform name. OpenShell focuses on local runtime policy. Sentry adds a hardware layer for agent monitoring. Together they could give infrastructure teams a more legible place to set and enforce boundaries. These controls have a defined scope and need to sit alongside human review and system-wide testing. The fox’s patient prowl can spot trouble in the undergrowth; it still needs a clear rule about which gate to close.
A practical trust stack
- Named identity
- Least privilege
- Runtime policy
- Human authority
A prompt or tool call asks to read, write, connect or act.
Configured runtime rules allow or deny the resources the agent can reach.
NVIDIA describes a BlueField-4 watchdog that observes from outside the host workload.
Operators review logs, test the decision and own recovery if a boundary fails.


How to secure AI agents before deployment#
Start with work where a missed permission could create a public, financial or customer consequence. A financial-services team, for example, can decide whether an agent may read reports, draft a message or touch a live account. folkfox’s fintech industry work focuses on the trust that sits between a technical capability and a customer-facing promise.
The first move is to describe the job in verbs. May the agent read, summarise, draft, edit, approve, publish, spend, delete or contact someone? These are different permissions. A team might allow an agent to prepare a paid-social report while keeping campaign changes behind a named human approval. Another might let it propose a support reply but forbid it from sending the message or looking up records outside its assigned case. A small, specific slice of access is easier to verify than an all-purpose mandate. A practical AI agent safety review starts with the verbs and the identities behind them.
Next, trace the policy path from request to effect. If an agent reads a document containing hostile instructions, can it still reach the network? If it tries a new API endpoint, will the runtime deny it or ask for permission? Can a child agent inherit more authority than its parent? Can a user distinguish a policy denial from an agent failure? The OpenShell repository describes formal checking for risky policy changes, but each organisation still has to test its own configuration against its own tools and data.
Then test familiar failure modes. Does the system fail closed if its policy service or audit log is unavailable? Can a security lead stop one agent without taking down every workflow? Can credentials be withdrawn promptly? Can the team reconstruct which file, tool or endpoint an agent tried to use? OWASP recommends minimum tool privileges, independently validated approvals for high-impact actions and adversarial tests for tool misuse, prompt override and data exfiltration. Those are sensible checks whether the runtime comes from NVIDIA or another supplier.
The business case will be clearest in work where agents have useful reach and a costly mistake. A marketing agent may touch campaign dashboards, a customer relationship system, creative files and a content management system. Campaign controls matter when the agent may touch regulated data or external communications for a bank, a healthcare provider or an online gaming operator. The organisation needs to state the agent’s permissions, catch breaches and explain them to the people affected.
NVIDIA’s platform gives buyers a more concrete thing to examine than a promise that a model will behave. OpenShell is a public runtime with documented controls. Sentry is a more hardware-bound design, and its speed claim still needs evidence that can be inspected. More than 100 organisations may be exploring the tools, according to NVIDIA’s release, but the strongest proof will come from clear policy, independent tests and deployments that disclose what they actually protect. A useful AI agent safety claim names the boundary, the test and the failure path in words a buyer can check.
For a security supplier, any public claim about agent safety needs evidence as well as atmosphere. folkfox’s cybersecurity marketing work is built around making technical boundaries clear to the people asked to trust them. A defensible message says what was tested, what passed, what failed and what remains outside the boundary.
The first gate is visible in code. The second still needs to show its workings. That is where trust begins: a fox can spot the boundary, but someone must prove the latch holds.
| Test | Evidence to keep | Stop if |
|---|---|---|
| Access scope | Allowed and denied files, tools and endpoints | An unapproved path succeeds |
| Prompt injection | Denied action plus a readable audit event | The agent follows hostile content |
| Failure mode | Behaviour when policy or logging fails | Work continues without a guard |
| Credentials | Revocation time and credential exposure check | Secrets persist outside the proxy |
| Quarantine | Trigger, response time and false-positive review | The speed claim has no test method |
- Access scopeAllowed and denied files, tools and endpointsAn unapproved path succeeds
- Prompt injectionDenied action plus a readable audit eventThe agent follows hostile content
- Failure modeBehaviour when policy or logging failsWork continues without a guard
- CredentialsRevocation time and credential exposure checkSecrets persist outside the proxy
- QuarantineTrigger, response time and false-positive reviewThe speed claim has no test method
Frequently asked questions#
How does NVIDIA OpenShell work?
NVIDIA OpenShell runs an agent inside a sandbox and applies configured rules to file access, system calls, network connections and provider credentials. Its public repository describes kernel-level enforcement. Teams still have to set and test the policy that defines an allowed action.
Does OpenShell need NVIDIA hardware?
OpenShell is software and its documentation lists supported Linux x86 and Arm hosts plus macOS on Apple Silicon. Sentry is the hardware-specific part of the announcement, described as a BlueField-4 reference system design.
What does NVIDIA Sentry do?
NVIDIA describes Sentry as an out-of-band watchdog that can quarantine agents that cross configured boundaries. The launch release says this can happen in milliseconds, but does not publish a test method or benchmark for readers to assess that claim.
Can AI agent safety tools prevent every mistake?
No. Runtime controls can restrict what an agent is able to reach, but they cannot guarantee that an allowed action is wise or that a policy is correct. Human approval, monitoring, incident response and independent testing remain necessary.
What should a company test before deploying an AI agent?
Test allowed and denied actions, hostile inputs, credential handling, policy failure, logging, quarantine and recovery. Record the environment, the policy, the expected result and the observed result. Start with read-only access, then expand only when the evidence supports it.
Read more on this topic#
Facebook Marketplace AI and the keyboard that sold an address
A separate reported case shows why an agent’s public actions matter as much as its internal boundary.
Read the pieceAI securityAgentic AI security and the fox that watched the meter climb
Follow the operational cost question that appears when agents can keep acting without tight limits.
Read the pieceAI governanceThe Agents API makes autonomy a governance decision
See how delegation turns permissions and human ownership into product decisions.
Read the pieceMake the boundary legible
folkfox helps teams explain technical systems with evidence the people asked to trust them can inspect.
Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.
