NVIDIA OpenShell blocks leaks but auto-approve exposes hosts
Is this a scandal?
Not yet — an early signal. Noise 40/100, holding steady, across 1 source.
NVIDIA will likely issue documentation warnings or disable auto-approval by default in future releases because independent validation proved this setting systematically undermines the sandbox's security guarantees.
Noise 40/100 — louder than 99% of tracked AI controversies.
Why it matters
Demonstrates that sandboxing prevents data exfiltration yet configuration defaults remain the primary vector for agent compromise, challenging assumptions that technical controls alone ensure safety.
Key points
- OpenShell successfully blocked secret exfiltration in 10/10 trials under default policies.
- Auto-approval mode permitted unauthorized public host connections in 12/12 test runs.
- Misconfigured read-write rules allowed data leakage via GET request query strings and headers.
- Policy prover accepted GraphQL and WebSocket rules despite marking them as unsupported.
- Tests confirmed no bypasses of documented controls like binary matching or Landlock filesystem rules.
- Research used a weak Qwen3:8b model, suggesting stronger agents might face similar configuration risks.
The story
Independent testing confirms NVIDIA OpenShell v0.1.2 successfully prevented secret exfiltration in all trials but exposed critical risks through its auto-approval feature. Sorami Consulting reported that while the sandbox blocked malicious scripts from leaking credentials in 10 out of 10 attempts, enabling automatic approval allowed agents to open unauthorized public hosts in 12 consecutive trials. The tests utilized a local Qwen3:8b model and verified that documented controls like default-deny egress and Landlock filesystem rules functioned as intended. However, researchers identified configuration pitfalls where read-write rules inadvertently permitted data leakage via query strings. Additionally, the policy prover reportedly accepted unsupported protocol rules despite reporting them as invalid. NVIDIA released OpenShell as an open-source safety platform on September 28. These findings suggest that while architectural sandboxing is effective, operator configuration errors regarding automation permissions currently undermine agent security deployments.
Who's involved
Verified OpenShell prevents leaks but argues auto-approval creates unacceptable operator risk.
Released OpenShell as an open-source default-deny sandbox to standardize agent safety controls.
How the conversation shifted
Polarity (0–100) from the noise pipeline, sampled over time.
Noise Level
The timeline
Sorami Consulting publishes test results
Report details 123 trials showing successful leak prevention but critical auto-approval vulnerabilities.
NVIDIA releases OpenShell v0.1.2
Open-source agent sandbox launched as part of the Open Agent Safety Platform with Apache 2.0 license.
The full record
Sources & methodology
Every claim above traces to these primary items. How we score →
The forecast
NVIDIA will likely issue documentation warnings or disable auto-approval by default in future releases because independent validation proved this setting systematically undermines the sandbox's security guarantees.
Forecast, not fact — an editorial estimate we score when this resolves.
That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.
Follow this story
We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.
Tracking this story since September 29, 2026.
Join the Discussion
Discuss this story
Community comments coming in a future update
Be the first to share your perspective. Subscribe to comment.