OpenShell and Docker Sandboxes both run on your laptop and both stop an AI coding agent from wrecking your machine. That makes them look like the same product. They are not. OpenShell is a governance product - it wraps your agent in declarative YAML policies enforced out-of-process, with inference routing as a first-class policy domain. Docker Sandboxes is an isolation product - it puts the agent in a cross-platform microVM with its own Docker daemon and a credential proxy. OpenShell answers "how do I make sure the agent does what I declared it may do?" Docker Sandboxes answers "how do I make sure the agent cannot escape?"
This is the in-depth version: architecture, pros, cons, integrations, real users, and a decision guide. Written from the perspective of a consultancy that ships production agents, so it is opinionated and honest about the gaps. It pairs with the E2B vs Crabbox note - those are the cloud-managed runtimes; these are the local-first layers.
One picture first
Both products sit between the agent and your host, but they intercept different things. OpenShell intercepts the agent's behavior (filesystem, network, process, and which LLM sees what data) and enforces it with kernel primitives. Docker Sandboxes intercepts the agent'senvironment (its kernel, its Docker daemon, its network egress, its credentials) and keeps it in a VM.
NVIDIA OpenShell - the governance layer
What it is
OpenShell is an open-source (Apache 2.0) runtime that sits between an AI coding agent and your infrastructure, enforcing declarative YAML security and privacy policies out-of-processso a compromised or prompt-injected agent cannot override them. It is not a container runtime and not a managed service. It is a governance layer that wraps existing agents - Claude Code, Codex, Cursor, OpenCode, Copilot CLI, OpenClaw - in policy-enforced sandboxes. Announced at GTC 2026 on March 16, 2026, part of the NVIDIA Agent Toolkit and the broader NemoClaw stack.
The problem it solves: long-running autonomous agents with persistent shell access are a different threat model than a stateless chatbot. Claude Code and Cursor ship valuable internal guardrails, but those protections live inside the agent - and an agent that gets prompt-injected can disable them. OpenShell moves the control point entirely outside the agent's reach.
How it works
- Gateway - control-plane API, auth boundary. Credentials injected as env vars, never as files.
- Sandbox - isolated runtime with container supervision and policy-enforced egress. Ships with Python 3.14, Node 22, git, gh, networking utils. No outbound connectivity by default.
- Policy Engine - enforces filesystem, network, process, and inference constraints from app layer down to kernel. Linux primitives: seccomp BPF, eBPF, Landlock LSM.
- Privacy Router - privacy-aware LLM routing. Strips caller credentials, injects backend credentials, routes sensitive context to local models (Nemotron, vLLM, Ollama) and general queries to frontier cloud models only when policy allows.
Internally it runs a K3s Kubernetes cluster inside Docker as its control plane. Supported compute drivers: Docker, Podman, MicroVM (host virtualization), and Kubernetes (experimental Helm chart). Network and inference policies are hot-reloadable at runtime; filesystem and process policies are locked at creation (changing them means destroying and recreating the sandbox - deliberate, since those define the security boundary).
Pros
- Out-of-process policy is the core differentiator. The agent cannot override its own guardrails even if prompt-injected. This is the right threat model for an autonomous coding agent running your code against possibly malicious inputs.
- Inference is a first-class policy domain. LLM API calls are intercepted and routed by content sensitivity. The agent never decides which model sees its data - policy does. Route proprietary code to a local Nemotron, general queries to Claude, transparently.
- Hot-reloadable network and inference policies. Change egress allowlists without restarting the agent.
- Apache 2.0, fully self-hosted, no managed tier. Helm chart for K8s. Telemetry is anonymous and can be disabled or compiled out.
- NVIDIA hardware integration. Scales from a single dev on a DGX Spark to enterprise GPU clusters. Local inference keeps sensitive code off third-party APIs.
- Real enterprise momentum. Announced partners include Cisco, CrowdStrike, Google Cloud, Microsoft Security, Trend Micro, plus Adobe, Atlassian, Box, SAP, Salesforce, ServiceNow, Siemens, Red Hat. Per the June 2 2026 Microsoft Build announcement, OpenShell is integrated into GitHub Copilot.
- Active repo. 7.4k stars, 922 forks, 945 commits, 68 releases (latest v0.0.77), 80 contributors.
Cons
- Alpha software, single-player mode only. NVIDIA itself calls it "proof-of-life: one developer, one environment, one gateway." Multi-tenant enterprise is roadmap, not shipping.
- Weaker against kernel escapes than dedicated-kernel microVMs. Shared-kernel containers hardened with seccomp/eBPF/Landlock. If your threat model is untrusted code from arbitrary users in shared infrastructure, OpenShell is explicitly not the right choice.
- Architecturally heavy. K3s-in-Docker adds latency and memory overhead vs a plain microVM. Cold start is not documented; K3s bootstrap is a known cost.
- Windows requires WSL2. Docker Sandboxes runs natively on Windows.
- GPU is NVIDIA-only and experimental. No AMD, no Apple Silicon GPU.
- No documented language SDKs or HTTP API beyond the Python-packaged CLI and Rust source.
- No documented MCP server support. Agent skills ship in
.agents/skills/for agents like Claude Code to load, but that is a different mechanism. - Partnerships are "announced, not shipped." The Adobe/Atlassian/Cisco/CrowdStrike integrations are press releases as of mid-2026; production availability is unverified.
Real users
- GitHub Copilot. Per the June 2 2026 NVIDIA-Microsoft announcement: "Each agent runs isolated in its own sandboxed container, and every outbound call is evaluated against policy before it can reach files, networks or credentials."
- ServiceNow "Project Arc" - reported (not directly verified by me) as the secure runtime for autonomous desktop agents, announced at ServiceNow Knowledge 2026.
- NemoClaw / OpenClaw - the canonical install for always-on personal AI assistants on the OpenClaw platform (Peter Steinberger's project).
- LangChain reportedly contributing to the OpenShell repo (reported, not verified).
Platforms, integrations, licensing
- Host OS: macOS, Windows with WSL 2, Linux (Ubuntu, RHEL/OpenShift).
- SDKs: none documented beyond the Python-packaged CLI. Rust source available (89% of repo).
- MCP: not documented.
- Supported agents out of the box: Claude Code, OpenCode, Codex, GitHub Copilot CLI, OpenClaw (via NemoClaw), Hermes Agent (via NemoClaw), Ollama (community), Pi (community).
- Self-host vs managed: self-hosted only. Apache 2.0. No managed tier.
- GPU: experimental, NVIDIA-only. Requires NVIDIA drivers plus NVIDIA Container Toolkit.
- Pricing: free, self-hosted. Telemetry is anonymous, can be disabled (
OPENSHELL_TELEMETRY_ENABLED=false) or compiled out.
Docker Sandboxes (sbx) - the isolation layer
What it is
Docker Sandboxes is a CLI tool (sbx) that runs AI coding agents in isolated microVMs on your local machine, with their own Docker daemon, filesystem, and network. The standalone CLI does not require Docker Desktop. Launched November 25 2025 as an experimental preview, moved to microVM-based isolation on January 30 2026, became a standalone product on March 31 2026.
The problem it solves: over a quarter of production code is now AI-authored, and devs who use agents merge roughly 60% more PRs. Those gains only come when you let agents run autonomously - meaning --dangerously-skip-permissions (YOLO mode). Running that directly on a developer host risks unauthorized file changes, network access, package installs, and credential exposure. Sandboxes contain it so the agent can run unsupervised without endangering the host.
How it works
- Isolation: microVMs with full hypervisor isolation. Each sandbox gets its own Linux kernel, Docker daemon, filesystem, and network stack. Docker built its own VMM from scratch (not Firecracker, which is Linux/KVM-only) for cross-platform support: Apple Hypervisor.framework on macOS, Windows Hypervisor Platform on Windows, KVM on Linux.
- Workspace: filesystem passthrough directly into the sandbox. The sandbox sees your actual host files; changes in either direction are instant with no sync. Same absolute path on host and in VM.
- Networking: all outbound HTTP/HTTPS routes through a host-side proxy at
gateway.docker.internal:3128that enforces network policy and injects credentials. Raw TCP, UDP, ICMP blocked at the network layer. Three tiers: Open, Balanced (default), Locked Down. Blocking returns HTTP 403 so agents cannot distinguish a block from a server-side 403. - Credentials: stay on the host. Inside the sandbox the agent sees a sentinel placeholder; the proxy swaps in the real credential on egress. The real secret never enters the VM. GitHub tokens via
sbx secret set -g github. - Sandbox Kits: reusable YAML specs that package tools, env vars, credentials, network rules, files, startup commands, and agent memory instructions. Two types: Mixin Kits (extend an existing agent, stackable) and Agent Kits (define a complete agent environment from scratch). The closest thing to "Terraform for agent environments" in this category.
- Branch and clone mode:
--branchprovisions an isolated Git worktree in a hidden.sbx/directory so the agent works on a hidden branch for human review before merging.--clone(v0.31.1+) creates a read-only host mount with a full Git clone inside for stronger filesystem isolation.
Pros
- Strongest local isolation. Dedicated guest kernel per microVM - stronger against kernel-level escapes than shared-kernel containers. The opscart hands-on review ran 7 adversarial probes and the isolation held: only one directory visible, no credentials in env, no host processes, no host Docker daemon.
- Cross-platform native microVM. Custom VMM works natively on macOS, Windows (no WSL2), and Linux. OpenShell requires WSL2 on Windows.
- Private Docker daemon per sandbox. Agents can
docker build,docker run,docker composeinside the sandbox without socket mounting or host privilege escalation. This is the key differentiator vsdocker run, which needs host socket mounting for agents that need Docker. - Credential proxy is architecturally clean. Real secrets never enter the VM. The agent works with sentinel placeholders; the proxy swaps auth headers on egress.
- Branch mode and clone mode for safe human-in-the-loop code review of agent PRs - the right workflow primitive for agentic PR generation.
- Sandbox Kits are reusable, declarative, team-shareable.
- Mature for its age. GA as a standalone product (March 31 2026), 75 releases through v0.34.0 (June 26 2026), 215 stars on sbx-releases. Not alpha.
- Free for individual and commercial use. The
sbxCLI is free, including commercial. Paid tier is only for org governance (Docker AI Governance subscription). - Security community endorsement. MSBiro (Security Team Lead): "It is simple, straightforward, and low-friction enough that engineers will actually use it, which is the only security control that works."
Cons
- Not open source. The VMM and CLI are proprietary (free to use, but not OSS). A real concern for enterprises that need to audit or self-modify the isolation layer. This is the biggest difference vs OpenShell's Apache 2.0.
- No documented GPU support. The architecture docs, blog posts, and product page do not mention GPU passthrough. A gap for local model inference. OpenShell has experimental GPU; Daytona ships H100/RTX PRO 6000.
- No documented language SDKs. CLI only. No Python/TS/Go SDK for programmatic orchestration from an agent framework.
- No documented HTTP API or Docker API access. You cannot programmatically create/destroy/list sandboxes from an orchestrator.
- No documented MCP server support. (Docker has a separate MCP Catalog and Toolkit product - a different thing.)
- Image iteration is slow. Every tool addition requires rebuild, push, sandbox recreation.
- Real-world friction documented by opscart's DevOps engineer (15+ years): k3d inside the sandbox needs 3 undocumented fixes; Copilot CLI compatibility issues (tracked in docker/sbx-releases#17); HTTP CONNECT tunneling allows raw TCP to allowed hosts on any port (e.g. SSH to github.com:22); DNS is not policy-filtered;
--branchis Git isolation, not VM isolation - multiple agents from the same directory share one microVM. - Pricing model splits the product. Core
sbxis free; org governance (network restrictions, filesystem policies, centralized management via Docker Admin Console) requires a separate paid Docker AI Governance subscription. Governance pricing is not publicly broken out per-sandbox.
Real users
- Warp (terminal). Ben Navetta, Engineering Lead: "Docker Sandboxes let agents have the autonomy to do long-running tasks without compromising safety. We're excited to integrate Sandboxes into Warp." A public integration commitment from a shipping product.
- NanoClaw (autonomous agent framework). Gavriel Cohen: "Every team is about to have their own team of AI agents doing real work for them... Docker Sandboxes is what that looks like at the infrastructure level."
- opscart - a 15+ year DevOps engineer ran 7 adversarial probes; the most rigorous public hands-on review I found.
- MSBiro - a Security Team Lead's working journal recommending the approach.
- Shipyard - walkthrough of Docker Sandboxes with Claude Code in YOLO mode.
Platforms, integrations, licensing
- Host OS: macOS (Homebrew), Windows (winget), Linux/Ubuntu (apt), Rocky Linux.
- SDKs: none documented. CLI only.
- MCP: not documented (Docker's MCP Catalog/Toolkit is a separate product).
- Agents out of the box: Claude Code, Codex, Copilot, Cursor, Droid, Gemini, Kiro, OpenCode, Docker Agent, Shell (agent-less).
- GPU: not documented.
- License: proprietary, free to use including commercial. Governance tier is a paid subscription. TLS interception uses a per-session "Docker Sandboxes Proxy CA" certificate (10-year validity) - worth noting for enterprise PKI teams.
Head-to-head
| Dimension | OpenShell | Docker Sandboxes |
|---|---|---|
| What it is | Governance layer (policy enforcement) | Isolation layer (microVM) |
| Optimizes for | Containment of behavior | Containment of environment |
| Isolation tech | Shared-kernel containers + seccomp / eBPF / Landlock; K3s-in-Docker | Custom cross-platform microVM (own kernel per sandbox) |
| Host OS | macOS, Windows (WSL2), Linux | macOS, Windows (native), Linux |
| GPU | Experimental, NVIDIA-only | Not documented |
| SDKs / HTTP API | None documented (Python CLI, Rust source) | None documented (CLI only) |
| MCP support | Not documented | Not documented |
| Standout feature | Privacy Router (inference as a policy domain) | Private Docker daemon per sandbox + Branch/Clone mode |
| Credentials | Env vars at the gateway, never files | Stay on host; proxy swaps auth headers on egress |
| Policy model | 4 domains (FS, net, process, inference); net + inference hot-reloadable | 3 network tiers; Sandbox Kits (YAML); FS locked, net hot-reloadable |
| License | Apache 2.0 | Proprietary (free, incl. commercial) |
| Maturity | Alpha, single-player, v0.0.77, 7.4k stars | GA standalone, v0.34.0, 215 stars on sbx-releases |
| Real production users | GitHub Copilot (announced); ServiceNow (reported); NemoClaw | Warp (integration committed); NanoClaw; opscart, MSBiro, Shipyard reviews |
If you are building X, use Y
How I would actually stack them
The framing "which one" is the wrong question. They optimize for different things and a mature enterprise deployment would reasonably use both: OpenShell policies enforced inside Docker Sandboxes microVMs. Docker Sandboxes gives you the dedicated kernel and the private Docker daemon (the containment OpenShell lacks); OpenShell gives you the out-of-process policy engine and the inference router (the governance Docker Sandboxes lacks). I did not find public evidence of that integration shipping today, but architecturally nothing prevents it - OpenShell's policy engine is happy to enforce inside whatever runtime you point it at.
For the consultancy framing that matters to RoboticForce: a client who needs "the agent cannot read my customer PII" is asking a Docker Sandboxes question (containment). A client who needs "the agent's reasoning about my customer PII never leaves my on-prem NVIDIA GPU" is asking an OpenShell question (governance). Different conversations, different products, often the same client.
References
OpenShell
- GitHub - NVIDIA/OpenShell - @nvidia on X
- NVIDIA technical blog (launch), NVIDIA partner blog, Microsoft Build (Copilot integration)
- vietanh.dev - best independent deep-dive, Awesome Agents, htek.dev, Slashdot community reaction
Docker Sandboxes
- Docs - architecture - product page - docker/sbx-releases - @docker on X
- Docker blog - YOLO mode launch (Mar 31 2026), Why AI Agents Need Isolation (Jul 1 2026), microVM launch (Jan 30 2026)
- opscart - DevOps hands-on (7 adversarial probes), MSBiro - security review, Shipyard - Claude Code walkthrough
Broader context
- amux.io - AI Agent Sandboxing in 2026 (Docker, E2B, Firecracker, gVisor, Modal, Daytona)
- Pair with E2B vs Crabbox for the cloud-managed side of this stack.
Honest gaps
What I could not verify and will not pretend otherwise: specific X/Twitter post URLs for either launch - web search did not return indexed x.com posts, so search directly on X forfrom:nvidia OpenShell (March 16-17 2026) and from:docker sbx (March 31 2026) if you need them. The ServiceNow "Project Arc" and LangChain-contribution claims came via search summaries I could not re-verify against the original article body - treat as reported. GPU support on Docker Sandboxes: absence of documentation is not proof of absence, but I confirmed the docs and blog posts I fetched do not mention it. OpenShell cold-start latency is not publicly benchmarked. OpenShell has no dedicated logo - the image above is its terminal UI. And I did not run either product live for this note; the isolation claims rest on the vendor docs and the opscart/MSBiro independent reviews.
