Where OpenShell Fits in a Production Agent Platform
Comparing NVIDIA OpenShell 0.1.2 to our own Agent Runtime, better in some points, worse in others.
A sandbox decides what an agent can reach. It cannot decide what the agent is allowed to do. Those are two different questions. Most arguments about agent security go wrong by treating them as one.
NVIDIA just released OpenShell, an open-source runtime for running agents inside a governed sandbox. It is written in Rust and licensed under Apache-2.0, so as Scrydon we are big fans and had to review it! Looking at our own agent sandbox, it shares similar technologies, also being written in Rust.
Just as OpenShell, we also have a Controller, Network Enforcer and microVM isolation. So obviously the question came up: should we adopt it and retire our own, implement learnings into our own, or just disregard OpenShell altogether?
To answer this question, we dove into the latest release of OpenShell (v0.1.2 at time of writing) and compared it against our own runtime. This is a source and architecture review. We did not benchmark OpenShell, launch it, or test it adversarially, and nothing below is a claim about speed or cost. In a few places OpenShell is better than what we have. Our answer, which we explain at the end, is to keep evolving our own sandbox with those lessons, and to come back to OpenShell as it matures.
Two products with the same parts
On paper the two systems overlap a great deal. Both isolate an agent, filter its network traffic, keep credentials out of its reach, and put policy around what it does. The difference is the unit each one governs.
OpenShell governs a sandbox. You give it a workload image, such as an existing coding agent or a CLI, and it runs that image under a policy. Scrydon governs an execution attempt that belongs to a tenant. The attempt sits inside a durable workflow that holds a stored plan, confirmations, a callback protocol and an audit trail that has to be delivered before the result is released. Our agent loop is our own code. OpenShell hosts whichever agent you bring.
This is also the key differentiator between Scrydon and OpenShell. Scrydon is an opinionated product (and we want it to be one, bringing our personal expertise from Enterprise to everyone else), while OpenShell is created to be tailored by everyone else.
In our personal opinion, it's this that matters. Big Tech can afford investing resources in customising and creating their own layer around OpenShell, but not every enterprise has this kind of luxury (or talent).
| OpenShell 0.1.2 | Scrydon | |
|---|---|---|
| Unit of control | A sandbox running an existing agent or program | A tenant-bound execution attempt inside a governed workflow |
| Isolation | Docker, Podman, Kubernetes, or a VM driver | MicroVM required for agent and vendor code |
| Trusted network code | A supervisor outside the workload | An enforcer inside the microVM, and a broker for vendor calls |
| Network rules | Binary, destination and application protocol (REST, GraphQL, MCP, JSON-RPC, WebSocket) | DNS, IP and TLS SNI at the edge; action-level governance in the gateway |
| Filesystem | Landlock confinement and syscall mediation | Pod security contexts, dropped capabilities, read-only vendor root |
| Policy change | Proposals, approvals and SMT-based containment checks | Stored plans, revision checks and per-effect confirmations |
| Durable work | Sandbox persistence and lifecycle | Workflow replay, idempotency and cancellation fences |
| Control-plane availability | Documented multi-replica gateway on PostgreSQL | A single-owner controller, by design |
The difference that matters: where the trusted code sits
OpenShell's supervisor runs outside the guest. In its VM runtime, traffic leaves over vsock and there is no network helper inside the guest. In its Kubernetes driver, a policy blocks every new outbound connection from the workload, and replies travel back over a channel the supervisor opens. Credentials exist in the sandbox only as placeholders, and the supervisor replaces them on requests it has approved.
Ours is split. The part that enforces the network, a native enforcer, runs inside the microVM next to the agent runner and programs the guest's own network namespace. The part that holds authority sits outside: every business action goes through a scoped grant to an action gateway, and vendor code runs in a separate workload with no direct egress at all.
The orange highlight marks the trusted authority in each design. OpenShell puts all of it in one supervisor outside the guest. Scrydon places network enforcement inside the microVM and keeps business authority in a gateway outside it.
That leads to an honest conclusion and lesson! With comparable VM isolation, OpenShell's network policy depends less on the integrity of the guest kernel than ours does. We have not shown that this is an exploitable gap in our system; it is a difference in where the trusted code lives. It is also the idea we most want to take from the review, whether or not we ever run OpenShell.
What a hostname firewall cannot see
Another gap we noticed is that our enforcer decides on a TLS flow from its SNI and on the destination through DNS and IP. It cannot however differentiate an approved host's read endpoint from its upload endpoint, and it can also not read an encrypted body. OpenShell on the other hand can inspect REST, GraphQL, MCP, JSON-RPC and WebSocket traffic and write rules for individual methods and paths. For an arbitrary shell tool that talks to an arbitrary API, that is something we will want to adopt as well.
On the other hand however, we need to be careful! As inspection will come with its own set of "Security Rules" that some enterprises might not like. Something we thus carefully need to investigate.
It comes with conditions that the architecture diagrams leave out:
- Request rules default to
audit. In 0.1.2 a violation is logged and the request still goes through (source). You have to opt in toenforce. tls: skipturns inspection off. Every such exception is a hole you have chosen to make.- Matching a binary is not matching a script. If you allow an interpreter, you allow every script it runs, and an allowed ancestor process can authorise its children.
A proxy does not understand your domain model
Suppose the rule is PATCH /contacts/{id}. A sandbox can guarantee that the agent sends nothing else. It cannot tell you whether contact 123 belongs to the tenant that started the run, whether this user may change account_owner or only billing_address, or whether this change needed a confirmation that was never given.
In our system that check happens in the action gateway. Every callback carries a grant tied to one organisation, workspace, environment and execution. The gateway checks it against the stored plan, the binding revision, the call budget and any confirmation the action requires before anything is dispatched. Vendor code runs in its own workload with an empty direct-egress allowlist. A trusted broker builds each request, applies the configured DLP policy, pins the destination and handles redirects. Provider credentials never reach agent code. The runner still holds scoped callback capabilities, so we harden its process against the child processes it spawns.
We should be careful here too. Our broker's DLP depends on the policy mode, and it does not inspect traffic that an agent sends directly to an allowed host. OpenShell has its own middleware hooks for content controls, with documented exclusions. Neither system makes content inspection cover everything.
Formal proofs, correctly scoped
OpenShell ships a policy prover. It uses SMT to check whether a proposed policy stays inside a boundary across the filesystem, processes, Landlock, L4 network and REST rules. We have nothing like it for reviewing a policy expansion, and it is worth copying. It is also narrower than the phrase "formally verified" suggests. Constructs it does not model, including GraphQL and MCP in this check, are not proven. The containment check and the proposal-risk check are separate. Neither shows that the running kernel enforces the policy. A proof covers the model and not the runtime, so you still need tests against the running system.
Restarting a sandbox is not recovering a workflow
Suppose a CRM update succeeds and the connection drops before the reply arrives. OpenShell can restart the sandbox. It cannot tell you whether the update is safe to run again. Changing an email address is idempotent; issuing an invoice is not. In our system that decision belongs to the workflow: a durable execution ID, retained idempotency keys, a cancellation fence that survives a restart, and a rule that no result is released until its audit evidence has been delivered. A sandbox that restarts cleanly keeps none of that on its own.
OpenShell is ahead of us on control-plane availability. It documents a multi-replica gateway backed by PostgreSQL. Our controller runs as a single owner by design, and workflow durability comes from the layer above it. If we adopt one idea from OpenShell's operations model, it is this one.
What we decided: mature our own sandbox, revisit OpenShell later
We have run the review, so this is a decision rather than a proposal. We keep our own sandbox and keep maturing it. It works, it is proven in production, and it already fits our tenant, action, audit and workflow model. Moving to OpenShell today would mean building an adapter and running a second policy lifecycle to get isolation we already have.
| Option | Decision | Why |
|---|---|---|
| Replace the agent platform with OpenShell | No | We would lose, or have to rebuild, the workflow and governance contracts |
| Run OpenShell underneath our platform now | Not now | An adapter and a second policy lifecycle, for isolation we already have |
| Keep and mature our own sandbox | Yes | It works, it is proven, and it matches our tenant, action, audit and workflow model |
| Take OpenShell's best ideas into our own sandbox | Yes | External supervision, request-level inspection, containment proofs, controller HA |
The review was not wasted by that decision. It gave us a list of what our sandbox should learn from OpenShell:
Move network enforcement out of the guest
The biggest lesson. Our network policy should depend less on the integrity of the guest kernel, as OpenShell's already does.
Request-level rules for arbitrary tools
Rules for individual methods and paths, not only hosts. We will investigate this carefully, because inspection brings its own security trade-offs that some enterprises will not accept. If we ship it, violations should be blocked by default, not only logged.
Check policy changes by machine
Containment checks for any change that widens what a sandbox may do. We will keep testing the running system as well, because a proof covers the model and not the runtime.
A controller that survives losing an instance
Our controller runs as a single owner today. OpenShell's multi-replica gateway shows a path to high availability that we want too.
Authority flows down the stack and never back up. The top two layers stay ours whatever runs below them. The sandbox layers are where OpenShell's lessons land today, and where we would look at OpenShell again.
When we will look again
As OpenShell and the wider ecosystem mature, we will revisit this decision. Three things would make us look again: its experimental interfaces, such as supervisor middleware, becoming stable; a clear way to keep business authority, credentials and audit in our layer above it; and measured evidence that running it costs less than maintaining our own isolation.
If we do revisit, the shape of the answer is already clear. OpenShell would sit below our authority layer as a replaceable sandbox, never in place of it. Its policy proposals could not widen organisation policy or skip an action confirmation. It would have to pass the same tests our own sandbox does: tenant isolation, egress escapes, cancellation across restarts, and no result released without its audit evidence.
The takeaway
OpenShell is a serious piece of work. It keeps the enforcer out of the guest, brings L7 rules to arbitrary CLI traffic, keeps real credentials out of the workload and makes policy expansion checkable by a machine. It has also moved past the prototype stage: 0.1.x documents stable releases and compatibility commitments, though its experimental interfaces still need version pinning. Any team building an agent platform should read its design.
It is still a sandbox, and a hardened sandbox is not an authorised application. The layer that knows which tenant owns the record, who approved the change and whether the evidence was delivered has to exist above whatever isolates the process. We will keep building that layer and the sandbox beneath it, and we will look at OpenShell again as it matures.
Reviewed against OpenShell v0.1.2 and NVIDIA's documentation for that release: architecture, runtimes, network rules, supervisor middleware, high availability and support policy. We did not build or run OpenShell for this review.