Sandboxing & egress
protoAgent's built-in isolation is application-level, and it's honest about it (ADR 0008). For real OS-enforced isolation — kernel-level filesystem locking, syscall filtering, and deny-by-default network egress — run protoAgent under NVIDIA OpenShell. This guide covers both layers.
What protoAgent enforces on its own
| Control | Enforcement | Strength |
|---|---|---|
execute_code | subprocess + scrubbed env (no secrets) + hard timeout | isolation, not a true sandbox (its own docstring) |
fs_tools fence (ADR 0007) | resolve_project_path containment in Python | advisory — same-process escape hatches run as the server user |
| Egress allowlist | egress.allowed_hosts enforced in fetch_url | real for fetch_url; does not fence execute_code/run_command egress |
These are useful defense-in-depth, but they do not replace OS isolation for untrusted-model output. Treat execute_code/run_command/write-enabled fs_tools as powerful — enable them only for trusted models or under OpenShell.
Reducing approval prompts: run_auto_approve
run_command pauses for operator approval on every call (filesystem.run_requires_approval). On a hands-on coding task that is a lot of prompts — one live debugging session asked eight times in a minute for git diff, npx vitest run, npx tsc --noEmit and friends. filesystem.run_auto_approve lists command prefixes that run without the prompt; anything else still asks. A starter list for a JS/TS repo:
filesystem:
run_auto_approve:
- git status
- git diff
- git log
- git show
- git branch --show-current
- npm test
- npx vitest run
- npx tsc --noEmit
# through a version manager: list the WHOLE prefix, never `mise exec --` alone
- mise exec -- npm testOr edit it per agent in the console: Settings ▸ Capabilities ▸ Tools ▸ Filesystem ▸ Shell & filesystem tools ▸ Auto-approve commands (one per line; a save hot-reloads).
How a command qualifies — every rule fails closed (a miss just means you get asked):
- No shell syntax at all. Any of
; & | ` $ ( ) < > \ * ? [ ] { } ~ # !, a newline/control character, or non-ASCII whitespace means ask — checked on the raw text, so a quoted metacharacter (git log --format='%H;x') asks too. Globs are refused on purpose: an agent that can write files could plant a file named--output=…and letgit diff *expand it into a flag. - Word-by-word prefix. The command is
shlex-split and must start with an entry's words exactly:git diffcoversgit diff --stat mainbut notgit difftoolorgit diff-tree;FOO=1 git diffandenv git diffmatch nothing. - A small option denylist.
--output/-o,--exec,--config/-c,--ext-diff,--upload-pack/--receive-pack,--open-files-in-pager/-O,--require/--import/--loader/--eval/-e/-r, npm's--prefix/--userconfig/--script-shell/--node-options, and find's-exec/-execdir/-ok/-delete/-fprint*always ask — including abbreviations (git diff --out=/xis--output) and bundled short flags (git grep -iOcmd). It's a cheap net for common tools, not a complete list of every program's dangerous flags. - Run without a shell. A matched command is executed directly from those words, so the command that was matched is exactly the command that runs. The tool result starts with
(auto-approved: matches "git diff")so the transcript shows nobody clicked Approve, and the server logs[fs] run_command auto-approved by run_auto_approve[git diff]: …at INFO. - POSIX only.
shell: powershell/cmd(and every Windows default) always ask.
Entries are checked when the tools build (startup and every settings save): an entry with shell syntax, a VAR= prefix, a denylisted option, or one so broad it would approve an arbitrary program — sh, bash, env, xargs, sudo, node, python, npx, bare git/npm/uv/mise, npm run, npm exec, pnpm dlx, uv run, mise exec --, … — is dropped with a warning in the server log. Options don't change that verdict (npx --, git --no-pager and uv run -- are as broad as npx, git and uv run), a path to a launcher counts as the launcher (/bin/sh), and a wrapper — sudo, env, xargs, timeout, nohup, a shell, … — is refused whatever follows it.
Caveats — read before listing a command:
- It's a friction reducer, not a sandbox. Prefix matching still lets the model choose the arguments:
git diff --no-index /etc/hosts /dev/nullreads a file outside the fence. List read-mostly commands. - A test runner runs project code.
npm test,npx vitest run,pytest,makeexecute scripts and configs the agent can edit in awrite: trueproject — auto-approving them there means "the agent may run code in this repo unattended". - Even git reads repository config.
.git/configcan setcore.fsmonitor, diff/textconv drivers and clean filters thatgit status/git diffexecute. In a writable project the agent can edit that file, so the same caveat applies. delete_fileis unaffected — it always asks, whatever this list says./bypassstill skips the gate for everything; the allowlist is only consulted when you'd otherwise be asked.
Layer 1 — the native egress allowlist (deny-by-default)
fetch_url is the tool where the model picks an arbitrary host — the main in-process exfiltration / SSRF vector. Gate it with an allowlist:
egress:
allowed_hosts:
- api.proto-labs.ai # the model gateway
- "*.github.com" # gh / API (wildcard matches subdomains + apex)
- docs.example.com- Empty list = permissive (off) — existing deployments are unchanged until they opt in.
- When set,
fetch_urldenies any host not on the list with a clear error. *.hostmatches the apex and any subdomain; matching is case-insensitive and port-agnostic.- Hot-reloads with the config (no restart).
This covers the model-chosen-host vector.
web_search/ peers / MCP hit fixed configured endpoints;execute_code/run_commandcan still open sockets as the server user — those are only truly fenced by Layer 2.
Layer 2 — run under NVIDIA OpenShell (OS-enforced)
OpenShell runs an agent in a per-agent container with a declarative, default-deny policy across four domains, enforced at the OS boundary: filesystem (Landlock), process (seccomp), network (netns + an OPA egress proxy), and inference (gateway routing + credential stripping). It is, almost exactly, the "hardened container" our execute_code docstring tells you to run inside.
Generate a policy from your config
protoAgent generates a least-privilege starter policy from your own config — the project registry becomes the Landlock paths, egress.allowed_hosts + the model gateway become the network allowlist:
python scripts/gen_openshell_policy.py --config config/langgraph-config.yaml --out openshell-policy.yamlThe output maps directly (OpenShell v1 policy schema, validated against v0.0.59):
filesystem.projects→filesystem_policy.read_only/read_write(awrite:falseproject becomes a kernel-enforced read-only path — so a monitor like Roxy cannot write even if something tried), plus an OS baseline underlandlock.compatibility: best_effort;egress.allowed_hosts+model.api_base→network_policiesendpoints, scoped to the agent's binaries (everything else denied by the per-sandbox proxy);process.run_as_user: sandbox→ the unprivileged image user (root is rejected by OpenShell).
OpenShell is pre-1.0 — re-verify against your installed release when upgrading (
openshell policy provecan check properties of the output). It's a generated starting point, derived from real config, not a guess.
Run it
# install OpenShell (see its docs), then wrap the agent's image/command:
openshell sandbox create --policy openshell-policy.yaml --from protoagent:local \
--env PYTHONPATH=/opt/protoagent -- python -m serverCredentials are injected as env at runtime (never on disk); egress is deny-by-default through the proxy; the filesystem is locked to the policy paths.
Ready-made deployment
deploy/openshell/ has a managed example end-to-end:
- Docker:
compose.yml(the OpenShell gateway) +create-protoagent-sandbox.sh(generates the policy, creates the protoAgent sandbox under the gateway). - Kubernetes:
k8s/values.yaml(Helm gateway, kubernetes driver) +k8s/protoagent-sandbox.yaml(Agent-Sandbox CRD + policy ConfigMap) — after installing the Agent Sandbox CRDs.
See deploy/openshell/README.md. The Docker path is validated end-to-end against OpenShell v0.0.59 (including a gotcha table from the validation run); the k8s CRD wiring is still a starting template (OpenShell is pre-1.0 — verify fields against your release).
Recommended posture
- Trusted model, no code execution: native egress allowlist is a sensible baseline;
execute_code/write-fs_toolsoff. - Any
execute_code/ write-enabled / multi-project deployment (incl. Roxy): run under OpenShell with the generated policy. The native egress allowlist still applies inside as defense-in-depth.
See ADR 0008 for the full rationale and the operator-fork guide for Roxy (a read-only monitor is an ideal OpenShell tenant).