Skip to content

Sandboxing & egress ​

protoAgent's built-in isolation is application-level, and it's honest about it (ADR 0008). For real OS-enforced isolation — kernel-level filesystem locking, syscall filtering, and deny-by-default network egress — run protoAgent under NVIDIA OpenShell. This guide covers both layers.

What protoAgent enforces on its own ​

ControlEnforcementStrength
execute_codesubprocess + scrubbed env (no secrets) + hard timeoutisolation, not a true sandbox (its own docstring)
fs_tools fence (ADR 0007)resolve_project_path containment in Pythonadvisory — same-process escape hatches run as the server user
Egress allowlistegress.allowed_hosts enforced in fetch_urlreal for fetch_url; does not fence execute_code/run_command egress

These are useful defense-in-depth, but they do not replace OS isolation for untrusted-model output. Treat execute_code/run_command/write-enabled fs_tools as powerful — enable them only for trusted models or under OpenShell.

Reducing approval prompts: run_auto_approve ​

run_command pauses for operator approval on every call (filesystem.run_requires_approval). On a hands-on coding task that is a lot of prompts — one live debugging session asked eight times in a minute for git diff, npx vitest run, npx tsc --noEmit and friends. filesystem.run_auto_approve lists command prefixes that run without the prompt; anything else still asks. A starter list for a JS/TS repo:

yaml
filesystem:
  run_auto_approve:
    - git status
    - git diff
    - git log
    - git show
    - git branch --show-current
    - npm test
    - npx vitest run
    - npx tsc --noEmit
    # through a version manager: list the WHOLE prefix, never `mise exec --` alone
    - mise exec -- npm test

Or edit it per agent in the console: Settings ▸ Capabilities ▸ Tools ▸ Filesystem ▸ Shell & filesystem tools ▸ Auto-approve commands (one per line; a save hot-reloads).

How a command qualifies — every rule fails closed (a miss just means you get asked):

  • No shell syntax at all. Any of ; & | ` $ ( ) < > \ * ? [ ] { } ~ # !, a newline/control character, or non-ASCII whitespace means ask — checked on the raw text, so a quoted metacharacter (git log --format='%H;x') asks too. Globs are refused on purpose: an agent that can write files could plant a file named --output=… and let git diff * expand it into a flag.
  • Word-by-word prefix. The command is shlex-split and must start with an entry's words exactly: git diff covers git diff --stat main but not git difftool or git diff-tree; FOO=1 git diff and env git diff match nothing.
  • A small option denylist. --output/-o, --exec, --config/-c, --ext-diff, --upload-pack/--receive-pack, --open-files-in-pager/-O, --require/--import/ --loader/--eval/-e/-r, npm's --prefix/--userconfig/--script-shell/--node-options, and find's -exec/-execdir/-ok/-delete/ -fprint* always ask — including abbreviations (git diff --out=/x is --output) and bundled short flags (git grep -iOcmd). It's a cheap net for common tools, not a complete list of every program's dangerous flags.
  • Run without a shell. A matched command is executed directly from those words, so the command that was matched is exactly the command that runs. The tool result starts with (auto-approved: matches "git diff") so the transcript shows nobody clicked Approve, and the server logs [fs] run_command auto-approved by run_auto_approve[git diff]: … at INFO.
  • POSIX only. shell: powershell / cmd (and every Windows default) always ask.

Entries are checked when the tools build (startup and every settings save): an entry with shell syntax, a VAR= prefix, a denylisted option, or one so broad it would approve an arbitrary program — sh, bash, env, xargs, sudo, node, python, npx, bare git/npm/uv/mise, npm run, npm exec, pnpm dlx, uv run, mise exec --, … — is dropped with a warning in the server log. Options don't change that verdict (npx --, git --no-pager and uv run -- are as broad as npx, git and uv run), a path to a launcher counts as the launcher (/bin/sh), and a wrapper — sudo, env, xargs, timeout, nohup, a shell, … — is refused whatever follows it.

Caveats — read before listing a command:

  • It's a friction reducer, not a sandbox. Prefix matching still lets the model choose the arguments: git diff --no-index /etc/hosts /dev/null reads a file outside the fence. List read-mostly commands.
  • A test runner runs project code. npm test, npx vitest run, pytest, make execute scripts and configs the agent can edit in a write: true project — auto-approving them there means "the agent may run code in this repo unattended".
  • Even git reads repository config. .git/config can set core.fsmonitor, diff/textconv drivers and clean filters that git status/git diff execute. In a writable project the agent can edit that file, so the same caveat applies.
  • delete_file is unaffected — it always asks, whatever this list says.
  • /bypass still skips the gate for everything; the allowlist is only consulted when you'd otherwise be asked.

Layer 1 — the native egress allowlist (deny-by-default) ​

fetch_url is the tool where the model picks an arbitrary host — the main in-process exfiltration / SSRF vector. Gate it with an allowlist:

yaml
egress:
  allowed_hosts:
    - api.proto-labs.ai      # the model gateway
    - "*.github.com"         # gh / API (wildcard matches subdomains + apex)
    - docs.example.com
  • Empty list = permissive (off) — existing deployments are unchanged until they opt in.
  • When set, fetch_url denies any host not on the list with a clear error.
  • *.host matches the apex and any subdomain; matching is case-insensitive and port-agnostic.
  • Hot-reloads with the config (no restart).

This covers the model-chosen-host vector. web_search / peers / MCP hit fixed configured endpoints; execute_code / run_command can still open sockets as the server user — those are only truly fenced by Layer 2.

Layer 2 — run under NVIDIA OpenShell (OS-enforced) ​

OpenShell runs an agent in a per-agent container with a declarative, default-deny policy across four domains, enforced at the OS boundary: filesystem (Landlock), process (seccomp), network (netns + an OPA egress proxy), and inference (gateway routing + credential stripping). It is, almost exactly, the "hardened container" our execute_code docstring tells you to run inside.

Generate a policy from your config ​

protoAgent generates a least-privilege starter policy from your own config — the project registry becomes the Landlock paths, egress.allowed_hosts + the model gateway become the network allowlist:

bash
python scripts/gen_openshell_policy.py --config config/langgraph-config.yaml --out openshell-policy.yaml

The output maps directly (OpenShell v1 policy schema, validated against v0.0.59):

  • filesystem.projects → filesystem_policy.read_only / read_write (a write:false project becomes a kernel-enforced read-only path — so a monitor like Roxy cannot write even if something tried), plus an OS baseline under landlock.compatibility: best_effort;
  • egress.allowed_hosts + model.api_base → network_policies endpoints, scoped to the agent's binaries (everything else denied by the per-sandbox proxy);
  • process.run_as_user: sandbox → the unprivileged image user (root is rejected by OpenShell).

OpenShell is pre-1.0 — re-verify against your installed release when upgrading (openshell policy prove can check properties of the output). It's a generated starting point, derived from real config, not a guess.

Run it ​

bash
# install OpenShell (see its docs), then wrap the agent's image/command:
openshell sandbox create --policy openshell-policy.yaml --from protoagent:local \
  --env PYTHONPATH=/opt/protoagent -- python -m server

Credentials are injected as env at runtime (never on disk); egress is deny-by-default through the proxy; the filesystem is locked to the policy paths.

Ready-made deployment ​

deploy/openshell/ has a managed example end-to-end:

  • Docker: compose.yml (the OpenShell gateway) + create-protoagent-sandbox.sh (generates the policy, creates the protoAgent sandbox under the gateway).
  • Kubernetes: k8s/values.yaml (Helm gateway, kubernetes driver) + k8s/protoagent-sandbox.yaml (Agent-Sandbox CRD + policy ConfigMap) — after installing the Agent Sandbox CRDs.

See deploy/openshell/README.md. The Docker path is validated end-to-end against OpenShell v0.0.59 (including a gotcha table from the validation run); the k8s CRD wiring is still a starting template (OpenShell is pre-1.0 — verify fields against your release).

  • Trusted model, no code execution: native egress allowlist is a sensible baseline; execute_code/write-fs_tools off.
  • Any execute_code / write-enabled / multi-project deployment (incl. Roxy): run under OpenShell with the generated policy. The native egress allowlist still applies inside as defense-in-depth.

See ADR 0008 for the full rationale and the operator-fork guide for Roxy (a read-only monitor is an ideal OpenShell tenant).

Part of the protoLabs autonomous development studio.