Skip to content

0042 — Fleet supervisor & unified console (background agents, in-place switching) ​

  • Status: Accepted (v0.31.0). Member launch now sets PROTOAGENT_HOME=<workspace> (not the retired PROTOAGENT_CONFIG_DIR) per ADR 0065; the fleet model is otherwise unchanged.
  • Date: 2026-06-09
  • Builds on: ADR 0004 (per-instance scoping), 0027 (plugins), 0040 (bundles), 0041 (workspaces & tiered stores).

Context ​

ADR 0041 made each agent a first-class, isolated workspace — its own config dir, its own instance.id-scoped data, its own port. But a workspace is still a separately launched server with its own console. The operator wants a fleet: many named agents on one machine, switchable from one console, each running in the background with its session intact — flip from Ava to Roxy and back, and Ava's chat (and her in-flight background work) is right where you left it.

Crucially, the hard half is already done. Each agent's chat history is in its own ~/.protoagent/<id>/checkpoints.db (ADR 0004/0041), so "Ava's session continues" is just reconnecting to Ava's thread — and it survives a restart (resume from checkpoint). What's missing is (a) something that keeps the processes alive so background activity (schedules, an in-flight agent loop) continues while you're away, (b) a console that switches which agent it's viewing in place, and (c) a way to spawn new agents from a starter type.

Decision ​

A fleet supervisor + a unified, proxying console, over persistent background agents, with archetype-driven creation. Recommended topology:

        ┌──────────── Hub (:7870) — the only UI ────────────┐
        │  Unified console + Supervisor + Reverse proxy      │
        │  [ Ava ▾ ]  switch = swap the proxy's active target│
        └───┬─────────────────┬─────────────────┬────────────┘
   supervises headless agent backends (--ui none), each isolated:
     ┌──────▼───┐      ┌───────▼───┐      ┌──────▼────┐
     │ ava :7901│      │ roxy :7902│      │trader:7903│   ← persistent bg processes
     │ scoped   │      │ scoped    │      │ scoped    │      (own data + live session)
     └──────────┘      └───────────┘      └───────────┘
  • Hub — a lightweight server on the main port that serves the one console, runs the supervisor, and reverse-proxies the active console's API / A2A / SSE to the selected agent backend. The operator only ever opens the hub.
  • Agents — ordinary workspace servers launched headless (--ui none, API+A2A only — lighter, no console each), supervised by the hub, each scoped per ADR 0041.
  • Switching is view-only — selecting an agent swaps the proxy target; the agent never stops, so its background work keeps running and its session is live.
  • Start/stop from the control plane — the hub can launch, stop, and report status for each agent; a stopped agent's history persists and resumes when restarted.
  • Archetypes — a bundle (ADR 0040) carries an archetype: block (label/icon/blurb); the new-agent picker offers each known bundle as a starter type, and "create" runs the ADR 0041 workspace new --bundle flow.

Design details ​

A. The supervisor (control plane) ​

Owns agent-process lifecycle, built on workspaces.manager.run_exec (ADR 0041): for each workspace it can start (spawn python -m server --ui none with the workspace's config-dir + instance + port), stop (SIGTERM → reap), and report status (running / stopped, pid, port, uptime). A small registry (in workspace.yaml + live process table) tracks it. Exposed as a CLI (python -m server fleet up/down/ls/status) and the API below.

B. Control-plane API (served by the hub) ​

  • GET /api/fleet — every workspace + live status (running/stopped, port, the active one).
  • POST /api/fleet/{name}/start · POST /api/fleet/{name}/stop.
  • POST /api/fleet/{name}/activate — set the console's proxy target.
  • POST /api/fleet — create from an archetype ({name, bundle}) → workspace new --bundle.
  • GET /api/archetypes — known bundles' archetype: metadata for the new-agent picker.

C. The reverse proxy ​

The hub forwards the active console's requests (/api/chat, /a2a, /api/events SSE, plugin views) to the active agent's loopback backend, streaming SSE through unbuffered. Auth: agents bind loopback and trust a shared hub token (the hub injects it); the browser only ever talks to the hub (same-origin), so no per-agent CORS/token sprawl.

Amendment (v0.35.0, #883/#900): the slug proxy also forwards WebSocket upgrades, not just HTTP/SSE. The original proxy stripped Upgrade/Connection, so a member's plugin view that opened a live WS (e.g. agent_browser's viewport) loaded over HTTP but its socket showed "Disconnected" behind the hub. proxy.forward_ws resolves the slug → member, opens a client WS (carrying the bearer + subprotocols), and pumps frames both ways — so WS plugin views traverse the hub like HTTP does.

Amendment (#1607): WS proxying is now refused for REMOTE members (close 1008), narrowing the v0.35.0 amendment above. The hub's default-deny auth is an HTTP middleware (Starlette BaseHTTPMiddleware skips non-HTTP scopes), so the @app.websocket proxy route runs with no hub auth. For a host/local peer that's benign — the hub attaches no stored credential, and a browser WS can't carry one — but for a remote member the hub would attach that member's stored bearer, lending an unauthenticated caller a ride into its authenticated sockets (e.g. a terminal plugin's PTY). Until the WS caller can be authenticated at the hub, a remote member's live sockets don't traverse the hub — use delegate_to / a direct connection for remote live views. Host + local-peer WS (the original #883 use case) are unchanged. Same change also makes the proxied SSE token (/api/events) hub-signed (the console fetches /api/sse-token with host:true), because the hub's auth middleware validates the proxied stream before forwarding — a member-signed token 401'd at the hub for every non-host member on a bearer-gated hub.

Amendment — ADR 0113 D6 (#3648): the #1607 refusal above is lifted; remote members' WebSockets traverse the hub again, under one rule: the hub never attaches the remote's stored bearer to an upgrade on its own. A presented ?token= (or a server-to-server caller's Authorization: Bearer) is authenticated at the hub on the tier ladder without the open-mode shortcut (credential_tier). If it is operator, it is swapped for the remote's stored token in the slot it was presented in; anything else closes 1008, so an open hub never swaps one in. A browser handshake must also come from the hub's own origin, the desktop webview (tauri://localhost, http://tauri.localhost) or the A2A_ALLOWED_ORIGINS allowlist, since the HTTP Origin check never sees a WS scope. A socket with no credential (ticket-based plugins such as agent_browser's ?ticket= and the terminal plugin's in-band ticket, both minted over the authenticated HTTP proxy) is passed through with no Authorization, and the remote checks the ticket itself. A remote registered without a stored token is still refused. The hub neither attaches nor forwards a credential of its own, and the loopback-only fleet token (ADR 0089) never leaves the machine. In-band frames are the plugin's responsibility: the hub pumps them opaquely. The terminal plugin ≤0.9.1 put the hub's bearer in its first frame when a ticket mint failed, so remote terminal views need terminal-plugin ≥0.9.2. An unauthenticated caller still gets nothing through the hub it couldn't get by dialling the remote directly, which is the property #1607 protected.

Amendment — A2A URL-based multi-tenancy: the slug proxy makes each member an independently-addressable A2A tenant at /agents/<slug>/a2a, which is exactly the A2A URL-based routing pattern — one hub server, many agents, each at its own URL prefix. For a peer to discover and route to a member correctly, the member's agent card must advertise its own tenant URL, not the hub root. Previously a member inherited the hub's A2A_PUBLIC_URL, so its card advertised the hub root ({hub}/a2a): every member's card collided on one URL, and a card-based dispatch landed on the hub agent instead of the member (the "wasted hub route" — jon's LLM brokering a message meant for matt). Fix: the supervisor spawns each member with A2A_PUBLIC_URL = {hub_public_url}/agents/<slug>, so _a2a_card_url() yields {hub}/agents/<slug>/a2a — the member's own proxied route. A member is now a first-class, discoverable A2A tenant reachable directly (transport-only proxy; the hub agent's loop is never invoked), while still living in the hub process — no separate container or port. The hub token remains the single auth boundary in front of all tenants. See A2A endpoints § Fleet multi-tenancy and Agent card § supportedInterfaces.

D. Session continuity (≈ free) ​

Each agent's checkpoints/goals/memory are already instance.id-scoped (ADR 0004/0041) — so switching back loads that agent's thread, and it survives stop→restart (resume from checkpoint). The supervisor keeping the process warm is what additionally keeps background activity (scheduler ticks, an in-flight agent loop) running while you're away.

E. The switcher UI ​

The console topbar agent name becomes a dropdown: the fleet list with status dots + ports, a one-click switch (swap active), start/stop affordances, and "+ New agent" → the archetype picker (cards from GET /api/archetypes) → name → create + start + activate.

F. Archetypes (additive — bundle metadata) ​

A bundle manifest gains an optional archetype: block (label, icon, blurb, optional persona/accent). It rides the settled bundle shape — load_bundle already returns it, no schema change. pm-stack is archetype #1 ("Project Manager"); every future bundle that adds the block becomes a starter type for free.

Amendment (2026-08) — naming. pm-stack (then labelled "Project Manager") bundled project_board and portfolio; it is now portfolio-manager-archetype, and the single-repo Project Manager archetype ships from project-manager-archetype (ADR 0100 amendment, #2895 — "stack" retired, archetype repos are <name>-archetype). The archetype system itself is ratified in ADR 0100.

G. Keep-alive policy ​

Resource-bound: default keep-N-warm (recently-active agents stay running; the rest are stopped and resume from checkpoint on switch), with an opt-in run-all for small fleets. Configurable in the hub.

H. Native desktop — the Tauri shell is the hub ​

On the desktop app (apps/desktop, Tauri), the shell plays the hub. It already spawns + supervises the server as a sidecar (the parent-death watchdog / autostart machinery), so it takes the supervisor + proxy roles directly and the fleet becomes GUI-first — the CLI's workspace/fleet/plugin verbs become panels over the same control-plane API, so native and headless never diverge.

Roles on native

  • Process management — the shell spawns each agent as a sidecar (python -m server --ui none per workspace, via workspaces.run_exec / the fleet API), tracks pid+port, reaps them on quit (reusing its server-sidecar machinery), and applies the keep-N-warm policy (a laptop won't run a big fleet hot).
  • Proxy — the in-app webview renders one console; the shell (or a thin in-process hub it sidecars) routes the active console's chat / A2A / SSE to the selected agent. Loopback + a shell-held token — no per-agent surface is exposed.
  • Contract — the panels drive the same /api/fleet + /api/archetypes the CLI uses.

Panels (React, apps/web)

  • Onboarding — first run becomes "create your first agent": an archetype picker (cards from GET /api/archetypes — Basic + every installed bundle's archetype:) → name → POST /api/fleet (create + start). Replaces the single-agent setup wizard.
  • Switcher — the topbar agent name → a dropdown (fleet list + status dots + "+ New agent").
  • Fleet manager (Settings → Agents) — GET /api/fleet rows with start/stop/remove + "+ New."
  • Per-agent config — the existing Settings drawer (model / plugins / secrets), now scoped to the active agent's workspace (config is per-PROTOAGENT_CONFIG_DIR).

Build seam (who builds what)

  • Server: the fleet supervisor (slice 1, shipped) + the control-plane API & reverse proxy (slice 2) + run_exec (the launch primitive).
  • Tauri shell: the agent-sidecar spawn/reap + proxy plumbing (recommended — it's already supervising a sidecar), or delegate both to a thin Python hub it sidecars. Either works.
  • React: the onboarding/archetype picker, switcher, and fleet-manager panels — all over the control-plane API.

Net: the desktop just renders the control plane the CLI already drives — one model, two front-ends.

I. Remote fleet members — agents on other machines ​

The slices above manage local agents: the supervisor run_execs a subprocess on this host, tracks it by pid, and the proxy forwards to 127.0.0.1:<port>. But the model already generalizes to remote protoAgents on other machines — the only thing that's intrinsically local is process spawning. Everything else is address-based:

  • Each agent is already an independent A2A endpoint (the a2a field is a URL). Nothing makes it loopback-only — point it at https://host:port and agent↔agent delegate_to works cross-machine today.
  • The reverse proxy is address-based — /active/* forwards to "the active agent's address." Localhost now; a remote host:port needs no console change.
  • The host already self-registers as a first-class agent (host: true) — a remote member is the same idea with remote: true and a non-local address.

A remote member is registered, not spawned. Add it by URL + token (not create+launch):

  • Agent kind — supervisor.status() gains a remote: true member kind: {name, kind: "remote", url, running, a2a: <url>/a2a}. No local pid.
  • Status — health is a poll of the remote's GET /api/runtime/status (+ the hub token), not the local _alive(pid) probe. Unreachable ⇒ running: false (same UX as a crash).
  • Lifecycle — you can't kill a remote process. Two tiers: (a) register-only — the remote runs itself; the hub bookmarks + proxies + monitors it, no start/stop; (b) remote hub — if the other machine runs its own fleet hub, start/stop proxies to its control-plane API (turtles all the way down).
  • Transport/auth — real network now: TLS + a per-remote token (reuse the existing A2A / console bearer-gate). The proxy attaches the token on forward.
  • Console management (#1608/#1609). A remote is added from Settings ▸ Agents by URL + optional token — the only way to register a token-gated peer (discovery can't carry a credential) — and edited in place via PATCH /api/fleet/remotes/{ident} (supervisor.update_remote: url/token/name, id + slug stable so open windows survive; token:"" clears, re-runs the same SSRF-egress/collision checks as add). An offline remote reads "unreachable" (not "stopped"); a member that 502s (box offline) or 401s (bad token) gets a targeted boot-gate recovery — Return to host / update its token — and a proxied member's 401 is kept off the hub's global token prompt (it's the member's credential, not the hub's).

Relationship to the delegate registry (ADR 0025). protoAgent already connects to remote agents — delegate_to over a2a/openai/acp (Settings → Integrations). That's the delegation axis (hand a remote a task). The fleet is the management axis (switch your console into an agent, see its status, control its lifecycle). Remote fleet members are the convergence: a registered remote could surface in both — delegate to it as a peer and focus the console on it. A future cleanup is one registry feeding both views.

Frontend is mostly free — the panels render /api/fleet; a remote agent is just an agent with a remote address + remote: true. Show a "remote" tag (like the host tag), gate start/stop on the lifecycle tier, and the switcher's address swap already routes through the proxy. The only real new UI is the "add remote agent" form (URL + token) alongside "+ New."

Status: the register-only tier shipped (#839, slice 6 below), and remote credentials are now paired rather than pasted. ADR 0113 covers that: the remote shows a one-time code, and the hub claims a per-hub device token the remote can revoke. The same ADR routes delegates to a remote through the hub (one registry, one token), re-enables remote WebSockets without the hub lending a credential, and advertises on mDNS only from a reachable bind. Remote-hub start/stop (tier b) is still future work.

Amendment (2026-09) — the fleet deck (#3466). The fleet gained a terminal: bare protoagent fleet / protoagent top opens a Textual deck over the running hub — the roster with the console's presence words, member detail, a fleet-wide work feed, member management (the REST-only create / remove / rename / remotes / order were promoted into ops/fleet.py for it, ADR 0075 D2), every hub on the box (--all: heartbeats under every box root, instance roots with a fleet, listeners by port, peers; attach, or bring a stopped hub up), and conversations with members — steering, HITL answers, delegation cancel — over the same A2A / /api surface as the console. It supersedes ADR 0075's "chat is proto's job" for the fleet (amendment there) and fixes the rule for a shell beside a running hub: live hub first, badged disk fallback (0075 open question 2). Ports are box-global and every box root's heartbeats are inputs — what a shell "sees" is no longer what it inherited from PROTOAGENT_HOME. Guide: the protoagent command → the fleet deck.

Options considered ​

  • Separate consoles per agent (ADR 0041 slice-4 simple) — navigate to each agent's own console. Ships fast, but no single home, no in-place feel, no shared background view. Rejected as the end state.
  • In-place hot-swap via a hub proxy (this). Slickest UX (one console, swaps in place), supports background agents + start/stop + a single auth boundary. Chosen.
  • Eager run-all vs keep-N-warm. Run-all is simplest but doesn't scale; keep-N-warm bounds resources while history always persists. Chosen (configurable).
  • Archetype source: bundle metadata vs a curated registry. Bundle metadata is self-describing and needs no second list to maintain. Chosen.

Consequences ​

  • A real fleet OS — one console over many persistent, isolated, archetype-spawned agents, with lifecycle control. The capstone of 0040/0041.
  • Resource cost — N warm agents ≈ N processes + model connections; the keep-N-warm policy bounds it, and history persists regardless of warmth.
  • The hub is now critical — it owns the console, the proxy, and the supervisor; it must fail safe (an agent crash shows as "stopped", never takes the hub down).
  • Security — agents bind loopback + trust a hub token; the browser only touches the hub. No new external surface.
  • Backward compatible — a single python -m server (no hub) is unchanged; the fleet is opt-in via the hub.

Implementation slices ​

  1. Supervisor — start/stop/status over run_exec, a process table, fleet CLI.
  2. Control-plane API + reverse proxy — /api/fleet + active-target proxying (chat/A2A/SSE).
  3. Switcher UI — the agent-name dropdown (list, status, switch, start/stop) over the API.
  4. Archetypes — GET /api/archetypes + the "+ New agent" picker → create-from-archetype.
  5. Keep-alive policy — keep-N-warm + resume, lifecycle polish.
  6. Remote fleet members (§I, shipped — #839, register-only tier) — a remote: true member registered by URL (+ optional bearer, stored hub-side and attached by the proxy): remotes.json beside fleet.json, status via a TTL-cached agent-card probe (refreshed off the event loop), _target_for_slug resolves remote slugs to URLs, Discover's "Add to this fleet" + POST/DELETE /api/fleet/remotes. Composes with the delegate registry (ADR 0025) — the same remote can be a member and a delegate_to target. Remote-hub start/stop (tier b) remains future work.

Slices 1–3 deliver the core (switch between running agents in one console); 4 adds the starter-type creation flow; 5 makes it scale; 6 extends the fleet across machines.

Part of the protoLabs autonomous development studio.