ADR 0006 — Observability & the Self-Improving Flywheel
- Status: Accepted (2026-06-01) — all 4 slices shipped (flywheel: measure → persist → surface → advise)
- Date: 2026-06-01
- Deciders: Josh Mabry; protoAgent maintainers
- Tags: observability, cost, tracing, metrics, a2a, optimization, flywheel
- Supersedes / Superseded by: —
Accepted. We want to measure what the agent actually costs and how long it takes — per LLM call, per tool, per turn — then turn that signal into a loop: measure → analyze → optimize → measure. The bones exist (Langfuse, a Prometheus module, an audit log, token capture, the outbound
cost-v1DataPart), but the LLM half is half-wired (therecord_llm_callmetric is defined and never called), prompt-cache tokens and USD cost aren't captured, there's no historical store, and none of it surfaces in the console. This ADR records the current-state map, the flywheel target, the alignment with Workstacean'scost-v1A2A extension, and a sliced roadmap. Slice 1 (the measurement foundation) ships with this ADR.
1. Context & Problem Statement
To refine cost and latency we first have to see them, reliably, at the right granularity. An audit of what protoAgent captures today:
What exists (good bones)
| Capability | Where | State |
|---|---|---|
| Langfuse tracing (session + tool spans, trace-id on audit lines) | tracing.py | Working; graceful no-op when unconfigured. Env-only until #3017 — see the amendment below |
Prometheus /metrics (per-fork namespaced) | metrics.py, server.py | Endpoint up; LLM metrics defined but unused |
| Per-tool latency + success → JSONL + Langfuse + Prometheus | graph/middleware/audit.py | Working |
Per-LLM-call token capture (on_chat_model_end, stream_usage=True) | server.py _run_turn_stream, graph/llm.py | Working (input/output only) |
Outbound cost-v1 DataPart on the terminal A2A artifact | a2a_handler.py (COST_MIME, _cost_payload) | Working (usage + durationMs; costUsd omitted) |
The gaps (what blocks the flywheel)
metrics.record_llm_callis defined but never called (metrics.py). The*_llm_calls_total,*_llm_latency_seconds,*_llm_tokens_totalseries are dead — the LLM half of any dashboard is empty. Onlyrecord_tool_callfires (fromAuditMiddleware).- No USD cost.
costUsdis deliberately omitted; there's no pricing table. Every consumer recomputes from tokens + its own rates. - No prompt-cache tokens.
usage_metadata.input_token_details(cache_read / cache_creation) isn't read. Prompt caching is on — so we pay for the cache and can't see the savings or the hit ratio. - No per-LLM-call latency. Only whole-task wall-clock duration is computed (from
created_at→updated_at). - No historical store. Prometheus is live-scrape-only (and missing the LLM data);
AuditMiddleware's session stats die on restart. Nothing inside protoAgent can answer "what was expensive/slow last week." - The operator console shows none of it — no tokens/cost/latency anywhere.
- No feedback loop. Nothing turns telemetry into action.
2. The flywheel (target)
┌────────── measure ──────────┐
│ per-call tokens (incl cache)│
│ per-call + per-tool latency │
│ USD cost, model, turn rollup│
└──────────────┬───────────────┘
▼
persist + aggregate (local store: per-turn / per-call rollups)
▼
surface (operator console: $/turn, p50/p95, cache-hit %)
▼
optimize + feed back (flag expensive/slow turns; prove the levers —
prompt cache, tool deferral, compaction, model routing; feed
signal into learned-skills / memory / operator recommendations)
▲
└──────────── measure again ───────────┘The levers already exist (prompt cache, tools.deferred, compaction, routing/fallback). What's missing is the measurement that proves they work and the loop that points them where they're needed.
3. Decision
Instrument at the seams the audit found; don't rebuild the backbone. Align outbound telemetry with the fleet so the data is useful beyond protoAgent.
- Capture is per-LLM-call, in
_run_turn_stream'sastream_eventsloop — the only place with per-call visibility (a task may hit the model N times). Add cache tokens + per-call latency + the model name to what's already captured there. - Cost is computed in-process from a pricing table (
pricing.py), cache- aware, and emitted ascostUsd— making protoAgent the source of truth for its own cost rather than every consumer re-deriving it. - Wire the dead metrics: call
record_llm_callper call; extend it with cache tokens + cost; add the missing series. - Align with Workstacean's
cost-v1A2A extension (see §5) — emit the Anthropic-shaped cache fields it already expects, populatecostUsd, and declare the extension URI in the agent card so the integration is explicit, not incidental. - Persist + surface + close the loop in later slices (§4).
4. Ranked Plan (slices)
✅ Measure (foundation) — ships with this ADR. Capture cache tokens + per-call latency + model in
_run_turn_stream; addpricing.py→costUsd(base input/output rates, fleet-consistent; cache-discounted cost deferred until gateway token semantics are validated); wirerecord_llm_call(extended with cache + cost series). Per-call Langfuse generation spans already come from the LiteLLM gateway callback — we deliberately don't add a manual shim that would bypasstrace_session's nesting (guarded bytest_no_legacy_shims_exist). Emit cache fields +costUsdoncost-v1and declare the extension URI in the card. (fixes 1–4)✅ Persist & aggregate — shipped.
telemetry_store.py(TelemetryStore) writes one row per turn leg (tokens incl. cache, cost, duration, LLM/tool call counts, model, outcome) — a HITL park/resume is two legs sharing one task id, and each owns its spend (#3001). Rows come from both turn drivers through the shared writerserver/turn_telemetry.py::record_turn: the A2A executor's terminal hook, and the non-streaming_chat_langgraphbehind/v1/chat/completions,/api/chat, and the pluginHOST.invoke()seam. Instrumenting only the first is how the second's spend stayed invisible for so long (#3000) — a new turn surface must route through that writer or it is not measured. A third producer joins them:AcpClient.promptwrites one row per CLI coding-agent run under acoder:<delegate>key and anacp:<delegate>model label, since those dispatches happen outside any turn and so reach no terminal hook (#3015). They carry zero tokens and zero cost — the coder bills its own subscription — so whole-instance cost is unaffected while turn counts, success rate and latency percentiles now span both populations; the per-model split is where they stay separate. Instance-scoped (ADR 0004). Read via/api/telemetry/summary(totals, success rate, cache-hit ratio, p50/p95 latency, per-model split) +/api/telemetry/recent. Survives restart; no TTL (history is the substrate),prune()available. (fixes 5)✅ Surface — shipped. A System ▸ Telemetry dashboard (
apps/web/src/telemetry/TelemetrySurface.tsx): summary cards (cost, turns, success rate, cache-hit %, p50/p95 latency, tokens, tool calls) + a by-model table + a recent-turns table, reading/api/telemetry/*. Functional-first (theme-consistent, no charts yet — a follow-up). (fixes 6)✅ Flywheel / feedback — shipped (advise-only).
/api/telemetry/insights- a Telemetry Insights panel: flags turns whose cost/latency ≥ 5× the rolling median for their own model (#3015 — the store holds several populations that aren't comparable: a CLI coding-agent run recorded under
acp:<delegate>is minutes where a gateway chat turn is seconds, and against one shared median every coder run cleared 5× by construction and filled the list), and proves the levers we can measure from the per-turn store — prompt-cache hit % + estimated USD saved (pricing.cache_read_savings_usd), plus model-mix (routing visibility). Read-only: it surfaces signal, the operator decides — no autonomous config changes. Levers needing extra per-turn signals (tool-deferral schema-token savings, compaction, detailed routing) are explicitly listed as not yet measured rather than faked — a follow-up that adds those signals can light them up. (fixes 7)
Slice 4b (per-turn signals) then made two of those levers real: the telemetry row records the actual model(s) used per turn (
model= primary,models= distinct set), so routing — incl. aux/fallback models — is proven per turn rather than stamped from the configured lead; andToolDeferralMiddlewareemits*_llm_tools_deferred_totalto Prometheus, proving the deferral lever live. Compaction is likewise proven viaCountingSummarizationMiddleware(subclasses langchain'sSummarizationMiddleware, counts each real compaction →*_compactions_total). With routing, deferral, and compaction all measured,insights.unproven_leversis now empty — every optimization lever the agent has is observable.- a Telemetry Insights panel: flags turns whose cost/latency ≥ 5× the rolling median for their own model (#3015 — the store holds several populations that aren't comparable: a CLI coding-agent run recorded under
Why advise-only (not auto-optimize). Letting telemetry change config automatically (auto-enable deferral, auto-downgrade model) is higher leverage but needs guardrails and risks surprising regressions. We start by making the signal trustworthy and visible; auto-optimization is a deliberate future step, not a default.
Priorities (per the kickoff): cost visibility ($) and latency breakdown lead; Slice 1 makes both real. The console surface (Slice 3) follows so the numbers are visible, then the loop (Slice 4).
5. Alignment: Workstacean's cost-v1 A2A extension
protoAgent already emits a cost-v1 DataPart; Workstacean already consumes one. This ADR closes the remaining gaps so they meet the same contract.
- Extension URI:
https://proto-labs.ai/a2a/ext/cost-v1(protoWorkstacean/src/executor/extensions/cost.tsCOST_URI;docs/extensions/cost-v1.md). protoAgent will declare it in the agent card'scapabilities.extensions, which is what gates Workstacean's cost interceptor (ExtensionRegistry.interceptorsFor(card)). - DataPart MIME:
application/vnd.protolabs.cost-v1+json(already matches). usageshape (Workstaceanlib/types/cost-v1.tsCostArtifactUsage):{ input_tokens, output_tokens, cache_creation_input_tokens?, cache_read_input_tokens? }— Anthropic-shaped. We don't emit the cache fields yet; Slice 1 adds them.costUsd— Workstacean's interceptor uses ourcostUsdwhen present and only falls back totokens × MODEL_RATESwhen it's absent. Emitting it makes us authoritative and sidesteps cache-discount mismatch (theirMODEL_RATESinlib/types/budget.tscarries input/output only, no cache tiers).- Rates:
pricing.pymirrors the structure + overlapping values of Workstacean'sMODEL_RATESand adds Anthropic cache multipliers (cache-read ≈ 0.1× input, cache-write ≈ 1.25× input). Both sides agree on the base rates; protoAgent's emittedcostUsdis the cache-accurate number. - Consumer fields (
docs/extensions/cost-v1.md): the interceptor records aCostSampleand publishesautonomous.cost.{actor}.{skill}for the planner's cost/confidence ranking and the fleet cost-per-outcome dashboard. So better protoAgent telemetry directly improves fleet-level planning — the flywheel isn't only local.
Note: Workstacean's cost store is observational, in-memory (last 200 samples/key), explicitly not billing. protoAgent's Slice 2 local store is the durable, queryable half on our side; the two are complementary.
Amendment — prompt-cache pricing resolved (#3003, 2026-08-23)
Decision 2 above said cost is computed in-process and cache-aware. The implementation shipped without the cache half and left a note in pricing.py deferring it "until the gateway's cache-token semantics are validated end-to-end — different gateways disagree on whether input_tokens already includes cached reads." The deferral outlived its reason and the code drifted from this ADR: every cached token was billed at the full input rate, roughly tenfold its real cost, on stores running a 73% cache-hit ratio.
The gateways do disagree at the raw provider layer, and they are reconciled before the shape we consume. LangChain's UsageMetadata is documented as "a standard representation of token usage that is consistent across models" in which input_tokens is "the sum of all input token types" and input_token_details.cache_read / .cache_creation are subsets of it. Every producer in the tree reads that shape, so no provider branch is needed.
Resolved as:
pricing.cost_usdtakes LangChain-shaped (cache-inclusive) usage and does the disjoint split internally — full input rate for uncached prompt tokens, ×0.1 for cache reads, ×1.25 for cache writes, output at the output rate. One place knows the arithmetic.- The stored
input_tokenscolumn is cache-EXCLUSIVE, normalised once in the shared writerserver/turn_telemetry.py::record_turn.input_tokens + cache_read_input_tokens + cache_creation_input_tokensis the turn's true prompt size and the three are disjoint, so a consumer summing them can't double-count.total_tokensre-adds the cache components, so it still means every token the turn moved. - The
cost-v1A2A extension is unchanged. ItsinputTokensis a fleet wire contract (protolabs-a2a 0.3.0) that consumers read with OpenAI-ish cache-inclusive semantics; normalising at the store writer rather than at the producers keeps that contract intact. cache_hit_ratioiscache_read / (input + cache_read + cache_creation). Rows written before this change carry the old cache-inclusive column, so their denominator double-counts the cached reads and the ratio reads slightly low. Not backfilled: quietly rewriting recorded history is worse than a documented seam, and the error is conservative.
Amendment — a peer's cost-v1 bills to the calling turn (#3016, 2026-08-23)
We emitted cost-v1 on every terminal artifact and then ignored it in our own client: plugins/delegates/adapters.py::A2aAdapter.dispatch read _extract_text(result) and returned a bare str, so a delegation to a protoAgent peer spent real money that appeared in no row on the calling side. A hub fanning work out to members reported only its own thinking, and ADR 0055 multi-team orchestration was uncountable by construction. Of every delegation kind this is the one where the number is already computed, already correct, and already on the wire — an ACP coder's spend genuinely isn't observable from here (#3015), a peer's is.
Where a peer's spend lands: the calling turn, tagged — not a row of its own. This follows #2872, which settled the same question for subagents: a delegated model call bills to the parent turn while carrying a tag, and the tag excludes it from the parent's context-window fill. One row per turn keeps "what did this turn cost" answerable without a join, and the tag keeps the delegated share separable. A peer row is {…tokens…, cost_usd, model: "peer:<delegate>", peer: "<delegate>"}.
The mechanics, and why each is what it is:
- Read at the terminal artifact.
tools/a2a_parse.py::_extract_costtakes the cost-v1 payload off the terminal artifact's metadata map keyed by the extension URI (protolabs-a2a 0.3.0 — the shapea2a_impl/executor.py::_terminal_partsemits), falling back to the terminal status message's metadata. It sits beside_extract_text— the wire reader belongs with the other wire readers, in a layer core surfaces can import without reaching intoplugins/. - Attributed through the existing
usagecustom-event lane (#2872,server/chat.py'son_custom_eventbranch → the executor's accumulator).Adapter.dispatch's-> strreturn type is a stable interface across three adapters and does not change; nothing is stashed on the adapter either, sinceADAPTERSholds one instance per type process-wide and last-call state on it would race across a concurrent fan-out. - The peer's
costUsdis used verbatim, never re-derived. The peer knows which models it routed to; we don't. A peer reporting tokens but nocostUsdcontributes tokens and no cost — a visible undercount beats an invented number. modelcarries apeer:marker, not a model name. cost-v1 has no model field, andmodelsexists to prove which model actually ran (Slice 4b). The prefix keeps "what did this turn spend on peers" answerable from the stored row — it is the only durable trace a delegation leaves, since thepeertag itself is a stream-only routing hint. Because a marker is not a model, every field that names the model that ran picks fromtools.a2a_parse.drop_peer_markers(models), never from the raw list: the row'smodelcolumn and theturn.usagebus event (server.turn_telemetry), and the fleet trace export'smeta.model(observability.trace_export), which the lab consumes as the row's teacher model. One deliberate exception:acp:<delegate>(#3015, andserver.chat._acp_drive_turnbefore it) stays in themodelcolumn, because unlike apeer:marker it is the model that ran, in the only sense we can observe from this side. One helper because a fourth reader is a matter of time, and a marker leaking into any of them is the same defect. Normally the lead's own call is first anyway, but a provider that reports no usage leaves the marker leading the list — and a per-model breakdown (or a training row) that sayspeer:orbisis simply wrong. The rawmodelslist keeps its markers everywhere.- Silent degradation is the contract. A peer that emits no cost-v1 — any non-protoAgent A2A agent — behaves exactly as before: no row, no cost, same text. So is a payload we can't read: reading and dispatching are both inside the best-effort guard, because a raise would propagate through
registry.dispatch, which records the delegate as failing before re-raising — discarding an answer already in hand and red-flagging a healthy peer over a bad number. Telemetry must never break a delegation (#2872's rule, and it applies to the parse half too).
Known limits, deliberately left rather than papered over:
- A peer delegation counts as one
llm_callshowever many calls the peer really made — cost-v1 reports no call count. - A background delegation (ADR 0050) bills nothing — and that takes a deliberate flag, not the absence of one.
asyncio.create_taskCOPIES the spawning context, so the detached job inherits thedelegate_totool body's LangChain run context and can reach the spawning turn's stream: left alone it would bill that turn whenever the peer answered before the turn closed and drop the row when it answered after — the same delegation writing two different rows depending on the peer's latency. A contextvar set at the top of the job's own coroutine (mark_delegation_detached) makes the exclusion deterministic. Detached spend belongs to the later turn the reply is drained into, not to the turn that spawned it; billing it there properly means carrying the peer's row through the background job's completion into that turn, which is a separate change. - Only spend that reached a terminal artifact is billable here. A peer leg that ends without one — a HITL park (#2943's
input_requiredleg, which emits a status message and no artifact) or a failed turn — put no cost-v1 on the wire, so the caller cannot bill it. It is not lost: the peer records that leg in its own store as a row of its own. The consequence is that a HITL delegation chain undercounts on the calling side by the peer's pre-park spend. Closing it means emitting cost-v1 on the park and failure paths too — an emitting-side change, and a separate one. - Nothing is deduplicated across dispatches. Correct today because each billed return path corresponds to one leg the caller just caused (a resumed task's terminal artifact is REPLACED in place under
{task_id}-answer, so it carries the resumed leg only), and the "already finished, nothing to resume" branch returns without billing. A peer that accumulated across legs would double-count.
6. Consequences
Positive
- Real per-call/per-turn cost + latency, with cache visibility — the data the flywheel runs on, and proof that the existing optimization levers work.
- Fleet alignment: one
cost-v1contract; protoAgent's numbers feed Workstacean's planner + dashboards directly. - Reuses the existing backbone (Langfuse/Prometheus/audit) — incremental, not a rewrite.
Negative / costs
- A pricing table is a maintenance surface — model rates drift; it must track the gateway (documented; falls back to a
defaultrate, never crashes). - Cache-token fidelity depends on the gateway surfacing
prompt_tokens_details/input_token_details(OpenAI-compat exposescached_tokens; Anthropic cache-creation may not round-trip through every gateway). Captured best-effort, defaulting to 0. - Per-call instrumentation adds a little work to the hot streaming loop — kept cheap (counters + a contextvar), and all of it no-ops when Langfuse/Prometheus aren't configured.
7. Alternatives Considered
- Lean entirely on Langfuse/Prometheus, no local store. Good for live ops, but Langfuse is opt-in/external and Prometheus is scrape-windowed — neither gives protoAgent a queryable history for the flywheel. Local store (Slice 2) complements them, doesn't replace them.
- Let consumers compute cost (status quo). Keeps protoAgent rate-free but every consumer re-derives cost, no one sees cache savings, and the console can't show
$. Rejected — emittingcostUsdonce, cache-aware, is the right source of truth. - A new bespoke telemetry DataPart. Rejected —
cost-v1already exists and is consumed fleet-wide; we conform to it rather than fork it.
8. Related
- ADR 0004 — Multi-Instance Data Scoping — the Slice 2 store scopes per instance via the same helper.
- ADR 0005 — Tool Pollution —
tools.deferredis a lever Slice 4 will prove (schema-token savings). - Cost & trace propagation, Wire Langfuse + Prometheus.
- Code:
server.py(_run_turn_stream),a2a_handler.py(cost-v1),graph/middleware/audit.py,tracing.py,metrics.py; newpricing.py. - Fleet:
protoWorkstacean/src/executor/extensions/cost.ts,protoWorkstacean/lib/types/cost-v1.ts,protoWorkstacean/docs/extensions/cost-v1.md.
9. Amendments
Slice-2 and Slice-4 decisions revised after the fact. Each is dated and cites the issue that forced it; the sections above are left as originally written so the change of mind stays legible.
Amendment — Langfuse credentials are configurable (#3017, 2026-08-23)
The state table above says tracing is "working; graceful no-op when unconfigured". Both halves were true and together they hid a hole: tracing.init read LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY from the environment and nowhere else, and nothing in the shape the fleet actually deploys — a member launched by the desktop app as protoagent-server --port … --ui none — puts those variables in its environment. So on every fleet member the deep-trace half of this ADR degraded gracefully to nothing, with no way to turn it on short of editing the app bundle's launch environment. The live PM measured 0 trace_ids across 336 turns and 5,000 model calls over a month.
The trace tree is not decoration on the SQL rollup: trace_id, the a2a.trace caller propagation (so a delegation nests under its caller) and trace_session all exist to feed it, and the console's _resolve_trace_url_template deep-link path can never resolve without it. With that half dark, the SQL rollup was the only observability an operator had — the condition under which the defects in #3000–#3006 went unnoticed as long as they did.
tracing.{enabled,host,public_key,secret_key} now sit in LangGraphConfig, editable at Settings ▸ Tracing, with the two keys declared in config_io.SECRET_PATHS so they live in the untracked secrets.yaml (the Langfuse "public" key authenticates the ingest client server-side — it is a credential here, not a browser token). The environment still wins: a complete LANGFUSE_{PUBLIC,SECRET}_KEY pair is used as-is, including when tracing.enabled is false, because a container deploy has no config file to flip. tracing.enabled is the fallback toggle, not a kill switch.
Settings ▸ Tracing, not Box ▸ Telemetry beside telemetry.enabled, and the difference is the whole acceptance rather than a taxonomy preference. telemetry.* is host-scoped box config, so the console files it under a hostOnly section that renders on the host console alone; the four tracing fields are agent-scoped credentials, and the deployment shape this amendment exists for — a member launched --ui none, seen only through the hub's slug-scoped window — is exactly where a hostOnly section is dropped. Filed there, the setting would have been as unreachable from the console as it already was from the environment. The Telemetry surface instead carries a QuickSetting onto the same four fields (ADR 0048's chip-is-a- shortcut, same fields, same save path), so the two halves of this ADR still meet in one place on the host console.
And the surface now states its own state: /api/telemetry/recent reports tracing_enabled, and the console's Trace column reads off rather than — when tracing is disabled. A column of dashes reads as "these turns weren't traced" — the failure mode that let a month of dark tracing pass for normal.
Amendment — env keys never follow a config host (#3039, 2026-08-23)
The amendment above made tracing.host config, and left the host resolving as one env_host or cfg_host or _DEFAULT_HOST chain regardless of which layer answered for the keys. docker-compose.yml passes LANGFUSE_HOST=${LANGFUSE_HOST:-} — set and empty for an operator who exports only the key pair — so cfg_host won that chain, and the deployment's credentials left the process as a Basic auth header aimed wherever tracing.host said, carrying prompts, tool IO and span bodies with them. tracing.host is not a secret field, so it carries none of the handling public_key/secret_key get, and config reaches an instance through more paths than deployment env does: snapshot import, a fork's committed YAML, any write to langgraph-config.yaml. Before #3017 the host was env-only and config could not redirect deployment credentials; that property is restored.
The block is directional, and the direction is the whole decision. Env keys resolve to LANGFUSE_HOST/LANGFUSE_URL or _DEFAULT_HOST and stop. Config keys still fall back to env_host (cfg_host or env_host or _DEFAULT_HOST). Env is the more trusted layer — only whoever starts the process sets it — so an env host aiming config-owned keys was never the hole, and refusing that fallback would break a shape this repo actively tells operators to run: compose persists /sandbox/config while config/langgraph-config.example.yaml says the two keys are secrets that belong in Settings ▸ Tracing, i.e. host-in-env + keys-in-Settings. The fleet has the same shape — graph/fleet/supervisor.py spawns members with full_env = {**os.environ, **env}, so every member inherits the hub's LANGFUSE_HOST while its keys come from its own per-agent Settings. Mirroring the block would have sent those members to host.docker.internal:3001, an address that resolves only inside compose, and taken their tracing dark to fix nothing.
A host that loses is named, not dropped. Both discard paths worked before the upgrade, so without a line the operator's only other signal is a Trace column that quietly stops filling — precisely the failure this ADR's #3017 amendment exists to remove, one layer down. resolve_credentials prints the ignored value, why it lost, and where the traces went instead, beside the initialized from {source} -> {host} line that already says which layer answered.
Snapshot import lists tracing.host in CAPABILITY_KEYS, so an import plan names it alongside filesystem.allow_run: "the config you accepted chooses where your prompts and tool IO get shipped" is consent-shaped, not config trivia (ADR 0071 D1 — show it, don't neuter it).