Starter tools
The tools tools/lg_tools.py::get_all_tools() binds — what an agent can do before you install a single plugin. Most of them are conditional: they appear only when their backing store or config flag is present, which is why two instances of the same build legitimately show different tool lists.
- What ships, and what turns each one on → At a glance
- A tool you expected isn't there → Why a tool isn't bound
- The exact shape of one tool → the reference sections below
Where the agent's tools come from
get_all_tools() is one of five sources. The final toolset is assembled in graph/agent.py, in this order:
| Source | Example tools | Bound when |
|---|---|---|
| Core — this page | web_search, memory_recall, schedule_task | per-tool gates below |
Plugins (register_tools) | read_note, docs_search, delegate_to, github_get_pr | the plugin is enabled — see Plugin tools |
| Subagent delegation | task, task_batch | subagents are included in the build (Subagents) |
| Filesystem fence | read_file, edit_file, run_command | filesystem — on by default, fenced to workspace |
| MCP servers | <server>__<tool> | mcp.enabled, or a plugin's managed server |
Two things then happen to the whole assembled set, not just the core part:
tools.disabled/tools.hiddendrop named tools — including plugin, MCP, delegation and filesystem ones (tools.disabled: [run_command]really does remove shell access).- If deferred disclosure is on, everything outside a small base set is hidden from the model's per-call schema list until it searches for it. The tools are still bound and callable; only the model's view is trimmed.
At a glance
41 tools in twelve groups. Each group's heading says what makes it appear.
General — always bound
| Tool | What it does |
|---|---|
current_time(timezone="UTC") | Wall-clock time in an IANA timezone. |
calculator(expression) | Arithmetic over an AST — never eval(). |
web_search(query, max_results=5) | DuckDuckGo text search. No API key. |
fetch_url(url, max_chars=8000) | Fetch a URL, return cleaned plain text. |
Asking the operator — always bound, lead agent only
Both pause the turn via a LangGraph interrupt() (A2A input-required) and resume with the answer. Hard-denied to subagents; auto-answered on autonomous turns so nothing deadlocks.
| Tool | What it does |
|---|---|
ask_human(question) | One free-text question. |
request_user_input(title, steps, description="") | A Back/Next form wizard — text, number, boolean, choice cards. |
Rendering — always bound
| Tool | What it does |
|---|---|
show_component(component, props, title="") | Render a table / keyvalue / timeline widget inline in chat instead of a markdown blob. |
Skills & curation — always bound
| Tool | What it does |
|---|---|
load_skill(name) | Expand one <available_skills> entry into its full procedure. |
list_skills() | Every indexed skill — name · source · confidence · description. |
save_skill(name, description, body, tools=None, provenance_reason="", source_session_id="") | Create a new skill. Additive-only: refuses to overwrite. |
recent_activity(limit=30, window_hours=168) | Read-only digest of recent turns + a telemetry rollup. |
list_skills / save_skill / recent_activity are defined in tools/self_edit_tools.py (_build_curation_tools); load_skill stays in tools/lg_tools.py.
Memory & knowledge — bound when a KnowledgeStore exists
Built by default; drop the whole group with middleware.knowledge: false. See knowledge.
| Tool | What it does |
|---|---|
memory_ingest(content, domain="general", heading=None, memory_kind=None, subject=None, delivery_policy=None, expires_in_days=None) | Store text you already have — optionally typed (what it is, when it enters the prompt, how long it lasts). |
knowledge_ingest(source, domain="general", title=None) | Fetch + extract + chunk a URL or file — the only path to a YouTube transcript or a PDF. |
memory_recall(query, k=5, domain=None, memory_kind=None, delivery_policy=None, include_superseded=False) | Search memory; returns cited matches (optionally the superseded history too). |
session_search(query, limit=5, surface="") | Search prior session transcripts by content and return expandable session ids. |
recall_session(session_id) | Expand one <prior_sessions> line into that session's full summary. |
memory_list(domain=None, limit=10, memory_kind=None, delivery_policy=None, review_state=None) | Most-recent-first listing, with each chunk's #id and typed-memory / review / expiry tags. |
memory_stats() | Per-domain chunk counts. |
forget_memory(chunk_id, reason="") | Hard-delete exactly one chunk by id. |
Scheduling — bound when a scheduler backend exists
Built by default; drop with middleware.scheduler: false (or SCHEDULER_DISABLED=1). See Schedule future work. Defined in tools/scheduler_tools.py (_build_scheduler_tools).
| Tool | What it does |
|---|---|
schedule_task(prompt, when, job_id=None, timezone=None) | Persist a future turn — cron expression or ISO datetime. |
list_schedules() | This agent's jobs (never another agent's). |
cancel_schedule(job_id) | Cancel one job by id. |
wait(seconds, then) | End the turn now, resume in seconds with then as the instruction. Lead agent only. |
Tasks — bound when a TaskStore exists
Built by default. The agent's in-process planning board, mirrored to the console Tasks panel. Defined in tools/scheduler_tools.py (_build_task_tools).
| Tool | What it does |
|---|---|
task_create(title, description="", priority=2, issue_type="task") | Open an issue. priority 0 (highest)–3; type task|bug|feature|chore|epic. |
task_list(include_closed=False) | List the board. |
task_update(issue_id, status="", title="", description="", priority=-1, issue_type="") | Change fields; status is open|in_progress|blocked|deferred|closed. |
task_close(issue_id, reason="") | Close as done or won't-do. |
Inbox — bound when an InboxStore exists
Built by default (absent only if the store fails to open).
| Tool | What it does |
|---|---|
check_inbox(priority_floor="next", limit=10) | Pull pending inbound messages posted to POST /api/inbox and mark them delivered. |
Goals — goal.enabled and at least one plugin verifier
goal.enabled defaults to true, but the three goal tools also need a registered plugin verifier — with none, only list_verifiers binds. See Goal mode and ADR 0028. Defined in tools/goal_tools.py.
| Tool | What it does |
|---|---|
list_verifiers() | Every verifier registered here. Bound if goals or watches are on. |
set_goal(condition, check, check_args=None, max_iterations=None) | Set this session's standing goal, ground-truthed by a plugin verifier. |
update_goal_plan(plan) | Carry a running plan across goal iterations. |
abandon_goal(reason) | Declare the goal unachievable and stop the loop. |
Watches — watches.enabled and at least one plugin verifier
An independent axis from goals — a watch is verifier-only and moved by an external process. Defaults to true, same verifier requirement. See Watches and ADR 0067. Defined in tools/scheduler_tools.py (_build_watch_tools).
| Tool | What it does |
|---|---|
create_watch(condition, check, …) | Poll a condition on a cadence; optionally run a follow-up prompt when it trips. |
list_watches() | Every watch — id · status · condition · verifier. |
update_watch(watch_id, …) | Adjust a live watch without losing its observation history. |
clear_watch(watch_id) | Remove one watch. |
Introspection & onboarding — config-gated
| Tool | What it does | Bound when |
|---|---|---|
show_config(section="", offset=0, limit=0) | Read the agent's own effective, merged config, secrets masked. | a config is available (always, in the server) |
onboard_project(repo, name=None, write=None) | Clone a repo from any git host into the onboarding root and register it as a managed project. | onboarding.enabled (default on; refuses until root + allow are set) |
register_local_project(path, name=None, write=None) | Register a directory already on disk as a managed project — under the onboarding root directly, outside it only on the operator's approval card. | onboarding.enabled (default on; refuses until root is set) |
Opt-in singles
| Tool | What it does | Bound when |
|---|---|---|
edit_soul(section, content, mode="replace", reason="", source_session_id="") | Rewrite one section of the agent's own SOUL.md. | legacy lead opt-in, or bounded auto self-improvement |
update_skill(name, description, body, reason, tools=None, source_session_id="") | Replace an editable skill and archive its outgoing version. | auto self-improvement on a private/layered store |
delete_skill(name, reason, source_session_id="") | Delete an editable skill and archive its outgoing version. | auto self-improvement on a private/layered store |
set_config(updates) | Change the agent's own operational config — models, routing, plugin settings. Lead agent only. | tools.self_config_enabled: true (default off) |
search_tools(query="", limit=10) | Load deferred tools by capability. | tools.deferred.enabled: true (default off) |
The self-editing tools (edit_soul, update_skill, delete_skill, set_config) are defined in tools/self_edit_tools.py.
Why a tool isn't bound
In rough order of how often each one bites:
- It's a plugin tool and the plugin is off.
read_note,docs_search,delegate_to,github_*,run_workflow,show_artifactare not core. Checkplugins.enabledand the console's Plugins ▸ Installed panel, whose search matches tool names — the fastest way to answer "which plugin ships this?". - Its backend isn't there. No
KnowledgeStore→ nomemory_*. No scheduler → noschedule_taskand nowait. See the group headings above for the exact flag. - Goals/watches are on but no verifier is registered.
goal.enabled: truealone binds onlylist_verifiers;set_goalandcreate_watchneed a plugin-contributed verifier and are omitted until one exists. This is the surprising one — the flag is on and the tool still isn't there. - It's on the denylist.
tools.disabled(off, still toggleable in the console) ortools.hidden(off and not offered in the UI at all). Both sweep the fully assembled set. Seetools. - The subagent's allowlist doesn't name it. Subagents get an explicit list, not the full set — and
ask_human/request_user_inputare hard-denied even when a fork lists them (Subagents). - Deferred disclosure is hiding it. With
tools.deferred.enabled, the tool is bound but its schema is withheld untilsearch_toolssurfaces it.
The console's Settings ▸ Capabilities ▸ Tools panel lists what this instance actually ended up with, and GET /api/tools is the same view over HTTP.
General
current_time
async def current_time(timezone: str = "UTC") -> strCurrent wall-clock time in the given IANA timezone (e.g. "UTC", "America/New_York", "Asia/Tokyo").
2026-04-17T13:23:42.644606-04:00 (America/New_York)
Human: Friday, April 17 2026, 13:23:42 EDTUnknown timezones return "Error: unknown timezone 'Not/A_Zone'. …" — never raises.
calculator
async def calculator(expression: str) -> strEvaluates a numeric expression by walking an AST. Does not call eval().
| Supported | Example |
|---|---|
+ - * / | 1 + 2 * 3 |
// floor div | 10 // 3 |
% mod | 10 % 3 |
** power | 2 ** 10 |
Unary - | -5 + 3 |
| Parens | (1 + 2) * 3 |
Rejected with an error string: names (__import__, any identifier), calls (abs(-5)), attribute access ((1).__class__) — anything that isn't pure arithmetic.
Success: "2 ** 10 = 1024". Division by zero: "Error: division by zero".
web_search
async def web_search(query: str, max_results: int = 5) -> strDuckDuckGo text search via the ddgs package. No API key. max_results is clamped to 1–10.
3 result(s) for 'LangGraph tutorial':
1. LangGraph Introduction — https://langchain.com/langgraph
LangGraph is a framework for building...Network failures, rate limits and import errors come back as "Error: …" strings, so the model can read them and degrade rather than losing the turn.
fetch_url
async def fetch_url(url: str, max_chars: int = 8000) -> strFetches a URL and returns cleaned plain text.
- Scheme must be
http://orhttps://—file://,javascript:,ftp://are rejected. - SSRF guard is on by default: with no allowlist configured, a host that resolves to a private / loopback / link-local / cloud-metadata / reserved address is refused. Setting
egress.allowed_hostsflips it to deny-by-default — only listed hosts (wildcards allowed) pass, and the configured model gateway is auto-included. Seeegress. - Response body capped at 2 MB before parsing; text truncated at
max_charswith a…[truncated]marker. - HTML: scripts, styles, nav, footer and noscript stripped; prefers
<main>/<article>. - Non-HTML (JSON, text, CSV) is decoded and returned as-is.
[200] https://example.com
Example Domain
This domain is for use in documentation examples...User-Agent is protoAgent/0.1 (+https://github.com/protoLabsAI/protoAgent).
Asking the operator
ask_human
def ask_human(question: str) -> strPause the turn, ask the operator one question, continue with their answer (HITL, ADR 0003). Issues a LangGraph interrupt(), which checkpoints the graph at the call site; A2A callers see the task move to input-required carrying the question and resume it by sending a follow-up message with the same taskId.
Lead agent only — hard-denied to subagents (the interrupt is resumed by the lead turn's runner, and a subagent's graph has no checkpointer to resume one).
On an autonomous turn (scheduler / inbox / webhook / background) nobody is watching, so rather than parking the task forever the runtime auto-answers with a "no operator — proceed" sentinel (bounded; it force-completes past the budget). Prefer proceeding with a stated assumption there. Use this for a decision you genuinely must wait on — never for narration.
request_user_input
def request_user_input(title: str, steps: list[dict], description: str = "") -> strAsk for structured input via a form dialog and continue with the submitted fields as a JSON object. Same mechanics and same lead-only/auto-answer rules as ask_human.
steps is a list of form steps; more than one renders as a sequential Back/Next wizard (step indicator, required fields gate Next/Submit) and the last step submits every step's answers together. Each step is {"schema": <JSON Schema draft-07 of that step's fields>, "title"?: str, "description"?: str} — at least one step with fields is required (an empty steps is rejected).
Field types, per property in a step's schema.properties:
- text / number / boolean —
{"type": "string" | "number" | "integer" | "boolean"}; add"format": "textarea"for multi-line. - single-choice cards —
{"type": "string", "oneOf": [{"const": "pg", "title": "Postgres", "description": "Durable, multi-writer"}, …]}. Each option is a selectable card with its label and description. A bare"enum": [...]renders as a plain dropdown instead. - multi-choice cards — wrap the options in an array:
{"type": "array", "items": {"oneOf": [...]}}; the value comes back as a list.
Mark fields required with the step schema's "required": [...]. For one free-text or yes/no question, use ask_human.
Rendering
show_component
def show_component(component: str, props: dict, title: str = "") -> strRender structured data as a typed widget inline in the chat (ADR 0051) instead of a markdown blob. Data-only and safe — no code execution — and the console renders it through an extensible component registry that plugins can add to.
component | props |
|---|---|
table | {"columns": ["A","B"], "rows": [["a1","b1"], …]} |
keyvalue | {"items": [{"label": "Credits", "value": "183k"}, …]} |
timeline | {"steps": [{"label": "Buy hauler", "state": "done|active|todo", "detail": "…"}, …]} |
title is an optional heading. An unknown component returns an error naming the valid ones.
The fourth registered type, code-ref (a pointer into a project file that opens the console's code pane — ADR 0112), is not built here: it is emitted only by the filesystem tool show_code(project, path, line, end_line?, note?), which checks the fence, the secret-like-name deny list and the line range first. show_component refuses it. show_code is bound only while the code pane toolset is on (filesystem.code_pane, default off).
Plugins can add kinds of their own (registry.register_component, validated by the plugin's own schema — ADR 0051); show_component builds none of them. The artifact plugin's artifact-ref is one: every artifact create/revise leaves a chip that opens the Artifact panel on that exact version.
Rule of thumb: a data shape (table / metrics / steps) → this tool; a generated visual (chart, diagram, bespoke HTML/React/SVG) → an artifact, which renders generated code in a separate sandboxed panel. A mermaid artifact about the code can also carry links — its nodes and sequence messages then open the exact lines in the code pane (validated like show_code; see the artifact plugin's diagramming-code skill).
Skills & curation
load_skill is the runtime half of progressive disclosure; the curation tools back the scheduled /dream (memory consolidation) and /distill (workflow → skill) passes (ADR 0054). They read STATE at call time and self-gate, which is why they're unconditional here. The update/delete tools are separately guarded by the unified self-improvement policy.
load_skill
def load_skill(name: str) -> strThe <available_skills> block in the agent's context lists each skill as a name plus a one-line summary (ADR 0060); this returns that skill's complete body. name must match a <skill name="…"> exactly. An unknown name returns the available names (capped at 40, then a pointer to list_skills) rather than an error.
list_skills
def list_skills() -> strEvery skill in the index — name [source · confidence] — description — so a distill pass extends instead of duplicating. Read-only.
save_skill
def save_skill(name: str, description: str, body: str, tools: list[str] | None = None,
provenance_reason: str = "", source_session_id: str = "") -> strCreate a new skill. Additive-only: it refuses if the name already exists and never overwrites. Saved with source=distilled and curator-managed confidence, so a mistaken capture decays and self-cleans rather than accumulating. description is required — it's how the skill gets matched. A self-improvement pass can supply provenance_reason; the producing session is recorded automatically.
update_skill
def update_skill(name: str, description: str, body: str, reason: str,
tools: list[str] | None = None, source_session_id: str = "") -> strAvailable only when the self-improvement review and skills facet are both auto. Replaces a user or learned skill, preserving user-facing metadata and existing tool hints when none are supplied. Bundled/commons skills are read-only, and flat shared stores receive no automatic skill writers. The outgoing file is copied unchanged under skills/.history/<slug>/ before the live artifact changes; a JSON sidecar records the mutation session and reason.
delete_skill
def delete_skill(name: str, reason: str, source_session_id: str = "") -> strUses the same gate and read-only rules as update_skill. The complete outgoing artifact is archived before deletion, providing the rollback copy named in the tool result.
recent_activity
def recent_activity(limit: int = 30, window_hours: int = 168) -> strRead-only digest of what the agent has actually done: the Activity feed (time · origin · trigger · text) plus a telemetry rollup (turns, tool calls, LLM calls, cost, success rate, by model) over the window. limit is clamped to 1–200. Returns a "nothing to consolidate" line when both sources are empty.
Memory & knowledge
See Memory & the knowledge store for the model behind these, and Ingestion for the pipeline knowledge_ingest drives. The memory tools are defined in tools/memory_tools.py (_build_memory_tools) and bound by get_all_tools() when a knowledge store is present.
memory_ingest
async def memory_ingest(
content: str,
domain: str = "general",
heading: str | None = None,
memory_kind: str | None = None,
subject: str | None = None,
delivery_policy: str | None = None,
expires_in_days: int | None = None,
) -> strStore a chunk of text you already have — preferences, environment facts, decisions worth recalling later. domain is a logical bucket ("preferences", "context", "general", …); heading is an optional short label that doubles as a stable de-dupe key.
The typed-memory arguments (ADR 0108 D4) are all optional: memory_kind says what the chunk is ("profile", "standing", "fact", "decision", "note", "episode", "reference"; inferred from the domain when omitted), subject what it's about, and delivery_policy when it enters the prompt — "always" (every turn, the same promotion as domain="hot"), "retrieved" (on a relevant query; the default when omitted) or "on_demand" (only through memory_recall). When knowledge.hot_write_confirm is on the tool refuses always-on writes — domain="hot" ordelivery_policy="always" — and tells the model to ask the operator.
expires_in_days (ADR 0108 D7, 1–3650) gives a volatile fact a shelf life: the row is kept but leaves delivery once it lapses. Every memory the agent stores starts review_state="pending" until the operator confirms it in the Memory inspector, and carries the session it was written in as its source — stamped from the graph state, not passed by the model, so memory_recall can cite src: for it later.
Returns "Stored chunk 17 in 'preferences'.", or an error string when the store is unavailable.
knowledge_ingest
async def knowledge_ingest(source: str, domain: str = "general", title: str | None = None) -> strRuns the full ingestion pipeline over a source the agent doesn't have the text of yet: it pulls the URL or file, extracts text, then chunks and embeds it. Handles web articles, YouTube transcripts, PDFs and text documents, and — when a config with a gateway is available — audio, video and image sources via STT/vision.
This is the only path that gets a transcript or decodes a file; web_search + fetch_url won't. When a background manager is present (ADR 0050) a slow source is detached as a background job instead of blocking the turn. Filing a source under domain="hot" is an always-on write, so with knowledge.hot_write_confirm on the tool refuses it the same way memory_ingest does — before anything is fetched.
memory_recall
async def memory_recall(
query: str,
k: int = 5,
domain: str | None = None,
memory_kind: str | None = None,
delivery_policy: str | None = None,
include_superseded: bool = False,
) -> strTop-k search over the store (FTS5, LIKE fallback), one match per line, each citing its provenance — domain, stored date, namespace:
[preferences] coffee: Operator's preferred coffee is a Gibraltar with oat milk.
[context] lab: Primary lab is Snickerdoodle in Spokane.domain scopes the search to one bucket — use it to separate the agent's own record from inherited or imported knowledge (a domain like claude-import is another codebase's history, not this agent's actions). memory_kind and delivery_policy narrow it to one typed-memory classification (ADR 0108 D4); delivery_policy="on_demand" is the only way an on-demand memory surfaces. include_superseded=True (D7) also returns rows a newer revision replaced, each tagged [superseded] — the audit history for "what did we believe before?". Returns "No matches." when nothing clears the threshold.
session_search
async def session_search(query: str, limit: int = 5, surface: str = "") -> strFull-text search over persisted prior-session summaries when the relevant session id is unknown. The disposable FTS5 index is synchronized lazily, so existing histories become searchable without migration and session persistence never depends on index health. Results are relevance-ranked, credential-redacted, capped at 20, exclude the active session, and carry an excerpt plus id for expansion with recall_session.
surface optionally limits results to chat, a2a/other, activity, palette, or background. Query text is converted to literal terms; raw FTS operators are never executed.
Matching is stemmed (FTS5 porter), so "audit log" finds a session about "audit logs" and "rotation" finds one about "rotating" — recall doesn't depend on reproducing the stored wording. Stemming handles inflections, not derivations: "retained" and "retention" have different stems and do not match each other. There is no semantic/vector search over sessions, so a query that shares no word stem with the stored text will miss — which is why the automatic <prior_sessions> digest still earns its place where it is enabled (ADR 0108 D9): it tells the agent a session EXISTS without requiring it to guess the wording first. It is not unconditional — context.prior_sessions: off turns it off entirely, and goal-driven turns suppress it — so on those turns this search is the ONLY way back to a past session, wording guess and all.
recall_session
async def recall_session(session_id: str) -> strExpands one entry of the auto-injected <prior_sessions> digest into that session's full persisted summary — messages and final output, reasoning-stripped, capped at ~6 000 chars (ADR 0069). The digest itself carries only one attributed line per prior session (id · timestamp · surface · topic · message count), so this is the on-demand path to the content. Errors cleanly on an unknown or malformed id.
memory_list
async def memory_list(
domain: str | None = None,
limit: int = 10,
memory_kind: str | None = None,
delivery_policy: str | None = None,
review_state: str | None = None,
) -> strMost-recent-first listing, filtered by domain, memory_kind, delivery_policy and/or review_state when given. Each row carries the #<id> that forget_memory takes, plus kind= / policy= / review= / expires= tags. review_state="pending" lists what still awaits the operator's verdict — confirmation is theirs to give, never the agent's. Useful for "what did I log today?".
memory_stats
async def memory_stats() -> strPer-domain chunk counts plus a total — the sanity check that an ingest actually landed.
forget_memory
async def forget_memory(chunk_id: int, reason: str = "") -> strHard-deletes exactly one chunk by the id memory_list shows. The forgetting half of a /dream pass: use it on a fact that is stale, superseded or duplicated, ideally after ingesting the corrected version first.
This is a real delete, not a supersede — automatic fact consolidation marks replaced rows invalidated_at and keeps them for audit, but an explicit forget is operator intent and removes the row. No bulk or wildcard form exists, by design. reason is recorded for the audit trail.
Scheduling
schedule_task
async def schedule_task(prompt: str, when: str, job_id: str | None = None, timezone: str | None = None) -> strPersist a future invocation; the agent receives prompt as a fresh turn when it fires.
when is either a 5-field cron expression ("0 9 * * 1-5" = weekdays at 9am) or an ISO-8601 datetime ("2026-05-01T15:00:00" = once). Backends auto-detect which. timezone is an optional IANA zone for interpreting when. job_id defaults to <agent_name>-<uuid> — you need it later for cancel_schedule.
Returns "Scheduled job <id> next at <iso>.", or "Error: …" on a malformed when.
Prompts must be self-contained: the agent has no memory of the scheduling moment when the task fires, so write a fresh turn ("review last week's pipeline incidents and post a summary"), not a back-reference ("do that thing we discussed").
list_schedules
async def list_schedules() -> strThis agent's scheduled jobs — one per line with id, next fire, schedule and prompt preview. Multi-agent isolation is real: each agent only sees jobs it created. "No scheduled jobs." when empty.
cancel_schedule
async def cancel_schedule(job_id: str) -> strCancel by id. Returns "Canceled <id>." or "Error: no such job <id>." Cross-agent cancellation is blocked — gina-personal cannot cancel gina-work's jobs even when they share a sqlite path.
wait
async def wait(seconds: int, then: str) -> strYield the turn and resume later instead of busy-polling a status tool (ADR 0053). Calling wait ends the current turn (via WaitYieldMiddleware) and schedules a one-shot resume seconds from now; when it fires the agent is re-invoked with then as its instruction, in the same conversation thread (history intact). This is how you run long-horizon "do X, wait, do Y" work without burning the recursion budget on a poll loop.
then is required and self-contained — it's the agent's only context on resume, so it must name the work and the entities ("Dock NOVAHAUL-5 at X1-UC87-K93, sell the ore, accept the next contract"). seconds is clamped to ≥ 1; pass the ETA a status tool gave you and wait the full duration in one call, since under-waiting just wakes early to wait again.
One pending wait per thread (#1702): a new wait supersedes any still-pending wait for the same session (a stable wait:<session> job id → cancel-then-add), so repeated waits can't stack into a pile of wake-ups that all fire into the thread. A schedule_task job uses its own id and is untouched. Every scheduling is logged — [wait] thread=… in Ns (superseded a pending wait) → resume: … — so a stacking loop is visible.
Lead agent only (subagents are bounded by max_turns and no allowlist names it). The yield is durable across restart. For an absolute time or a recurring cadence use schedule_task — wait is for "yield for a bit, then pick this back up".
Tasks
The agent's own planning board — an in-process SQLite issue tracker, mirrored to the console Tasks panel. Tasks are attributed to the session that created them.
task_create
def task_create(title: str, description: str = "", priority: int = 2, issue_type: str = "task") -> strOpen an issue and return its id. priority is 0 (highest) to 3 (low); issue_type is one of task / bug / feature / chore / epic. An invalid value returns "Error: …" rather than raising.
task_list
def task_list(include_closed: bool = False) -> strOne line per issue — [status] id (pN, type) title. Open issues only unless include_closed=True. "No issues on the board." when empty.
task_update
def task_update(issue_id: str, status: str = "", title: str = "", description: str = "",
priority: int = -1, issue_type: str = "") -> strChange any subset of fields. status is open / in_progress / blocked / deferred / closed. Leave a string field empty — or priority at -1 — to keep it unchanged.
task_close
def task_close(issue_id: str, reason: str = "") -> strClose an issue as done or won't-do, with an optional reason.
Inbox
check_inbox
async def check_inbox(priority_floor: str = "next", limit: int = 10) -> strPull pending inbound messages (ADR 0003) — webhooks, external systems and sister agents that posted to POST /api/inbox — and mark them delivered.
priority_floor selects the tiers: now (now only), next (now + next, default) or later (everything pending). now-priority items have already fired an Activity turn on arrival; next and later wait for this call, so the agent decides when to surface them. Returns the items one per line, or "Inbox empty.".
Goals & watches
Both surfaces are plugin-verifier only. The tools hardcode type="plugin", so an agent cannot open a shell, test or data goal on itself — those stay operator-only via /goal and the operator API. A verifier name looks like <plugin-id>:<name>.
list_verifiers
async def list_verifiers() -> strEvery verifier registered on this instance: the core types (command, test, ci, data, llm, plugin) for awareness, then the plugin-contributed checks with their <plugin-id>:<name> identifier and description. Only the plugin checks are usable by set_goal / create_watch. When none are registered it says so explicitly — which is the signal that goal mode is on but has nothing it can verify.
set_goal
def set_goal(condition: str, check: str, check_args: dict | None = None,
max_iterations: int | None = None) -> strSet a standing goal for this session. The agent is re-invoked toward condition until the plugin verifier named by check passes; check_args is declarative data the verifier reads (e.g. {"min": 1000000}).
An unknown check is rejected up front, listing the registered verifiers — without that guard the goal would be created but could never pass, spinning to the iteration cap and finishing unachievable. Also errors when goal mode is off or there's no active session.
update_goal_plan
def update_goal_plan(plan: str) -> strRecord or refresh the running plan for this session's active goal (what's done, what's next, what failed). It's persisted and fed back into the next continuation prompt, which is how a coherent plan survives across iterations. A harmless no-op when goal mode is off or no goal is active.
abandon_goal
def abandon_goal(reason: str) -> strFlag the active goal unachievable and stop the loop. The goal finishes unachievable after the turn — unless the verifier finds it already met, which wins. No-op when no goal is active.
create_watch
def create_watch(condition: str, check: str, check_args: dict | None = None, run_prompt: str = "",
watch_id: str | None = None, interval_s: float | None = None,
expires_in_s: float | None = None, stall_after: int | None = None,
repeat: bool = False, on_change: bool = False) -> strPoll condition on a cadence, ground-truthed by the plugin verifier check; when it's met, run run_prompt (if given) as a follow-up turn in this session. Many watches run in parallel — that's the point: a deploy, a CI run and a metric are three watches, not one goal.
By default a watch is a tripwire: it fires once and is done. Two flags make it a standing monitor:
repeat— keep watching after it fires, firing each time the condition becomes true again (so a latching condition likecredits >= 1Mdoesn't spam).on_change— fire whenever the checked value moves, whatever the condition says. Use it to track something rather than wait for it. Impliesrepeat.
Three knobs shape lifetime and cost; a watch with none of them polls at the default cadence until something clears it:
| Knob | Effect |
|---|---|
interval_s | Seconds between checks for this watch — a floor, never faster than the global cadence. |
expires_in_s | Give up this many seconds from now; the watch finishes expired. Relative on purpose — a model asked for an absolute timestamp guesses, and a guess in the past expires the watch on its first tick. Must be positive. |
stall_after | After N consecutive checks with unchanged evidence, fire the stall signal (the watch stays active) — how you notice a deploy that's wedged rather than slow. |
watch_id defaults to a slug of the condition; pass one to hold two watches on the same condition. Set expires_in_s on any repeating watch unless you really mean forever.
list_watches
def list_watches() -> strEvery watch for this agent — id · status · condition · verifier — or a note when there are none.
update_watch
async def update_watch(watch_id: str, condition: str | None = None, run_prompt: str | None = None,
interval_s: float | None = None, expires_in_s: float | None = None,
stall_after: int | None = None, clear_deadline: bool = False,
repeat: bool | None = None, on_change: bool | None = None) -> strAdjust a live watch, passing only what changes. Use this instead of clear-and-recreate — recreating resets the stall history and starts the evidence over.
expires_in_s is measured from now, as in create_watch. Because None already means "not supplied", removing an expiry entirely needs its own flag: clear_deadline=true (passing both is an error). A finished watch can't be edited — set a new one.
clear_watch
def clear_watch(watch_id: str) -> strRemove a watch by the id list_watches shows. Reports whether it existed.
Introspection & onboarding
show_config
def show_config(section: str = "", offset: int = 0, limit: int = 0) -> strReturns the agent's own effective, merged configuration — the same view GET /api/config serves — as JSON. Pass a section as either a top-level key ("model", "mcp", "filesystem", or any plugin section such as "project_board") or a dotted nested path such as "project_board.projects" to keep the output small; omit it for the whole document, which falls back to a section index when it's too large to render at once. Exact top-level keys win before dotted parsing, so existing sections whose names contain dots remain addressable. An unknown top-level section returns the available names plus a near-match suggestion; an unknown dotted path names the missing segment and the keys available at the last resolved object.
Selected dicts, lists, and oversized strings can be paged with offset and limit. Dict pages use sorted keys, so repeated calls reconstruct the same value deterministically. Paged responses are JSON envelopes with section, pagination (offset, actual limit, returned, total, next_offset, has_more) and value. If a selected value is too large for the 12k transport safeguard, show_config returns an explicit first page instead of cutting the JSON; continue with offset=<next_offset> until next_offset is null. When a single child is itself too big to embed in a page, it comes back as a __truncated__ pointer carrying the deeper read_with path (and shape metadata, never the value) — so paging a parent never dead-ends on one oversized child, and nothing is silently dropped.
Selectors are dot-separated. A key that itself contains a dot (or is empty) is escaped in the path — a literal dot is written \. and a literal backslash \\ — so read_with pointers resolve back to exactly that child instead of re-splitting through the middle of its name. Exact top-level keys are still matched before any dotted parsing, so a section whose name contains a dot stays addressable as-is.
Read-only. It never writes, and it binds whenever a config is available. Drop it with tools.disabled: [show_config] like any other core tool.
Why it exists. An agent's own config/langgraph-config.yaml lives outside every filesystem fence, so the file that says how the agent is wired was the one file it couldn't open — making a misconfiguration indistinguishable from a bug. In the incident that prompted it (#2540), an agent spent two sessions diagnosing a board bound to the wrong repo — checking paths on disk and on GitHub, reading the plugin's defaults from source — before its operator found the answer in one grep. Reading a plugin's source tells you what it does unconfigured; this tells you what your instance set.
Secrets are masked, not inherited as safe. config_to_dict blanks secret-typed schema fields and plugin-declared secrets, which is the right bar for the token-gated operator API — but this output lands in the model's context and the chat transcript, so the tool applies its own pass on top. mcp.servers[].env and [].headers are free-form string maps that routinely hold tokens; every value in them is masked, along with any key that names a credential, at any depth. Masked values read as «redacted», so the agent still learns that a credential is set — a blank stays blank, because masking one would claim a token is present when none is.
onboard_project
async def onboard_project(repo: str = "", name: str | None = None, write: bool | None = None, github_repo: str = "") -> strClone a git repository into the configured onboarding root and register it in the managed projects registry, so the filesystem tools, the GitHub plugin and the project board can reach it. repo accepts any git host: a bare owner/repo (GitHub), gitlab.com/owner/repo, https://host/owner/repo(.git), git@host:owner/repo(.git) or ssh://git@host/owner/repo. git gets the URL as given, so ssh keys and credential helpers apply; credentials in a URL are masked in the result. github_repo is the older name for repo and still works. name defaults to the repo name; write overrides the operator's default access mode. The registry's github binding is filled only for a github.com source.
The source must match an onboarding.allow glob on its host/owner/repo (so gitlab.com/acme/* admits GitLab, and a github.com/* glob does not); local paths, file://, ext::-style transports and anything starting with - are refused before git runs. An existing checkout is reused untouched (no fetch), with its drift from the tracking branch reported.
register_local_project
async def register_local_project(path: str, name: str | None = None, write: bool | None = None) -> strRegister a directory that is already on disk — a checkout the operator made, or one the agent created — in the same registry, without cloning. path must be absolute (~ is expanded) and is judged by where it resolves (symlinks followed). Inside onboarding.root it registers directly. Outside it — or anywhere, when no root is set — the turn parks on an operator approval card for that one folder — Allow read-only / Allow read-write / Deny, the operator's choice winning over write — that /bypass and "allow for session" never skip; a denial is a plain result the agent can act on. Some places are refused with no card (filesystem root, home, system and credential directories — the full list is under onboarding); onboarding.approve_outside_root: false restores the flat refusal. allow does not apply — nothing is fetched. The github binding comes from the directory's origin remote when that is GitHub; default_branch from origin/HEAD, else the current branch. Idempotent: an already-registered path is reported, not re-written.
Both tools are present unless onboarding.enabled: false, which removes them entirely. With enabled on but root/allow unset they refuse by naming the setting to change — that is how the agent tells the operator what to configure.
Self-configuration
set_config
async def set_config(updates: dict) -> strGuarded, off by default. Bound to the lead agent only when an operator sets tools.self_config_enabled: true; bounded subagents never receive it. updates is a flat dict of dotted keys → values ({"project_board.coder": "proto"}), applied through the same ops.config.set_config write path as protoagent config set and the console's PATCH /api/config (ADR 0075 D2 — one op, thin adapters). The change persists and reloads the agent, and the operator is notified on the bus.
Scope is operational settings only. The tool refuses the whole write — never partially — when any key falls in the trust surface, because those settings decide what the agent is allowed to do rather than how it behaves:
| Refused | Why |
|---|---|
filesystem | allow_run (shell) and the ADR 0007 project fence |
operator | allowed_dirs / project_dir — the operator fence |
egress | the ADR 0008 network fence |
plugins | enabling a plugin runs its code in-process as the agent (ADR 0071) |
delegates · runtime · acp | each entry names an executable the host spawns (command/args, often permissions: auto) |
auth · security · mcp | operator credentials and out-of-process capability |
soul | persona has its own guarded path (edit_soul) |
tools | including self_config_enabled itself — no widening its own fence |
Plugin sections can't be enumerated in advance, and any plugin may grow a key naming a binary, so below the section the fence refuses a key that names a program to run, in any section:
- a key segment that is, or has as a
_/-/camelCase token, one ofcommand,cmd,args,argv,binary,bin,exe,executable,interpreter,entrypoint—coder.command,local_gate_cmd,rh_bin,binary_path,browserArgs; - a plugin setting its manifest marks
spawns: true— for names the rule above can't see, likeffmpeg_path. Read from every installed plugin, enabled or not; if that lookup fails, plugin-section writes are refused rather than guessed at.
Nested values are checked too: {"some_plugin": {"command": …}} is the same write as {"some_plugin.command": …}. *_path keys in general are not refused — most name data (brand_kit_path), which is the agent's to repoint. The line is that an agent may choose among the executables its operator provisioned (project_board.coder: proto) but never define one.
Any secret-typed key is refused outright and never echoed back: the write path would faithfully route it into secrets.yaml, which is right for the CLI and the console and wrong here, since it would put a live credential in the turn transcript.
So an agent can retune how it runs without being able to change what it may do. Batching a denied key with legitimate ones refuses everything, so nothing can be smuggled through alongside a valid write.
Persona
edit_soul
async def edit_soul(section: str, content: str, mode: str = "replace",
reason: str = "", source_session_id: str = "") -> strGuarded, off by default (ADR 0081). Bound to the lead agent when an operator sets soul.self_edit_enabled: true. The one exception is the built-in, policy-bounded self-improve reviewer: it receives the tool only when the master, post-goal review, and persona modes are all auto. Other bounded subagents never receive it. It lets the agent durably refine its own persona by rewriting one markdown section of SOUL.md — section matches case-insensitively (a missing one is created), mode is replace or append.
Scope is persona only — identity, voice, values, temperament — never operating doctrine (ADR 0079). Guardrails: one section per call (it can't blow away the file), a 64 KB cap on the whole persona (it rides in the system-prompt prefix every turn), and empty / no-op / invalid-mode edits are refused with an error string.
Every edit goes through write_soul, so the outgoing persona is snapshotted to soul-history and restorable from Settings ▸ Identity. The change takes effect on the agent's next turn: the server injects its graph-reload as a callback so the compiled graph rebinds atomically — the current turn is unaffected — without tools/ importing server/. Where no callback is wired (subagent, eval, script), the save still lands and applies on the next natural reload.
Never silent. Every accepted edit publishes a persona.self_edited event (section, mode, new revision, producing session, and optional reason) on the event bus, so the operator sees an identity change in the console even when it happens on an autonomous turn — and it leaves a trail if a prompt injection ever drove one. That transparency guardrail came out of ADR 0081's due diligence against prior art: Hermes keeps SOUL.md operator-only; OpenClaw invites unguarded self-edit and treats the soul as a prompt-injection surface; Letta added a read-only persona guard after unconstrained self-edits degraded identity.
Progressive disclosure
search_tools
def search_tools(query: str = "", limit: int = 10) -> strAdded only when tools.deferred.enabled is on (ADR 0005). With deferral on, the model sees a small always-on base set plus this meta-tool each turn; everything else stays bound and callable but its schema is withheld until the agent searches for it. ToolDeferralMiddleware reads the backticked names out of the result and binds those tools on subsequent turns.
Keyword-matches deferred tools by name and description, returning - \name` — purposelines. An emptyquerylists everything;limitis clamped to 1–50; a query that matches nothing falls back to the full list rather than a dead end. Override the always-on base withtools.deferred.keep—search_tools` itself is always kept, since without it nothing could be loaded back.
Plugin tools
Everything else the agent can call comes from a plugin (register_tools), not this registry. First-party plugins that ship in-tree:
| Plugin | Tools | Default |
|---|---|---|
notes | read_note / write_note / append_note over one shared agent-global markdown notebook, plus the Notes console panel | on |
docs | docs_search / docs_read over protoAgent's own documentation | on |
artifact | show_artifact — generated HTML/React/SVG in a sandboxed panel | on |
craft | (no tools — ships engineering slash-command skills) | on |
delegates | delegate_to(target, query) — route a sub-task to another agent or endpoint over a2a / openai / acp, managed and hot-swappable from the console (ADR 0025). An acp delegate drives a CLI coding agent over ACP (ADR 0024). Replaced the retired peer_consult / peer_list / code_with. See Delegates + CLI coding agents | built-in |
workflows | run_workflow + Workflow Studio | off — plugins.enabled: [workflows] |
friction | record_friction / friction_review / resolve_friction, the Friction console view, and the /friction operator command. Open friction is projected into <working_state> (ADR 0079) so the agent observes its own backlog. Auto-capture logs tool errors, and a shell command only when it duplicates a bound first-class tool (cat → read_file, grep → search_files, ls → list_dir); git, tests and builds are never friction. Config: friction.escape_hatch_exempt, friction.issue_repo (see the Friction log guide) | on |
cowork | (no tools — ships the knowledge-work skill pack: docx / xlsx / pptx / pdf deliverables through execute_code, /daily-brief, drop-folder watches via the cowork:folder_changed verifier, schedule, consolidate-memory, writing-voice, /setup-cowork. ADR 0083) | on, and turns on execute_code — off with plugins.disabled: [cowork] |
agent_browser | 17 browser tools over the native agent-browser CLI: browser_open / browser_back / browser_forward / browser_reload, browser_snapshot (accessibility tree with @eN refs) / browser_get_text / browser_get_html / browser_get_value, browser_click / browser_fill / browser_type / browser_press / browser_hover / browser_eval, browser_screenshot / browser_pdf (page→PDF, fenced to the plugin's capture dir) / browser_close — plus the web-browse skill, two workflows, and the drivable Browser panel. See Browser automation | off — plugins.enabled: [agent_browser], and the agent-browser binary on PATH |
execute_code | execute_code — a Python interpreter the document skills run in: a child process with a scrubbed environment and a hard timeout (isolation, not a true sandbox) | on while cowork is on — plugins.disabled: [execute_code] always wins |
coder, orgchart, telegram, hello | — | off |
Others — including the GitHub plugin (github_get_pr, github_get_issue, github_list_issues, github_get_commit_diff and the Review API tools) — live in their own repos and are installed from a git URL:
python -m server plugin install <git-url>Installed plugins are pinned in plugins.lock and manageable from the console's Plugins ▸ Installed panel. See Plugins and the plugin registry.
Adding your own
Ship a plugin. register_tools is the supported path: it survives upstream re-sync, hot-reloads, and can carry its own config, secrets, Settings and console views. Editing get_all_tools() still works, but it is a core edit that conflicts on every upstream merge.
The tool itself is the same either way:
from langchain_core.tools import tool
@tool
async def my_tool(required_arg: str, optional_arg: int = 5) -> str:
"""First line becomes the LLM's summary of the tool.
Args:
required_arg: What this argument is. The LLM reads these docstrings.
optional_arg: Optional, with a sensible default.
"""
try:
result = await do_the_thing(required_arg, optional_arg)
except Exception as e:
return f"Error: {e}"
return f"Success: {result}"Two conventions every tool here follows, and yours should too:
- Never raise into the turn. Return
"Error: …"and let the model read it, retry with different arguments, or degrade. A raised exception costs the turn. - Don't hand-roll
subprocess. Build ontools/shell.py::run_command(async; handles timeout/kill, missing-binary → structured error, env merge, stdin/cwd) ortools/gh_cli.pyforghspecifically.
If you're inside a tool body and need the current session, read it from injected graph state — current_session_id() is empty there (the tool runs in a different execution context than the middleware).
See Write your first tool for the walkthrough, and Build a plugin for the packaging.
Related
- Configure subagents — tools are allowlisted per subagent
- Configuration —
tools.disabled,tools.hidden,tools.deferred - Environment variables — SSRF allowlist vars affect
fetch_url; scheduler backend selection lives there too - Eval your fork — the eval harness exercises these end-to-end
- Schedule future work — the firing model and multi-agent isolation behind the scheduler tools