Skip to content
You are reading documentation for unreleased main. Read the 0.1 version.

Tool Skills

Tool Skills are modular prompt fragments that teach an agent how to use a particular tool or MCP server. Each skill lives as one page in a dedicated container of the knowledge backend (a Confluence space, or a container of the engine’s own pages), is loaded into every node’s in-memory skills.Registry, and is kept current on every node by that node’s skill sync: a walk of the container at boot and whenever the applied configuration moves it, a read of the one page that changed when a page changes, and a periodic walk behind both (see Keeping every node current).

Two-tier consumption:

  • A short summary (≤240 bytes) always appears in the per-phase prompt catalogue so the executor knows the skill exists.
  • A rich body (≤32 KB) loads on demand via the load_tool_skill(key) builtin, which is always available in Execute and Sub-agent.

This keeps the prompt prefix small (no eager body inlining) while still making rich skill content available the moment the agent decides it needs it.

Skills are also enforced by default: the engine blocks calls to the tools a skill’s trigger covers until the body has been loaded in the current session. Mark orientation / hint-grade content required: false to keep it advisory — see Required skills.


The engine hardcodes no tool-specific guidance — no per-tool prompt constants, no per-MCP-server hint sections in the prompt builders. All of that prose lives in the knowledge backend (Confluence pages). Adding a new MCP server or tweaking the guidance for an existing one is a page edit — no engine code change, no release.

This covers the how-to half of tool decoupling. The structural half — the engine’s phase contracts, sub-agent guard, and builtins not naming concrete tools either — is covered by Tool Capabilities: prompts describe a tool by its capability and the LLM picks the match from its catalogue, while the engine derives any runtime tool classification from MCP annotations rather than a tool-name list.


walk: boot, apply, periodicpage read: page webhook

tool_skill_page_changed(broadcast)

page read

(operator authors / edits)

Knowledge backend: the Tool Skills container(Confluence space TS)

skill sync(one loop per node)

every other node's skill sync

skills.Registry(in-memory, one per node)

summary in per-phase catalogue

load_tool_skill(key) builtin(LLM-driven, always available)

LLM sees the rich body onlywhen it decides it's needed

No database row. The registry is in-memory and belongs to the node rather than to a configuration epoch, so an apply keeps it (an apply refreshes the skill_variables map, and re-points the sync when it moved the skills source). A restart walks the container again.

No code defaults. The engine ships zero skill prose. An empty container → empty registry → just the tool catalogue. Operators seed it by publishing markdown: with their own assistant over /operator/mcp on the native backend, or crewlet confluence import on Confluence (see below).

One skill sync per node, reading the single-homed knowledge backend (see Knowledge System). A walk and a single-page read apply the same admission test (skills.AdmitPage): the page lives in the configured container and its leading YAML frontmatter declares a trigger:. A page with no frontmatter, or frontmatter with no trigger, is an ordinary page and is skipped quietly; a page that declares a trigger and does not parse is reported (skill_page_undecodable) and skipped. A previously admitted page that stops passing the test (deleted, moved out, or edited into a non-skill) is dropped, never left serving its last-good body.

A page is the registry’s identity, and a key is what a model asks for. Every change names a page, so the registry records what each page holds and derives the key-addressed catalogue from that. A page whose key was edited leaves nothing behind under the old key, and reading the one page that changed ends exactly where a walk of the whole container would. Two pages declaring one key are an authoring error: the page with the lower id (on Confluence, the older page) is served on every node, the other is logged as skill_key_duplicated, and it takes the key the moment the first page gives it up.


The registry is per node and its content is a wiki’s, so each node runs one sync loop that owns it. Everything that can change the registry is a request to that loop, and the loop does one read at a time, so a later read always sees later state than an earlier one.

What happenedWhat the node’s sync does
The node bootsWalks the container. On the native backend it waits until this node’s page projection has caught up first, because a walk over a projection that is behind is a partial set and the registry replaces wholesale.
An apply moves the skills source: a different knowledge backend (connecting Confluence, or disconnecting it), a different skills container, or a different Confluence instanceRetires the previous source’s skills at once (tool_skills_retired) and walks the new one. The registry never serves a source the applied company no longer names: the old container’s pages are ordinary pages now, and a disconnected wiki’s skills describe a stack the company stopped running. A company moved onto the native knowledge base by a revision serves its native skills after the node restarts, because the native backend starts with the node; until then the node logs tool_skill_source_unreadable.
An apply turns tool skills off (knowledge.backend: none, or skills_container: "")Retires every skill and walks nothing (tool_skills_off).
An apply that leaves the source as it wasNothing, unless the last walk failed or this node could not read the source until now (a token put back, a Confluence integration that starts this time). An apply is the gesture that fixes a credential, so it walks again then; and a source that was unreadable is walked even after a walk that succeeded, because every page change that arrived while it could not be read was dropped.
A Confluence page webhook, on the node that won the deliveryReads that page only and updates it in the registry (tool_skill_page_synced), then publishes tool_skill_page_changed on the event stream. A page event in any other space costs nothing unless the page currently holds a skill (a page that just moved out of the skills space is announced from the space it moved to).
tool_skill_page_changed from a peerReads that page itself and updates it. Webhook deliveries are a fleet-wide consumer group, so exactly one node parses each one, and this broadcast is how every other node hears about it. A node whose applied company names a different container ignores it.
A Confluence page trashed or deletedReads it back like any other change, on the node that heard the webhook and on every peer, and drops it on the wiki’s answer that the page is gone. A removal is never acted on without that read: a page restored from the trash announces nothing, and a trash delivery or a nudge that reached a node after the restore would otherwise drop a skill the wiki still serves until the next periodic walk.
A single-page read failsWalks the whole container instead (tool_skill_page_read_failed), which is the path that knows how to back off.
On the native backend, a skill page commits in this node’s page projection (created, edited into or out of a skill, trashed, restored or purged)Walks the container, a read of the node’s own store, and announces it: tool_skill_page_changed with walk: true, since the applier knows the container moved but not which page. Every node that holds data applies the page log itself and hears the change without it; the announcement is for a node that holds no data, which runs no applier and would otherwise hear of the edit only on its periodic walk. Several data nodes announce one edit, and a hearer coalesces them into one walk. A node whose company moved its skills to Confluence by an apply keeps its page projection running, and a native commit there walks nothing.
tool_skill_page_changed with walk: true, on a node that holds no dataWalks the container through a data node, which answers only once its own copy answers requests — level with its logs, or drained and close to their ends.
Every 10 minutes, with a fifth of that as jitter either wayWalks the container. This bounds everything the event path misses (a broadcast a node did not hear, a webhook the wiki never delivered, and the changes no subscribed Confluence event names: a page moved into or out of the space, a page restored from the trash) at 12 minutes, the longest a jittered interval can be.
A walk failsKeeps what the registry held, logs tool_skill_sync_failed (with attempt and retry_in), and retries after 5 seconds, doubling per consecutive failure up to the 10-minute interval. A source that stays broken then costs one walk per interval, the same as a healthy one.

A single-page update and a walk cannot disagree about a page, so the periodic walk never moves guidance under a seat that already read it; it only catches up what the event path missed.

A walk or a page read that finds the container as the registry was already built from it changes nothing and logs nothing above debug (tool_skills_unchanged, tool_skill_page_unchanged). Almost every periodic walk finds exactly that, and installing it again would repeat every warning a registry change raises (skill_key_duplicated, skill_variable_unresolved, skill_page_undecodable, and the trigger audit’s skill_trigger_* lines) once per interval on every node. Those are logged when a walk or a read changes something; the trigger audit and the variable check also run on every config apply. A walk that recovers after failures still logs tool_skills_synced with recovered, because the failures before it were logged as errors.


Skills are authored as markdown files with YAML frontmatter — the same format the backend pages round-trip to via the per-backend codec (below). The repository ships ten bundled examples under examples/tool-skills/.

---
key: skill:code_runtime
trigger:
tool: run_sandbox
phases: [execute]
required: false
title: Running code work in the sandbox
summary: |
For real code work, call the run_sandbox tool: a coding agent
implements in an isolated checkout and pushes a branch or opens a pull
or merge request.
---
When a task needs you to implement or modify code, run tests, or
reproduce a bug, call `run_sandbox` with a concrete brief …
FieldRequiredDescription
keyyesStable identifier. By convention mcp:<server>, tool:<name>, or skill:<descriptive>. Used for idempotency on re-import, prefix-cache ordering, and as the argument to load_tool_skill.
triggeryesA discriminated expression; see Triggers below.
phasesyesList of turn phases this skill is catalogued in: any subset of execute, review, subagent.
titleyesRendered as the header when the body loads.
summaryyes≤240-byte one-liner. Always inline in the per-phase catalogue, regardless of whether the LLM ends up loading the body. Keep it tight.
requiredno (default true)Load-before-use safeguard. When true (the default), the engine rejects calls to the tools the trigger covers until the LLM has loaded this skill via load_tool_skill in the current session. Set false for advisory orientation / hint content. See Required skills below.

Exactly one of these fields:

TriggerFires when
tool: <name><name> is in the phase’s tool surface
mcp_server: <server><server>’s tools reach the seat: a shared server (the default) reaches every agent seat, and a shared: false one only a seat that declares it under mcp_env, its own or its unit’s
any_of: [...]Any sub-trigger fires (logical OR)
all_of: [...]Every sub-trigger fires (logical AND)

Examples:

# Single MCP server
trigger:
mcp_server: github
# Several tools that share guidance
trigger:
any_of:
- tool: query_episodes
- tool: confluence_search
- tool: refresh_memory
# Composite: only when an agent has both atlassian AND slack
trigger:
all_of:
- mcp_server: atlassian
- mcp_server: slack
  • Body: 32 KB UTF-8 (MAX_SKILL_BODY_BYTES). Looser than a prompt-prefix cap because bodies load on demand, not on every turn. Skills bigger than this almost always want to be split.
  • Summary: 240 bytes UTF-8 (MAX_SKILL_SUMMARY_BYTES). The summary appears on every turn for every role triggering the skill — keep it tight.

Both caps are enforced at file-parse time and at registry-load time as a small defence against accidental or malicious prompt-injection prose. The caps bound the source bytes; a skill that references skill variables (below) can render slightly longer.


A skill body is the same static text for every company, but some guidance needs a per-org fact the body can’t know — the canonical case is the org’s tracker/knowledge-base base URL, so an agent shares a human-clickable link instead of what its tools hand back (mcp-atlassian under cloud_id auth returns api.atlassian.com/ex/… gateway URLs colleagues can’t open, and a tool result’s self link is a REST endpoint).

Skills are therefore parameterized. A skill writes ${name} where it needs an operator-supplied value, and the founder sets the value once in company.yaml:

skill_variables:
jira_base_url: "https://nimbus.atlassian.net" # values support ${ENV}
confluence_base_url: "https://nimbus.atlassian.net/wiki"

The bundled platform_mentions skill uses them:

- **Work item:** `${jira_base_url}/browse/{ISSUE-KEY}`
- **Page:** `${confluence_base_url}/spaces/{SPACE}/pages/{page-id}`

(See Confluence / Jira for where each base comes from.)

The engine substitutes the variables at render time, everywhere a skill’s summary / title / body reaches the LLM — the per-phase catalogue, the load_tool_skill result, and the required-skill block message. The rule lives in the knowledge-base-authored skill page; only the value comes from config. The engine carries no integration-specific code — skill_variables is generic, so the same mechanism serves a repo base URL, a support email, a runbook root, or any other per-org fact a skill wants to name.

Substitution grammar (deliberately narrow):

  • Only the braced identifier form ${name} is substituted, and only when name is a declared variable. A bare $name, a literal $$, the agent-facing single-brace placeholders skill bodies use ({project-uuid}, {page-id}), and an unknown ${other} are all left byte-for-byte unchanged. This is why skill prose dense with shell / regex / currency $ is safe, and why an unchanged variable map keeps catalogue summaries byte-stable for prompt prefix caching.
  • Keys must be identifiers (^[A-Za-z_][A-Za-z0-9_]*$), validated at config load so the operator key-space matches the render grammar exactly (a key like base-url is rejected up front rather than silently never matching). This identifier-only grammar also means a ${name} can never name a config path or traversal — there are no paths in the flow, just a flat operator map.
  • A variable that resolves to empty (unset ${ENV}) is dropped, so its ${name} reference renders as the literal ${name} — visibly broken and debuggable, never a silently malformed value.
  • Because that literal only ever shows up inside an LLM prompt (which no operator reads), the registry also emits a skill_variable_unresolved warning whenever a registered skill references a ${name} missing from the map, checked whenever a walk or a single-page update registers a skill, and re-checked for every skill on each config apply, so a dropped variable surfaces in the logs immediately.

What this fix is — and is not. Skill variables are enforced-reading guidance: the required-skill guard guarantees the rule and the correct base URL are in the agent’s context before it can post, but the engine deliberately does not rewrite MCP tool results (that would be content-mangling middleware around a tool stack it stays agnostic to). If your MCP server itself returns unshareable links in its tool results (e.g. mcp-atlassian under cloud_id auth builds every result link from the gateway CONFLUENCE_URL), the deterministic complement is to fix that in the server so its results carry the human-readable base — the prompt-layer rule then remains as defense-in-depth and continues to cover links the agent composes itself (Jira browse links are always composed: Jira tool results only carry REST self-links).

Trust + secrets. Values are set by the Tier-B config operator; a Confluence skill author (lower trust) can only reference keys, never define values, and an undeclared reference renders inert. Because resolved values render into prompts — and, like any prompt content, into the durable event store and dashboard — treat a skill_variables value exactly as you would a policies or backstory string: do not reference secrets. The map is an explicit allowlist of the handful of values the operator deliberately exposes, which is why it is safer than letting skills reach into arbitrary config.

skill_variables is company-wide and hot-reloadable: changing a value on PUT /config re-renders every skill on the next turn.

${crewlet_base_url} is the engine’s own. It carries where people reach this deployment’s dashboard, from integrations.dashboard_base_url resolved like any ${VAR}, so a skill links to a work item as ${crewlet_base_url}/#/work/{KEY} and to a page as ${crewlet_base_url}/#/knowledge/pages/{page-id} with nothing for an operator to keep in step. The name is reserved — a company declaring it is refused, since two sources for one address is how a link comes to point at the deployment the company used to run on. It is never taken from integrations.public_base_url, which is where vendors reach the deployment and, behind a public listener, a socket that serves no dashboard. With no dashboard address set it is left undefined, so a skill renders it literally and the registry warns skill_variable_unresolved.


A skill is catalogued in a phase’s prompt when its declared phases includes that phase and its trigger matches the active surface:

PhaseTool surface seenNotes
Executeevery first-party tool plus whatever the executor has activated, and the names of the MCP servers whose tools reach the seatThe executor’s live surface, so a skill for a tool it discovered mid-run is catalogued from the next round. run_sandbox is on it only for a seat whose role.sandbox is enabled and whose executor is not itself a coding agent in agent mode, so a tool: run_sandbox trigger reaches exactly the seats that can launch a run.
Reviewempty (Review has no domain tools)MCP-server-keyed skills still appear when an operator scopes a skill to the Review phase (lists it in the skill’s phases), even though Review has no domain-tool surface.
Sub-agentthe delegated task’s allowlistSame matching as Execute, against the worker’s narrower surface. subagent is the wire name of the worker phase

execute, review and subagent are the only names a skill’s phases: accepts; an unrecognised one is a parse error rather than a skill offered in no phase.

Within a phase, catalogue entries are sorted in alphabetical key order so the prompt prefix stays byte-stable across turns (LLM prefix-cache stability).

Catalogue position: skills land immediately before the ## Available tools block in Execute — one conceptual section, how-to then names — and immediately after the review header in Review, so guidance precedes the evidence it is meant to be weighed against. Review’s catalogue has its own header: it asks the reviewer to weigh the work against the conventions listed and names no loader, because the reviewer’s only tool is its submission.

Every rendered catalogue is recorded. Each prompt that carries a catalogue — the executor’s, the reviewer’s, and each delegate worker’s — publishes one knowledge_read with via: skill_injected, naming the page behind every skill it listed, so “which skill pages reach our agents” counts the pages a catalogue shows as well as the ones a model loads.


Catalogue entries carry only the summary. The full body reaches the LLM via load_tool_skill(key) — a built-in tool registered alongside lookup_colleague / use_skill / etc. The LLM calls it when the summary’s hint suggests detail it needs:

load_tool_skill(key="mcp:github") → returns the full body as a tool result

Available in Execute (always-on: it is a builtin, so it needs no activate_tool promotion) and Sub-agent (same as Execute). Returns an error with the list of registered keys if the key doesn’t exist, so the LLM can recover.

Review intentionally doesn’t have it — Review’s contract is the decision enum, not domain action.

The engine deliberately does not auto-inject skill bodies. All loads are LLM-driven so the prompt prefix stays small and the model’s tool-call log is the complete record of what guidance it consulted.

Each load publishes a skill_used naming the page the skill was read from (source_page_id, source_container) — a key lives inside a page and moves when the page is edited, so the page is what a load is attributed to — and, beside it, a knowledge_read with via: skill_loaded against the same page.


Required skills — load-before-use enforcement

Section titled “Required skills — load-before-use enforcement”

A skill is practices for the tools its trigger names — and in practice models sometimes skip the load_tool_skill call and use the tool without reading them. Skills are therefore enforced by default (required: true): the guard enforces the load in code, not in prompts. Every tool call in Execute and Sub-agent sessions passes through the engine’s dispatch gate; a call to a tool the skill’s trigger covers is rejected until the LLM has successfully called load_tool_skill(key) in the same session:

Tool 'create_pull_request' is gated behind a required tool skill you
have not loaded in this session:
- 'mcp:github': Read / search / review code + issues + PRs. Author code via the sandbox …
Call `load_tool_skill(key='mcp:github')` to read the required
practice(s), then retry 'create_pull_request'. Loading is needed once
per session; after that the tool works normally.

The blocked call never executes; the LLM loads the skill on the next round and retries.

Session scope, not turn scope. “Loaded” means the body is in this model’s context, and the executor, the reviewer and each worker are separate message histories — so a skill loaded by one is genuinely not in front of the others, and each session loads it itself. A round-cap extension continues the same session and keeps the load; a self_iterate starts a fresh session and therefore a fresh guard, because its context started over too. A suspended executor is the exception in the other direction: it is re-entered as the same message history, possibly on another node, so its loaded keys ride the saved conversation and the rebuilt guard starts from them.

The exempt set is about deadlock, not policy. load_tool_skill itself, the discovery meta-tools, and the phase submitters are never blocked, whatever a trigger says: gating the unlock would make a session unrecoverable, gating discovery would add rounds without protecting anything, and gating a submitter would brick the phase. A misauthored trigger can cost a phase some tools; it can never cost the phase. Every block also publishes a phase.tool_skill_blocked event (agent, phase, tool, skill keys, turn id) so operators can see agents attempting to skip required practices.

Mark a skill required: false when you want its content available as a pure hint — catalogued exactly as before (summary inline, body on demand) but never blocking anything:

required: false

Controlling the cost — keep enforced triggers narrow, broad bodies tiny

Section titled “Controlling the cost — keep enforced triggers narrow, broad bodies tiny”

Because enforcement derives from the trigger, the per-session cost of an enforced skill is governed by two knobs the author controls:

  • Trigger width decides which calls wait. Prefer exact tool names — the tools the practices actually apply to. The bundled skill:platform_mentions triggers on the write tools only (posting broken mention markup is visible to humans; Atlassian/Mattermost reads never wait on it) — its leaves name the write tools of mcp-atlassian and the version-pinned mcp-server-mattermost==0.5.1 release (jira_add_comment, confluence_create_page, mattermost_post_message, …), which is exactly why the example org pins the server versions it can: an unpinned upgrade could rename a tool out from under the gate. tool:refine_skill triggers on the single refine_skill builtin (no other tool waits on skill-refinement conventions).
  • Body size decides what each wait costs. An enforced mcp_server trigger gates every tool on that server, so its body must be a one-pager: the bundled mcp:github orientation is ~100 tokens, making the once-per-session load before any GitHub work near-free. If a server-wide skill’s body grows past orientation size, split it — keep the tiny server-wide overview and move the detailed practices into a narrow tool-triggered skill.

Enforcement gates exactly the tools the trigger names:

  • tool: <name> — that one tool.
  • mcp_server: <server> — every tool served by that MCP server.
  • any_of / all_of — the union of their leaves’ tools. An all_of skill only gates when its full trigger matches the session’s surface (a role with only one of the two servers is not the skill’s audience), but once matched, tools from each named surface are covered.

Session scope — why “per phase”, not “per turn”

Section titled “Session scope — why “per phase”, not “per turn””

“Loaded” is tracked per LLM session: one phase’s message history. The executor, the reviewer and each worker run on separate message histories, so a body loaded by a worker is not in its parent’s context — and the same applies to every self_iterate iteration (each starts a fresh phase session). This is deliberate: the entire point of a required skill is that the practices are in the context window of the model actually making the call. Round-cap extension loops continue the same message history, so loads carry across extensions without re-loading.

  • Catalogued required skills carry a visible (required — load before use) marker plus a one-line enforcement note in the ## Tool skills section, so the model learns the contract up front; the block message is the recovery path, not the discovery path. Review renders required skills unmarked — it has no domain tools and no load_tool_skill, so nothing is enforced there.
  • Engine plumbing is exempt (load_tool_skill itself, activate_tool, list_mcp_server_tools, submit_work, submit_review) — a misauthored trigger cannot brick a phase.
  • The guard only arms when the session can satisfy it: load_tool_skill is an Execute always-on and rides along on every worker surface that can reach it. A surface without it (e.g. a custom always-on override that removed the loader) disables enforcement for that session rather than soft-locking the LLM, with a skill_guard_disabled_no_loader warning.
  • A failed load_tool_skill call (wrong key, registry error) does not unlock anything.

An enforced skill costs the session one extra tool round plus the body tokens, only in sessions that actually use a covered tool — cheap when triggers are narrow and bodies are sized to their blast radius (the tokens are the point: the practices end up in context before the call). When several enforced skills cover the same tool, the block message lists every missing key and the LLM can load them all in a single round of parallel load_tool_skill calls. Most bundled examples ship enforced: narrow practices like skill:platform_mentions gate only their exact tools, and the server-wide enforced skills (mcp:github / mcp:gitlab) keep their bodies to ~100-token orientation one-pagers so the per-session load is near-free. Three bundled skills are advisory (required: false): skill:code_runtime (advice on briefing run_sandbox, not markup correctness, triggered on that tool so it reaches exactly the seats offered it, whichever code host they push to or none), and skill:getting_unstuck / skill:retrieval_research — each carries a server-wide mcp_server: atlassian leaf in its trigger, and enforcing advisory practice prose at that width would gate every Jira and Confluence read, inverting the trigger-width rule above.


crewlet confluence import is a unified publisher: it routes every .md file by frontmatter and publishes both tool-skill pages and general knowledge docs in one pass — a file with trigger: ⇒ a Tool Skill (this page, → the Tool Skills container); otherwise ⇒ a knowledge doc whose container is its parent directory and title is its first # H1. The two land in different containers with different encodings; everything below is the skill side.

crewlet confluence import company.yaml

The positional config is the Tier B company YAML — the importer reads the backend credentials from its confluence: block (the Tier A bootstrap has no such block and is rejected). Walks every .md file under examples/ (or any path you pass, recursively), and for each skill file encodes it in the backend’s page format and creates a page in the Tool Skills container. Idempotent: pages are matched by their crewlet-skill-key-<key> label — rename-stable, and skipped unless you pass --update.

Useful flags (see the CLI reference for the full per-command tables):

  • --dry-run — print what would be created/updated without making page writes.
  • --update — overwrite existing pages (Confluence keeps the prior version in page history for rollback).
  • --prune — after publishing, delete import-managed skill pages in the space whose source .md is gone (e.g. a renamed or removed bundled skill). Only touches pages the importer itself published — identified by the crewlet-skill marker + per-key label that no local file claims — never user-authored pages or knowledge docs. Combine with --dry-run to preview deletions.
  • --space TS — target a different Tool Skills space than the default TS (skill files only; knowledge docs take their container from their parent directory).
  • --create-space — auto-create any target space that doesn’t exist (requires space-admin on the bot account). Without it, a missing space fails the pre-flight with remediation rather than publishing half a tree.

Publish, then run — two commands, in that order:

Terminal window
crewlet confluence import company.yaml ./skills/ # publish the pages
crewlet run -config crewlet.yaml -company company.yaml

The importer reads the backend’s credentials from the Tier B company YAML, not from the Tier A bootstrap, so it works before a node is configured at all. Running it first on a fresh deployment means the engine’s own boot-time sync picks up the pages that are already there.

Open the page in your browser, edit, save. The Confluence page webhook reaches one node, which reads that page back into its registry and tells the rest of the fleet, and each of them reads the page into its own. The next agent turn on any node sees the new body. No restart, no CLI invocation, no deploy.

Every offer and every load is a recorded knowledge_read (skill_injected and skill_loaded), kept per company day in the replicated usage domain. A skill page’s page answer carries skill_loaded_by — each seat, how often it loaded the body and how often a phase offered the summary — and the dashboard’s Agent skills screen lists every skill with the seats it reached over thirty days. A skill offered constantly and never loaded is a summary that answers the question on its own, or one nobody’s work matches.

If you suspect a webhook was missed during a long outage:

crewlet confluence resync company.yaml

Runs the engine’s own walk and admission test against a throwaway registry and prints the container’s page count, the skills it admitted and any page that declares a trigger and did not parse (which also makes the command exit non-zero). resync is skills-only: knowledge docs are never loaded into a registry, so there is nothing to resync for them. It does not reach into a running engine, and a running engine does not need it to: every node walks its container every 10 minutes (give or take a fifth, so a fleet does not walk in lockstep), so a missed webhook is caught up within 12 minutes without a restart.


The shape is one idea: a machine-readable YAML frontmatter block at the top of the page (edit to change binding metadata: key, trigger, phases, title), followed by the guidance rendered as a normal page body (edit to change the prose). When the skill sync reads a page back, it parses the YAML and flattens the body HTML to plain text for the LLM. The conversion is intentionally lossy on formatting (bullets and headings flatten to text-with-newlines) because the body’s only consumer is an LLM, not a human reader. Operators who want exact source-text fidelity should keep the .md files in version control and re-run the import when they change; an existing page is updated in place.

Every synced skill records the backend page it came from in skills.Skill.SourcePageID and SourcePageVersion. The page id is the registry’s identity for the skill (every later change to it is addressed by page), and the version is provenance. Confluence stamps its integer page version; a backend without one leaves SourcePageVersion at 0.

Each skill page combines a leading YAML frontmatter code macro (the small box at the top of the page) with the markdown body rendered to Confluence storage XHTML. A walk lists every page in the space through the content API (Client.PagesIn, which pages to exhaustion and refuses a space too large to be a skills container) and decodes the leading code macro back into frontmatter (confluence.DecodeSkillPage); a single-page read does the same for one page id and reports the space the page lives in now. Confluence’s auto-generated space home page and other non-skill content declare no trigger, so admission skips them as ordinary pages. The crewlet-skill label is the importer’s provenance marker, read only by -prune; a skill page written by hand in the space loads without it.


SettingDefaultDescription
knowledge.skills_containerTSConfluence space key the engine watches when Confluence is the knowledge backend — read by the skill sync, the searcher’s result exclusion and the parser’s routing exclusion alike.
-space <key> (CLI flag)the config fieldPer-invocation override of the Tool Skills space for crewlet confluence import / resync (skill files only — knowledge docs take their container from their parent directory).
CREWLET_TOOL_SKILLS_SPACE (env var)—The flag default for those two commands, so an operator running them repeatedly does not retype the space. Nothing else reads it.

The field is three-valued, and the empty string is an answer. Absent takes the reserved default TS; a named space takes that space; an explicit skills_container: "" turns tool skills off — no sync, no routing exclusion, no search exclusion, and an import that meets a skill file refuses rather than filing it as prose. The off switch exists because the default reserves a real space key: a company whose ordinary work space happens to be TS would otherwise have it silently dropped from every knowledge search and every routing decision, with no way in the config to say otherwise.

The engine reads the config field and only the config field. The environment variables are flag defaults for the operator commands and nothing more: a fleet whose nodes each read a container out of whoever’s shell started them would disagree about which one holds the skills, and the symptom is agents on one node following guidance the others have never heard of. A routing decision belongs in the versioned document describing the company.

The Tool Skills container holds engine-managed scaffolding, not general knowledge. Crewlet does not maintain a synced knowledge index, so there is nothing for the container to pollute; just don’t add it to the knowledge read scope (knowledge.scope) and skill pages won’t surface in the ## Relevant knowledge query-time search — the searcher drops the skills space from results wholesale as well, since a skill page is machinery rather than knowledge.

The container is also excluded from notification routing. Webhooks for tool-skill pages still drive the in-memory registry update via the engine’s skill-sync callback, but both transports short-circuit the recipient-routing path (set_notification_excluded_spaces / set_notification_excluded_projects) so engine-managed pages don’t surface as notification_undeliverable warnings or emit spurious notification_skipped events. Page edits in the Tool Skills container have no human or agent recipient by design — only the engine consumes them.


What happensEffect
Knowledge backend unreachableA walk that cannot enumerate the container completely replaces nothing (a partial walk must not silently delete skills): the registry keeps what it holds, which at boot is empty, and the node logs tool_skill_sync_failed with the attempt number and when it retries. Retries start at 5 seconds (sized for a wiki whose API comes up seconds after the engine, and a rate-limit refusal) and double to the 10-minute walk interval, so the ordinary boot race heals itself and an outage costs one walk per interval. Applying the configuration again retries at once.
Webhook lost, or a page moved or restored (events the engine does not subscribe to)Every node’s periodic walk catches it up within 12 minutes (a 10-minute interval, jittered by a fifth). crewlet confluence resync is a read-only diagnostic that prints what the container holds, and does not reach into a running engine; see Drift recovery above.
A node misses a peer’s tool_skill_page_changed (a broker reconnect, or tool_skill_nudge_unavailable at boot)That node converges on its periodic walk; the node that heard the webhook, and every node that heard the broadcast, already have the change.
This node cannot read the skills sourceLogged as tool_skill_source_unreadable with the reason: no integrations.confluence.token (the skills space is read with the organization’s credential), a Confluence integration that did not start, or a revision that moved the company onto the native knowledge base on a node already running the native tracker without it — the runtime’s halves are fixed by the first company that started it, so that change needs a restart. (On a node running no native backend at all, the apply that moves the company onto one starts it, knowledge base included.) The node keeps serving what it last read from that source (nothing, for a source it never read) and walks nothing until an apply or a restart fixes the reason; the apply that does walks the source at once.
Skill body over the 32 KB capThe page fails validation and is skipped with skill_page_undecodable, which names the page and the cap; other skills load normally.
Page is missing the leading YAML frontmatter block, or its frontmatter names no triggerNot a skill: it is counted as an ordinary page and skipped without a warning, and a page that was previously a skill is dropped from the registry. A page that declares a trigger and does not parse is logged as skill_page_undecodable and dropped the same way.
Two pages declare the same keyThe page with the lower id (on Confluence, the older page) is served on every node, and skill_key_duplicated names both pages. Give the shadowed page a key of its own or remove it.
Misauthored skill bodyDegrades every agent on the matching tool/MCP surface on the next webhook tick. Use Confluence’s native page history to roll back, or re-push the source file with crewlet confluence import, which updates the existing page.
Chronic phase.tool_skill_blocked events on one skillThe model keeps trying the tool before loading the required skill. Recovery works (the block message names the key), but each block wastes a round — rewrite the catalogue summary so the load happens proactively, or narrow the trigger if the skill is over-scoped.
Required skill on a surface without load_tool_skillThe guard is not armed for that session instead of soft-locking it. Execute always carries the loader and a worker receives it whenever its parent can reach it, so this affects a worker whose parent lacks the loader and the reviewer, which has no domain tools to gate.
Skill references a ${var} missing from skill_variablesThe literal ${name} would only ever surface inside an LLM prompt, so the registry logs skill_variable_unresolved (skill key + variable + field) whenever a walk or a single-page update registers a skill, and re-checks every skill on each config apply.
Trigger names a tool that exists nowhere (e.g. an upstream MCP server renamed it)Trigger matching is exact-string, so the skill silently stops cataloguing and, if required, stops gating. The engine checks every skill against the current epoch’s tool surface (skills.Registry.Audit) after every change the skill sync makes to the registry (a walk or a single-page update, whatever asked for it) and whenever an epoch becomes current (at boot and on every config apply, which is where an MCP server is added or removed), always against the epoch serving at that moment rather than one an apply is still building, so the last change to either side is checked against the other: a partially live skill with a dangling tool name logs skill_trigger_partially_dangling (warning: near-certain name drift), while a skill whose whole trigger matches nothing logs skill_trigger_matches_nothing (info: plausibly authored for a stack this org doesn’t run).

  • Agent Runtime — where the per-phase prompt builders live; how the registry is threaded into the turn engine.
  • Turn Engine — phase contracts and the tool surface each one gets.
  • CLI Reference — full flag reference for crewlet confluence import / crewlet confluence resync.
  • Environment Variables — CREWLET_TOOL_SKILLS_SPACE (import/resync flag default; the engine reads knowledge.skills_container).

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.