Skip to content
You are reading documentation for unreleased main. Read the 0.1 version.

CLI Reference

Crewlet ships one command, crewlet — a single self-contained binary. Every subcommand below is served by it.


CommandDescription
crewlet run [config.yaml]Read Tier A bootstrap (positional, or -config; default ./crewlet.yaml), connect to DB, run engine; falls into unconfigured state if no active revision
crewlet validate [file.yaml]Validate a Tier A or Tier B YAML and print a summary (-json for located, classified problems and warnings); with no positional it checks both tiers via -config and -company
crewlet migrate [config.yaml]Apply pending schema migrations (Tier A file, default ./crewlet.yaml). Every process migrates on open, so this is a way to do it without starting one — -check reports pending work and exits non-zero without applying it
crewlet budgets show [config]Print each scope’s day, ISO week and month on the company clock — spend, ceiling (unlimited where none), the engine’s STATE (ok, near, refusing), when the window turns over and LAST REFUSED — read from a running node, because the counters are the fleet’s and not this file’s. There is no reset: a window’s allowance comes back when the window turns over
crewlet backup -dir PATH [config]Copy a running node’s store files and its stream estate into one verified directory on the engine’s host — the only way to copy either, since the store is locked to that process and the embedded broker binds no socket. See Backups & Restore
crewlet retention status [config]What each domain’s log is holding, what the trim concluded and which of the six terms is stopping it, every node’s position, and what this node costs to replace. Exits non-zero when any alarm is active, printing each one’s measurement and remedy on stderr — the hook for your own cron
crewlet retention snapshots [config]The per-node snapshot inventory: what each machine holds, per domain, how old and how large — or why it holds none. The question you ask when a join fails
crewlet retention ack -stream NAME -position NPublish an operator backup floor, for backup_floor: operator. It exists because the engine cannot see a copy that has left the host
crewlet retention evict <node> -confirm <node> [-op-id ID] [-force]Stop a node’s records applying on every log the trim counts nodes on — the tracker’s and the pages log — so the trim can pass a floor an absent machine is pinning. Refused, with nothing written, while the node still holds a live presence lease, or while its lease cannot be read (-force overrides both). Prints one line per log with its outcome and the watermark before and after; exits non-zero when not every log holds the record, printing under each unfinished log what finishes it — the command with -op-id where running it again can, and what to do instead where it cannot
crewlet retention readmit <node> -confirm <node> [-op-id ID]The inverse commit, on the same logs. Refused while the node has not applied every record up to the one just before the higher of the trim floor and the first surviving sequence of the tracker’s or the pages log, and the refusal prints the numbers. Nothing is written to either log on a refusal; a partial readmission is finished like an eviction
crewlet retention set-capacity <stream> <bytes> -confirm <bytes>Change a log’s byte ceiling, up or down. Runs inside a fleet-wide maintenance window and costs three restarts, because a log’s Tier A ceiling is only the value its stream is created with. A raise the broker has no room for is refused before the window opens, and one it refuses at the apply is reported as a refusal rather than as an unknown outcome
crewlet retention maintenance status|abandon|exclude -stream NAMEWhere that window stands, who has not acknowledged, and the two gestures that act on it
crewlet retention reanchor -stream NAME -confirm <created_at> [-force] [-discard]Adopt a recreated stream, or a broker restored from an older copy: move that one log to its next generation, declaring every position below it comparable and safely stale, and resume its applier with no restart. A recreated log is followed from its first surviving record, a restored one from its end, and one continuing in a generation only an evicted peer held from this node’s own checkpoint, that generation’s records void. A restored log holding records written after the restore that this node’s rows do not hold is refused unless -discard accepts that they are applied on no node
crewlet retention verify --restore -dir DIRRestore the newest artefact and open the copy. Exits non-zero past its cadence — the cron hook that turns a lapsed restore test into a failing check. Talks to no node
crewlet work purge <task-id> -project KEY -reason TEXT -confirm <task-key>Destroy a task and every row it produced, on every node. The one operation with no inverse, restricted to a person or an operator token. Its children move onto its own parent rather than being destroyed with it
crewlet objects status [config] [-json]Where the company’s files are kept — the object store’s backend, nats or an S3 bucket — read from a running node, the node that ran the collector’s last passes, and what each found: objects listed and deleted and unfinished uploads abandoned by the last collection; files named, missing and damaged by the last audit to finish, with those files listed. -json prints the fleet view’s objects block as the node answered it
crewlet fleet broker list [config] [-json]The fleet broker’s membership: each live node’s broker kind (member, leaf, client, or unknown for a kind this build does not know) beside how the JetStream metadata group counts it — read through a member — and, in words, every disagreement: a member gone for good that the group still counts in every election, with the command that removes it
crewlet fleet broker remove <node> -confirm <node> [-force], or -peer <peer> -confirm <peer>Stop the metadata group counting a member that is gone for good, through a live member’s system account — by node id, or by the peer id list shows for a voter no member can name. Refused while the node holds a live presence lease as a member; -force is for a member wedged in a way that still renews it
crewlet seats pause <handle> [-stop] [-reason TEXT]Pause an agent seat: it starts no new turn, its mail waits in order and its scheduled runs are skipped. -stop also ends the turn it is on at its next round. As the person the token is bound to
crewlet seats resume <handle>Lift the pause; what waited is delivered first, in order
crewlet schema [company|bootstrap]Print the JSON Schema for a config tier (editor autocomplete, CI, AI-assisted authoring)
crewlet config import <company.yaml>Load Tier B YAML, activate as a new company_config revision
crewlet config export [--revision <UUID>]Dump the active (or specified) revision as YAML to stdout
crewlet config showOne-line summary of the active revision
crewlet config revisions [--limit N]List recent revisions (newest first)
crewlet config diff <UUID> [-against <UUID|active>]Structural diff of two revisions — paths and values, always redacted on both sides
crewlet config activate <UUID>Re-point the fleet at a revision; re-activating the current one mints a new epoch, which is how a rotated secret takes effect
crewlet config sealEncrypt the active revision as one document under the Tier A keyring (one-time migration off plaintext-at-rest) — see Secrets
crewlet config rekey [-dry-run]Re-encrypt the active revision’s config document under the active key (master-key rotation)
crewlet secrets keygen [-key-id ID]Generate a fresh encryption-keyring key + the crewlet.yaml snippet to install it
crewlet secrets set <NAME>Store an encrypted secret in the secret store; the engine resolves ${NAME} from it ahead of the environment
crewlet secrets listList stored secret names + metadata (never values)
crewlet secrets unset <NAME>Remove a stored secret
crewlet secrets get <NAME> -revealPrint one stored value to stdout — break-glass, audited, CLI-only
crewlet secrets rekey [-dry-run]Re-encrypt stored secrets under the active keyring key
crewlet search eval [-store PATH]Measure the two-stage semantic search against the exact scan, on the vectors a store file actually holds. Ground truth is the exact scan’s own top-K, so nobody authors a judgement; exits non-zero below the floor for that corpus size
crewlet llm listEvery cli-agent provider the company declares, with its CLI, model and how it is signed in
crewlet llm doctor [KEY]Verify a subscription backend end to end — the CLI is installed, the login answers, a real completion returns, the CLI’s own shell is refused and its web tool reaches the network — and an anthropic entry against its model: the Models API’s record compared with the request shape the engine sends, and one real round that must come back as a tool call (-no-smoke stops before every real completion)
crewlet llm login <KEY>Establish the vendor’s own login for a provider: brokered interactively, -from-host to adopt one this machine already has, -capture-token to mint a headless token into the secret store (add -print-token to send it to stdout and store nothing), -token-stdin for one you already hold
crewlet llm status <KEY>Ask the CLI who it is currently logged in as
crewlet llm logout <KEY>Revoke locally and delete the provider’s credential files
crewlet llm export <KEY> [-secret-store]Pack the login into one portable blob — stdout, or the secret store under the name the engine restores from on a fresh host
crewlet llm import <KEY>Restore a bundle from stdin onto this host; refuses to overwrite a login that is already there
crewlet confluence import <company.yaml> <directory>Publish a directory of authored markdown into Confluence spaces — one space per directory, plus the tool skills the files themselves declare. Every target space is checked before a single page is written
crewlet confluence provision <company.yaml>Register the inbound Confluence webhooks and mint the credential each one carries. Cloud gets one token-bearing hook per event, Data Center one signed hook for all of them, and a re-run converges what is there rather than adding to it
crewlet confluence resync <company.yaml>Re-run the engine’s own tool-skill walk of the Confluence skills space against a throwaway registry and print what loads — a read-only diagnostic, not a way to change a running engine
crewlet slack provision <company.yaml>Create, update and install one Slack app per agent seat from the canonical manifest, minting each seat’s bot token and signing secret into the ${VAR}s its config points at. The install itself is an OAuth grant, so the run hands the operator one authorize URL per seat and takes the code back
crewlet jira provision <company.yaml>Report a Jira instance against the config: which account each seat’s own credential authenticates as, whether every project the org chart names exists and agrees about its lead, and — on Data Center — register the inbound webhook with a minted secret. Jira issues no credentials on a provisioner’s behalf, so this run reports far more than it changes
crewlet gitlab provision <company.yaml>Reconcile the config into GitLab: one service account per agent seat, membership, per-agent PATs minted into the config’s own ${VAR} references, and the group webhook. A re-run leaves a working token alone; -dry-run reports without touching anything, and a run that cannot record what it minted revokes it.
crewlet github provision <company.yaml>Report a GitHub deployment against the config — which account each seat’s own credential authenticates as — and register the inbound webhooks with a minted secret: one on the organization where the credential may, otherwise one per named repository. GitHub issues no credentials on a provisioner’s behalf, so this run reports more than it changes
crewlet mattermost provision <company.yaml>Create/update one Mattermost bot account per Mattermost-enabled agent, add it to the team + channels, mint its access token into an env file or the secret store. Keeps each bot’s display name in step with the company document. A re-run leaves a working token alone; -rotate mints fresh ones, -handles a,b narrows the run, -decommission disables departed seats’ bots. Runs a preflight first: system-admin role, EnableBotAccountCreation, EnableUserAccessTokens.
crewlet mattermost doctor <company.yaml>Check a Mattermost install end to end: reachability, the Site URL every browser inherits, a browser-shaped websocket upgrade, and one real authenticated socket per agent seat
crewlet --versionShow the installed version

Every command that reads the Tier B company document takes -config (default ./crewlet.yaml), and resolves its ${VAR} references the way the engine does: the secret store first, the process environment behind it. A command that read the environment alone would see an empty string for every value already rotated into the store — and for integrations.gitlab.signing_secret, empty is the signal to mint, so a re-run would replace a working webhook secret at the third-party app. With no bootstrap at that path, or one declaring no secrets.keys, the run resolves from the environment alone and says so on its first line. The one exception is the operator’s own credential (-admin-token / $GITLAB_ADMIN_TOKEN and its siblings), which is read from the environment only — see the secret store.

Every command except crewlet run logs at warn. They open a store, which logs a migration line per schema file and an open line per call — noise on a one-shot command whose stdout is meant to be piped, read or diffed. That is a default, not a ceiling: export CREWLET_LOG_LEVEL=debug (or info / error) to turn it up, which is exactly what a half-applied migration or a failing deploy gate needs, and CREWLET_LOG_FORMAT=json to change its shape for a collector. CREWLET_LOG_FILE is the third of them: it appends what the command logs to a file, beside stderr, so a CI step can keep a crewlet migrate in the same durable record as the node it is migrating for. It never touches what a command prints — stdout is still yours to pipe or diff. They are environment variables rather than flags on a dozen commands because they belong to the invocation — a CI step exports them once and everything it runs answers. A value this build does not recognise resolves to the default (warn, console): a bad log level must never be why an operator cannot run a migration, and it must not quietly change the default either. A log file is the exception, because a path is not an enum with a default: one that cannot be opened fails the command. crewlet run has its own -log-level / -log-format / -log-file / -debug flags, its own default of info, and reads the logging: block from its Tier A file — it ignores all three variables. Colour is CREWLET_LOG_COLOR / NO_COLOR everywhere — see Environment Variables.

crewlet run [<config.yaml>] [-company PATH | -import-company PATH] [-debug]
[-log-level LEVEL] [-log-format FORMAT] [-log-file PATH]
[-mode MODE]
[-roles ROLE[,ROLE...]] [-api-host HOST] [-api-port PORT]

Reads Tier A bootstrap and starts the agent engine.

Tier B is not read from a file at runtime. A running node serves the revision the fleet’s activation pointer names, so a Tier B file on this command line is only ever a way of getting a document into the store — and the two flags above are the two reasons to want that. The file is read against the rules a running company depends on, and against the admission rules only when it is actually written as a new revision (-company into an empty store, -import-company over a different company): a file that is already the active revision, or a bootstrap the store’s own company outranks, starts the node even when it carries a duplicate name. To change a running fleet with no restart at all, use crewlet config import, which goes through the node’s API. The path comes from the positional argument, or from -config, defaulting to ./crewlet.yaml. Naming it both ways is refused: the two would have to agree and nothing checks that they do. A leftover positional is refused too, rather than ignored — Go’s flag parser stops at the first non-flag token, so a command that took the path and kept going would silently boot from the default without ever mentioning the file the operator named.

If the default is missing and a config.yaml sits beside it, the error says so. config.yaml is the name this project’s own guides used to give the Tier A document, and the bundled example still carries it in examples/nimbus.config.yaml, so an operator with a file written against that guidance gets “no such file” about a name they never typed. It is a hint, not a fallback: silently loading a file nobody asked for is how a node boots from the wrong document on a machine that has both. Tier B is read from the company_config table in the store — if no active revision exists, the engine boots in the unconfigured state with the API still serving so an operator can bootstrap via crewlet config import or PUT /config. See Configuration concept doc.

FlagDescription
-config PATHTier A: this node’s broker, store and API (default ./crewlet.yaml)
-company PATHTier B bootstrap seed (default ./company.yaml): imported only when the store holds no company yet. Once one exists this file is ignored, loudly (company_seed_ignored at warn), so a restart with a stale file never reverts a live change. Absent at its default is fine — the node boots on whatever the store holds. A revision the seed writes is the node’s own: its created_by is the node’s id and its created_by_kind is node. On a deployment whose Tier A names api.auth.company_writers the seed is never imported, empty store or not, and says so the same way: the managing system writes the company.
-import-company PATHTier B to make the active revision now, over whatever the fleet is running. The deliberate “this file is the company again” gesture. Mutually exclusive with -company; both together is refused, because they ask for opposite things. Refused before the node starts when Tier A names api.auth.company_writers: the file presents no credential, and the managing system would replace it at its next reconcile.
-log-level LEVELdebug, info (default), warn or error. Overrides logging.level in Tier A, and only when actually given. A typo resolves to info — a bad log level must never be why a company will not boot.
-log-format FORMATconsole (default), text or json. Overrides logging.format in Tier A, and only when actually given. console is columns and colour for a person; text is slog’s key=value; json is one object per line for a shipper. A typo resolves to console.
-log-file PATHAlso write the log to this file, overriding logging.file.path in Tier A, and only when actually given. It is a second destination: stderr keeps every line. An explicit -log-file "" writes no file for one run, whatever the Tier A document says — unless that document also set logging.stderr: false, which would leave the node writing its log nowhere; that combination is refused by name. It moves the path only — the file’s shape, level and rotation caps stay Tier A’s. A path that cannot be opened fails the command rather than resolving to a default: a durable record an operator did not get, with nothing saying why, is worse than not starting.
-debugShorthand for -log-level debug; wins if both are given. It only ever raises — to quieten a file that sets logging.level: debug, pass -log-level info. It makes the engine verbose and not the embedded broker: nats-server’s own debug output is stream.debug in Tier A, off by default, because it is per internal-client rather than per event and the engine’s coordination reads produce a constant stream of it. The broker’s warnings and errors reach the log at either setting.
-api-host HOSTBind address, overriding api.host
-api-port PORTBind port, overriding api.port. With the ingress role it carries the whole HTTP surface (less the routes api.public.port moves to its own listener); without it, only the probes — /health and /ready — and a seats node’s agent-mode tool bridge where api.public.port is unset. 0 serves no HTTP at all — no dashboard, no REST, no webhook endpoint and no probe, so every integration goes deaf and no orchestrator can probe the node. That is why leaving the flag off is not the same as passing 0. Either flag is refused when it contradicts the file’s other listeners: 0 beneath an api.public.port, or a port another listener in the file binds on the same address.
-mode MODEmaintenance or seal: boot for a capacity window rather than for service. Both start the broker and no publisher — no seats, no duties, no schedulers, no writer on any surface, no barrier (a linearizable read is refused maintenance), no object collection or audit (each pins the estate with a barrier), and no eviction, readmission or reanchor — while the node keeps its presence lease, and the difference is that maintenance may write stream configuration while seal may not, which is exactly what makes a seal-mode acknowledgement evidence. Leave it off for a node in service; a node in either mode refuses to run a company.
-roles ROLE[,ROLE...]What this node runs, overriding node.roles: data (hold the company’s durable state), ingress (serve the HTTP API and its webhooks), seats (claim seat leases and run agents), workers (the company-wide singleton duties). Default: all four — one process running a whole company. An unknown name is rejected rather than dropped, because a typo would otherwise produce a node that runs nothing and reports itself healthy. The flag is applied after the file validates, so what the roles require of the file is checked again here: dropping data from a node whose store is not store.scratch is refused, naming the setting. See Running a Fleet.

The logging flags override the Tier A logging: block only when they are actually given: a flag carries its default whether or not anyone typed it, so applying them unconditionally would pin every node at info and make the file’s own setting dead on arrival. -log-file needs that distinction in both directions — its default is the empty string, which is also how an operator says “no file for this run”.

The four overrides — -roles, -api-host, -api-port and -log-file — are the fields whose right value depends on where the process is running rather than on what the company is: which job this node does, where its HTTP surface binds, and which path on this host its log lands on. Everything else in Tier A belongs in the file, where it can be reviewed — including the log file’s own shape and rotation caps, which describe the disk rather than the invocation.

Publishing knowledge and tool skills is its own command rather than a flag on run: an engine that published on every boot would rewrite a company’s knowledge base from whatever tree the deploying machine happened to have. Use crewlet confluence import, and crewlet config import to load the first company revision.

Press Ctrl+C for graceful shutdown — signals escalate in two tiers:

  1. First Ctrl+C: graceful. Every held seat quiesces, running LLM turns finish their rounds (there is no internal timeout, and drain_in_progress logs the in-flight count every 10 s), and turns still queued behind the concurrency limit are returned to the broker for prompt redelivery. The HTTP surface stays up for the whole drain and closes only once it completes: /health stays 200, /ready answers 503 naming the drain, and the dashboard and every read keep answering, so the drain can be watched there as well as in the logs. What it stops doing, from the first moment, is taking work: every webhook, config, secret and setup write, operator write and backup answers 503 with {"error": "draining"} and a Retry-After, because the drain waits on in-flight turns and a surface still accepting deliveries would keep making more.
  2. Second Ctrl+C: the process exits immediately. The first signal hands signal handling back to the operating system precisely so this works: in-flight turns are killed with their triggers unacknowledged, and the broker redelivers them once its ack window elapses.

There is no third tier, because the second is already an unconditional exit rather than something the engine has to be well enough to perform.

Under an orchestrator, SIGTERM follows the same tiers; the host’s grace period (k8s terminationGracePeriodSeconds, systemd TimeoutStopSec) is the SIGKILL backstop. See Graceful shutdown.

When piping output, prefer tee -i — Ctrl+C reaches the whole pipeline, and a plain tee dies on the first press, taking the drain logs with it (the shutdown itself is unaffected).


Manage the Tier B company configuration in the store. Every subcommand opens the store named by the Tier A bootstrap (-config, default ./crewlet.yaml), and decrypts what it holds with that file’s keyring.

Revisions are immutable and the activation pointer is append-only. Nothing here edits a revision: importing writes a new one, activating appends to the pointer. That is what makes re-activating an unchanged revision meaningful — it mints a new epoch every node is watching, which is how a rotated secret reaches a running fleet without a restart.

crewlet config import <company.yaml> [-config PATH] [-api URL] [-summary STR]

Validates the Tier B YAML and writes it as a new active revision, recording the previously-active revision as its parent_revision_id.

It reaches a running node. The store is exclusive to one process, so against a live engine this cannot open the database — and it no longer needs to: it detects the held store and goes through that node’s PUT /config instead, which stores the revision and activates it fleet-wide, so every node converges with no restart. -api URL names a node explicitly, which is also how this works from a machine that is not the node at all. This is the same routing crewlet secrets does for the fleet’s secret store, and for the same reason: both estates live inside the engine’s process.

Through a node, the write is a compare-and-set like every other. If another write activated first, the import says so and names the revision that won; when the node kept the document as an inactive revision it names that too, with the node’s own routes that compare it with what is live and make it live (GET /config/revisions/<UUID>/diff and POST /config/revisions/<UUID>/revert), because crewlet config diff and crewlet config activate open the store the running engine holds.

With the engine stopped it writes to this node’s own store and marks the revision active there, which the node publishes to the fleet at its next start. The line it prints says which of the two happened.

-summary is the audit note recorded with the revision (default imported from <path>). The revision history is the record of who changed what and why, so a fleet-wide write is worth a sentence; created_by is the token’s id when it goes through the API, and the invoking operator when it does not — either way the revision’s created_by_kind is operator, and that author travels with the revision to every node in the fleet.

It does not refuse because a revision is already active, and there is no flag to force it past one: the pointer is append-only, so an import chains a revision rather than overwriting one, and crewlet config activate takes you back to any earlier id.

Idempotent by content, the same rule the boot seed follows: a file that already matches the active revision imports nothing and says so, an edited one imports once. The audit note is imported from <path>, created_by is the invoking operator and source is file — all three are recorded for you and none is settable from the command line. (PUT /config takes an X-Summary header; the CLI has no equivalent.)

Like seal and activate, this writes the revision to this node’s store; the note it prints says what publishes it to a running fleet.

On a managed document — Tier A names api.auth.company_writers — the offline route presents no credential and is refused, naming the writers. Through a node, the node judges the token this command sends (CREWLET_API_TOKEN when set, else Tier A’s first token): one of the writers imports as before, and any other is refused with config_managed, which the command prints with the token it sent.

crewlet config export [-config PATH] [-revision UUID] [-redact]

Dumps the active revision (or -revision <UUID>) as YAML to stdout. It does not emit the stored bytes verbatim: the payload is decrypted with the Tier A keyring, decoded, and re-encoded as YAML, so an encrypted revision exports as the readable document rather than as its {__encrypted__: "enc:v1:…"} blob.

Without -redact this prints secrets in the clear. ${VAR} references export as themselves — they name a credential rather than being one — but a config that inlines a literal key exports that key. This is the one command that will do that, deliberately: it is what a genuine restore or migration needs.

-redact is what you want for anything you are going to share. It masks every field the config types declare as secret-bearing (structurally, not by pattern-matching the text), which is the share-safe dump. crewlet config show and crewlet config diff are always redacted and have no flag to turn it off.

crewlet config show [-config PATH]

Prints the whole active revision as redacted YAML — exactly crewlet config export -redact, with no way to ask for the unredacted form. Use it to read what the fleet is actually running; use crewlet config revisions for the id, author and audit note.

Errors with no revision is active; run `crewlet config import` when nothing is active.

crewlet config revisions [-config PATH] [-limit 20]

Lists recent revisions. The active revision is marked with *.

crewlet config diff <UUID> [-against <UUID|active>] [-config PATH]

Prints a structural diff of two revisions — one line per path that moved, with the value on each side. -against active (default) compares against the currently-active revision, and the direction reads as “what -against became”, so crewlet config diff <UUID> answers “what would reverting to this change”.

--- active
+++ 35ecd5ab-96ff-4d6d-b101-9de41d644832
~ providers.llm.main.model: "claude-opus-5" -> "claude-sonnet-5"
+ agents.roles[3].handle = "qa"
- integrations.github.enabled (was true)

Paths and values, not lines. The stored form is JSON produced by marshalling a struct, so re-ordering a map or adding a field with a default rewrites lines that mean nothing to a reader. The question an operator asks is “what changed about the company”, and paths answer it — this is the same differ GET /config/revisions/{id}/diff serves. A string value is quoted and other values are not, because "true" and true are different settings and a renderer that printed both bare would show a type change as no change at all. Every change, however many there are. GET /config/revisions/{id}/diff cuts its listing at 500 entries because a response body and a dashboard socket frame have a size budget — it reports changes_total beside the listing, so a reader knows what was left out. A terminal has no such budget and does have a pager, so this command prints the whole comparison: the reader best equipped to read a long diff was the one a shared cap kept it from.

Both sides are always redacted, and there is no flag to turn that off. A diff is what an operator pastes into a ticket or a chat thread to ask a colleague whether a change looks right, which is the single most likely way a credential leaves the machine. crewlet config export -revision <UUID> is there for the rare case that needs the real values, and it takes a deliberate act. A changed credential still appears: the revisions are compared as stored, so a rotated key is reported at its path as a change from "__redacted__" to "__redacted__", never with either value.

crewlet config activate <UUID> [-config PATH]

Re-points the fleet at a revision. Every node applies it on its next reconcile.

Re-activating the revision that is already active is not a no-op, and that is the point: the pointer is append-only, so it mints a new epoch. A node’s reconciler skips on the epoch it has applied, never on the payload, so the apply always runs — re-reading the secret store and rebuilding every provider, transport and MCP child that captured a resolved value. It is the documented way to make a rotated credential take effect on a running fleet.

On a managed document that gesture stays open — it changes nothing in the company — and activating any other revision is refused, because it presents no credential and replaces the document the managing system wrote.

crewlet config seal [-config PATH]

Encrypts the active revision under the Tier A keyring and writes a new active revision holding the whole config as one opaque {"__encrypted__": "enc:v1:…"} document — the one-time migration off plaintext-at-rest. ${VAR} references inside are kept verbatim and resolve at construction time. A no-op when the active revision is already sealed. Requires a keyring in crewlet.yaml (crewlet secrets keygen), and says so if there is none.

Like import, this writes the revision to this node’s store; the note it prints says what publishes it to a running fleet.

crewlet config rekey [-dry-run] [-config PATH]

Rotates the master key: re-encrypts the active revision’s config document under the current secrets.active_key_id, writing a new revision. The document is decrypted with whatever key sealed it (that key must still be in secrets.keys) and re-encrypted under the active key.

Run this and crewlet secrets rekey together. They are the two halves of one rotation — the company document and the secret store are sealed with the same keyring — and doing only one, then dropping the retired key, makes whatever the other still holds unreadable.

Workflow: crewlet secrets keygen -key-id <new> → add the new key to secrets.keys and set active_key_id: <new> while keeping the old key → crewlet config rekey and crewlet secrets rekey → once both succeed, drop the old key from crewlet.yaml.

-dry-run reports what would move by reading the key id off the envelope, decrypting nothing. Idempotent: a document already under the active key is skipped and says so. A plaintext revision is refused rather than silently sealed — “rotate the key this is under” and “start encrypting this at all” are different decisions, and the refusal points at config seal. Fails clearly, naming the key, if the document is sealed under one no longer in the keyring.


crewlet secrets keygen [-key-id ID]

Prints a fresh base64 32-byte encryption key plus a copy-pasteable crewlet.yaml secrets: snippet that references it via an environment variable (keeping the raw key out of the file). --key-id (default key-1) names the key; the id is stamped into every envelope the key seals, so keep it stable across restarts and pick a new id only when rotating. Key generation is always explicit — Crewlet never auto-generates a key, because a silently-generated key that isn’t captured makes every backup unrecoverable. See Secrets.

The remaining subcommands operate on the secret store — the company’s encrypted credentials, one sealed value per ${VAR} name, which the engine consults ahead of the process environment. All of them read the Tier A bootstrap (-config, default ./crewlet.yaml) for the keyring, and all of them need one: the store has no plaintext mode, so a config declaring no secrets.keys is refused with a pointer at keygen rather than silently storing plaintext.

Which store they reach depends on whether the engine is running, and the command says which it used. The rows live on the coordination KV so every node reads them, and on the default topology that KV is inside the engine’s own process — so a running node is written through its authenticated /secrets API, and a stopped one falls back to its own local table, which the engine migrates onto the fleet at its next start — replacing a name the fleet already holds only when the fleet’s value was written earlier (see Secret Store). The engine’s exclusive database lock is what tells the two apart, with a pid attached.

-api URL names the node to write through, for running the command from a machine that is not the node. The bearer token comes from CREWLET_API_TOKEN when set, and otherwise from the first api.auth.tokens entry in the Tier A config; the token’s id is recorded as the author of the write.

crewlet secrets set <NAME> [-value V] [-source STR] [-config PATH] [-api URL]

Stores one encrypted secret under NAME, which must be a valid environment-variable name — a name outside the ${VAR} grammar could never be read back, so it is rejected. Existing names are replaced.

The value comes from stdin by default (echo "$TOKEN" | crewlet secrets set GITLAB_TOKEN_SWE), or from an interactive prompt on a terminal. -value exists for scripted use, but an argv value is visible in ps and lands in shell history, so prefer stdin. -source records provenance (default cli); the provisioning CLIs stamp their own. The author is the Tier A token’s id when the write goes through a running node, and $CREWLET_OPERATOR/$USER when it goes to the local table.

A running engine picks the new value up at its next config activation or restart, not immediately — the command says so after each write, along with which of the two stores it wrote.

crewlet secrets list [-config PATH] [-api URL]

Prints one row per stored secret: name, sealing key_id, last-updated timestamp, who wrote it, and source. It names the store it read first, because a stopped node’s own empty table and a fleet with nothing in it look identical otherwise. Never prints a value — the listing drops the envelope on the way out and the route behind it has no value field at all. Works without a keyring on the fleet’s store, so an operator locked out of the key can still take inventory.

crewlet secrets unset <NAME> [-config PATH] [-api URL]

Removes the record and says whether one was there — “was not set” is the outcome a cleanup script wanted on its second run, not a failure. Afterwards ${NAME} falls back to the environment.

crewlet secrets get <NAME> -reveal [-config PATH] [-api URL]

Prints one decrypted value to stdout, with no trailing newline so it can be piped. -reveal is required — without it the command refuses and reads nothing. The route behind it (GET /secrets/{name}) is gated the same way, by an explicit ?reveal=true no crawl reaches by accident, answers Cache-Control: no-store, and logs the access against the operator its guard authenticated. Intended as break-glass for recovering a credential the upstream API will never show again, not as a scripting interface.

crewlet secrets rekey [-dry-run] [-config PATH] [-api URL]

Re-encrypts every stored secret not already sealed under secrets.active_key_id, and prints the names it moved. The per-record counterpart of crewlet config rekey — run both before dropping a retired key from secrets.keys, or records still sealed under it become unreadable. Each envelope names its sealing key, so mixed-key states decrypt correctly throughout. -dry-run lists what would re-encrypt without decrypting anything, reading the denormalised key_id instead.

The store is shared, so this is run once for the fleet, not once per node. A run through a node’s API sends the key id it expects and is refused with a 409 if that node seals under a different active_key_id — a silent success there would report a rotation the fleet did not make. It aborts rather than half-completing if any record cannot be opened with the keyring in hand, because retiring the old key on the strength of a partial pass is what makes a secret unreadable for ever.


crewlet validate [<file.yaml>] [-tier auto|company|bootstrap] [-json]
crewlet validate [-config <tier-a.yaml>] [-company <tier-b.yaml>] [-json]

Validates a config and prints a summary, reaching nothing: no broker, no store, no provider is dialled.

Two forms, because it has two jobs. Name one file and it validates that document. Name neither and it validates both tiers together, from -config and -company — which is what a CI step wants, and which reports both tiers’ problems rather than stopping at the first: an operator fixing a broker URL only to be told about their org chart on the next boot has been made to pay twice for one edit.

A leftover positional is refused, not ignored. Go’s flag package stops at the first non-flag token, so a command that took the file and kept going would silently validate the defaults instead and print a success line about files it never opened.

Validation is deep: it builds the Organization, so duplicate seat and unit names, bad cron expressions, invalid timezones, human seats missing a contact identity, and a knowledge scope with no backend behind it all fail here rather than at run time. It reads no environment: Tier B keeps ${VAR} references verbatim, so a config validates fully before any secret exists. For the same reason a company with no providers.llm at all validates, and its summary reports 0 LLM providers: every node applies it, and its seats hold their work until a provider is added (see A Company With No Model Provider).

FlagDescription
-tierWhich tier a positional file is. auto (default) reads the document’s keys, not its filename. The two tiers share no top-level key, so name/roles/units/providers mean Tier B and node/stream/store/coordination mean Tier A. A document that carries neither, or an equal count of both, is refused naming this flag rather than guessed at: guessing wrong reports every field of the file as invalid, and an operator reading that cannot tell it from a genuinely broken document.
-jsonEmit a machine-readable result on stdout instead of prose.
-config / -companyThe two-tier form. Ignored when a positional file is given.

With -json, the payload is {"valid": bool, "tier": str, "file": str, "problems": [...], "warnings": [...], "summary": {...}}. Each problem is located and classified, so an editor, CI job, or AI authoring loop can fix everything in one pass:

{
"valid": false,
"tier": "company",
"file": "company.yaml",
"problems": [
{ "path": "roles[1].llm", "segments": ["roles", 1, "llm"],
"kind": "unknown_value", "seat": "cto",
"message": "roles[1].llm: value not in the allowed set: \"nonexistent\" is not a configured provider: providers.llm has primary. A key that misses is not an error at run time: the seat falls back to another model and bills against it, so this is the only place the typo can be seen" }
],
"warnings": []
}
Problem fieldMeaning
pathThe authored path. Empty only for a failure that belongs to no place in the document, such as a file that is not YAML.
segmentsThe same path as an array of keys (strings) and list indexes (numbers). A map key can hold a dot, so read these rather than splitting path. null when path is empty.
kindmissing, out_of_range, conflict, shape, unknown_field, unknown_value, or invalid for anything this build does not classify: a closed set with a fallback, because a consumer branching on it must never receive an empty string and read it as a field somebody forgot.
messageThe full line, exactly as the prose output prints it.
seat / unitThe handle of the seat, or the name of the unit, it is about. Omitted when neither.
lineThe line in the file, for a failure the parser found (an unknown key, a value of the wrong shape). Omitted otherwise.

A rule that several seats or units break together, such as two seats sharing a name, is one message and one problem beside each of them, so problems can hold more entries than the prose output has lines. The prose output prints that message once, led by every path it applies to.

Each warning is a reference that resolves to nothing (a unit lead, a root seat’s unit, a manages entry, a GitLab access level): {"kind": "dangling_reference", "ref": "lead" | "unit" | "manages" | "gitlab_access_level", "path", "segments", "seat", "unit", "from", "to", "message"}. The engine runs a company with one, so a warning never fails validation; prose output prints each on its own warning: line. Warnings are reported for any company document that parses, valid or not. Both lists are always arrays.

Exit code is 0 when valid and 1 otherwise, in both output modes: a -json run that printed {"valid": false} and exited zero would pass every CI gate built on crewlet validate x.yaml -json || exit 1. In -json mode nothing is echoed on stderr, because the payload already carries every problem and a second copy is what makes a machine consumer’s log unreadable.


Applies pending schema migrations. Every process that opens the store migrates it, so this is never required — what it is, is a way to do it without starting a node.

Terminal window
crewlet migrate # uses ./crewlet.yaml
crewlet migrate /etc/crewlet/crewlet.yaml
crewlet migrate -check # report pending work, apply nothing
FlagDescription
configTier A YAML (positional, or -config; default ./crewlet.yaml). Name it once: a second positional, or a positional alongside -config, is refused rather than resolved — the two would have to agree and nothing checks that they do.
-checkList pending migrations and exit 1 if there are any; applies nothing. This is what a deploy gate calls, and a gate that reported pending work and exited 0 would stop nothing.

A node without data has nothing to migrate. Its store is store.scratch — created fresh, at the binary’s own schema, every time crewlet run starts — so migrate refuses it and says so. The offline config, secrets and search eval commands refuse a scratch store for the same reason: what they wrote into it would be gone at the next boot. Use the API of a node that holds data instead.

Rolling out N nodes at once means N processes opening the same database and racing to apply the same files. That race is safe — one transaction per file, with the version row written inside it, so a file is either fully applied and recorded or neither — but it is not what an operator wants to watch during a deploy, and a failure mid-rollout is a fleet in two schema states. Migrating once, deliberately, before anything starts makes the outcome one thing that either worked or did not.

Applying is done by opening the store, not by a second code path: a migrator the engine does not use is one that can disagree with it about what “applied” means. -check is the exception and has to be — it reads schema_migrations and creates nothing, because a command that migrated while answering “what would you migrate” could never answer it. A database with no schema_migrations table has applied nothing, which is what a fresh install looks like rather than an error.


Token-budget usage lives in the fleet’s coordination store: one counter per scope for the whole company, with a slot for each calendar window — the day, the ISO week and the month on the company’s clock — and surviving restarts. A counter each node kept privately would make an org cap of 500k into N × 500k.

This command talks to a running node, not to a file. That follows from where the counters live: on the default topology the coordination store is the engine’s own embedded broker, so there is nothing on disk to open — and opening it anyway would be worse than useless, because a second broker on the same store directory is accepted rather than refused, and two writers on one store is corruption rather than contention.

Terminal window
crewlet budgets show # usage per scope
FlagDefaultWhat it does
-urlthe api block of the config named on the command lineThe running node’s base URL. A wildcard bind (0.0.0.0, ::) becomes the loopback address, because a wildcard is not something anything can dial
-token$CREWLET_API_TOKEN, then the config’s first api.auth.tokens entryThe bearer token, sent on the read. A node with allow_anonymous_read off answers nothing without one

The environment wins over the config so an operator who exported a token deliberately gets that one. There is no token default on the command line: a token typed as an argument is in the shell history, in ps, and in any CI log that echoes the command.

The ceilings are not stored here — they come from the active company config (token_budget on the org, role.token_budget on a seat), so every process derives the same numbers without coordinating. Only the usage is shared.

show names the company clock the windows are cut on, then prints one row per window: the company’s day, week and month, then each seat’s.

Windows on the company clock: Europe/Berlin
SCOPE PERIOD WINDOW USED LIMIT STATE RESETS AT LAST REFUSED
org day 2026-09-23 2710450 3000000 near 2026-09-23T22:00:00Z -
org week 2026-W39 9120045 unlimited ok 2026-09-27T22:00:00Z -
org month 2026-09 31004188 unlimited ok 2026-09-30T22:00:00Z -
eng day 2026-09-23 102120 100000 refusing 2026-09-23T22:00:00Z 2026-09-23T07:29:51Z
…

A window no ceiling caps still shows its spend, with a LIMIT of unlimited. STATE is the engine’s own judgement, the one every surface shows: refusing when no charge fits — which every window that refused a round reaches, since the refused round is counted — near at nine tenths of the ceiling, ok otherwise. LAST REFUSED is when that window last turned a call away — a round whose charge it refused, or work turned away unsent because the window was already full (a turn’s next call, a parked delivery, a person’s question, a reflection pass), which is recorded as the gate’s refusal too — printed only while its STATE is refusing, and - otherwise: a window whose ceiling was raised after it refused reads ok or near and shows - there, although it keeps the refusal’s stamp (refused_at on GET /budgets) until its next admitted charge clears it. A refused round is counted like any other, because the vendor billed it, so a window that refused one reads USED past LIMIT by that round — the eng day above refused a 3 000-token round at 99 120 — and every later charge is refused against that figure. The next charge the scope admits clears the stamp, and so does the window turning over.

show refuses rather than printing zeros when the node reports it could not read the counter (durable: false on the query surface). A counter nobody could look at is not a counter that reads zero, and a table of zeros draws a company at 0% of its budget at exactly the moment nothing is known. Seats with no ceiling and no spend in any window are left out for the same reason in reverse: three permanent zero rows per seat bury the seats that matter.

There is no reset. A window’s allowance comes back when the window turns over — at local midnight, on Monday, on the 1st — rolled inside the charge that crosses the boundary, and room before then is made by raising the ceiling in the company config — which is a revision with an author, where a counter zeroed by hand left no record of who made the room or why.


crewlet backup [<config.yaml>] -dir <absolute path> [-url URL] [-token TOKEN] [-wait DURATION]

Copies what a running node holds — its own store file, the replicated estate’s file on a node with the data role, and every JetStream stream and coordination bucket — into one directory, and verifies each store copy before calling it a backup.

It writes to the engine’s host, not yours. -dir is resolved where the node runs; nothing is downloaded. A relative path is refused rather than guessed at, and a destination that already holds something is refused rather than merged, so each run gets a directory of its own.

Like budgets, this goes through the node rather than opening files, and here the reason is doubled. The store is locked to the engine’s process for the life of the handle and the driver does not support a second process on a database file — so no outside tool can read it, and copying it anyway is a torn copy, because committed data lives in the file and its -wal together. The stream estate is worse: on the default topology the broker is embedded in the engine and binds no socket, so there is no address to give the nats CLI. The one process that can reach both is the engine, and this asks it to.

The report names what it captured, per estate. Every node holds its own store whatever its node.roles, and a data node the replicated estate beside it. On the embedded topology the fleet’s streams come too: a member snapshots them from its own broker, and a leaf, whose broker keeps no stream, snapshots the same streams from the members across its leaf link. A node that dialled an external NATS cluster copies its store files alone and says so, naming the cluster as where the stream half is backed up, rather than presenting a partial copy as a backup.

The company’s files come too, as an objects line. On the default nats object store their objects are in the OBJ_crewlet_files stream the backup snapshots with the rest, and the line counts the objects the copy names that the stream holds (lost ones are listed beneath); on s3 each object the copy names is read from the bucket into objects/ and checked against its file’s size and SHA-256 as it is written, and the line gives their count and size. A file whose object the store answered it does not hold, or holds wrong, does not fail the backup: it is listed beneath, as PROJECT/path and the object’s key, to restore from an earlier backup or upload again. One the store could not answer about at all fails it.

-wait (default 30 minutes) bounds how long the command waits for the answer, not the copy: the engine finishes what it started, so a wait that expires leaves a good backup this command has already called a failure. It is set far past any plausible duration for that reason.

See Backups & Restore for what the directory contains, what the copy is a copy of, and how a restore uses it.


crewlet work purge <task-id> -project KEY -reason TEXT -confirm <task-key>
[-op-id ID] [<config.yaml>] [-url URL] [-token TOKEN]

The gestures on work items that belong to a person rather than to a seat. Everything else a company does to its tasks is done by seats through their own tools, and that is the design — an engine whose operator edits work by hand is one whose org chart is decoration. The exception is the operation no seat may perform.

Destroys a task and every row it produced: its own row, its comments, its body revisions, its checklist, its field values, its watchers, its relations, its dependencies, and every inbound reference and key alias that made it resolvable. A marker is written in their place, and every later record about that task is dropped for ever — which is what stops a redelivery months afterwards resurrecting any of it.

remove hides a task and delete stops later records about it. Neither takes the body out of a single database, which is why this verb exists: an erasure request and a credential pasted into a task description both need the rows gone.

The confirmation is the task’s key, not its id. The id is on the command line already, so repeating it confirms nothing; the key has to be looked up. The reason is required because it is the only thing that survives — the marker’s reason is the entire account of what used to be at that key for whoever reads it a year later.

Its children are moved, not destroyed. Each direct child re-parents onto the purged task’s own parent, or becomes a root when the purged task was one. Destroying the subtree would destroy work nobody confirmed; leaving it would leave every child pointing at an id that resolves to nothing.

The outcome is three-valued, like every write here. applied means the rows are gone on this node and the record is durable. pending means the record is durable and this node has not applied it yet — do not run it again, because a second gesture appends a second purge. unknown is the one to retry, and the operation id it prints goes back in -op-id so the retry cannot append a second record. A retry whose first purge did land answers applied, at the position that purge landed at, rather than finding the task already gone. The command mints that id before it asks and waits thirty seconds — past the three five-second waits a write can make — so a purge the node never answered, which may well have landed, still prints the -op-id to run it again with. Pass the id exactly as printed: it carries the instant it was minted, and a node whose operation ledger may have lost the first purge’s row since — to the ledger’s thirty-day sweep — judges the retry by it, answering unknown again rather than purging twice. An id the engine did not print is refused (op_id_invalid), and so is one altered on the way back — trimmed, spaced or over 128 bytes — since the broker would carry it as a different id.

What it does not reach: a node that is offline or evicted keeps its copy until it replays, adopts a snapshot, is replaced or is destroyed. There is no duration to state, and crewlet retention status names which nodes those are.

crewlet objects status [-json] [<config.yaml>] [-url URL] [-token TOKEN]

Where the company’s files are kept, and what the object store’s collector last found. It talks to a running node, for backup’s reason: the collector’s report is a record in the coordination store, which on the default topology is the engine’s own embedded broker and binds no socket. status reads the fleet view’s objects block (GET /fleet), which every node answers from that one record.

It prints the store the fleet’s files are in — the fleet’s own NATS bucket (OBJ_crewlet_files) or an S3 bucket with its endpoint, bucket and prefix — the node that ran the last passes of the object-collector duty, and one line for each pass:

  • Last collection — how many objects it listed in the store, how many it deleted for being older than a day with no row naming them, and how many uploads begun more than a day ago and never finished it abandoned. A collection that stopped judging says why (stopped judging: … — the node’s view of the estate was incomplete, so a row it could not read might name any object) beside what it had deleted before it stopped; one that stopped says what stopped it, and one that did not finish says incomplete. A sweep of unfinished uploads the store refused — on S3, an identity without s3:ListBucketMultipartUploads or s3:AbortMultipartUpload — is a line of its own beneath.
  • Last audit — from the last audit to run to its end: how many files the company’s rows name, how many the store does not hold (missing) and how many it holds at the wrong size or digest (damaged). An audit that failed after it is said beneath, never in its place. When any file cannot be read it lists them — the file, whether it is missing or damaged, and its object’s key; the first hundred, and how many more — to restore from a backup or upload again.

Before the collector has finished a pass — on a new fleet, in the minute or so before a data node first claims the duty — it says so. -json prints the block exactly as the node answered it, for a script. The exit status is non-zero when the node cannot say — its coordination store did not answer, the node runs no object store, or it answered a state this build does not know — and zero otherwise, files that cannot be read included: a store that lost bytes is an answer, and the objects_missing alarm is what pages for it.

crewlet fleet broker list [-json] [<config.yaml>] [-url URL] [-token TOKEN]
crewlet fleet broker remove <node> -confirm <node> [-force] [<config.yaml>] [-url URL] [-token TOKEN]
crewlet fleet broker remove -peer <peer> -confirm <peer> [-force] [<config.yaml>] [-url URL] [-token TOKEN]

The fleet broker’s membership. Two records say who its members are: every node advertises its broker kind on its presence lease — derived from its stream block, never from its roles (see Configuration) — and the broker’s metadata group, the raft group that places every stream and consumer, counts its voters for itself. Both verbs talk to a running node, and the node asks a member where it is not one: only a member holds the group, and a removal is answered only on its system account, which nothing outside its process can reach. They are clients of the routes in the API reference.

The question it answers is does the broker still count a member I have lost? One row per live node — what it advertises, its roles, how the group counts it (leader, current, offline, last heard 5m ago, or - for a node that is no voter) and its peer id where it is a voter — then one row per voter no live node is, and under the table every disagreement in words.

The group counts each voter by its raft peer id, which is derived from the node id. A member hears another’s name only from that server, so a voter whose survivors have restarted since it died is listed as (name unknown) by its peer id alone; a voter is matched to its node by peer id, never by the name the answering member happens to have heard.

  • DEAD MEMBER — a voter no live node is, or whose node came back as a leaf or a client. Every election and every create goes on counting it; the line prints the remove command for once it is not coming back — by -peer for a voter nobody can name, since it has no node id to give.
  • NOT COUNTED — a live node advertising a member the group does not count: still joining, or removed while it ran.
  • UNKNOWN — a node whose presence does not say what its broker is, which a capacity seal counts as a member.

A group no member could report is said to be unread, never shown as empty. On a fleet whose broker is an external cluster it says so and lists nothing: that membership is the cluster’s own operator’s. -json prints the node’s answer as it came.

Stops the metadata group counting a member once the group has committed the change — printing which member’s system account carried it and the voters that remain. Name it by node id, which reaches it by the peer id the id hashes to whether or not any member still remembers its name, or by -peer with the peer id list shows, for a voter nobody can name. Refused while its node holds a live presence lease as a member (or without saying what its broker is): it is still running, and a running member removed from the group rejoins it as a voter at its next restart, so stop it first. -force removes it anyway, for a member wedged in a way that still renews its lease. A node alive under the voter’s name as a leaf or a client is removed without -force: the member it was is gone for good. The member being removed never carries its own removal — if it is the only live member, the command says so.

The whole removal is bounded by one wait: the carrying member’s commit budget and a round trip. A carrying member that does not answer ends it as an outcome nobody knows — it may have committed the change first — rather than another member proposing it again; list reads the group.

Only the group’s leader answers a removal, and a member removed because its host died was often that leader — so the removal asks again every second while the survivors elect another, and a removal made during an election waits it out rather than failing. A group that answers nobody for the whole wait (two minutes) has lost its quorum and can change nothing about itself, and the refusal says so: bring enough members back for a quorum, then ask again. A removal the node did not answer may still have been committed; list reads the group, and asking again for a member already removed is refused as not a member.

crewlet seats pause <handle> [-stop] [-reason TEXT]
[-request-id ID] [<config.yaml>] [-url URL] [-token TOKEN]
crewlet seats resume <handle>
[-request-id ID] [<config.yaml>] [-url URL] [-token TOKEN]

Pause an agent seat, or resume it — the same pause_seat and resume_seat the dashboard’s buttons call, over the same route (POST /operator/act/{tool}), so a pause from a terminal and one from a profile screen are one gesture with one record. It acts as the person your token is bound to (contact.crewlet_operator_id on their seat); a token nobody bound is refused unbound, and the refusal names the seat to bind it on.

A paused seat starts no new turn, its incoming mail waits on its inbox in order, and its scheduled runs are recorded skipped_paused rather than sent. The turn it is on finishes first unless -stop is given, which ends that turn at its next round; a stopped turn is not run again. -reason is one line (at most 500 characters), shown on the seat and in the feed. Pausing a paused seat changes nothing — except that -stop adds the stop — and resuming a seat that is not paused changes nothing. See Agent Runtime § Pausing a seat.

The outcome is applied — the pause is the fleet’s record, and the node holding the seat carries it out from its own copy within about a second — or unknown, when the node could not confirm the write. unknown prints the request id to send again with -request-id, so the retry is the same request. The id is a UUIDv7 — the node reads the instant it carries to decide whether it can still vouch for a retry — and the command mints one when -request-id is not given; an id of another version is refused invalid_request_id.

crewlet retention status|snapshots|ack|evict|readmit|set-capacity|maintenance|reanchor|verify
[<config.yaml>] [-url URL] [-token TOKEN]

What the state log is holding, why it is not shrinking, and the gestures that change it. A group rather than nine top-level verbs, because several of them stop a machine writing and that should not sit at the same level as version.

Every verb but one talks to a running node, for backup’s reason: the register they read and write is a coordination bucket on a broker embedded in the engine, which binds no socket, so there is no address any other tool could be given. The exception is verify --restore, which reads an artefact off disk on purpose — see below.

A domain this node refuses leads, before anything else about it, with the refusal’s own sentence — and for a wrong_stream, which finding is behind it (recreated, ahead_of_log, log_diverged, generation_passed):

NOT READY pages: wrong_stream (recreated) — pages's rows are keyed to the stream created at … and the broker's CREWLET_PAGES_LOG was created at …
WRITES REFUSED tracker: log_truncated — peer node-4 stands at sequence … of CREWLET_TRACKER_LOG, which ends at …

Then the blocking term, in prose, because it is the answer to the only question anybody runs this for:

Nothing is being trimmed on tracker: the newest complete backup is 3 days old
(backup_max_age is 24h).

Then one row per registered domain — its stream, generation, replay protocol, both ends, bytes, ceiling, reserve, headroom and trim floor — the six terms with their state and detail, one row per node per domain, and this node’s own replica line. TRIM FLOOR is - where the trim has concluded nothing about the domain’s generation yet — right after a reanchor, until its first tick on the adopted stream — and unreadable where the floor register could not be read; the watermark block says which instead of listing terms. Each node row carries its position’s generation (GEN), and a position from a generation the log has left prints left gen N (or ahead gen N) in place of a lag, since its sequence compares with nothing the log holds now; a node that reported its log diverged at its checkpoint is marked LOG DIVERGED. RESERVE is the top of the ceiling the tracker and pages logs keep for gate records (the gate reserve), - on the vector changelog, which keeps none; HEADROOM is what is left of the rest, the ceiling ordinary writes are refused at — so a log at 0% still takes an eviction.

A term that does not apply to a domain prints n/a rather than 0: the vector log has no wake feed, and an absent term is a different fact from one that permits nothing. Where a log does have one, its feed_ack_floor is that log’s own feed — never another domain’s.

It exits non-zero exactly when this node has an active alarm, on the report’s own rule rather than a second one in the CLI — the shell script watching this exit code would otherwise be watching the definition nobody maintained. Each alarm’s measurement and remedy go to stderr, so a cron capturing stdout for a dashboard still gets the reason in its own mail.

-domain <name> narrows the log and watermark blocks to one domain.

Its own verb rather than a block of status, because the repository is per node: which of my machines can donate, and how old is what they hold is a disk question. A node with no artefact still gets a row, carrying the reason — sole_node, lagging, unhydrated, deferred, insufficient_space, ahead_of_log, log_diverged or failed — because the absence is the answer to “why did the join fail”. A node that holds a current artefact carries no reason at all: recent is the skip a healthy node takes, so it is never published as one.

crewlet retention ack -stream CREWLET_TRACKER_LOG -position 918100000

Publishes an operator backup floor. Under backup_floor: operator the trim follows this rather than the engine’s own copies, because a backup is not a backup until it leaves the host and the engine cannot see that it has. Both flags are required: an acknowledgement moves the floor the trim deletes against, so there is no value to guess.

crewlet retention evict node-4 -confirm node-4

An absent node pins the applied term for ever — its position never advances, so nothing above it can be deleted. That is deliberate for a node that is coming back; eviction is the gesture for one that is not.

-confirm repeats the node id, the same shape the other destructive gestures use. The trim counts nodes per log, so the eviction is a record on every log it counts nodes on — the tracker’s log and the pages log — and each answers on its own line with its three-valued outcome, or not written and the reason that stopped it:

$ crewlet retention evict node-4 -confirm node-4
evict node-4 (operation 01a0cd85-735a-7294-9d3e-38998abd698c.evict-node-4)
tracker: applied at CREWLET_TRACKER_LOG 918280002
pages: applied at CREWLET_PAGES_LOG 4410
node-4 stays COUNTED for about a minute, so a live node is certain to have read its own tombstone before the trim passes it.
tracker trim floor 918100000 → 918100000
pages trim floor 4100 → 4100

The command prints the watermark before and after and the instant the eviction takes effect: the node stays counted for about a minute, so a live one is certain to have read its own tombstone before the trim passes it.

The node you ran it on writes every log itself: every data node holds the whole estate, so one command reaches every log.

The gesture is judged once, before either log is written. A node that still holds a live presence lease is refused with 409 eviction_refused — it is still reaching the fleet and almost certainly running, and an eviction would drop everything it writes and take its copy out of service. Stop it and wait for its LIVE column in crewlet retention status to read no, or pass -force for a node wedged in a way that still renews its lease — the refusal prints both. A node that cannot read the presence leases at all answers 503 eviction_unjudged; -force takes the eviction past that as well, since the leases are all the judgement reads, and the refusal says so. A run that already passed -force is never advised it again. A node id no node could run under is refused 400 invalid_gate before anything is judged.

When not every log holds the record, the command exits non-zero, and under each log it did not finish it prints what to do. Where the same command can finish that log — an unknown outcome, a node still catching up — it names it, with -force carried over from a forced run:

pages: unknown — the record may or may not be on the log
its outcome is unknown: the same gesture under the same operation id answers from this log's own ledger if the record landed, and writes it if it did not
The gesture has not reached every log. Run it again with -op-id 01a0cd85-… to finish it: a log that already holds the record answers from its own rows and is not written twice.

Run it again with that -op-id. Each log’s record is published under an id derived from it, the sign, the log and the node, and each log’s snapshot reads that id’s ledger row before anything is decided — so a log whose record is already the gate in force answers at the position it has, and only the missing log is written, however long after the first run the retry comes. A fresh id would be a second eviction rather than this one finished, and an id that is not in the engine’s grammar is refused. Where the same command cannot finish a log straight away it prints what has to happen first, and then the same -op-id that finishes the gesture once that is done: a log_full log — full past even the reserve its gate records may use, since a log full only for ordinary writes still takes them — needs crewlet retention set-capacity:

pages: not written (log_full) — pages: statelog: unavailable (log_full): …
CREWLET_PAGES_LOG is full to its broker ceiling, past even the reserve kept there for gate records: raise its ceiling, which is the only thing that makes room for it — then the same gesture under the same operation id finishes it, where a fresh one would write every log that already holds the record again
crewlet retention set-capacity CREWLET_PAGES_LOG <bytes> -confirm <bytes>, then run this again with -op-id 01a0cd85-…

Not the same command now — it is refused the same way until the ceiling moves — and not a fresh one afterwards: once there is room, this gesture’s own -op-id is what finishes it, because a fresh id is a second eviction that writes the tracker’s log again and re-dates the eviction there. A node that is itself evicted cannot write, so it prints -url <that node> with the same -op-id; a wrong_stream log needs a reanchor first, and then the same -op-id; and a superseded operation — an eviction retried after a readmission took the node back — is a new gesture, without -op-id. The indented sentence is the node’s own and names no flag, because the dashboard renders the same answer; the line under it is this command’s rendering of the node’s actions (see the gate answer).

The command mints the operation id before it asks, and waits a minute for the answer — past the fifty seconds the node takes at most to answer one gesture (twenty to judge it and thirty to write every log), which it finishes even if the connection drops. So a gesture that got no answer at all still prints the -op-id that finishes it — and so does one whose answer the node did not write: a reverse proxy’s 504 page, any status carrying no engine error code, or a 200 cut off part way through, each of which says nothing about what the node did. Pass it back exactly as printed: the node refuses (op_id_invalid) an id that is not in the engine’s grammar — a UUIDv7 carrying its mint instant — and one over 128 bytes or holding anything but visible ASCII, since the broker would carry that as a different id.

readmit is the inverse commit rather than a delete, so the eviction’s whole history survives a replay. It is refused while the node is below a trim floor — and the refusal prints its position beside the floor, because that inequality is the reason:

$ crewlet retention readmit node-4 -confirm node-4
crewlet: the node answered 409: readmission_refused
statelog: node-4 may not be readmitted: its last position in tracker (reported 2031-04-02T03:14:00Z) is 1200 and records below 9000 may already be gone from the tracker log (published floor 9000, first surviving sequence 8800) — it has to catch up before the fleet counts it again
start node-4 if it is not running: it catches up on its own, …
its SEQ in `crewlet retention status`, at the domain's own GEN, says when it has caught up, and `crewlet retention snapshots` whether a peer can donate one; then run this again

The comparison is the one the node’s own write fence makes: its last published position in each log that claims identity — the tracker’s and the pages log — against the higher of the published floor and the log’s first surviving sequence, and the node must have applied every record up to the one just before it. A node that has never published a position is judged as holding nothing; one whose position is from a generation the log has since left is refused, since nothing it holds compares with anything the log still has. The judgement is made once, for both logs, before either is written, and nothing is written to either on a refusal. A readmission that reaches one log and not the other is finished with -op-id, exactly as an eviction is.

What to do is let the node catch up, which it does on its own: start it if it is not running, and it replays what the log still holds or adopts a peer’s snapshot where it does not (crewlet retention snapshots says whether any peer can donate). Readmit it once its SEQ in crewlet retention status has reached one less than the higher of that domain’s TRIM FLOOR and FIRST — the refusal’s own floor, first_seq and generation are the numbers it compared, and right after a reanchor the domain shows - in TRIM FLOOR until the trim’s first tick on the adopted stream, so the bound is FIRST alone. The node’s GEN column says whether its position is at the log’s current generation: one from a generation the log has left prints left gen N in place of a lag and is refused whatever its SEQ reads. The position is a heartbeat old, so a node that has only just caught up can be refused once more; run the command again.

A readmission that cannot be judged is refused as well, with nothing written: 503 readmission_unjudged, when the positions register, a published floor or a log’s first surviving sequence could not be read. The node’s position is not what is wrong there, so it is not told to catch up: the command is asked again, on this node or through another (-url <that node>). And either gesture on a node in a capacity window is refused 409 not_publishing before anything is judged; its line points at crewlet retention status, which leads with the window while it is open, since nothing is written to any log until the fleet is back in normal mode.

The refusal is the truth about the node rather than the thing keeping your data safe. Readmitting a node below the floor used to succeed, and it put back the pin the eviction had lifted: the trim, counting that node again, stops advancing. Its writes were never the danger — its own fence refuses any write that assumes a history it does not hold.

crewlet retention set-capacity CREWLET_TRACKER_LOG 8589934592 -confirm 8589934592

Changes a log’s byte ceiling, raising or lowering it. A log’s Tier A ceiling (stream.tracker_log_max_bytes, stream.tracker_vectors_max_bytes, stream.pages_log_max_bytes, stream.usage_log_max_bytes) is not a live setting, only the value its stream is created with. A resize is decided against the usage the log is at, and a publisher makes that a moving quantity, so this runs with the whole fleet in a maintenance mode and costs three fleet-wide restarts, two more per retry. The procedure is documented once, in the retention guide; this command’s help prints the five steps.

-confirm repeats the byte count, and the target is then fixed for the life of the operation: a verify compares the observed ceiling against it, so a target that could move would make a mismatch unreadable.

A raise the broker has no room for is refused before the maintenance window opens, naming what the raise would reserve, what the broker has left, the limit it is held to and the field that sets it — so an operator learns it in one round trip rather than after three fleet-wide restarts. A raise the broker refuses at the apply is reported as a refusal rather than as an unknown outcome, with the same numbers, read at that moment rather than carried from the window’s opening: an operation spans restarts, and what matters is what the broker had when it said no. A limit this node could not read is not a refusal — a node that could not hear has not been told no — and the operation goes ahead for the broker to decide.

-i-have-excluded-all-publishers is required only on stream.type: nats. There the engine does not run the broker and cannot establish who else holds a connection to it, so the assertion is yours in your own words rather than a check that quietly proves nothing.

A target whose ordinary ceiling is at or below what the log already holds is refused before the window opens, naming both numbers and the least target that would do: that ceiling would refuse every ordinary append the moment it applied. On the tracker and pages logs the ordinary ceiling is the target less its gate reserve, a sixteenth; on the vector changelog it is the whole target. A target under a gibibyte is refused too — Tier A’s floor for every log, and the one the reserve is sized against. A target under the current ceiling and above the usage is accepted, and is how a log gives a reservation back. A raise the broker cannot reserve is refused before the window opens too, naming what it reserves and what the broker has left, wherever the node can read that limit: a lone embedded node, or a NATS account’s own limit. A clustered member cannot, and there the broker refuses such a raise when the window applies it.

Run from a node in normal mode it refuses outright, naming the restart. Run from maintenance it opens, baselines and applies; run from seal it collects the barrier, seals, verifies and confirms. It prints the phase it reached and the next gesture, never the transition table.

crewlet retention maintenance status -stream CREWLET_TRACKER_LOG
crewlet retention maintenance exclude -stream CREWLET_TRACKER_LOG -node node-4 -confirm node-4
crewlet retention maintenance abandon -stream CREWLET_TRACKER_LOG -confirm capacity-01J...

status is the one page an operator can see why a fleet is still excluded: the operation, its phase and attempt, every participant’s baselined incarnation against what it acknowledged as, which acknowledgements are missing, any unresolved write attempt, and any admission still blocking activation. On a fleet with no window it says so — no window and a window in phase opened are different facts with different next steps.

exclude is your assertion that a participant’s process is stopped and holds no outstanding request. It is the only thing that waives an acknowledgement; an eviction does not, because that is about whose records apply and this is about whose process is running.

abandon changes what the operation is trying to reach and never the barrier it must cross. From opened it clears outright — no request was issued. From anywhere else the fleet still has to restart into seal, because a paused coordinator’s request is outstanding whether or not a person has read a status page. -confirm repeats the operation id from status.

crewlet retention reanchor -stream CREWLET_TRACKER_LOG
crewlet retention reanchor -stream CREWLET_TRACKER_LOG -confirm 2031-04-02T03:00:00.418226517Z

The recovery for a stream that was genuinely recreated, or for a broker restored from an older copy. Run without -confirm it prints the live stream’s own created_at, and which case it is, and refuses: the confirmation means I looked at the thing I am re-anchoring, and a verb that read the value and fed it straight back would be confirming against its own output. Paste the instant back exactly as printed — it is the same one a wrong_stream refusal names as the broker’s.

The three cases are where the log is followed from:

  • recreated — the stream is not the one this node’s rows are keyed to, so it is followed from its first surviving record;
  • restored — it is the same stream, ending below this node’s checkpoint or written past it since, so the rows already hold every record the copy kept and it is followed from its end — one below the generation record the verb appends, so a record written while it ran is not applied past it either — and none of them is applied again. If the log also holds records the rows do not — written after the restore, by a node whose rows were the copy’s age — they would be applied on no node, so the verb names the newest of them and refuses unless you pass -discard. The alternative to discarding them is to keep them: do not re-anchor this node, replace its rows with a peer’s instead;
  • abandoned — it is the same stream and holds every record the rows are missing, but it continues in a generation only a peer the fleet has since evicted held; it is followed from this node’s own checkpoint, in the generation after the evicted peer’s, and every record of the generation it skips is void.

A same-stream log that reaches the checkpoint, holding there the record the checkpoint names, and continues in no abandoned generation is none of these, and the verb prints that there is nothing to re-anchor rather than a command to run. The answer names the case and the sequence the checkpoint went to — and, with -discard, the newest record it discarded.

-discard is its own flag rather than part of -force: -force says the fleet cannot be asked who is most caught up, which says nothing about what a log holds, and forcing past an unreadable register must not discard writes on the same keystroke.

It moves only the log you name: that domain’s checkpoint goes to the next generation on the live stream — the one after every generation the domain has used, an evicted peer’s included — every other domain’s stays where it is, and the domain’s applier resumes on the node with no restart.

For the tracker and the knowledge base it refuses while any peer has already re-anchored the stream, naming the peer — that peer’s rows are the fleet’s history in the new generation, and this node adopts its snapshot instead. A peer that is gone for good is released by evicting it first (crewlet retention evict): an evicted peer counts toward neither this rule nor the next. Otherwise only the most caught-up node on the stream its rows came from may run it, among the peers holding history the log does not — a peer whose checkpoint record the log holds is on the log and is not weighed — and the refusal names the peer further along; -force overrides that rule, and an unreadable positions register, and never a re-anchored peer. Nor does it ever open a generation another node already opened: if a peer got there first, the verb reads its record back, names it, and commits nothing.

For a recreated log it does not recover records that were on the old stream and were never applied here, and the refusal says so. See Re-anchoring a recreated or restored log.

crewlet retention verify --restore -dir /var/backups/crewlet [-cadence 720h]

The one verb here that talks to no node. It restores the newest artefact under -dir and opens the copy — the point is to establish that the artefact alone is enough, and running it through a running engine would be asking the thing under test to test itself. It writes nothing to the live store and takes no lock on it, so it is safe beside a running node.

It prints what the artefact holds and exits non-zero past its cadence, which defaults to 30 days. Put it in cron: a lapsed restore test that exits zero is a paragraph in a runbook nobody read.

A directory with no manifest is debris rather than a partial backup — the manifest is written last — and an artefact naming no domain position cannot be verified whatever else it contains, because a restore replays from that sequence.

See Retention for the six terms, the snapshot repository and the join runbook.


crewlet schema [company|bootstrap] [-o PATH]

Prints the JSON Schema for a config tier — company (Tier B, the default) or bootstrap (Tier A) — to stdout, or to PATH with -o/--output.

The schema is generated from the config types the loader itself uses, so it cannot drift from what the engine accepts, and because every config type forbids unknown keys it carries additionalProperties: false throughout — a mistyped field is flagged rather than silently ignored.

It is deliberately a subset of the validator, never a superset. A schema can express structure — key spaces, types, closed sets, ranges, patterns — and not the cross-field rules the validator enforces. The invariant is one-directional and tested: everything the schema rejects, the validator also rejects. An editor that red-underlines a config the engine would happily run teaches authors to ignore it.

So it accepts what the decoder reads, not merely what a field’s type suggests, and a test offers every field of both tiers each such value to hold the two to the same verdict:

  • An empty value — ~ or "" — anywhere the engine reads it as unset, which is everywhere but the company document’s root.
  • A number or a boolean in a text field, read as the text it was written as (name: 2024, org_webhook: false).
  • YAML 1.1’s switch words in a boolean field — yes, no, on, off, y and n, each in YAML’s three casings (yes, Yes, YES) — because the decoder reads them there. A quoted "true" is not one of them: it is text, and a boolean field refuses it.
  • A ${VAR} reference in Tier A wherever the engine resolves one into a value the field can take: anywhere in any text field, a patterned or closed-set one included, since the value is judged only once it has resolved; and as the whole value of a number or a boolean field, whose resolved text is read as if it were written there. A reference with other text around it in a number or a boolean is refused, by the schema and the engine alike.
  • A Tier B reference as the literal text it is. The company keeps a reference verbatim and resolves it only where a provider or transport is built, so a field with a pattern or a closed set — an enum, a handle, a unit id — refuses one as it would any other text that is not a value of it. The exception is a field whose consumer resolves a whole reference, which takes exactly one beside its own rule: today that is a Mattermost seat’s username. See Environment Variable References.

Every position that holds a credential carries the annotation "x-crewlet-secret": true — on the field itself for a single value (integrations.github.webhook_secret), on items for each member of a list (providers.llm.*.api_keys), and on additionalProperties for each value of a map (mcp_servers[].headers, a seat’s sandbox.env, and both levels down in mcp_env, whose values are per-server maps of variables). A map’s keys are never marked: they are names (GITHUB_TOKEN, Authorization), and only what they map to is a credential. A position without the keyword holds none, and the keyword’s only value is true.

It is read from the same declaration on the config types that drives redaction, and a test holds the two to one set in both directions: in the company schema, a position is marked exactly when GET /config, crewlet config show and crewlet config export -redact mask it. The Tier A schema marks its own by the same rule — the keyring’s secrets.keys[].material, api.auth.tokens[].token, stream.token and store.objects.s3.secret_access_key — though no read surface serves Tier A to redact. store.objects.s3.access_key_id is not marked: AWS treats an access key id as an identifier rather than a secret — it travels in every signed request’s Authorization header and appears in CloudTrail records and error messages — and only the secret access key proves possession. A consumer that prefers every value of that block to arrive by reference may still write a ${VAR} there; the mark only says where a literal is wrong.

It is an annotation, not a rule. The engine runs a literal credential in a marked position (it seals it at rest and masks it on every read), so the schema accepts one there, and a validator that does not know the keyword ignores it, as JSON Schema 2020-12 requires. What a consumer does with it is its own policy; the one the engine is built for:

  • Write a whole ${VAR} reference there, never the value. In the company document the reference is stored verbatim and resolved where the provider or transport is built — from the secret store first, then the environment — so a tool that renders a company (a Kubernetes operator from its Secrets, a CI job from its vault) pushes the value to the secret store with PUT /secrets/{name} and writes ${NAME} into the marked position. In Tier A the reference resolves from the node’s environment.
  • Treat a literal in a marked position as a finding — an editor warning, a linter failure, or a refusal in tooling that must never hold a credential in a document it versions.
  • Mask it when displaying a document, as the engine’s own reads do, unless the value is exactly one ${VAR}, which names a credential and carries none.

The mark lives in the published schema, and a consumer acts on it there. It does not survive being copied into a Kubernetes CustomResourceDefinition: a CRD’s schema accepts only its own x-kubernetes-* extensions (and no $ref), so a generator that builds one from fragments of these files has to strip x-crewlet-secret and keep the policy in its own code.

Both documents are checked into schema/; a test regenerates and compares them, so a config field added without a schema entry fails the build rather than leaving a stale file nobody opens.

Point your editor at it for completion, inline field docs, and typo squiggles:

# yaml-language-server: $schema=https://docs.crewlet.ai/schema/company.schema.json
name: "Acme AI"

See Authoring with an AI assistant.


crewlet search eval [-store PATH] [-config crewlet.yaml] [flags]

Measures the semantic half of knowledge search — the two-stage 1-bit retrieval described in Knowledge System — against the exact scan it approximates, on your own vectors.

There is no other number that answers this. The engine’s own quality gate measures the arithmetic on a seeded fixture and deliberately makes no claim about recall on a particular company’s documents: a 1-bit code keeps only each vector’s orthant, and how much that says about ranking is a property of your corpus’s distribution. Over a family of embedding-shaped generators the same arithmetic spans 0.29 to 0.98 recall.

Ground truth is the exact scan’s own top-K, so nobody authors a judgement — which is the step that otherwise makes an evaluation stop being run. The queries are held-out documents from the corpus itself, spread deterministically across it so two runs are comparable.

It reads a file, not a running node. The replicated estate is exclusively owned by the running engine, so pass the copy inside a backup — which needs nothing stopped and measures the same rows — or the node’s own file with the engine stopped. With neither -store nor a reachable Tier A file it has nothing to open.

Terminal window
$ crewlet search eval -store /var/backups/crewlet/2026-09-01/store-replicated.db
corpus 118432 sources, text-embedding-3-large at 3072 dimensions
measured 25 queries at depth 150 from 1200 candidates
first stage the semantic index: 128 of 1024 lists probed (generation 1099511744562)
index measured on 118432 sources: 0.9812 against a 0.9800 floor in its worst shape (container:task), 0 head miss(es)
recall 0.9761 (floor 0.9312 for this corpus size)
worst query 0.9467
head misses 0 (documents dropped from the exact top ten)
scan recall 0.9803 with 0 head miss(es) — the same search with the full scan as its first stage
narrowed source:page recall 0.9950 floor 0.9800 head misses 0 (25 of 25 scanned) — scan 0.9950, 0 head miss(es)
narrowed source:task recall 0.9772 floor 0.9368 head misses 0 — scan 0.9810, 0 head miss(es)
narrowed container:task recall 0.9967 floor 0.9800 head misses 0 (19 of 21 scanned) — scan 0.9967, 0 head miss(es)
narrowed container:page recall 1.0000 floor 0.9800 head misses 0 (4 of 4 scanned) — scan 1.0000, 0 head miss(es)
verdict the two-stage search recovers the exact ranking at the shipped depth, in every shape
window 8192 bytes a source (the corpus's opening): a semantic search sees each source's title and body up to it, a keyword search the whole body
past window task 2110 of 104208 sources (2.0%), 7.4 MiB of 196.0 MiB of text (3.8%)
past window page 9874 of 14224 sources (69.4%), 402.1 MiB of 518.6 MiB of text (77.5%)

The last three lines are printed on every run: for each corpus, how many sources and how much of their text lie past the window each vector was computed from — what a search by meaning cannot see, which a keyword search still reads (Knowledge System). That part of the report reads every source’s whole body.

FlagDefaultWhat it does
-store PATHfrom -configThe replicated database to measure. Overrides the Tier A file
-config PATH./crewlet.yamlTier A, read only for the replicated store’s path
-queries N25How many held-out documents to measure over. Each one is a full exact scan, which is what bounds the run
-limit N150The depth recall is measured at — the shipped returned depth
-candidates N1200The stage-1 candidate depth — the shipped pair
-model NAMEmost populatedThe embedding model to measure
-dimensions Nthe model’sThe width to measure
-probes Nthe index’s ownHow many lists of the semantic index to probe. 0 takes the count the index’s own training measured; a larger one shows what probing more lists would recall
-window BYTESthe smaller of 8 192 and the model’s own per-input boundThe bytes of each source the vectors were computed from, for the report of what lies past them. The store does not carry the company’s configuration, so state it when providers.embeddings.max_input_tokens or max_batch_tokens lowers the per-input bound below the model’s own — the embedding duty embeds the smaller of 8 192 bytes and that bound
-fitoffAlso print the corpus’s own mean pairwise cosine, which is the parameter the engine’s seeded fixture is fitted from
-metricsoffPrint one key value line per metric instead of a report, for a collector or a shell — the window report included, as search_eval_window_bytes and, labelled by source, search_eval_window_sources, search_eval_window_beyond_sources, search_eval_window_text_bytes and search_eval_window_beyond_bytes

It exits non-zero when the recall is below the floor for that corpus size, or when any document was dropped from the exact top ten, in any shape — so it can go in a schedule. Both conditions matter: an aggregate of 0.98 is compatible with losing exactly the documents that mattered, and a semantic-only document the first stage drops leaves the fused answer entirely.

Run it monthly, and after any change to providers.embeddings.model or providers.embeddings.dimensions — those are the two inputs that move the answer. A report naming more than one embedding space is a refill in progress: until it finishes, documents still on the old model are not in the candidate pool at all, because the scan filters on the model/width pair.

Below the floor, the remedy is decided in advance and in this order: raise the quantization over-fetch (measured free in latency — stage one is a full scan whose cost does not depend on how many candidates it keeps), then an 8-bit first stage, which ships in the same release that moves the model default.


crewlet llm list
crewlet llm doctor [KEY] [-no-smoke]
crewlet llm login <KEY> [-from-host | -capture-token | -token-stdin |
-username U -password-stdin] [-home PATH]
[-print-token]
crewlet llm status <KEY>
crewlet llm logout <KEY>
crewlet llm export <KEY> [-secret-store]
crewlet llm import <KEY> # bundle on stdin

The operator side of a subscription LLM backend: a providers.llm entry of type: cli-agent drives a vendor’s own CLI under the operator’s Pro/Max plan instead of an API key, and the login that makes that work is established here rather than in the config document. KEY is the providers.llm key; commands that take one and are given none act on the only cli-agent provider when there is exactly one. doctor also examines every type: anthropic entry; the other subcommands are cli-agent only, since an API entry has no login to broker, list or export.

list is the inventory — provider key, CLI, model and how it is signed in. doctor is the one to run before a company’s first turn: it checks the CLI is installed, that something signs it in, and — unless you pass -no-smoke — that a real completion comes back, which is the only check that catches a plan whose quota is exhausted, or a profile whose text_paths no longer find the answer in what the CLI prints. The same run measures two claims the profile makes rather than trusting them: that the CLI’s own shell is refused (it runs on the engine host, so a denial that stopped working is a seat reading whatever the engine user can read) and that its web tool reaches the network (the one local tool every profile deliberately keeps on). Both are believed only on evidence a model cannot invent — the current clock, read by the tool.

Both commands answer “signed in” the same way, from the environment the CLI is actually given rather than from the configuration: credential files a login wrote, a headless token, an api_keys value under auth.mode: api-key, a variable auth.mode: inherit-env forwards, or — for a CLI whose profile reads its provider’s key from its own environment (hermes, pi, opencode) — a credential-named variable in cli.env. doctor’s sign-in line names every one it found, and an entry with none is a no sign-in problem naming the route to take — unless the smoke test was answered, which proves the CLI authenticates some way the engine does not hand it (an endpoint that takes no key, a credential in its own configuration). A token auth.mode: api-key removes from the child is not counted. ${VAR} references and the variables inherit-env forwards are resolved in the process running crewlet llm — the secret store first, then that shell’s environment — so run it with the engine’s environment to get the engine’s answer.

llm list’s SIGN-IN column is the first route doctor would name: credentials (files a login wrote), token (a headless token), api key (under auth.mode: api-key), inherited token or inherited key (forwarded by auth.mode: inherit-env), environment (a cli.env key), or none.

On an anthropic entry doctor checks the two things that make every call on it a 400 while the config validates clean. The request is shaped from the model by a capability table compiled into this build, so doctor reads the model’s record from the vendor’s Models API (GET /v1/models/{id}, under the id the table reads — a Bedrock or dated spelling is the model it spells) and compares the thinking types, the effort levels and the output ceiling with what the entry sends. A disagreement that puts a field the model refuses on the wire is a problem, naming what to change: a claude_model for an alias the table does not know, the row a claude_model names, or — for a model the table reads itself — a Crewlet whose table matches the API. One where the table is merely more cautious than the model is a note. A gateway that does not serve /v1/models (a 404, 405 or 501, or a 200 that is not a model record) is reported as not served and is not a problem; any other failure of that read is. That read is made on one key and benches nothing, whatever it is answered: /v1/models is not the route a phase calls, and a gateway that refuses a key there (or rate-limits that route alone) while its messages route takes it would otherwise leave the round below with every key cooling. Then, unless you pass -no-smoke, it sends one real round in the shape a phase sends — a tool offered, no tool choice forced, the instruction naming it, the entry’s own thinking, streamed, at effort low — and certifies that a call to that tool came back. The Models API read bills nothing, so it runs under -no-smoke too. doctor exits non-zero when any entry has a problem, and an entry that does not build at all (a ${VAR} model that resolved to nothing, a dial its resolved model refuses) is reported as one beside the others. openai entries are not examined.

login has four shapes because the vendors do:

ShapeWhen
(no flag)Broker the vendor’s own interactive login and keep the result
-from-hostAdopt a login this machine already has, e.g. from running the CLI by hand
-capture-tokenMint a headless token into the secret store — the only shape a remote code sandbox can use, because a token is one scoped revocable variable and credential files never leave the engine host
-capture-token -print-tokenThe same mint, written to stdout and stored nowhere — for an operator whose secrets live in somebody else’s manager. It refuses to run on a terminal: this is a credential, and a token in a scrollback outlives the command, while a screen-share or a shell history outlives the scrollback. Pipe it or redirect it. The two-step alternative (-capture-token, then secrets get -reveal) writes the token into the store on the way past, which is precisely what this avoids.
-token-stdin / -username U -password-stdinStore a credential you already hold, where the CLI genuinely has one

export packs a login into one portable blob so another host can come up already authenticated. -secret-store writes it to the secret store instead of stdout; without it the blob goes to stdout in the clear, which is what you want when piping into your own secret manager and never what you want in a shell history.

import is the other half, and it reads the bundle from stdin — a credential on argv is visible in ps and lands in shell history, and the natural spelling is a pipe anyway:

Terminal window
crewlet llm export claude | ssh other-host crewlet llm import claude

It refuses to overwrite an existing login, loudly: a host that has been running holds the fresher refresh token, and restoring a boot-time blob over it is how a fleet logs itself out. crewlet llm logout <KEY> first if you mean to replace it.

The engine also restores a bundle itself, at boot, from cli.auth.credential_bundle or the conventional CREWLET_LLM_CLI_<KEY>_CREDENTIALS variable — and only into an empty credential directory, by the same rule. That path authenticates a fresh container before its first turn rather than after an operator remembers a command; import is for the hosts that are already up. See environment variables.


crewlet mattermost provision <company.yaml> [-admin-token TOKEN]
[-secret-store | -env-file PATH | -print]
[-rotate] [-dry-run]

Creates (or updates) one Mattermost bot account per Mattermost-enabled agent seat — a role whose mattermost.bot_token is a whole-value ${VAR} reference — adds it to the configured team and channels, and mints its personal access token into the exact ${VAR} the YAML references.

Unlike the Slack provisioner there is no app manifest, no local ledger and no OAuth click: Mattermost is its own directory, so a seat is found by looking up a deterministic username and the reconcile is stateless. The one manual prerequisite is a system-admin personal access token (-admin-token, or $MATTERMOST_ADMIN_TOKEN) — creating bot accounts and minting their tokens both require system-admin rights, and an admin must first enable personal access tokens in System Console → Integrations.

A bot hears only what it has joined, so the channel list is not a convenience: a bot that exists, authenticates and is in no channel is an agent that never wakes, and the failure is silent on both sides. A channel that does not exist is a note rather than an abort — half a fleet joined and the run stopped is a worse state than every bot joined to the channels that do exist and a line saying which did not.

A plain re-run does not rotate a working token. Mattermost returns an access token’s value once, so minting every run would revoke the credential every bot’s websocket is currently authenticated with — an operator adding a tenth seat would take the other nine down. A bot is left alone when the variable holding its token still has a value and the account still has a token under this tool’s description (crewlet-<handle>). -rotate mints for every bot regardless, retiring the previous one after recording the new.

FlagDescription
-admin-tokenSystem-admin personal access token. Falls back to $MATTERMOST_ADMIN_TOKEN. The bots’ own tokens are what this run mints, so it cannot bootstrap itself from them.
-secret-store / -env-file PATH / -printWhere minted credentials go — exactly one, and there is no default. See the secret store. -print writes export VAR=… lines and, when a run rolls back, unset VAR for each — the stream is meant to be sourced, and a comment is a no-op to a shell, so an operator who piped it into source would otherwise keep a revoked token exported.
-rotateMint a fresh token for every bot, including bots whose current one still works. Restart the engine afterwards.
-dry-runPrint the plan and touch nothing. It is the same plan the run uses.

A bot this run created is rolled back by revoking every token on it — nothing else has ever minted there. On a bot that already existed, only the token this run minted is revoked. See Mattermost integration.


crewlet mattermost doctor <company.yaml> [-admin-token TOKEN] [-config PATH]

Checks a Mattermost install by exercising what actually breaks, in the order it breaks — and no operator credential is required: the seat tokens already in the config are what the engine authenticates with, so they are the honest thing to check with, and minting an admin token to find out whether a company works is a step that exists only to be skipped. Pass -admin-token (or export MATTERMOST_ADMIN_TOKEN) to run the shared checks as somebody else.

CheckWhat it catches
GET /system/ping, unauthenticatedWrong URL, no route from here. First and without a credential, because a bad token must not make a healthy server look dead — the two have completely different remedies.
ServiceSettings.SiteURL vs integrations.mattermost.urlThe setting with no error message: Mattermost accepts a websocket only from a client whose Origin matches SiteURL, so a mismatch silently costs every human live updates while agents keep working. See The Site URL. A path-only difference is reported separately — the socket is fine, but the server builds its absolute links from its own value.
A websocket upgrade sent with a browser’s OriginWhat every human’s live feed does, including a reverse proxy that drops Upgrade. The Origin carries scheme and host only, because that is the exact string the server compares.
The configured teamChannels are team-scoped, so a team that does not resolve is a company where no bot can be placed.
Per seat: its own credential, a real socket, its channelsA revoked token, a disabled bot, a ${VAR} that never reached this deployment, or a bot in no channel — which hears only direct messages while its account looks perfectly healthy.

Each seat is checked with its own credential, because “the server accepts sockets” and “this bot wakes” are different questions and only the second delivers a message. A seat that fails early is not asked the later ones: one whose token did not resolve is never dialled, and one whose credential is refused is never asked about its channels — reporting those as separate failures would send an operator after faults nobody observed. The same rule governs the run as a whole: an unreachable server, an unreadable server config or a missing credential stops the checks and says so, because one failing line with nothing after it otherwise reads as “one thing is wrong” when it means “nothing else was even asked”.

Read-only. Exits non-zero when any check fails, so it drops into a deploy script.


crewlet gitlab provision <company.yaml> [-admin-token TOKEN]
[-secret-store | -env-file PATH | -print]
[-public-url URL] [-mode group|instance]
[-rotate] [-decommission]
[-token-expiry-days N] [-dry-run]

Idempotent reconcile from company config to GitLab state. For each agent seat that declares a GitLab credential under mcp_env.gitlab as a whole ${VAR} reference, it ensures a service account exists (username <username_prefix><handle>), adds group and project membership, mints a personal access token into that variable, and — with -public-url — registers the group webhook at <url>/webhooks/gitlab. The positional argument is the Tier B company YAML; its integrations.gitlab.provisioning block supplies the group, access levels, and token scopes. Human seats are resolved, never created.

A dry run says what it would do to the signing secret. It is the most consequential thing a run can do — replacing the key a working hook signs with fails every delivery in flight until the new value reaches the engine — so the plan states which of untouched / reused / minted / rotated will happen, and into which ${VAR}. That decision is made by the same function the real run uses, so the two cannot disagree.

A run that recorded something says what is left to do. Where the values went and what still has to happen are different questions, and only one of the three sinks answers “source a file”: -env-file needs sourcing and a restart, -print needs the values moved before the terminal closes, and -secret-store needs the current revision re-activated (crewlet config activate) so the running engine rebuilds its secret snapshot — it needs no file, which is exactly why a report that stopped at “recorded in the encrypted secret store” read as finished. A run that changed nothing prints no follow-up.

A plain re-run does not rotate a working token. GitLab returns a personal access token’s value once, so a provisioner cannot verify that what it recorded last time still matches — and minting every run is an outage, because the engine is running with the old value: an operator adding a tenth seat would revoke the nine credentials the other agents are authenticating with. A seat is left alone when the variable holding its token still has a value and the account still has a usable token under this tool’s name (crewlet-<handle>), and both halves are checked because either alone is wrong — a recorded value whose token was revoked leaves an agent 401ing for ever, and a live token nobody wrote down cannot be deployed. -rotate mints for every seat regardless, retiring the previous one after recording the new: never before, or a failed record leaves the seat with nothing.

FlagDescription
-admin-tokenOperator credential — a top-level group Owner PAT with api scope on GitLab.com, or an admin PAT self-managed. Falls back to $GITLAB_ADMIN_TOKEN. The seats’ own tokens are what this run mints, so it cannot bootstrap itself from them.
-secret-store / -env-file PATH / -printWhere minted credentials go — exactly one, and there is no default: a run with nowhere to put what it mints creates live credentials at the third-party app and prints none of them. See the secret store.
-public-urlThis deployment’s public base URL; the group webhook is registered at <url>/webhooks/gitlab. Overrides integrations.public_base_url for this run, and registration is skipped only when both are empty — a hook pointing at the wrong host is worse than none, because the instance then reports a healthy integration.
-rotateMint a fresh token for every seat, including seats whose current one still works. Restart the engine afterwards.
-decommissionDelete service accounts whose seats have left the config. Scoped twice: the username must start with provisioning.username_prefix (never empty — it defaults to crewlet-) and the account must be a member of this company’s group, because either alone is too broad. That group scan is used in both modes and it is deliberate: every seat is made a member of provisioning.group whatever created it, so a managed account of this company is always in the group — while the instance’s own service-account listing also holds every other company’s accounts on the box, which a shared prefix would then sweep. Only the DELETE route differs by mode; sending one down the wrong route answers “no such account” for one that is still live. An account the instance refuses to delete because it is not a service account is reported rather than aborting — that refusal is GitLab catching what the scan should not have proposed, so it is a signal about the prefix.
-mode group|instanceWhere service accounts are owned, for this run only; it defaults to integrations.gitlab.provisioning.mode, which is what the engine’s own reconcile reads. group (the default, and all GitLab.com offers) creates them under provisioning.group; instance creates them on the instance itself, which needs an instance-administrator PAT and buys an account that is a member of nothing until this run adds it — so one company’s seats can span several top-level groups, and an account survives its group being deleted. Everything downstream is identical: memberships are added the same way, and the token API is user-scoped on both. An unknown value is refused before the config is loaded.
-token-expiry-daysLifetime minted onto each token. Omitted, no expires_at is sent and the instance’s own policy applies — GitLab.com caps personal access tokens at a year regardless, which is the instance enforcing its policy rather than this tool choosing one. Nothing in Crewlet renews a credential on a schedule, so a lifetime nobody renews is an outage with a date on it.
-dry-runPrint the plan and touch nothing. It is the same plan the run uses.

An account this run created is rolled back by revoking every token on it — nothing else has ever minted there. On an account that already existed, only the token this run minted is revoked: sweeping it would take an administrator’s own token with no way to tell that it had. Both go through a detached context, because the failure is often the cancellation itself.

See GitLab Integration — Provisioning for the permission matrix and the full walkthrough.

crewlet jira provision <company.yaml> [-secret-store | -env-file PATH | -print]
[-public-url URL]
[-recreate-webhook] [-dry-run]

The Atlassian tracker’s reconcile, and it is a different shape from its peers: Jira issues no credentials on a provisioner’s behalf. A Cloud API token is created by the person it belongs to at Atlassian’s own account site, and a Data Center personal access token can only be minted for the calling user — so a command that offered to provision accounts would be printing instructions dressed as actions. What it does instead is the three things Jira genuinely allows, each answering a question that is otherwise invisible until an issue reaches nobody:

  1. Which account each seat’s credential authenticates as. It calls /myself with every seat’s own token, read from mcp_env.atlassian or mcp_env.jira, and reports the seats that resolved and the seats that did not. A seat with no account id receives no Jira events at all — the routing gate drops every target that names it — and nothing else in the engine says so. Human seats are never probed: they hold no tool credential and are reached by contact.atlassian_account_id or by email.
  2. Whether every project the org chart names exists, and whether Jira’s own project lead agrees with the org chart’s. A key with a typo in it is a routing gap that produces no error anywhere; a disagreement about the lead is reported and never failed, because a human manager owning the project while an agent triages it is an ordinary arrangement.
  3. The inbound webhook, on Data Center only, at <public-url>/webhooks/jira — subscribed to exactly the events the parser routes, delivering the whole body, signed with an HMAC secret. On Cloud the step is skipped with a note rather than attempted: a dynamic webhook there belongs to an app, so the endpoint refuses an API token however privileged it is, and reporting that 403 would send an operator to rotate a credential that is fine. Cloud events arrive through the Forge app instead.

A working webhook secret is never replaced. If integrations.jira.webhook_secret already resolves, that value is registered as-is. Minting on every run would be an outage: the engine holds the old secret, so the instance would start signing every delivery with a key nothing can verify. A secret that resolves to nothing is minted into the ${VAR} the config points at and recorded in the sink you chose; -recreate-webhook forces a rotation, which invalidates the secret every other deployment of this company holds.

FlagDescription
-secret-store / -env-file PATH / -printWhere a minted webhook secret goes — exactly one, and there is no default. See the secret store.
-public-urlThis deployment’s public base URL; the webhook is registered at <url>/webhooks/jira. Overrides integrations.public_base_url for this run, and registration is skipped only when both are empty — a hook pointing at the wrong host is worse than none, because the instance then reports a healthy integration that delivers into the void.
-recreate-webhookDelete and remake the hook to mint a fresh secret. The only recovery for a secret that was lost, because the value cannot be read back off the hook — and destructive for every other deployment holding the old one.
-dry-runRead the instance and report; register nothing. The sink is not opened, so it prompts for no passphrase.

A hook somebody else registered is reported and never touched: an instance may carry integrations that are none of this run’s business, and taking over the first one found by name would break one.

The run refuses outright when the org credential in integrations.jira.token is rejected — nothing else it reported would be trustworthy.

See Jira Integration — Provisioning for the full walkthrough.

crewlet github provision <company.yaml> [-secret-store | -env-file PATH | -print]
[-public-url URL]
[-recreate-webhooks] [-dry-run]

The hosted code host’s reconcile, and it is the same shape as Jira’s for the same reason: GitHub issues no credentials on a provisioner’s behalf. There is no API that creates a user, and the API that once minted a token for somebody else was withdrawn in 2020 — so a command that offered to provision accounts would be printing instructions dressed as actions. What it does instead is the two things GitHub genuinely allows:

  1. Which account each seat’s credential authenticates as. It calls GET /user with every seat’s own token, read from mcp_env.github under whichever key that seat’s tools use (GITHUB_TOKEN, GITHUB_PERSONAL_ACCESS_TOKEN, GH_TOKEN, or an Authorization header), and reports the seats that resolved and the seats that did not. A seat with no login receives no GitHub events at all — the routing gate drops every target that names it — and nothing else in the engine says so. Human seats are never probed: they hold no tool credential and are reached by contact.github_login.
  2. The inbound webhooks, at <public-url>/webhooks/github, subscribed to exactly the events the parser routes and delivering JSON — GitHub’s own default is form-encoded, which nothing here can decode. One hook on the organization where the credential may register it, covering every repository in the org including ones created later; otherwise one per repository named in provisioning.repos. org_webhook: true turns a credential that cannot into a failed run rather than a silent fallback, because an operator who asked for the org arrangement must not quietly get the other one. admin:org_hook is the scope, which a fine-grained token cannot carry at all.

A working webhook secret is never replaced. If integrations.github.webhook_secret already resolves, that value is registered as-is. Minting on every run would be an outage: the engine holds the old secret, so GitHub would start signing every delivery with a key nothing can verify. A secret that resolves to nothing is minted into the ${VAR} the config points at and recorded in the sink you chose; -recreate-webhooks forces a rotation, which invalidates the secret every other deployment of this company holds.

One unhookable repository does not stop the rest. A list will contain one that was renamed, archived, or made private to a team this credential is not in; each is reported with what is wrong and the others are still hooked. GitHub answers 404 for both “absent” and “invisible to this credential” — deliberately, so a probe cannot enumerate what exists — so the report says both rather than sending an operator to look for a repository that is right there.

FlagDescription
-secret-store / -env-file PATH / -printWhere a minted webhook secret goes — exactly one, and there is no default. See the secret store.
-public-urlThis deployment’s public base URL; the hooks are registered at <url>/webhooks/github. Overrides integrations.public_base_url for this run, and registration is skipped only when both are empty — a hook pointing at the wrong host is worse than none, because GitHub then reports a healthy integration that delivers into the void.
-recreate-webhooksDelete and remake every hook to mint a fresh secret. The only recovery for a secret that was lost, because the value cannot be read back off a hook — and destructive for every other deployment holding the old one.
-dry-runRead GitHub and report; register nothing. The sink is not opened, so it prompts for no passphrase.

A hook somebody else registered is left alone: matching is on the delivery URL, never on a name, because GitHub gives a webhook no name at all and an organization carries hooks other integrations put there.

The run refuses outright when the credential in integrations.github.token is rejected — nothing else it reported would be trustworthy. That token is optional for the engine (without it, thread activity degrades to the payload’s author and assignees) and required here, because there is no degraded form of registering a webhook.

See GitHub Integration — Provisioning for the full walkthrough.

crewlet slack provision <company.yaml> [-secret-store | -env-file PATH | -print]
-public-url URL
[-config-token TOKEN] [-ledger PATH]
[-handles a,b] [-reinstall]
[-no-install] [-dry-run]

Every other third-party app’s provisioning runs unattended. Slack cannot, and that is not an implementation gap: installing an app into a workspace is an OAuth grant, and OAuth exists precisely so that a person decides. So the run creates and updates the apps by itself, then hands the operator one authorize URL per seat and takes the code back — and where there is nobody to ask (-no-install, -dry-run), it prints the URLs and stops rather than pretending.

For each seat whose integrations.slack credentials are whole ${VAR} references it builds the canonical manifest (internal/slack is the single source of truth for the scopes and events), creates or updates the app, records the signing secret into the variable signing_secret points at, and — after the click — the bot token into the one bot_token points at. A seat whose credentials are literals is one an operator manages by hand: it is reported and skipped, never rewritten.

An unchanged manifest is not pushed. The manifest methods are Slack’s slowest rate class, roughly one request a minute, so a company of seven re-running would otherwise spend minutes achieving nothing. The ledger holds a fingerprint of the last manifest Slack accepted, and a re-run compares against it.

One seat’s failure does not cost the others. A mistyped code paste or a refused manifest is recorded against that handle, the remaining seats still provision, and the command exits non-zero naming what failed. Everything completed is durable — the ledger is written after every mutation — so a re-run resumes.

The app-configuration token is persisted before it is used. Slack’s rotation is single-use in both directions: the call that returns a new refresh token invalidates the one it was given. A run that rotated and then failed to record the result would lock the operator out of their own apps, so the pair is written to the ledger first, and a still-valid access token is reused rather than rotated again. For the same reason the ledger’s pair beats -config-token and $SLACK_CONFIG_REFRESH_TOKEN, which is the reverse of the usual rule: a value left in a shell export is dead the moment the first run used it, and preferring it would trade the only live pair for a retired one on every run after. The flag and the variable are a bootstrap for a ledger that holds nothing.

A recorded app that no longer exists is replaced, not reported kept. Its manifest fingerprint still matches, so without a probe the seat reads as healthy while the bot token in its ${VAR} authenticates as nothing. Every run validates each recorded app id first — one call, no write — and reads app_not_found / invalid_app_id / invalid_app as gone; a permission refusal is not an absence and never triggers a replacement, because an app this credential may not touch still exists and replacing it would leave two.

A code that belongs to another app is refused and records nothing. One authorize URL is printed per seat and they look alike; pasting the wrong one would mint a colleague’s bot token into this seat’s variable, and the seat would post as them with nothing reporting it.

FlagDescription
-public-urlPublic HTTPS base URL of this deployment. Required unless integrations.public_base_url is set, which it overrides for this run: every app’s Events API request URL and OAuth redirect URL are built from it, so an app created without one delivers nowhere and cannot be installed. A public_base_url written as a ${VAR} this process cannot resolve counts as unset, and the refusal says so.
-secret-store / -env-file PATH / -printWhere the minted bot token and signing secret go — exactly one, and there is no default. See the secret store.
-config-tokenThe operator’s app-configuration refresh token, from api.slack.com/apps → Your App Configuration Tokens. Falls back to $SLACK_CONFIG_REFRESH_TOKEN. Read from the environment alone, never from the secret store: it is the operator’s credential rather than the company’s.
-ledgerThe app ledger (default slack-apps.json beside the company document). It holds the client secrets Slack returns only at creation, so it is written 0600 — gitignore it like .env.
-handles a,bOnly provision these seats. Worth having against a method that allows about one request a minute.
-reinstallRedo the OAuth install even where a token is already recorded. Required for a scope change to take effect — a bot token carries only the scopes it was minted with — and destructive: the new install revokes the token every running node is authenticating with.
-no-installCreate and update the apps and record the signing secrets, then print the authorize URLs instead of asking for codes. For a non-interactive run.
-dry-runPrint the plan and check every manifest through apps.manifest.validate, which writes nothing: no app created, no manifest pushed, no install run, and the sink is not opened either, so it prompts for no passphrase. That check is why a dry run touches the network at all — apps.manifest.create is Tier 1, roughly one request a minute, so a malformed manifest discovered from the create costs a minute per seat and leaves the seats before the bad one already created. Validating needs a config token, so a dry run that cannot get one prints the plan and says the manifests were not checked; getting that token may rotate it, which is the one write a dry run makes and has to.

Run the API server first, publicly reachable at -public-url: Slack verifies each app’s request URL with a url_verification challenge, which the edge answers unconditionally — it has to, because during provisioning the signing secret does not exist yet and a verified handshake would be impossible.

See Slack Integration for the full walkthrough.

crewlet confluence import <company.yaml> <directory> [-space KEY] [-prune]
[-config PATH] [-dry-run]

Publishes a tree of authored markdown into Confluence. One walk, two destinations, decided by the file: a file whose frontmatter declares a trigger: is a tool skill and goes to knowledge.skills_container with the leading code block the engine parses back out; everything else is a knowledge doc, published as prose into the space its parent directory names, titled by its first # H1.

The routing is the FILE’S, not the directory’s, because a skill is identified by what it declares — an operator who files one under ENG/ still means a skill, and publishing it there as prose would put an instruction meant for one phase of one turn into every seat’s context.

Every target space is checked before a single page is written. A typo in a directory name would otherwise be discovered half way through, leaving an operator to work out which pages landed. The importer never creates a space: that names a container the whole company then works in, and it is not this command’s guess to make.

A page that already exists is updated in place, matched by title within its space. Confluence has no external-id field, so a page somebody renamed in the UI is orphaned and a re-import creates a second one. That is the backend’s limitation, reported rather than worked around: a marker page or a label pressed into service as an identity would be this tool inventing a second answer to a question Confluence already answers, and the two would disagree the first time somebody moved a page.

Nesting and labels come from frontmatter. Frontmatter may also declare a parent: — the title of a page in the same space to nest this one under, which is the one thing a flat directory of files cannot say about a wiki that has trees in it — and labels:, the author’s own page labels. The plan is ordered parents-first so a parent: naming a page published by the same run resolves; a cycle stops the walk naming the files, and a parent nobody publishes is a note and a page at the space root, because a doc nobody can read is worse than a doc in the wrong place. An existing page is never re-parented — where a page sits is something people move deliberately, and a run that dragged it back every time would be fighting them with no way to say so. Labels are attached on every run, not only on create, because the server call is idempotent and a label an author adds to a file that already publishes has to reach the page somehow; a label that will not attach is a note, not a page failure.

Provenance is a different question, and it is recorded. Every skill page this command writes gets the global label crewlet-skill. That says nothing about which page a file belongs to — it says only that the importer wrote it, which is a fact no field on the page carries and which only the writer can know. -prune is the one caller that needs it: an orphaned page is deleted only if this tool published it, because a lead who authored a skill by hand in the wiki has no local .md and would otherwise lose their work on the next import. A label that cannot be written is a note, not a page failure — the page is published and correct, and what is lost is the ability to prune it later.

Page failures are isolated: a restricted page or one 403 does not cost the other forty. The run reports what failed and exits non-zero.

FlagDescription
-space KEYPublish tool skills into this space instead of knowledge.skills_container. Empty reads $CREWLET_TOOL_SKILLS_SPACE, then the config field. Skill files only — a knowledge doc takes its space from its parent directory.
—Frontmatter on a knowledge doc may declare parent: (the title of a page in the same space to nest under) and labels: (the author’s own, lower-cased and de-duplicated because that is what Confluence stores). See below.
-pruneAfter publishing, delete skill pages in the skills space that carry the crewlet-skill label and whose key no local file publishes any more. Three conditions, all required: in the skills space, labelled, and parsing as a skill whose key this run’s tree does not publish — the label protects a hand-authored page, the parse protects an ordinary page filed in the same space, and the key comparison is what makes a renamed skill a delete-and-create rather than a silent duplicate. The orphan set is derived by subtraction, so a prune that cannot enumerate the space deletes nothing and fails the run: a partial read would make the orphan set larger and delete live pages. The set is taken from the plan, not from the writes that landed, so a page whose update happened to 403 is never deleted as an orphan of itself. A page that declares a trigger and does not parse has an unknown key and is reported rather than deleted.
-configTier A config naming this node’s store and keyring, for resolving the ${VAR}s in the company’s confluence: block.
-dry-runPrint the plan and write or delete nothing.

A company that has turned tool skills off (knowledge.skills_container: "") has no space for a skill file to go to: a tree containing one stops the walk naming both the setting and -space, rather than filing an instruction meant for one phase of one turn into a space every seat searches.

See Confluence Integration.


crewlet confluence provision <company.yaml> [-secret-store|-env-file PATH|-print]
[-public-url URL] [-recreate-webhooks] [-dry-run]

Registers the inbound webhooks that make Confluence push page and comment events to this deployment, and mints the credential each one carries into the ${VAR} the company document already points at.

The two deployments are registered differently, because they authenticate differently. On Cloud the run creates one hook per event at <public-url>/webhooks/confluence/<event>?token=…: Cloud signs nothing and honours no registration field for a header, so a shared token in the URL is the whole authentication, and one hook per event is how the engine knows which event fired at all (a Cloud payload does not say). On Data Center it creates one signed hook at <public-url>/webhooks/confluence covering every event, verified by HMAC over the body against integrations.confluence.webhook_secret.

A re-run converges rather than adding: hooks are found by the crewlet: name prefix this engine registers under, so one whose address has moved is re-pointed rather than duplicated, and one an operator registered by hand is never touched. A hook that is already correct is left exactly as it is, and one event this instance refuses is reported without stopping the rest.

FlagMeaning
-secret-store / -env-file PATH / -printWhere a minted token or secret goes. The same three sinks every provisioning command takes; a run with none of them refuses rather than minting a credential it cannot record.
-public-url URLThis deployment’s public base URL, which the hooks are registered against. Overrides integrations.public_base_url for this run, and registration is skipped only when both are empty.
-recreate-webhooksDelete and remake every hook with a fresh token or secret. Destructive across deployments: the previous value stops working everywhere else this company runs, so it is the recovery for a leaked or lost credential rather than part of an ordinary run.
-dry-runRead and report; register nothing.

See Confluence Integration — Webhooks.


crewlet confluence resync <company.yaml> [-space KEY] [-config PATH]

Runs the engine’s own tool-skill walk of the Confluence skills space against a throwaway registry and prints what admitted, so you can see what a running node’s next walk will see. The read-only diagnostic beside the importer, and it exists because a page that fails to admit is invisible: the only symptom is guidance that never appears in an executor prompt.

TS holds 12 page(s): 9 skill(s), 3 ordinary page(s).
deploy Cutting a release
incident-review Running a post-incident review

A page that declares a trigger: and does not parse is printed separately and exits non-zero. That is the case worth failing on: somebody wrote a trigger and got the rest wrong, and counting it as an ordinary page is exactly how a skill goes missing unnoticed.

-space overrides knowledge.skills_container, for checking a space the company document does not name yet.

Skills only. Knowledge docs are searched live at query time and never loaded into a registry, so for them there is nothing to resync.

It does not reach into a running engine, and a running engine does not need it to. A page edit reaches every node through the page webhook, and every node also walks its skills space every 10 minutes (give or take a fifth, so a fleet does not walk in lockstep), so a change the webhook path missed is applied within 12 minutes without a restart. See Keeping every node current.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.