Skip to content
You are reading documentation for unreleased main. Read the 0.1 version.

API Endpoints

The Crewlet API is served by any node with the ingress role (api.port > 0 in the Tier A config). By default a node has every role, so it runs alongside the agents in one process; a node with -roles data,ingress serves only these routes and reaches the rest of the fleet over the event stream. The routes below are identical either way.

A node without the ingress role binds api.port too, and serves only its two probes there — /health and /ready — plus a seats node’s agent-mode tool bridge, unless api.public moves the bridge to its own listener: see Probes on a node without ingress.

There is nothing to install for it — it is compiled into the binary, along with the dashboard it serves and the WebSocket that is the dashboard’s data plane (see WS /ws/stream).


Every route below is served on api.port, unless the node’s Tier A sets api.public.port. Then the routes outside parties call — the ones that authenticate by a provider signature or a signed per-run token, never an operator credential — move to that listener and are served nowhere else:

PrefixWho calls itWithout api.publicWith api.public
/webhooks/*Vendors (deliveries) and a person’s browser returning from a vendor’s app flow (/webhooks/slack-oauth, /webhooks/github-app)api.portapi.public only
/otlp/{token}/v1/{signal}A sandbox exporting a coding run’s telemetryapi.portapi.public only
/mcp/{token}A sandbox calling its seat’s tools (agent mode)api.portapi.public only
Everything else — /health, /ready, /, /dashboard, /favicon.ico, /static/*, /ws/stream, /query/*, every REST read, /config, /secrets, /setup, /operator/*, /backup, the /work/* and /fleet/* gesturesOperators, the dashboard, the CLI, orchestratorsapi.portapi.port only

A public route asked of api.port answers 404 with the code no_route and a hint naming api.public, since whoever asks there is on the operator’s own network, most likely a vendor configured with the wrong port. Every other route asked of the public listener answers exactly what a path nothing serves answers — the same 404 no_route, the same generic hint — so the published socket says nothing about the admin routes behind it. It requires no operator token and decides before the guard runs, so a guarded route there is that same 404 whatever credential is sent. The probes are deliberately not public: /health describes the node to whoever runs it, so point every health check at api.port. A node without the ingress role serves its probes on api.port and /mcp/{token} on api.public when it is set, beside the probes otherwise — and nothing at all when its api.port is 0. See Deployment → Exposing webhooks without the admin API.


Three deadlines bound a request on every route below, and each covers a phase the others do not. None is configurable — they are properties of what the surface is for, not of a deployment.

PhaseBoundWhat it stops
Request line and headers10 sA connection opened and left silent — the cheapest denial there is against a listener.
Reading the request body30 sA client that dribbles: a body under every size cap, delivered a few bytes at a time. The size cap and this are different failures, and a cap alone stops only the first. 30 s carries the largest body any route accepts — a 25 MiB webhook delivery — at roughly 7 Mbit/s sustained, far below what any forge, CI runner or operator workstation delivers.
Between keep-alive requests60 sA client that completed one request and then went quiet while holding its connection slot.

A body that does not arrive inside its deadline fails the read like any other truncated delivery: 400 with unreadable_body (logged as webhook_body_unreadable on a webhook path). There is deliberately no whole-request timeout: it would have to be large enough for the largest body on the slowest link, which makes it no bound at all on a small one. Long-lived responses — WS /ws/stream above all — bound themselves.


A node that has been told to stop (SIGTERM, or Ctrl+C once) keeps serving HTTP for the whole of its drain, and closes its listener only once the drain has completed. What changes from the drain’s first moment is which requests it will still take:

RequestDuring a drainWhy
GET /health200, status: "shutting_down"Liveness. An orchestrator that could not reach the node would kill it in the middle of the turns the drain exists to finish.
GET /ready503, reason: "draining"Readiness: takes the node out of rotation so traffic moves to a peer.
Every other read (GET, HEAD, OPTIONS): the dashboard, the REST reads, /query/*, /ws/streamServedA read starts nothing, and it is how the drain is watched.
/mcp/{token} and /otlp/{token}/v1/{signal}ServedThey carry the tool calls and spans of coding runs that started before the drain. A detached run outlives the turn that started it, so the drain never waits on one, and refusing these would shorten no drain and only break a run mid-flight.
Every /webhooks/* route, whatever its method503A delivery is new work, and one of the two GET landings acts: the GitHub App return seals a credential and writes a config revision, and an install arrival asks the reconcile loop for a pass. The Slack OAuth landing only renders a page and is refused with the rest, because a per-route carve-out is what refusing by default avoids.
Every other write: /config, /secrets, /setup, /backup, the /work/* writes, POST /operator/mcp, POST /operator/act/{tool}503Each one starts work or changes the company the drain is leaving. Refusing by default is what keeps a write route added later from slipping through a drain.

/operator/mcp is the one route the by-method rule splits, because it is mounted for every verb: its POST — every JSON-RPC call, reads included — is refused, and its GET server-to-client stream is served like any other read. Its DELETE, which ends a session, rides the default with the writes; the session dies with the listener a moment later either way. /mcp/{token} is not split, because the whole prefix is served: a coding run’s tool calls are the one thing on this listener the node must not break.

A refusal is 503 with a Retry-After of 30 seconds, long enough for a load balancer following /ready to have moved traffic to a peer, and a body the CLI and the dashboard both render:

{
"error": "draining",
"detail": "this node is draining for a shutdown: the turns already running finish, and nothing new is started here",
"hint": "retry against another node, or once this one has restarted; /ready answers 503 for as long as the drain lasts"
}

A write still needs its token first: an unauthenticated write answers 401 whether or not the node is draining. And a request that was already running when the drain began is not interrupted by it; it is cut only if it is still running five seconds after the listener starts to close.

A node with a public listener drains both by this one table — its webhooks are refused on api.public and its sandbox endpoints served there — and closes api.public first, then api.port, each with its own five seconds, so /health answers until the very end even while an open bridge session holds the public listener for its whole grace.


Every answer the API writes other than a success is JSON with an error code: the auth guard’s 401 invalid_token, each route’s own refusals, and a request nothing on the node serves — 404 no_route for a path (one this node’s build does not have, or one its configuration leaves unmounted, such as POST /work/{id}/purge on a company with no native tracker) and 405 method_not_allowed, with an Allow header, for a served path under another method. Both of those carry a detail naming the method and path.

The CLI and the dashboard rely on it. A non-2xx answer with no code was written by something in front of the node — most often a reverse proxy’s read timeout — and says nothing about what the node did, so a write that meets one is reported as unknown, with the operation id to finish it under, rather than as refused. An answer with a code is the node’s own, and a refusal from the node means nothing was done.


MethodPathDescription
GET/healthLiveness + the engine-health envelope (see below). Stays 200 through a drain (see During a drain); use /ready to steer traffic
GET/readyReadiness for a load balancer: 503 while draining, before the first config revision applies, or on a shed or stuck posture, and 200 otherwise. A 503 names why in reason: draining, unconfigured, shed or stuck, in that order of precedence. It never reads the fleet’s presence or alarm counts, which decide nothing here. A node without the ingress role answers a different question on the same route — whether it is doing its work — see Probes on a node without ingress
GET/agentsList agent roles, each merged with live state from the in-memory projection (including the in-flight live_call). Human seats are excluded — they appear only in /org with "kind": "human"
GET/agents/{id}Single agent — role, the live overlay (incl. live_call), and llm_history: the seat’s finished phases newest first, capped at 50. {id} is the seat’s handle, which is what every roster row carries as its id; a role name is accepted too
GET/agents/{id}/memoryDurable memories (personal, episodic, counterparty, synthesized skills). Same {id} — the handle resolves to the derived agent id the diary is keyed by
GET/agents/{id}/memory/episodes/{episode}One episode whole — its ask and its account complete, which the memory listing carries the openings of (see below)
GET/orgThe company’s charter and its seat and unit tree, in an explicit public shape that carries no contact identity, email, credential or deployment setting (see below). Human seats appear with "kind": "human"
GET/toolsRegistered tools, each tagged with the source that registered it — builtin or mcp:<server> (see Where a tool comes from) — plus its behavioural annotations, where it delivers, and its input_schema (see below)
GET/eventsRecent engine events from the event store (limit caps at 400; keyset-paged, see below)
GET/events/{event_id}Single event incl. payload — inside the 30-day history (event_history_seconds) like every other read of the log, so a link to an older event answers not_found: every node floors its lookup at the serving node’s horizon, whatever the retention sweep of the node holding a copy has or has not reached (see coverage)
GET/events/trace/{trace_id}All events in one trace, oldest first, capped at 500
GET/tokens/breakdownThe token-spend rollup by phase / model / provider entry / worker / seat — the live 24 hours, or any window of up to 90 company days from the replicated usage domain (see below)
GET/tokens/seriesThe same spend with a time axis — one bucket per company day or ISO week, split into bands (see below)
GET/agents/activityEvery seat’s turns over a window of company days — counts, the first-pass rate over reviewed turns, turn-duration quantiles and a day-by-day series (see seat_activity)
GET/schedulesConfigured role/unit schedules + next-run + recent dispatch ledger
GET/fleetEvery live node, its roles and labels, seat ownership, singleton duties, per-node config epoch, and where the company’s files are kept with what the object store’s collector last found. Always needs a token — it describes the deployment rather than the company, and the dashboard locks the screen that draws it (see below)
GET/sandbox-runsEvery detached sandbox run the engine still holds, read from the durable run record in the coordination store (see below)
GET/budgetsToken caps, the durable shared counter they are enforced against — per calendar window — and which scopes are being refused (see below)
POST/backupCopy this node’s store and stream estate into ?dir= on the engine’s host. Always needs a token — it writes every credential the company holds to a path the caller names (see below)
GET/fleet/brokerThe fleet broker’s membership: every live node’s broker kind as its presence advertises it, the JetStream metadata group as a member reports it, and every disagreement between the two — a member gone for good first among them. Always needs a token (see The broker’s membership)
POST/fleet/broker/remove/{node}Remove a member from the metadata group through a live member’s system account. ?confirm= repeats the node id; refused while the node holds a live presence lease as a member unless ?force=true. Operator-only
POST/fleet/broker/remove-peer/{peer}The same removal, naming the voter by the raft peer id GET /fleet/broker shows — for a voter whose name no member has heard. ?confirm= repeats the peer id. Operator-only
GET/integrationsEvery inbound surface, how it is wired, whether a signing secret is present, and what has arrived through it (see below)
GET/accessWho can reach the company through this engine and as whom: the API token LABELS the guard accepts (never a value), the person each one acts as, every human seat with its contacts and the state of its binding, and the auth posture — including company_writers, the tokens that alone may change a managed company document. Always needs a token (see below)
GET/credential-poolEvery providers.llm entry, each key it rotates through by variable name, and which of them a vendor is refusing and until when — this node’s pools beside the fleet’s cooldown ledger (never a value). Always needs a token (see below)
GET/backupsWhat the fleet has backed up: each owner’s newest point as the trim reads it, and every backup a person asked a node for, failures included. Always needs a token (see below)
GET/mcp-serversWhat each configured MCP server did on each live node — started, failed, tools served and the first failure — read off every node’s presence heartbeat, beside what the configuration declares (never a credential). Always needs a token (see below)
GET/workThe company’s own tracker: a filtered listing of work items, plus the last key number minted per project. Served only where tracker.backend is native — a company on Jira gets 404 unknown_query, not an empty board (see below)
GET/work/retentionWhat the state log is holding, what the trim concluded and which term is stopping it, every node’s position, and what this node costs to replace. Operator-only, reads included (see below)
POST/work/retention/ackPublish an operator backup floor, for backup_floor: operator
POST/work/retention/evict/{node}Install the eviction gate on a node on every log the trim counts nodes on — the tracker’s and the pages log — so the trim can pass a floor it is pinning. Refused 409 while the node holds a live presence lease; answers per log (see below)
POST/work/retention/readmit/{node}Lift it on every one of those logs — the inverse commit rather than a delete. Refused 409 while the node still lacks records a trim floor lets the log delete
POST/work/retention/capacityDrive a log’s byte-ceiling change as far as this node’s mode allows
GET/work/retention/maintenanceWhere that window stands and what is holding it
POST/work/retention/maintenance/abandonChange what the operation is trying to reach, never the barrier it must cross
POST/work/retention/maintenance/excludeRecord that a participant’s process is stopped and holds no outstanding request
POST/work/{id}/purgeDestroy a task and every row it produced, on every node. The one operation with no inverse: ?confirm= repeats the task’s KEY, ?project= names the container the record arbitrates under, ?reason= is required and is the only account of the task that survives — carried WHOLE on the one line the project’s lead is told, so a reason that would not fit that 600-byte line is refused 400 before anything is destroyed, rather than cut, and ?op_id= is how an unknown outcome is retried without appending a second purge — pass back the op_id a previous answer returned, unchanged: it carries the instant it was minted, and a node whose operation ledger may have lost the first purge’s row since — to the ledger’s thirty-day sweep — answers the retry unknown again rather than purging twice. An id that is not in the engine’s grammar is refused with 400 op_id_invalid: it carries no instant, so no node could tell whether it already ran. So is one over 128 bytes or holding anything but visible ASCII — a space included — since the broker carries the id in a header that trims its ends and rewrites a line break (see the three retention gestures that write for the whole rule). The answer carries the write’s outcome, its op_id and the position the record holds, which is absent for unknown: that outcome has none. Operator-only, and absent rather than 503 on a build with no tracker, where it answers 404 no_route like any path the node does not serve
GET/work/retention/reanchorThe live stream’s own created_at, which a reanchor’s confirmation has to echo, and the case a reanchor would answer
POST/work/retention/reanchorAdopt a recreated stream, or a broker restored from an older copy, at the next generation
GET/work/viewsOne container’s view strip: the six every container has without anybody saving one, and whatever was saved beyond them. ?container= takes the query grammar’s own spelling (workspace, project:ENG, unit:engineering, person:ana) — a project key is upper-cased and a unit is resolved to its id where the chart gave it one, so a team’s strip is one strip under either of its spellings and ?viewer= is whose personal views and pins order the strip — your own seat, or operator-only for anybody else’s, the same scope rule as /work/my-work; absent is the shared strip, which needs no credential
GET/work/views/savedEvery saved view ?viewer= can see, in EVERY container — the shared ones and their own personal ones, pinned first, each row carrying the container it lives in. What the dashboard’s view inventory and its sidebar’s pinned group read, because a view saved on a project board is in no workspace strip. ?viewer= takes the same scope rule as /work/views; ?counts=true counts each pinned view in its own container
GET/work/catalogueThe company’s vocabulary: the task types a create may name and the workspace’s custom-field declarations. ?archived=true also lists what was retired. The types are the EFFECTIVE set — the six this build ships plus whatever the company declared, a declaration replacing a builtin of the same slug
GET/work/projectsEvery project work is filed into, with its task_counts — {todo, active, done, closed}, one per status group, read from maintained columns and never aggregated per poll; there is no open, which folded the waiting work into the started work, so a caller wanting every unfinished item adds todo and active — its target_date (the day, YYYY-MM-DD on the company’s clock, the lead means it to be finished; omitted when none is set), its last_change (when the project’s work last changed and who changed it, ABSENT for a project nothing has been filed into), its chart-owned unit and its lead. ?q= narrows by a word in the key, the name or the purpose and ?unit= to the projects one unit owns — named by the unit’s id or by its name, in any case, since a stored unit carries whichever was current when the row was written. ?archived= SELECTS a set rather than widening one — false (the default) for the live projects, only for the retired ones alone, true for both — so “what did we retire” is a query rather than a caller’s own filter over a wider answer. ?sort= orders the whole selected set before the page is taken: one of key, name, unit, todo, active, done, closed, last_change, target, each optionally with a leading - for descending, defaulting to key, with the key breaking every tie; a project with no last_change or no target_date sorts last in both directions. An archived or sort value that is neither is a 400 naming the parameter and what it accepts. ?limit= caps at 200, which is also the default. The answer carries a census — {active, archived}, the same question under the same q and unit MINUS its archival term — so a caller that selected one set can still tell an empty set from an empty company; total is the census of the mode that was asked for. A set read, so it carries complete and its incomplete beside the read level
GET/work/projects/{key}One project in full: the six statuses with their labels, groups and descriptions; the effective types; the custom fields grouped by which type they apply to, required first, with the workspace ids this project shadows named; its tags; its default assignee, lead and owning unit. ?for_type= narrows the fields to one type plus the ones that apply to every type. Unknown key answers 404 naming the nearest three; a for_type the company does not file answers 400 bad_params, not 404 — the project is there and the argument is what to change
GET/work/filesOne project’s files, in path order: ?project= (required), ?folder= to narrow to the paths under one folder, ?removed=true to include removed files, ?limit= (default 200, max 1000) and ?after=, the next the previous page answered. An unknown project is 404 no_project, never an empty page (see below)
GET/work/files/{project}/{path...}One file’s bytes, streamed, with its version as the ETag. 503 content_unavailable where the file is listed and its content could not be read right now (see below)
PUT/work/files/{project}/{path...}Write a file: the request body is its content, Content-Type its type, If-Match the version it replaces and ?op_id= how an unknown answer is retried. Always needs an operator token — the write is attributed to the person (see below)
DELETE/work/files/{project}/{path...}Remove a file. Same credential, If-Match and ?op_id= as the write
GET/work/activityThe activity feed — one durable row per applied commit, quiet ones included, at any age with no live/archive boundary to cross. Ordered by the COMPOSED LOG POSITION rather than by any clock, so ?since= and ?cursor= are both positions written <stream>@<generation>:<sequence> — which is what lets a cursor span a reanchor with no gap and no repeat. ?task= (by key, id or a FORMER key), ?container=, ?kinds=, ?actor=, ?assignee=, ?notified=, ?from=/?to= (RFC3339, bounding the AUTHORED instants), ?limit= ≤200. ?q= is an escaped LIKE over the excerpt and is REFUSED unless it names a task, or a project and a since inside 90 days. Each record carries fields — what MOVED, as {"<field>": {"from": …, "to": …}}, and for a SET (tags, watchers, muted, collaborators, each link kind, blocking, a catalogue’s or tag set’s declarations, a person’s favorites) the members that joined and left, as {"<field>": {"from": "", "to": "", "added": […], "removed": […]}}, each sorted and every member whole — for every kind and not only the ones about a task: a project reconcile names the purpose, unit or epoch that changed, a view save the query parameters, a priorities write the order before and after, and a dependency the item it now waits on. A task’s own row draws on twenty-eight names: title, status, assignee, priority, project, type, tags, due, due_all_day, start, estimate, points, reporter, watchers, muted, collaborators, parent, routing_unit, archived, removed_with, waiting_on, linked, duplicates, page, blocking, checklists, fields and body. Values are the STORED form (a status slug, a whole RFC3339 instant, an item’s id) rather than a rendering, because every node writes the row identically and a rendering would depend on the reader’s zone and the company’s live vocabulary. Nothing is cut: an ordered list (priorities, pinned_views, checklists) is carried whole, free text past 600 bytes (a project’s purpose, say) is described as <N> bytes, and every field a change moved is on the row — only a notification card stops at 32, saying how many it left off (fields_omitted). The two largest are MARKED rather than carried: body is <N> bytes on each side (empty where there was none) and never the prose, and checklists is <list>: <done> of <total> done per named list, plus (<n> promoted) where an item became a sub-item. fields names each custom value by its SLUG — resolved against the project’s catalogue by the node applying the change, which is why a NOTIFICATION carries every other delta and not this one — with a choice as its option’s slug, a multi-valued field’s members joined with /, a value past 150 bytes as <N> bytes, and a count of any whose field the project no longer declares. The ANSWER also carries keys, an id-to-item-key map naming the tasks those deltas point at — waiting_on, linked and duplicates but never page, which names a knowledge-base page; the blocking mirror; a person’s priorities queue; and the two scalars that name a task, parent and removed_with — resolved on the answering node: a delta records another task by its ID, because a key belongs to that task’s own row and a history row is written once and never repaired. An id this node holds no row for is absent rather than empty, and a renderer falls back to the id
GET/work/my-workEverything one person is expected to look at, in seven bounded lists: priorities in the stored order, assigned, asked_of_me (each with the literal call that answers it, whether it is open, and the decision it carries when it asks somebody to choose — see Asking for a decision), checklist_items (which live on other people’s tasks and no assignee filter reaches), collaborating, watching_recent and unblocked_recent. ?handle= is whose, and it defaults to the caller’s own seat — see Whose record a personal question answers for. Naming somebody else’s handle is operator-only
GET/work/inboxOne person’s inbox: the notices a change wrote to them, each naming the ONE reason of eighteen it found them under, whether it asks something or merely informs, whether it arrived only because nobody better was found, and their own read and snooze marks — plus the person behind an operator’s token (actor_seat), the comment and turn the change came from, and the ask it is about, read as it stands now. Same scope rule as /work/my-work. ?unread=, ?primary_only=, ?snoozed= — exclude (the default: a snooze means not now), include or only, anything else refused naming the three — ?reasons= (comma-separated, refused naming the eighteen); every one of them narrows the SCAN, so a page is full whenever the scope holds 50 notices and next_cursor is never a cursor past an empty page, ?limit= ≤50, ?cursor=, and ?since= — a LOG POSITION written <stream>@<generation>:<sequence>, which is what seen_through reports back, never a bare sequence: the comparison is on the packed `(generation << 40)
GET/work/people/{handle}One human’s own state: their inbox (unread, read, snoozed, and which snoozes are now due), the order they mean to work in and who set it, and their pinned views. Same scope rule as /work/my-work, with the handle always named here because it is the path: your own seat’s needs no credential, anybody else’s is operator-only. A person nobody has written yet answers the EMPTY state with held: false, not a 404 — every human starts this way and the first write is what creates the record
GET/work/{id}One item with its description, thread, history and links. {id} is either the key (ENG-42) or the id — a person holds the first and every internal link the second
GET/work/commentsOne page of an item’s thread, walking back from the newest: ?item= (key or id), ?cursor=, ?limit= (default 20, at most 50) — see work_comments
GET/work/turnsOne page of the agent turns charged to an item, newest first: ?id= (key or id), ?cursor=, ?limit= (default 20, at most 50) — see work_item_turns
GET/pagesThe company’s own knowledge base: a filtered listing. Served only where knowledge.backend is native
GET/pages/{id}One page with its body, comments, revision metadata, children and ancestor breadcrumb. {id} is the id, or CONTAINER/Title — the title matches the way the fleet CLAIMED it, so case and runs of whitespace are ignored and ENG/deploy runbook reaches a page called “Deploy Runbook”
GET/containersEvery knowledge container this node knows about, with how many pages each holds. The engine materialises one per space: the org chart names, plus the two reserved ones, on every config apply; each carries chart_epoch, the activation its name and purpose were last written from (Unix milliseconds), so a configuration activated earlier never overwrites them
GET/viewerWho is asking. The presented credential’s operator id, whether it is an operator one, and the seat that binds it — a human seat naming that id in contact.crewlet_operator_id. Three distinct states, and a caller must tell them apart: no credential at all, a credential no seat claims, and a bound one. An unbound token is an ordinary state, not an error — the remedy is a line of company configuration, so the id is answered with no seat rather than refused. acts names the tools /operator/act would serve this caller: the catalogue’s writes for a bound token, and an empty list for anybody else. config_writer is whether this credential may change the company document, and config_managed_by who may when another system manages it — empty when nothing does, and for an anonymous caller
GET/stream/snapshotDashboard initial-state bundle, served from the in-memory projection (REST fallback for the WebSocket)
WS/ws/streamLive dashboard stream — agents, events, LLM invocations, health
GET/dashboardDashboard shell (/ redirects here; /static/{path} serves its assets)
POST/webhooks/jiraReceive Jira Data Center webhooks (Cloud arrives via /webhooks/forge)
POST/webhooks/slack/{handle}Receive Slack Events API deliveries for one seat’s app
GET/webhooks/slack-oauthOAuth install landing page for crewlet slack provision
POST/webhooks/githubReceive GitHub webhooks — HMAC-SHA256 over the raw body
POST/webhooks/github/{handle}The same route addressed to one seat, which is where that seat’s own GitHub App delivers
GET/webhooks/github-appLanding page for the per-agent GitHub App flow: converts the one-time creation code, or reports an install (see below)
POST/webhooks/gitlabReceive GitLab webhooks
POST/webhooks/confluenceReceive Confluence Data Center webhooks, HMAC-signed. Cloud arrives on the two routes below instead
POST/webhooks/confluence/{event}Receive one Confluence Cloud event, authenticated by the shared token the registered URL carries — X-Crewlet-Token first, ?token= as the fallback, because Cloud honours no registration field for a header
POST/webhooks/datadogReceive a Datadog monitor alert, authenticated by a constant-time comparison of X-Crewlet-Token. Datadog signs nothing, so the token is the whole check: an unset one answers 503, and one shorter than 26 characters answers 503 too — see Datadog
POST/webhooks/forgeReceive Forge events (FIT-verified)
POST/otlp/{token}/v1/{signal}Engine-fronted OTLP receiver for sandbox telemetry (per-run token in the path)
GET POST DELETE/mcp/{token}The tool bridge: one running seat’s tool surface, served over streamable-HTTP MCP to a coding agent in agent mode. Per-run token in the path; all three verbs because that is what the transport uses
GET POST DELETE/operator/mcpThe company’s own tracker and knowledge base, served over MCP to your AI assistant. Always needs a token — it files and moves work (see below). Absent where the company runs neither native backend
POST/operator/act/{tool}The same catalogue’s writes, one tool per request, as the person your token is bound to — the dashboard’s write surface. Refused unbound to a token no seat binds and to a disabled guard’s caller (see below). Absent where /operator/mcp is

Auth. Writes and every /config, /secrets and /setup route require Authorization: Bearer <token>. Reads (GET / HEAD outside those three) serve without one unless api.auth.allow_anonymous_read: false is set, at which point they need the same token — /ws/stream included, and it accepts ?token=… too since browsers cannot set headers on a WebSocket. Only there: a token in the query string of any other route authenticates nobody, because a URL lands in proxy logs and browser history. Never guarded either way: /health, /ready, /webhooks/*, /otlp/*, /mcp/*, and the dashboard shell (/, /dashboard, /static/*). See Configuration § Auth.

A token that is present and wrong is refused even where reads are open. Sending a credential says you meant to be somebody, so /ws/stream answers 401 rather than quietly serving you as anonymous — which is how a revoked token goes on appearing to work. The practical consequence is that a stale token in a browser breaks a dashboard that would have connected with none at all; the dashboard detects that and offers to forget it (see below).

GET /ws/stream without an Upgrade header answers 401 for a refused credential and 426 Upgrade Required for an accepted one. That pairing is a contract, not an accident: a browser is told nothing about why a WebSocket handshake failed — no status, and no close code, because a connection that never opened sends no close frame — so the dashboard re-asks over plain HTTP to tell “your token is wrong” from “the engine is down”. Without it a reader holding a stale token sees “retrying” for ever.

The guard is always mounted, whether or not Tier A is present. An API built without api.auth configuration has no token, and a route that needs one is therefore refused rather than served: reads work, every write and the whole of /config, /secrets and /setup answers 401. There is no way to start a process that serves those writes without a guard in front of them.

Every /webhooks/* route fails closed. They are exempt from the bearer token because each verifies its provider’s signature instead — so a route whose secret is not configured has nothing to verify with, and answers 503 with Retry-After rather than accepting the delivery. The sender retries and the delivery flows once the secret is set; nothing is discarded, and nothing unsigned is ever recorded, published, or shown on the dashboard.

Plus the four always-guarded surfaces: /config/*, /secrets/*, /setup/* and /operator/* — /operator/mcp and /operator/act. /setup is guarded on its READS as well, and deliberately: the list of which credentials a company has not configured yet is a map of what to attack.

Every response the API writes carries four headers, set before its status line (so a 304 carries them as well as a 200):

HeaderValue
Content-Security-PolicyThe policy for what the response is (below)
X-Frame-OptionsDENY
X-Content-Type-Optionsnosniff
Referrer-Policyno-referrer

The policy depends on what was served:

ResponseContent-Security-Policy
The dashboard shell (/dashboard), /favicon.ico and every /static/* assetdefault-src 'self'; script-src 'self'; style-src 'self'; img-src 'self' data:; font-src 'self'; connect-src 'self'; object-src 'none'; base-uri 'none'; frame-ancestors 'none'; form-action 'self' https:
The landing pages /webhooks/github-app and /webhooks/slack-oauthdefault-src 'none'; img-src data: (a page’s mark is inlined, since /static/ is not served on the public listener these pages move to), then style-src and script-src naming the sha256 hash of each page’s own inline block ('none' where a page has none), then base-uri 'none'; form-action 'none'; frame-ancestors 'none'
Everything else: JSON, plain text, the redirect from /, a 404 or a 401default-src 'none'; frame-ancestors 'none'; base-uri 'none'; form-action 'none'

These matter because the operator token the dashboard stores lives in the browser’s storage for this origin, and the two landing pages are unauthenticated pages on that same origin that render values from their query string. A policy is per response, so each page carries its own: the dashboard runs only the bundle it was built into, and a landing page runs only the style and script the engine wrote into it. No response may be framed by another site.

form-action on the dashboard allows https: as well as 'self' for one flow: creating a seat’s GitHub App posts the app manifest as a form to the code host, which is github.com or the GitHub Enterprise Server base the company configures. A reverse proxy in front of the engine should pass these headers through unchanged; one that adds its own Content-Security-Policy produces two policies, and a browser enforces both.

Read-side handlers live in the internal/api package (one module per domain — agents, events, tokens, org, fleet, access, mcpstatus, sandbox_runs, budgets, integrations, webhooks, dashboard, health); webhooks and /config/* keep a stable external contract, while the read/stream surface is free to evolve since the dashboard is its only consumer.

/config/* — live config management (auth-gated)

Section titled “/config/* — live config management (auth-gated)”

All /config/* routes require Authorization: Bearer <token> matching one of the tokens listed in Tier A api.auth.tokens. See the Configuration concept doc for the full auth model.

A managed document refuses every other writer. When Tier A lists api.auth.company_writers, only those token ids may change the document: PUT, PATCH, every per-entity PUT, a revert and the ?dry_run=true check of each answer any other credential 403 config_managed, before anything is read or stored, and /setup’s connect, disconnect and GitHub App routes answer the same refusal before they touch the secret store or a vendor. Reads and POST /config/reload stay open to every credential. The body names who manages the document and what to do:

{
"error": "config_managed",
"managed_by": ["gitops"],
"detail": "the company document is managed by gitops, and the credential \"founder\" may read it but not change it",
"hint": "change the company at its source, the system that writes it here with the token gitops: an edit made directly would be overwritten at its next reconcile. A leaked credential can still be rotated: write the new value with /secrets and POST /config/reload. To take the document back, remove api.auth.company_writers from every node's Tier A and restart"
}

403 rather than 401, because the credential was accepted and a different one is the managing system’s to hold, not the caller’s to retry with. See Managed configuration.

Read-only:

MethodPathDescription
GET/configActive revision, redacted (full JSON; ?format=yaml for YAML)
GET/config/revisionsPaginated history (newest first), metadata only
GET/config/revisions/{id}Single revision including its payload
GET/config/revisions/{id}/diff?against=<uuid|active>Structural diff
GET/config/referencesEvery ${VAR} the active document names, each with the config path of the field that names it, plus the revision they were read from

The dashboard reads four of those facts over the query channel rather than these routes — config, config_audit, config_diff and config_entities, each operator-gated for the same reason the prefix is. The reference index has no query of its own; the Secrets screen reads it over REST beside /secrets.

Why the reference index is a route and not a client-side scan. It answers “what breaks if I remove this credential”, which is the question in front of an operator about to delete or rename a secret: the config keeps ${VAR} pointers, so a removed row leaves every pointer at it resolving to the empty string and the surfaces holding one start refusing deliveries with nothing naming the row that went away. Deriving it from GET /config in the client would mean a second copy of the ${VAR} grammar, and the engine has already paid for that twice: a looser pattern once displayed a literal secret unmasked, and another once minted a live credential into a variable nothing reads. It would also miss references the read masks: GET /config shows a credential only when it is one whole ${VAR}, so "Bearer ${TOKEN}" arrives as "__redacted__". The index is built from the unredacted document and answers names and paths only, never a value. The path is the operator’s own spelling (roles[0].integrations.slack.bot_token), the same one a validation failure reports, and a name with several readers appears once per reader. It carries the document’s own ETag, because the index changes exactly when the revision does.

Full-document write:

MethodPathDescription
PUT/configReplace the active revision. The body is JSON or YAML, read the same way whatever Content-Type says. Requires a revision summary: an X-Summary header, or a top-level _summary key in the body. Conditional via If-Match / If-None-Match, see below. ?dry_run=true checks it and stores nothing, see Dry runs
OPTIONS/config204 with Allow and Accept-Patch: application/merge-patch+json
PATCH/configMerge one or more sections into the active revision, see below. ?dry_run=true checks it and stores nothing
POST/config/reloadRe-publish the active document unchanged, so every node re-applies and re-reads the secret store. See below
POST/config/revisions/{id}/revertCreate a new active revision whose payload equals revision {id}

A JSON Merge Patch (RFC 7396): send only the sections you are changing, in the shape the document already has.

The registered media type is application/merge-patch+json; plain application/json and an absent Content-Type are accepted too, since every example here sends one of those. Any other patch format is 415 with an Accept-Patch header naming what would have worked — notably application/json-patch+json, an RFC 6902 list of operations, which is a different format this surface does not serve. Editing one list member is what the per-entity routes are for. A patch format that can address a list member does not replace them: a patch addresses by structure, and a seat’s position in a unit’s list is not its identity — so an index-addressed edit rewrites a different seat the moment anything above it moves.

Terminal window
curl -X PATCH https://engine.example.com/config \
-H "Authorization: Bearer $TOKEN" -H "X-Summary: raise the executor round cap" \
-d '{"turn_engine": {"max_tool_rounds": 32}}'
  • Deep merge. {"providers": {"llm": {"main": {"model": "claude-opus-5"}}}} changes that model and leaves the provider’s type, its keys and every other provider alone.
  • null deletes. {"integrations": {"gitlab": null}} removes the section — without it a config surface can only add.
  • Arrays replace. RFC 7396 cannot address a list element, so roles: [...] in a patch replaces the whole roster. Editing one seat is what PUT /config/roles/{handle} is for; inventing a list syntax here would give two answers to one question. What a replacement does not remove is a field this build cannot represent: see Fields a newer build wrote survive every write.
  • Unknown keys are refused, not ignored. A patch is the edit least visible in a diff, so a typo that silently changes nothing is the worst outcome available, because the caller believes they changed something. That holds whatever the key is set to, null included: deleting a key this build does not know is refused rather than ignored, because the write carries it back and the caller would be told a deletion landed that did not.
  • Validated as the whole document it produces. A section that is fine alone is still refused when it leaves the company invalid.
  • Same summary rule and same If-Match as PUT /config, and a 409 when nothing is active: a patch is defined against a document, and building a company out of one section is not what this route is for.

If-Match matters more here than on the full write. A PUT carries the caller’s whole intended document; a PATCH is merged against whatever is active at that instant. See Concurrent writes for what the engine does and does not guarantee.

Every write that stores a revision (PUT, PATCH, a per-entity PUT, a reload and a revert) answers 201 with the revision, its epoch, and what the engine makes of the document it stored:

{
"revision_id": "3f1c0f0e-8a52-4d3b-9d7e-2b6f3f0c9a41",
"epoch": 42,
"warnings": [
{
"kind": "dangling_reference", "ref": "manages",
"path": "roles[0].manages[1]", "segments": ["roles", 0, "manages", 1],
"seat": "ceo", "unit": "", "from": "CEO", "to": "Ghost",
"message": "seat \"CEO\" manages \"Ghost\", which is neither a seat nor a unit, so the entry manages nobody. Correct the entry or add a seat or unit with that name"
}
],
"derived": {"seats": [...], "units": [...]}
}
  • warnings is what the engine will run but a person should know about. Always a list, empty when there is nothing to say. Each has the same locators as a problem (path, segments, and the seat handle or unit name it is about, empty when neither), plus from and to as display text. Two kinds:
    • dangling_reference: a reference that resolves to nothing. ref says what carries it: lead (a unit’s lead), unit (a root seat’s unit:), manages (one manages entry, at the index it was written) or gitlab_access_level (a key under integrations.gitlab.provisioning.access_levels naming no seat).
    • admission: an admission rule the stored company breaks, with ref, from and to empty. A write that keeps one is refused, so only a reload or a revert of a company a newer peer admitted under other rules answers with one, one beside each entity the violation names.
  • derived is the hierarchy the engine derives from the document, in full: every seat in the engine’s own order with its effective unit, primary manager, managers, reports, automatic reports and onboarding chain, and every unit with its effective type, lead and channel (and whether each was inherited). Each seat and unit carries its authored path. The fields are the ones GET /org carries without paths; a client draws the hierarchy from this rather than deriving it again.

PUT /config?dry_run=true, PATCH /config?dry_run=true and PUT /config/{kind}/{id}?dry_run=true are the same request, checked in the same order, that store, activate and publish nothing. The dashboard’s organization builder sends one on every edit, and its Budgets screen one before every ceiling it saves, so a check is always exactly the write a save would send. An entity write needs its check more than the whole-document writes do: its caller never sees the rest of the document, so the whole-company validation behind the splice is the only place it learns that a seat fine on its own leaves the company invalid, or that a ceiling it raised now sits above the company’s (a warning, which only a check shows before the save).

Terminal window
curl -X PATCH "https://engine.example.com/config?dry_run=true" \
-H "Authorization: Bearer $TOKEN" -H "If-Match: \"$REV\"" \
-d '{"mission": "Ship the thing"}'

A valid check answers 200:

{"valid": true, "base_revision_id": "3f1c0f0e-8a52-4d3b-9d7e-2b6f3f0c9a41", "warnings": [], "derived": {"seats": [...], "units": [...]}}
  • dry_run is read before anything else, and takes exactly true or false, or nothing. Any other value (1, yes, an empty value, the parameter twice) is 400 invalid_query: the two readings of a guess differ by whether the fleet’s configuration changes.
  • No summary is needed, because nothing is stored to record one on. A _summary key in the body is still lifted out, so the document checked is the one the write reads.
  • base_revision_id is the revision the check was built on, and "" when nothing is active. A client whose draft was built on a different revision learns that the configuration moved without a second request.
  • Every other refusal is the write’s, in the write’s order: 409 no_active_revision for a patch with nothing to patch, 409 revision_advanced for a stale If-Match, 412 already_configured for If-None-Match: * on a configured company, and 400 with problems for a document the write would refuse.
  • A dry run needs the same token a write does.

A refused document (400 validation_error, 400 invalid_patch, 400 invalid_body) keeps error, detail (one line per failure) and hint, and adds problems: the same failures, located and classified, so a client puts each beside the field it is about without parsing the detail.

{
"error": "validation_error",
"detail": "roles[1].llm: value not in the allowed set: \"nowhere\" is not a configured provider: providers.llm has zulu. ...",
"hint": "the WHOLE document a write produces is validated, ...",
"problems": [
{
"path": "roles[1].llm", "segments": ["roles", 1, "llm"], "kind": "unknown_value",
"message": "roles[1].llm: value not in the allowed set: ...", "seat": "cto"
}
],
"derived": {"seats": [...], "units": [...]}
}
FieldMeaning
pathThe authored path in the whole document that was validated. For a per-entity write that is the document the entity was spliced into, and an entity body it cannot read is placed where that entity sits (roles[1].gaol for a typo in the second seat). "" only for a failure that belongs to no place in it, such as a whole document that is not YAML at all
segmentsThe same path taken apart: strings for keys, numbers for list indexes. A map key can hold a dot, so read these rather than splitting path. null when path is ""
kindmissing, unknown_value, out_of_range, conflict, unknown_field, shape, or invalid for anything this build does not classify
messageThe failure’s whole line, exactly as it appears in detail. A duplicate name is one line naming every entity and one problem beside each, so there can be more problems than lines
seatThe engine-derived handle of the seat the problem is about, when it is about one
unitThe name of the unit the problem is about, when it is about one
lineThe 1-based line in the text that was sent, for a failure the parser found. A patch’s failure found in the merged document names no line, because that text is the engine’s merge rather than anything sent

derived is present whenever the document parsed: a document with problems still has a hierarchy, and a person fixing a misspelled lead finds it in the chart it breaks. A body or patch that never became a document carries none.

No message repeats a credential. A document read from GET /config carries masks, which a write restores from the stored revision before validating, so the values a refusal judges are ones the caller was never shown: a message says what rule a value breaks and never the value, a fragment of it, or its length.

The /setup submissions that change the document answer their validation_error with the same problems and derived.

GET /config and every entity GET return an ETag — the active revision id, quoted. It is the token the write side takes, so a read-modify-write needs no second request to find it.

HeaderOnMeaning
If-None-Match: <etag>GET304 Not Modified when the document has not moved
If-Match: <etag>writesProceed only against that revision; 409 revision_advanced otherwise
If-Match: *writesProceed only if something is active; 412 on an unconfigured node
If-None-Match: */config writesProceed only if nothing is configured, on this node or anywhere in the fleet; 412 already_configured otherwise, naming the revision it lost to
If-None-Match: *entity PUTProceed only if that entity does not exist: the create-only write that adds an MCP server or an LLM provider. 412 entity_exists when one does; 400 conflicting_preconditions beside an If-Match

The bare revision id is accepted wherever an ETag is, unquoted, because this surface shipped that form before it had entity tags. If-None-Match: * is the only create-only precondition — about the company at /config, about the entity at an entity’s own address — and every If-Match value other than * is an entity tag, matched against the active revision and nothing else.

Independently of any header, every write names the revision it derived from as the new revision’s parent, and the activation is a compare-and-set on that parent — so a lost update is refused whether or not the caller sent a precondition. See Concurrent writes.

Nothing under /config is cacheable. Every response the surface writes carries Cache-Control: no-store: reads, 304s, writes, refusals and error bodies, and the 404 and 405 it answers for a path or method it does not serve. A body here is the company document, with its contact identities and the ${VAR} name behind every credential, and a stored copy would outlive the session and the token that read it. Revalidation still works, because it never depended on a cache: a client that wants a 304 sends If-None-Match with the ETag it kept.

Four collections, GET and PUT:

MethodPathDescription
GET/config/{kind}/{id}One entity, redacted, with an ETag. The body is the entity itself, so it goes straight back into the PUT
PUT/config/roles/{handle}Replace one seat, wherever it lives — root-level or inside a unit, at any depth. Every entity PUT takes ?dry_run=true, see Dry runs
PUT/config/units/{name}Replace one org unit
PUT/config/llm-providers/{key}Replace one named LLM provider; with If-None-Match: *, add one under that key
PUT/config/mcp-servers/{name}Replace one MCP server entry; with If-None-Match: *, add one after every server already declared

Any other method is 405 with an Allow header naming GET, PUT. There is no DELETE — removal is a full-document edit, for the reasons below.

Why these exist beside the whole-document write: PUT /config makes every edit a company-wide one. A founder renaming one seat’s goal sends back a document carrying every other seat, every provider and every integration, and a concurrent edit anywhere in it is theirs to lose. Editing one entity narrows what a write claims to have changed, which is what makes the revision summary mean something.

It is the same write underneath, and that matters more than the convenience: an entity PUT opens the active revision, splices the entity in, restores the masks the read showed against that same revision, validates the whole document, and stores a new revision. A change that would leave the company invalid is refused even when the entity itself is fine — a seat naming a provider that no longer exists is exactly the break a per-entity surface invites, because the caller never sees the rest of the document.

Four rules follow from that:

  • An unknown field is refused, not dropped. The entity body is read by the whole-document parser, JSON or YAML: gaol where goal was meant is 400 invalid_body with an unknown_field problem placed where the seat sits in the document (roles[1].gaol), with its line in the body. A decoder that ignored what it did not recognise would answer 201 and store the seat with its goal silently gone.
  • A plain PUT never creates. An id nothing carries is 404 no_such_entity, not a new entity: naming one that is not there is far more often a typo than an intent to add one. The intent is SAID with If-None-Match: * — the create-only write, for the two flat collections (mcp-servers, llm-providers): a taken id is 412 entity_exists rather than a replacement, the body’s identity must match the path as on any PUT, and the whole company is validated with the new entity in it. An added entity lands LAST in its collection’s order — a server after every declared server, a provider at the end of providers.llm_order, the order an unpinned seat resolves a provider in. A seat and a unit are 400 not_creatable — each has a place in the chart the path cannot name — and are added through PUT /config, which shows the whole thing. The id is looked up before the body is read, so a mistyped one is a 404 whatever the body holds.
  • The id in the path is the identity, and a PUT never renames. A body whose own identity disagrees with the path is 400 identity_mismatch, not a move: nothing that points at the old identity travels with the splice. A seat’s durable id is a UUIDv5 over (company name, handle), so a renamed handle strands that seat’s diary, onboarding marker and counterparty profiles behind an id nothing derives any more; a unit’s name is referenced by every manages: entry and root seat unit: that names it, and an MCP server’s name by every mcp_env block, a seat’s or a unit’s, keyed on it. For a role the check is on the derived handle, so a body that omits handle and changes name is refused too: that is a rename, just an accidental one. Send the identity back unchanged (changing a seat’s display name while keeping its handle is an ordinary edit); rename through PUT /config, where what has to move with it is visible.
  • The same summary and If-Match rules apply, and a node with no active revision answers 409 no_active_revision — there is nothing to splice into, and building a company out of one seat is not what this route is for.

/secrets/* — the company’s credentials (auth-gated)

Section titled “/secrets/* — the company’s credentials (auth-gated)”

All /secrets/* routes require Authorization: Bearer <token>, reads included, for the same reason /config does: the listing alone says which credentials a company holds and when each last changed. Every node serves them, because every node opens the fleet’s coordination store that holds the rows.

MethodPathDescription
GET/setup/integrationsWhat every integration this build can set up still needs, plus the address third-party apps reach this deployment on. public_base_url answers present and resolved separately — set and set to something are different facts, and a ${VAR} nobody exported is present: true, resolved: false with the variable named in reference
GET/setup/integrations/{kind}One integration’s requirement list and state
POST/setup/integrations/{kind}/inputsSupply or generate those values: credentials are sealed, the rest is patched into the company
DELETE/setup/integrations/{kind}Disconnect: remove what the integration holds at the third-party app, then its block
POST/setup/integrations/{kind}/provisionRun the third-party app’s provisioning pass: mint what it needs, register its webhook
POST/setup/integrations/{kind}/checkRun the same pass read-only, to see whether something fixed at the third-party app took
GET/setup/integrations/{kind}/runsThe passes THIS NODE remembers for one surface, newest first, ten at a time. A pass is executed by whichever node held the surface’s lease and is remembered in that node’s own process, so the answer carries scope saying as much — an empty list on a fleet where another node ran the pass is an honest answer to a question the reader did not mean to ask. It exists because nothing could name a run id: the route below answered one pass and was reachable only by a caller that had just started it
GET/setup/integrations/{kind}/runs/{id}One pass, as the node that executed it remembers it
GET/secretsEvery stored name with its key_id, updated_at, updated_by and source. Never a value
GET/secrets/{name}The same fields for one name. 404 not_found when it is unset
GET/secrets/{name}?reveal=trueBreak-glass. The decrypted value, Cache-Control: no-store, logged by name against the authenticated operator
PUT/secrets/{name}Store or rotate one value. The request body is the value, raw bytes, up to 64 KiB. ?source= records provenance (default api). 400 invalid_name when the name is not an environment-variable name
DELETE/secrets/{name}Remove one value. 200 either way, with {"removed": true|false}
POST/secrets/rekeyRe-seal every record not already under this node’s secrets.active_key_id, answering the names it moved. ?key_id= is refused with 409 when it names a different key

The name is an environment-variable name, and a write that is not one is refused. The store is keyed by the name a ${VAR} resolves through, so gitlab-token or my token would be sealed, listed and read by nothing at all, a success the operator only discovers when a provider fails to authenticate hours later. Letters, digits and underscores, starting with a letter or an underscore. The refusal comes before the body is read, so the name is what the answer points at. A name outside the grammar names no secret, so reading one answers 404 and removing one answers {"removed": false}.

The body is the value, not a JSON wrapper. A credential is arbitrary bytes — a PEM key has newlines, a token can hold anything — and an encoding step between the operator and the sequence the vendor compares is a 401 nobody can explain.

Reveal is opt-in on the wire, not merely in the CLI. Without ?reveal=true the route answers what a listing answers for one name, so a browser, a crawl or a link preview cannot pull a credential out by accident.

A node with no secrets.keys answers 503 no_keyring on every route that seals or opens, pointing at crewlet secrets keygen. The store has no plaintext mode; refusing is the only alternative to holding credentials in the clear.

crewlet secrets is the client for all of this — see the secret store for why the CLI goes through a running node rather than writing the KV itself.

There is no leader — any node’s API can write the config, and the coordination KV is the shared truth. Writes are not serialized by a lock, but a concurrent one is detected: the activation is a compare-and-set.

  • Every write on this surface reads the active revision, derives from it, and names it as the new revision’s parent. That parent is what the flip compares against, so a write that lost is refused with 409 revision_advanced — whether or not the caller sent If-Match, because the server knows what it read.
  • If-Match: <revision_id> is still worth sending: it is checked before any work is done, so a caller editing a revision that has already moved is told so without a document being built, validated and stored first.
  • A losing write’s revision is kept, and the 409 (or 412) names it as stored_revision_id. It is stored in the history, valid and inert, so the operator’s work survives as history they can revert to. Inert means on the node that served the write too: a revision becomes that node’s active one only once the fleet has taken it, so the node goes on serving what it served, never offers the loser to the fleet at a restart, and adopts whichever revision actually won. Unwinding it instead would mean a second write that can itself fail, on the path where something has already gone wrong.
  • A node’s active revision follows the fleet. Once a node applies the fleet’s epoch, the fleet’s revision is its active one, which is what its GET /config serves and what it boots on; its reconciler checks that on every tick and corrects a copy that says otherwise.
  • An unset pointer is not a race. A node seeded from a file holds a locally-active revision before it has published anything; refusing there would fail every config write on a fresh single-node deployment that had done nothing wrong.
  • A write built on nothing is a create. A node whose own store is empty answers 404 no_active_revision on GET /config, and it reaches that state while its fleet runs a company: it joined and has not reconciled yet, or its best-effort copy of the fleet’s pointer failed. A write there was derived from nothing, so its activation is a create-only compare-and-set: it lands only while the fleet has no activation. If-None-Match: * also consults the fleet’s pointer before anything is built, and answers 412 already_configured naming the revision the fleet is on. Without both, the dashboard’s create flow on such a node replaced the running company outright, which renames it, changes every seat id derived from the name and orphans all of their memory.
  • The boot publish is deliberately unconditional. Two nodes starting at once may both offer the revision they hold; both are legitimate, last-write-wins is the right answer, and every node converges. It is the edit path that must not lose a write.

On a 409, re-read /config and send the edit again.

  • 200 OK: a successful read, or a dry run that found the write valid ({"valid", "base_revision_id", "warnings", "derived"})
  • 201 Created: a write produced a new revision; the body is {"revision_id", "epoch", "warnings", "derived"} (see What a write answers). A per-entity write, a reload and a revert return this too: each created one revision.
  • 400 Bad Request: invalid_body, invalid_patch or validation_error, each with detail (the field path and what to change) and problems; summary_required when a write has neither an X-Summary header nor a _summary body key; invalid_query when dry_run is anything but true or false; identity_mismatch when a per-entity body renames what the path addresses
  • 401 Unauthorized: missing or invalid bearer token ({"error": "invalid_token"})
  • 403 Forbidden: config_managed when api.auth.company_writers names the document’s writers and this credential is not one of them, with managed_by, detail and hint — a write and its dry run alike; never a read or a reload
  • 404 Not Found: a revision that is not there, no_active_revision on a read before the first write, no_such_entity on a per-entity write naming an id the active revision does not carry, or no_route for a path under /config this surface does not serve
  • 405 Method Not Allowed: method_not_allowed for a /config path under a method it does not take, with Allow
  • 409 Conflict: revision_advanced (a stale If-Match, or a race with a concurrent writer) or no_active_revision (a PATCH or a per-entity write on an unconfigured node, or a reload)
  • 412 Precondition Failed: already_configured when If-None-Match: * meets an active revision, or no_active_revision when If-Match names a revision and none is active
  • 415 Unsupported Media Type: unsupported_patch_media_type when a PATCH body is a patch format other than a JSON Merge Patch, with Accept-Patch
  • 503 Service Unavailable: draining when the node has been told to stop, with a Retry-After — see During a drain

Recent revision metadata for the dashboard’s Configuration and Audit screens. A query, not a REST route — there is no GET /config/audit in this build; the screen asks the query channel for config_audit and gets the same revision records GET /config/revisions serves, as a bare array (the one question in the set that is not an object).

query config_audit { "limit": <N> }
ParameterDefaultRangeDescription
limit501..500Number of revisions to return, newest first. An out-of-range number is CLAMPED to the range, and a value that is not a number reads as the default.

Response (200 OK):

[
{
"revision_id": "11111111-1111-1111-1111-111111111111",
"parent_revision_id": "00000000-0000-0000-0000-000000000000",
"created_at": "2026-05-17T10:31:02.118431Z",
"created_by": "founder",
"created_by_kind": "operator",
"source": "api",
"summary": "add Designer role",
"is_active": true,
"activated_at": "2026-05-17T10:31:02.118431Z"
}
]

Payloads are NOT included — fetch a specific revision via GET /config/revisions/{id} for the full JSON.

created_by is a label and created_by_kind says what it names, on this answer, on GET /config/revisions and on GET /config/revisions/{id} alike:

created_by_kindWho wrote the revisioncreated_by
operatorA person, through a credential: an API token on /config or /setup, or the login running crewlet config import / crewlet config rekeyThe token’s id, or the login ($CREWLET_OPERATOR, else $USER)
nodeThe engine itself: a node seeding the store from its -company file at boot, or the reconcile loop’s own writes (removing a disconnected integration, recording a discovered site, reloading after sealing a credential)The node’s id for a seed; reconcile loop for the loop

Read the kind rather than inferring it from the label — the two name spaces overlap, and an operator token may be called anything. A revision reads the same on every node: the fleet’s activation pointer carries its author, so a node adopting it records the origin’s author rather than its own (see Control Plane). A kind a newer engine adds arrives as itself.

A structural diff of two revisions — one entry per path that moved, with the value on each side. ?against= names the other side and defaults to the active revision; the direction reads as “what against became”, so a bare call answers “what would reverting to this change”. The same answer serves the dashboard as the config_diff query.

{
"from": "00000000-0000-0000-0000-000000000000",
"to": "11111111-1111-1111-1111-111111111111",
"changes": [
{ "path": "providers.llm.main.model", "kind": "changed",
"from": "claude-sonnet-5", "to": "claude-opus-5" },
{ "path": "roles[3].handle", "kind": "added", "to": "qa" }
],
"changes_total": 2
}

kind is added, removed or changed, and from/to are absent on the side where the path does not exist — which is what makes the first two readable without consulting kind. Both sides are redacted, always: a rotated credential shows as a changed mask and never as either value.

changes_total is how many differences there are; changes is how many this answer carries. The listing is cut at 500 entries because a response body and a socket frame have a size budget, so render the total and say what was left out — a short listing read as the whole comparison is a caller believing nothing else moved. The two differ only on a document several times the size of the example company, whose 401 leaves all changing at once still fits. crewlet config diff writes to a terminal, which has no such budget, and prints every change.

changes is always a list: two identical revisions answer [] with a changes_total of 0, never null.


POST /config/reload: after a secret changes

Section titled “POST /config/reload: after a secret changes”

Takes no body and changes nothing. It stores a new revision carrying the same document and activates it, which advances the epoch and makes every node apply again.

That is the gesture a rotated credential needs, and no other route performs it. A secret lives in the company config as a ${VAR} pointer, resolved when a provider or a transport is constructed, from a snapshot taken at apply time. Writing a new value with PUT /secrets/{name} therefore changes nothing in a running process: the pointer is already correct, so there is no patch to make, and with no activation there is no apply and no refreshed snapshot. Re-activating an unchanged revision is exactly why the activation pointer is append-only rather than keyed on a revision id.

A new revision rather than a re-pointed old one, for the same reason a revert writes one: the history stays append-only, so “the credentials were reloaded at 04:12” is a fact somebody can find later. X-Summary names it; unset, it records reload configuration.

Answers 201 with the revision, its epoch, its warnings and its derived hierarchy (see What a write answers), 409 no_active_revision when nothing is configured, and 400 validation_error when the active document breaks a runnable rule of this build (a reload is an apply, so it re-publishes only a company every node can run; correct it with PUT or PATCH). A document that breaks only an admission rule, such as a duplicate seat or unit name a newer peer admitted, reloads: that is how a credential rotation still reaches a company carrying one. Its answer lists each violation as an admission warning.

The command-line equivalent is crewlet config activate <UUID> naming the revision that is already current.

Open on a managed document. A reload changes no byte of the company, so api.auth.company_writers does not refuse it: it is how a credential somebody rotated in an emergency reaches the running seats when the system that manages the document cannot know it has to.

A read never validates what it reads. GET /config, a revision read, a diff, the reference index, the entity reads and the prior a write restores its masks from all open a stored revision as it is, even one this build’s validator would refuse: a revision is valid under the build that wrote it, and a later build can leave one in the store that this build refuses. What validates is whatever would RUN a document, and to the rules its question needs:

  • A write (PUT, PATCH, a per-entity PUT, a /setup submission that changes the document) validates the whole document it produces against every rule, admission rules included. So a revision this build refuses is always readable, a corrected PUT or PATCH always replaces it, and a write that leaves it uncorrected is refused with 400 validation_error, even when the write touched nothing near the problem.
  • A reload or a revert (including a /setup submission that only rotates a sealed credential, which reloads) validates what it re-activates against the runnable rules only. A revert to a revision that breaks one answers 400 validation_error naming the field; a revert to one that breaks only an admission rule is accepted, and each node logs org_admission_warning when it applies it. A revert to a revision sealed under a key this node does not hold answers 409 unreadable_revision.
  • Every one of them is also judged against the answering node’s own Tier A, with the rule the apply runs: a document this deployment cannot run — the engine’s own tracker or knowledge base on an embedded stream that keeps its streams in memory — answers 400 validation_error naming stream.store_dir, rather than being activated for every node to refuse a moment later and leave the fleet on its old epoch.

Fields a newer build wrote survive every write

Section titled “Fields a newer build wrote survive every write”

During a rolling upgrade an older node holds, byte for byte, documents a newer node wrote, including settings the older build has no field for. GET /config on the older node cannot show them, because its types cannot hold them, so a document read there and sent back never names them. Every write keeps them anyway: PUT, PATCH, a per-entity PUT, a reload and a revert all store a document built from the stored bytes, and carry back each key the writing build cannot represent that the write does not name.

  • Only keys the writing build cannot decode are carried. A key it knows and the write left out was removed on purpose and stays removed. A caller of the older build cannot name an unknown key at all (the strict reader refuses it), so no write through it can mean to remove one.
  • A list member is matched by identity, not by position: a seat by its handle and a unit by its name anywhere in the document, so a seat moved to another unit keeps its settings; an MCP server or a sandbox setup step by its name within its own list. An identity held twice in the stored document, or empty, matches nothing, so no seat’s setting reaches another. A member of a list with no identity (a schedule, for example) keeps nothing the write replaced.
  • A PATCH replaces only what it names. Everything it does not name is stored exactly as it was, so a list the patch leaves alone keeps every member’s settings, identity or not, and a schedule loses a newer build’s settings only when the patch replaces the seat or unit list that holds it.
  • A reload and a revert store the document exactly as it was stored, since neither changes it.
  • A renamed seat or unit is a new identity, so it keeps nothing of the old one’s unknown settings.

Connecting an integration means putting values in two places: a credential into the fleet’s sealed secret store, and everything else into the company document. /setup is the surface that does both, in the one order that is safe, so the dashboard never has to sequence it and never holds a credential across two requests.

Guarded in full, reads included, on the same terms as /config and /secrets: this surface answers with the names of the credentials a company holds, which of them are unset, and the pages at each third-party app an administrator would visit. That is a map of what to attack, and it is not something the anonymous-read posture opens.

On a managed document (Tier A api.auth.company_writers), a submission that would write a pointer into the document, a disconnect, and the begin of an agent’s GitHub App answer a credential the list does not name 403 config_managed — the same refusal /config answers — before a value is sealed, a teardown is queued or GitHub is asked for anything. A submission that only rotates a value the document already points at is not a change to it and is served, ending in a reload as always; so are the provisioning pass and its read-only check. A provisioning pass any operator starts may still record what it discovered into the document — where integrations.jira or integrations.confluence has no cloud_id, the Atlassian pass records the site’s cloud_id and site_url — and that revision is written as the node (reconcile loop), not as the credential that started the pass. A GitHub App creation carries the operator who began it in its signed state, so the callback records the app — its revision and its sealed key — as that credential, and asks again before it exchanges GitHub’s one-time code.

GET /setup/integrations/{kind} answers what that third-party app needs, whether or not the company has configured it:

{
"key": "datadog",
"configured": false,
"enabled": false,
"satisfied": false,
"inbound_path": "/webhooks/datadog",
"public_url": "https://engine.example.com/webhooks/datadog",
"requirements": [
{
"field": "webhook_token",
"label": "Shared token",
"kind": "secret",
"config_path": "integrations.datadog.webhook_token",
"secret_name": "DATADOG_WEBHOOK_TOKEN",
"required": true,
"mintable": true,
"help": "...",
"where": "...",
"vendor_url": "https://app.datadoghq.com/integrations/webhooks",
"blocks": "credential_missing",
"present": false,
"resolved": null
}
]
}

kind is one of secret, url, id, choice, text, handle, toggle. Each third-party app’s own package declares its list, so the surface serves a third-party app it has no screen for and the dashboard renders a third-party app it has no code for. A toggle is a JSON boolean in the document; a handle must name a seat this company has.

A secret’s credential is never echoed. Its value carries the field’s ${VAR} reference when the document holds one, because that is a name rather than a credential: it says which entry of the sealed store the field reads, it is already visible through GET /config to anybody this surface answers, and a client that could not see it would have no way to tell “this reads SHARED_TOKEN” from “type here to replace what is behind this field”. A document holding a literal in that position sends no value at all, and a composite such as https://${HOST}/hook is a literal for this purpose: it names a variable and carries an address beside it, so it is not a reference to anything.

present and resolved are the same two facts secret_present and secret_usable are, asked per field: written down, and actually usable in this process. resolved is null for a field the document leaves empty, since there is nothing to resolve. blocks names the reconcile finding that this input being absent produces, which is what lets a row reporting credential_missing offer exactly the fields that clear it.

mintable means the engine can generate the value, so nobody should be asked to invent it. No route here ever returns a credential.

Terminal window
curl -X POST https://engine.example.com/setup/integrations/datadog/inputs \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{
"if_match": "<revision_id>",
"values": {"enabled": "true", "route_to": "sre-lead"},
"generate": ["webhook_token"]
}'

generate is separate from values on purpose: a client that could send both under one key would eventually send a weak token by accident, and a field can be one or the other, never both.

Any field the engine reads through the resolver may hold a ${VAR}, not only the credentials: the Atlassian organization id, a site address, a cloud id and an account email are all read that way, so a company can keep them in the sealed store and name them here. present and resolved then say the two things separately, and a reference naming an entry that is not there reports resolved: false rather than passing for a working setting.

A field holding a reference also carries resolved_value: what that ${VAR} currently reads as. It exists because the links this surface describes are built out of values, and a reference is a name: the Atlassian API keys page is per-organization, so an organization id kept in the store would otherwise put the literal text ${ATLASSIAN_ORG_ID} in the path. Never present on a secret, on a literal (which is already the value), or on a reference naming nothing.

A secret field’s value may be a ${VAR} instead of a credential. Sent one, the route writes that reference into the config path and seals nothing, so a credential already in the store can serve several fields and rotating it is one write in one place. It needs no secret store in the answering process, because naming an entry is not writing one. Anything else in that position is a credential and is sealed under the field’s own name, a composite included.

What the route does, in this order:

  1. Refuses a stale base. if_match names the revision the requirement list was read against. A submission built on an older one is refused before anything is sealed, so a caller working from a stale page does not end up with a credential in the store that nothing points at.
  2. Seals every credential, under the name the third-party app declared or one derived as VENDOR_FIELD[_HANDLE]. The row records source: "setup".
  3. Patches the document with the non-secret values and, for a credential whose slot was empty, a whole ${VAR} pointing at the name from step 2. A slot that already holds a ${VAR} is written through, which is what makes rotating a credential a change to the store and not to the company. A slot holding a literal is 409 literal_in_config, naming the path: overwriting it would edit the company from a setup form and destroy a credential somebody put there on purpose.
  4. Activates, through the same merge, validation and compare-and-set PATCH /config performs. When the pointer needed no change (a rotation) it reloads instead, because a value written into the store after the last apply is invisible to every running seat until something activates. The answer says "reloaded": true when that is what happened.

Answers 201 {"revision_id", "epoch", "wrote_secrets", "reloaded", "state"}. Refusals: 400 invalid_input, 400 validation_error, 404 unknown_kind, 409 revision_advanced, 409 literal_in_config, 409 no_active_revision, 503 no_keyring.

Some third-party apps need something done at them, not just written down: a webhook registered, a signing secret minted and pushed. That is the third-party app’s provisioning pass, and POST /setup/integrations/{kind}/provision is what runs it.

The reconcile loop does this on its own. It runs the same provisioning function every few minutes with the sink and the public base supplied, because connecting an integration is the permission: a person named that one app and handed over an administrator credential for exactly this. This route is the same work on demand, for an operator who wants a pass to run now rather than at the next tick, and it holds a fleet lease under its own name so the two never overlap. Nothing in the dashboard calls it.

can_provision on a tool’s state says whether this build has a pass for it. needs_operator is present only for a third-party app whose pass still asks for a credential per run; no third-party app in this build does. GitLab’s group Owner token and Mattermost’s system-admin token are ordinary stored requirements now, sealed in the fleet secret store with a ${VAR} in the document like every other credential.

They used to be transient, asked for on every pass and dropped the moment it returned, on the reasoning that a one-time grant held permanently is a standing power. What that reasoning did not price is the disconnect: removing a service account needs the authority that created it, so with nothing held there was no way to take one away from here, and every account the engine created outlived the integration that created it. The credential is held so that it can be undone, and it is named in orphaned_secrets when an integration is disconnected, so an operator knows exactly what to revoke.

DELETE /setup/integrations/{kind} asks; it does not remove. It answers 202 and records the intent on the fleet row, and the reconcile loop removes what the integration holds at the third-party app (the webhooks it registered, and the accounts it created when asked) before the block leaves the company document.

That order is the whole design. The block carries the credential the teardown authenticates with, so dropping it first would strand every webhook and account with nothing left to authenticate a second attempt. Until the teardown succeeds the surface reports phase disconnecting, labelled Disconnecting, and a failure holds it there and retries rather than letting it drift back to looking connected.

{ "remove_seats": false, "force": false }

remove_seats is the console’s “also remove the accounts Crewlet created”. It defaults to false and is never inferred: the engine’s own webhooks come out either way, because nothing else uses them, but an account is a colleague at that third-party app with history attached. Mattermost bots are disabled rather than deleted, because deleting a Mattermost user takes its posts with it.

force drops the block immediately without waiting for the third-party app, answers 200, and is the way out of a teardown that can never succeed: a revoked credential, an instance that is gone. It is the operator saying they will remove what the third-party app holds themselves.

Either way the sealed credentials are named, not deleted, in orphaned_secrets: one an operator may be sharing with another deployment is not something a disconnect decides about on its own. crewlet secrets unset is the deliberate path.

“Either way” is newly true. The list was built only on the force path — the ordinary 202 returned no orphaned_secrets field at all — and it walked the company-level requirements only, so no seat’s own credential was ever named on either path: not the token a pass minted under the name that seat’s mcp_env points at, not Atlassian’s address slot beside it, not Slack’s per-seat bot token and signing secret. It is computed once now, before the request splits, which is the only moment it can be: every name is derived from a ${VAR} in the company document, and both paths end with that block gone. The dashboard shows the list rather than discarding the response and closing.

A credential a sibling surface still reads is not on the list. Jira and Confluence normally share one seat credential — Atlassian issues one API token per account, and the ordinary place for it is the shared mcp_env.atlassian block — so disconnecting one product alone used to name a token the other went on resolving. Following that list takes the product that stayed down. A sibling counts as a user unless it is itself disconnecting, which is what makes a whole card work: the dashboard takes the card’s surfaces in order, each request records its intent before the next is made, and the union the dialog shows names the shared credential exactly once.

GitHub’s per-seat App credentials are on the list too, and they are the ones nobody typed in: the engine converts a manifest and writes what GitHub returns, so an agent’s private_key and its App’s webhook secret exist without an operator having chosen either name. They also survive a disconnect by design — GitHub has no API for deleting an App registration, so the engine uninstalls the App and hands over a link to the page a person deletes it from, and the key stays valid for something that still exists. That makes this list the only place either value is ever mentioned; it named the company-level signing secret alone, so two sealed per-seat credentials stayed in the store with nothing telling the operator they were there.

Refusals: 503 surface_busy when a reconcile tick or an operator’s own pass is writing at this surface right now, which is the one refusal here that clears on its own and carries its own code for that reason: a caller that cannot tell a race from a fault treats both as terminal. The request waits a busy surface out for a few seconds first (a tick a moment from finishing is the common collision) and then names it; repeating the request is correct, because every step of a disconnect is idempotent. It used to answer internal_error, which a caller can only treat as terminal — the dashboard stopped at the first refusal, so a collision on the second of Atlassian’s three surfaces left the tool half disconnected. A busy surface is now waited out rather than skipped, because the order matters: the organization’s credential is what removes the accounts.

Refusals worth knowing: 409 requirements_outstanding names the fields still missing (a pass writes at the third-party app and must not run against a half-configured integration), 409 no_public_base_url when nothing has told the engine what address third-party apps reach it on, and 409 pass_in_flight when another pass for the same third-party app is already running. That last one is a refusal rather than a queue on purpose: minting twice is not something a retry should paper over.

POST /setup/integrations/{kind}/check runs the same pass with no sink, which is what makes it read-only: every vendor gates its registration and its minting on having somewhere to seal a credential, so a run without one reads and reports and writes nothing at the third-party app. It answers “did what I just fixed at the third-party app take”.

A check does get integrations.public_base_url, and the 409 no_public_base_url refusal above is the writing route’s alone. Withholding the address from a check made it report the wrong fact: a vendor handed no base reads that as this deployment has no inbound address and reports ingress_blocked owed by an admin — and a check records its findings through the same fold as everything else, so pressing Check on a healthy company wrote “every monitor that fires reaches nobody” into the live status row and flipped the card to Action required over a value that was already set. Supplying the base is not the permission to register; having a sink is.

Both record their outcome on the same fleet integration status the reconcile loop writes, through the same fold, so a pass run by hand and a tick that runs a minute later cannot disagree, and Settings › Integrations updates with no extra plumbing. A pass that failed is recorded too, as the loop records one: phase activating, actor engine, findings dropped, because a pass that failed did not observe anything.

recreate_webhooks re-registers with a fresh secret and is destructive across deployments: the previous secret stops working everywhere else this company runs. On GitLab it also rotates every seat’s token, which revokes the credential each agent is currently authenticating with.

No pass on this surface deletes anything. Decommissioning a service account whose seat left the configuration stays a command-line gesture, because a company mid-edit looks exactly like one that removed a seat.

Two third-party apps put an agent’s identity on the seat rather than on the company, because on both of them one app is one bot: Slack, whose credentials an operator pastes in, and GitHub, whose app the engine creates. Both carry a seats array in their tool state, one entry per agent seat.

Slack’s entries are a form: each agent has its own Slack app, so each has its own bot token and signing secret, and every entry carries its own requirement list, its own inbound_path and its own satisfied.

A submission for one of them names it:

{"seat": "sre-lead", "values": {"bot_token": "...", "signing_secret": "..."}}

Those write through the entity route rather than a merge patch, because a merge patch replaces an array wholesale and patching the roster to change one seat would delete every other one. The engine addresses the seat by its handle, which is its identity rather than its position, and everything the submission did not send stays exactly as stored.

Every agent seat is listed, configured or not: the list is what a screen renders a form from, so leaving out a seat with no app yet would leave an operator no way to give it one. Human seats are excluded, because a person’s Slack account is not something this engine holds a token for.

It does not create the apps. That goes through Slack’s app-manifest API, which authenticates with a configuration token Slack issues only by hand and which an organisation may decline to allow at all, so crewlet slack provision remains the automated path where those tokens are available. What this surface does is make a hand-created app usable without one: it takes the two values Slack shows on the app’s own page and seals them.

A GitHub seat’s entry carries no requirement list, and the empty one is deliberate rather than unfinished: nothing here is typed in. The app is created from a manifest, and GitHub returns its id, its slug and its private key once, to the engine, which seals the key and records the rest on the seat. What the entry answers instead is what is still outstanding for that agent:

{
"handle": "builder",
"name": "Builder",
"requirements": [],
"tier": "full_access",
"step": "install_app",
"action_url": "https://github.com/apps/acme-builder/installations/new",
"present": true,
"satisfied": false,
"detail": "the app exists and nothing has installed it, so it sees no repository and mints no usable token"
}

step and action_url are the two clicks, described under One agent’s own GitHub App. tier is the seat’s access tier, answered before the app exists because it is what the manifest asks for, and defaulting to read_only on a seat whose configuration is silent.

present is whether the seat has an app at all, and satisfied needs that app installed and its sealed key readable by this node. A ${VAR} naming a secret the store does not hold is the state that reads as configured everywhere else while the agent mints no token, so detail names the variable to set. It names the reference, never a key.

An unfinished roster does not hold the card open: the company block is what decides whether deliveries arrive, and a company running apps for three of its ten agents chose that.

A seat is never satisfied on an integration the company does not declare. The roster used to answer about the credential alone, so a seat whose ${VAR} resolved read as satisfied over a surface a disconnect had removed from the document entirely — with detail naming the config path the value sits at, which is a fact about YAML rather than a state. It says the seat is waiting for the integration to be connected instead.

Nor is it satisfied when the reconcile loop has a finding about it. What the document and the sealed store can establish is that a credential exists where the app looks for one — a real fact, and not the one a green row is read as. A key deleted at the third-party app leaves the pointer resolving perfectly while every call the agent makes is refused. The loop is the only thing that has asked the vendor, so a seat named by a finding on that surface’s status row comes back unsatisfied, carrying the loop’s own sentence as its detail. Advisory findings are skipped on their own verdict — a permission wider than the role asked for is a note on a working agent, and reporting it as a broken one would contradict the card’s own tag.

There is a window this cannot close: between a credential being destroyed at the vendor and the loop next looking, nothing anywhere knows. That window is integrations.check_interval_seconds (600 by default, floor 60), and POST /setup/integrations/{kind}/check is how to ask immediately.

GitHub is the other per-seat case, and it is not a form. A GitHub App is created by POSTing a manifest from a page carrying the operator’s own GitHub session, so this surface hands the dashboard what to submit rather than collecting values, and the two clicks that follow are a person’s. There is no server-to-server equivalent, which is why a reconcile pass cannot do this one alone.

POST /setup/integrations/github/app begins it, naming the seat:

Terminal window
curl -X POST https://engine.example.com/setup/integrations/github/app \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"seat": "senior-engineer"}'
{
"seat": "senior-engineer",
"tier": "review",
"action_url": "https://github.com/organizations/acme/settings/apps/new",
"manifest": {"name": "Acme Senior Engineer", "public": false, "...": "..."},
"state": "<signed token naming the seat>"
}

Those three values are what GitHub’s manifest flow takes: the browser POSTs the manifest as a form field to action_url, carrying state on that URL so the redirect can be tied back to the seat that started. action_url is the organization’s own app registration page whenever integrations.github.provisioning.org names one, because an app registered under a person’s account cannot be installed on the organization that owns the repositories.

Refusals: 400 bad_body, 400 seat_required, 409 no_active_revision, 404 no_such_seat, and 409 no_public_url when integrations.public_base_url is unset. The last one matters more than it looks: an app is created with its delivery, redirect and setup addresses baked in, and only a person at GitHub can change them afterwards, so creating one now would mean creating it again later.

GET /webhooks/github-app is where GitHub returns the browser, twice. It is unauthenticated because a redirect carries no engine credential; the state is what stands in its place, and it is validated before anything else happens.

  • With a code, the engine converts the manifest, seals the app’s private key and webhook secret first, then records app_id, app_slug and a ${VAR} pointing at the sealed key on the seat, through the same per-entity config route per-seat setup uses. installation_id is written as 0: the install is a second act. The page then links to the install. The seal comes first because GitHub returns those two values exactly once and reissues neither, so a failure after it costs a retry and a failure before it costs the app.
  • With ?installed=<handle>, it confirms the install and writes nothing. There is nothing to convert — installing is GitHub’s own act and returns no code — and nothing to check the query against either: an agent’s app is private, so it is installed from the organization’s own installations page, which sends back no state. An unsigned query is therefore never allowed to reach the company document. The reconcile loop adopts the installation on its next pass, having listed the app’s own installations — the reading that can tell a real id from a typed one.
  • With ?error=, it renders GitHub’s own error_description, which is the operator’s to read (they cancelled, or they may not create apps on that organization).

It answers 200 for a completion or an install, 400 for a refusal, a missing code, or a state or code the engine will not accept, and 503 when this process has no setup surface behind it. An error message here is always the engine’s own wording and never a quote of GitHub’s response body, because that body carries the app’s private key.

The roster says what is outstanding. A seat entry in a tool’s seats array carries three more fields where a per-seat app has tiers and clicks: tier is the seat’s access tier, step is what is left as a closed set of create_app and install_app, and action_url is where a person goes to do it. Two steps rather than one, because they are two acts minutes or days apart and an operator who has done the first needs to be told the second is left rather than shown the same button. action_url is empty for create_app, and that is not an omission: an app is created by POSTing a manifest, not by following a link, so the dashboard asks the begin route above for one and submits a form.

The engine reaches Atlassian over three of these keys: jira, confluence and the Forge relay. Each is its own block in the company document and its own entry here, and each is written on its own, which is what keeps a submission atomic on the thing it changes. A screen that presents them as one tool asks for one section per key and submits one request per section.

The Forge app id is the exception, and it rides with jira: one app relays both surfaces, so it is one value with two consumers. Listing it under both would be two forms writing one field.

/ws/stream is where the dashboard READS. State comes down it and questions go up it: the handshake snapshot carries every section a screen needs on first paint, subsequent pushes carry what changed, and anything fetched on demand — an agent’s LLM history, one event’s payload, a trace, a different spend window — is a query sent on the same socket and answered on it. The socket carries no write.

REST carries the rest, and it is two things the socket deliberately is not:

  • Writes. A change to the company’s work is POST /operator/act/{tool}, made as the person the dashboard’s token is bound to (see /operator/act); a change to the company document is PATCH /config; a credential is /secrets; an integration’s setup is /setup; a backup is POST /backup. A write answers with the position it landed at, and the reads it moved are asked again on the socket at that position.
  • Guarded reads the query registry does not answer — the secret names (/secrets), where each ${VAR} resolves from (/config/references) and an integration’s setup (/setup/integrations). They are credential-scoped surfaces with their own refusals, read through the dashboard’s one REST loader rather than mirrored onto the socket.

The REST read routes below are also a public read API, and GET /stream/snapshot is the fallback for a browser that cannot upgrade to a WebSocket (corporate proxies).

Every named read route is an adapter, never a second implementation: it resolves its path values and hands them to the same answer the socket’s query channel reaches. The generic form GET /query/{what}?a=b reaches the same answers by name and is what the socket’s own frames map onto, so the two can never drift. A question whose source this node lacks — an event log, a schedule ledger — is left unregistered rather than answered empty, so its route replies 404 with unknown_query, which tells it apart from the no_route a path nothing serves answers.

One shape, whether it came from the live projection or from the event store, and the field names are the same either way — a screen shows a live row beside a historical one, so two spellings of one event would render the two halves of one list differently:

{
"id": "…", "type": "chat_message_received", "source": "mattermost",
"timestamp": "2026-08-25T12:00:00Z", "category": "chat",
"summary": "…", "actor": "…",
"trace_id": "…", "span_id": "…", "parent_span_id": "…",
"failed": false,
"payload": { }
}

payload is present only where the event was fetched by id or by trace: a listing deliberately never selects it, because a page of events with every payload attached is the query that makes a live screen slow.

Every stored type reaches the activity feed — the snapshot’s events and each event frame — except auxiliary_spend, which is accounting rather than activity: what the auxiliary model cost, coalesced per key per flush. It is a row like any other in GET /events, a turn’s turn answer and a trace, and it reaches every spend figure through the rollups; it is never pushed as an event frame and never takes a row of the feed (see Event System). A read that is merged with the feed or drawn beside it asks feed_only=true on events and event_series, and gets the feed’s rows alone — what the dashboard’s event log, its Live strip and a seat’s latest events do, so a page scrolled on from the ring and the bars above it hold exactly what the ring would.

llm_history is the seat’s finished phases, read from the event store of every node that ever held the seat — placement moves a seat, and each node keeps the phases it ran — one row per agent_phase_completed, newest first, capped at 50, with the answer’s coverage saying which nodes it heard from (see Reading the fleet’s history). The call in flight is not in it; that is live.live_call, which comes from the projection, and the two are different sources on purpose: the store holds what completed, memory holds what is happening. A screen renders both with one renderer, so each history row carries the same fields a live one does — turn_id, phase, iteration, model, response, tool_executions, round_narration, partial_round, total_tokens — plus the envelope’s timestamp and failed. It is the stored record, so it also carries the phase’s price where its own CLI reported one (cost_usd), which the dashboard never reads (rule 19). A detached coding run is a row of its own, phase: sandbox, published when the run is collected: its tokens, its launch_id (a turn can launch two runs in one iteration, so the launch is part of that row’s identity), its report as response and its activity_transcript — see each run is published as a phase. A finished row also carries duration_ms, which a live one cannot: it is the engine’s own measurement of the phase, published on the record rather than reconstructed by pairing it with the agent_phase_started that shares its key. Zero means not measured — an agent-mode executor’s rounds ran inside a coding CLI’s own loop, in another process — never took no time. A phase that failed before it reached a provider also truncates to zero at this resolution; failed is what separates the two.

It measures the whole phase, including a suspend. A detached coding run parks the executor mid-loop and the phase is re-entered later — after a restart, possibly on another node — and the parked row carries the clock across with the rounds and the tokens, so the one record the phase publishes reports the run rather than the seconds spent collecting its answer.

A finished row carries the phase’s timeline too, and so does every other reader of the record — phases, turn, trace and event answer the stored payload verbatim: started_at (when this segment began — not published minus duration_ms, which on a resumed phase spans a coding run), rounds[] (one {round, started_at, duration_ms, model, input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, tool_calls} per provider call, the model’s half of the round only), each tool_executions[] row’s started_at, duration_ms, origin (builtin or mcp:<server>) and server, the phase’s cache_read_tokens / cache_write_tokens (a breakdown of input_tokens, never an addition to it), max_rounds / round_ceiling, a worker’s or a judge’s host_round, a resumed executor’s launch_id and the turn’s work_item, and steers[] — {round, note_id} for each person’s note the phase read. The live frame carries the same so far plus round_started_at and the running_call in flight. Each is absent where nothing measured it — a tool call nobody timed has no duration_ms, never a zero — so a reader treats absent as not recorded. The whole field list, and why each is measured where it is, is in what a turn records about its time.

An unreadable or absent event log costs the history and nothing else: the answer still carries the seat and its live state.

One frame, kind: "snapshot", with every section a screen needs on first paint. Three of them are derived from configuration rather than from anything that has happened, and they are present from the moment the socket opens — before a single turn has run:

SectionWhat it is
agentsThe company’s agent seats, each merged with its live overlay. Every seat in the company, not the ones this node runs, because the dashboard is a view of the company. Human seats are excluded — they have no turn, no phase and no spend; they appear in org with "kind": "human"
orgThe same public projection GET /org answers: the charter, the company’s resolved timezone and its token_budget, root-level roles and units nesting to any depth, with only the public fields of each — a seat’s resolved llm chain and tool_sources among them
toolsThe catalogue this node serves, each entry tagged with the source that registered it — builtin or the MCP server’s name. Empty on a node with no active revision, which has no catalogue yet
events, sandboxes, tokens, budget, healthThe live projection: what has happened

Every seat carries one seat-state vocabulary, computed by the engine and never by a client: activity is working, needs, stopped or idle, and stopped_reason is paused, unplaced, budget or provider on a stopped seat and null otherwise. Both keys are on every row, a seat no event has mentioned included. What each word means, which inputs it is read from and in what order they win is Agent States. Placement is the fleet’s lease table, read on every snapshot and on the five-second tick, so a seat a peer runs reads idle or working here rather than being left without a state — which the dashboard used to draw as offline, on every seat another node held — and a seat no node holds reads stopped/unplaced. A token meter report names every capped seat whether or not anything runs it, and moves a seat’s state only when one of its windows is refusing.

The three config-derived sections are re-sent on every config apply, as seats, org and tools pushes. Nothing else would correct them: a revision that adds, renames or removes a role produces no event a projection could learn from, and an overlay merge cannot express a row going away.

Live-state projection (api/stream + LiveState)

Section titled “Live-state projection (api/stream + LiveState)”

Every node that serves the API maintains an in-memory projection of every agent’s current state (internal/api/livestate.LiveState, owned by internal/api/stream.Service). It is fed by the same event stream the WebSocket fan-out consumes and read in O(1) thereafter — so /agents, /stream/snapshot, and the WebSocket handshake never re-derive state from a multi-query event scan on a request.

What the projection holds is what has happened: a phase, an in-flight call, a spend. Who the seats are, how they are organised, what tools they can reach and what they are scheduled to do all come from the company document instead — see the handshake snapshot.

What the projection computes, it also pushes. A dashboard mirrors it rather than re-deriving it: before this, every tab ran its own copy of the state machine below, its own sandbox tracking, and its own re-implementation of the spend aggregation, all applied to the raw event stream — three copies of server logic, three ways to drift, and a refresh that regularly disagreed with what had been on screen a moment earlier. Applying an event now yields a change set, and the changed agents’ overlays, the sandbox list, and the spend rollup go out as their own envelopes.

Crucially, the projection holds each agent’s in-flight LLM call — the latest agent_turn_progress (phase, round, model, accumulated response including the model’s reasoning, tool calls so far) — and surfaces it as a live_call field on the agent in every snapshot. Its response is built by the same function that builds the finished phase record’s, so the live row and the turn you expand afterwards are the same text rather than two assemblies of it — see Turn Engine § What streams during a turn. Each phase publishes an opening agent_turn_progress (round_num = -1) before its first provider call carrying the prompt, so the live row shows what the agent was asked while it is still answering. The field is otherwise ZERO-BASED, so the round a reader counts is round_num + 1, and -1 is neither round zero nor a missing value — it is the phase’s FIRST round, in flight and not yet answered, and the frame carries max_rounds so a row can say “round 1 of 24” from that moment. A surface drawing it resolves every case through lib/seats.ts’s roundOf (at least 1 while there is a call) and roundLabel, whose hint says when the first round has not come back; the roster once drew a bare dash, and the stepper a bare “Execute” for as long as a slow first answer took. agent_turn_progress is stream-only (never written to the event store); carrying live_call in the snapshot means a tab that refreshes or reconnects mid-call re-renders the live row immediately instead of waiting for the next progress event. State transitions in the projection are gated on the event timestamp so out-of-order delivery (measured, JetStream returns a failed (NAK’d) message behind never-delivered ones rather than replaying it from the head, and in a fleet several nodes publish into the event stream at once) can’t clobber newer state with an older event — including a final progress round that overtakes its own phase completion, which would otherwise re-open the row of a phase that had already finished.

Every one of those comparisons normalizes to a single ordering key first. The same instant reaches the API in more than one encoding (Z, +00:00, naive, a non-UTC offset), and as raw strings those order differently: …05Z sorts after …05+00:00 for the same moment, and 13:00+01:00 sorts after 12:30+00:00 while being half an hour earlier. Compared that way a straggler round resurrects a phase that has already finished, and a stale event clobbers newer state.

A failed phase is a first-class part of this. When a phase dies, the projection stamps its in-flight call failed, keeps it on screen rather than clearing it, and records the classified cause on the agent as last_error — so a seat that stopped can say why, and the call it stopped on is still there to read.

The agent detail page streams the same way: it paints from a agent query, then keeps itself current from the pushes. agent_phase_completed / agent_turn_completed envelopes append to the phase transcript live, and the agent’s own live_call — pushed on every round — is the in-flight row inside its turn card, replaced by the completed record when the phase finishes. Progress envelopes carry turn_id / phase / iteration for this correlation; they are stream-only and never persisted to the event store.

The spend rollup is maintained by the projection too. It HOLDS the spend records for its own 24-hour window — each phase’s, and each auxiliary_spend record of what the auxiliary model cost — the newest 24 000 of them, the WHOLE COMPANY’s since every projection is fed fleet-wide (about 2 600 turns a day; a company past that sees a rollup covering slightly less than a day rather than a wrong total) and folds them with internal/tokens, which folds the named windows’ company days into the same breakdown, so changing the window cannot change what a phase is counted as. The window is a ROLLING one, aged on the serving node’s clock and never on a record’s own stamp — every read leaves out what the window has aged past, and both an arriving record and the shared tick drop it:

  • a record stamped before the window is not counted, however late it arrives;
  • one stamped ahead, by a node whose clock runs fast, is aged from when it arrived — from the startup seed, for one read from history — so it leaves a day after it came and the cap takes it in that place, rather than moving the window or outliving it. It still shows the stamp it was published with, which is how the node with the wrong clock is found;
  • one whose stamp does not parse is kept (nothing can age it) and is the first the 24 000 cap drops.

The rollup’s since and until are the two instants the window was cut at. It ships in the snapshot and is re-pushed on the shared 5-second tick after any phase completed or any record aged out, so a client holding it stays live without a fetch — and a company that has gone quiet sees its spend leave the window as it ages rather than keep its last busy day — without a second implementation of the aggregation in the browser. (The bundled dashboard’s Spend screen reads named windows instead.) What a seat has spent is that rollup’s per-agent row: the projection keeps no second total of its own.

The projection is seeded from the fleet’s event stores when the process starts, after the broadcast subscription is attached and before the HTTP listener binds. The seed makes three bounded reads, side by side, and each read is bounded by the projection’s own limit:

  • the newest 400 persisted events, for the feed;
  • the newest 24 000 spend records — phases and auxiliary records — inside the 24-hour spend window, read in pages of 4 096 records a node, because a busy node’s whole day is more than one reply of the scatter carries (8 MiB) — asked whole, its reply was cut, and every node’s records were cut with it;
  • the newest three turns of every agent seat, for the seat’s last_turn and for a turn it left parked.

Each read goes to every live data node, through the same scatter the history queries use (see Reading the fleet’s history). Each data node’s store holds only what that node published — plus what any node holding no data handed it, see Deployment → The event store — so a seed that read this node’s store alone showed a restarted node only its own share of the company. Without any seed, every one of these surfaces started at this process’s boot: a restart, a deploy or a node joining a fleet showed an operator a company that had apparently done nothing, beside a store that said otherwise. The projection records which nodes answered, as a coverage, and a read that failed makes that coverage incomplete. An event that arrives both ways is recognised by its id and is listed and counted once, whichever arrives first. The stream can deliver it before the read, and the read can find a row that the publishing node wrote inline before the stream delivered it. History is ordered behind the live rows it predates, and a seat’s turn that the stream has already moved is left as the stream left it. A read that fails is logged as live_projection_not_seeded and costs that history, never the start-up. The reads run side by side, so a slow read does not use up the time budget of the others.

The paused seats are seeded from the coordination record, beside the reads above and before the bind: a pause taken before this process started is in no event it will hear, and a paused seat drawn as working is the one state a person pausing it must not be shown. A failed read is logged as seat_pauses_not_seeded; the pause itself is in force either way.

The running coding runs are reconciled against the durable run record. The record is read once before the listener binds and then every 30 seconds. The stream is lossy and in memory, so it cannot say which runs exist, and the record can. See the running-runs panel for which of the two wins when they disagree. A reconcile that changed the set pushes it as sandboxes, and a reconcile that changed nothing pushes nothing.

Single-shot bundle equivalent to the WebSocket handshake’s first envelope: every section a dashboard screen needs on first paint. Assembled entirely from the in-memory projection — no database round-trip on the hot path. Used as a fallback when the browser cannot upgrade to a WebSocket (corporate proxies, etc.).

Response

{
"health": { /* the whole health envelope described below */ },
"agents": [ { /* /agents row: activity (working | needs | stopped |
idle) + stopped_reason (paused | unplaced | budget |
provider, or null) + budget meter + live_call (the
in-flight LLM call, or null between turns) +
last_error (the phase failure that stopped this
seat, or null) + turn (the turn the seat is on,
or null) + last_turn (the newest turn it ended,
or null) + paused ({by, at, reason, stop_running}
while a person has the seat paused, or null) */ }, ... ],
"events": [ { /* recent event row, newest first — payload-free, plus
a `failed` boolean */ }, ... ],
"sandboxes": [ { /* in-flight detached coding run: turn_id, role,
agent_handle, agent_id, coding_agent, sandbox_id,
task, status (the run record's own word), started_at,
question, audience, work_item, owner, paused_at */ }, ... ],
"tools": [ { /* one catalogue entry — see The Tool Catalogue below */ } ],
"org": { /* /org payload */ },
"tokens": { /* the spend rollup — same shape as /tokens/breakdown */ },
"budget": { /* the live org-wide token meter, or null before any node has reported — see below */ },
"schedules": [ { /* configured schedule + computed next_run */ }, ... ]
}

Each events row is the payload-free feed shape — id, type, timestamp, source, actor, summary, category, trace_id, span_id, parent_span_id, topic — plus agent_id, the id of the seat the event concerns (the store’s own agent_id, absent for an event about no seat), which is what lets a client narrow its LIVE rows to one seat exactly as /events?seat= narrows the stored ones — and failed: true when the work the event reports did not succeed. It is true for an event carrying its own failed field (a phase or turn that died) and for an event type that is a failure (sandbox_run_failed, llm_unavailable, budget_exhausted, turn.guard_breach). Deciding it once, here, is what lets a dashboard mark failures without re-deriving them from a type list of its own.

The flag survives a restart: the event-store writer stamps a failed tag on those events, and the projection reads it back when it seeds its feed from the store at startup. The store’s listing (EventLog.List) deliberately never selects the payload column, so without the tag every historical failure would read back as a success.

One builder (App.health, internal/api/health.go) answers GET /health, the snapshot’s health section and the 5-second health push, and all three carry the whole envelope, so no two of them can disagree about whether the engine is healthy. There is no query for it: a screen reads the push. (The push once carried three fields while a stream query answered the rest, and five screens polled that query at cadences of their own, so the rail and the panel in front of it could disagree for fifteen seconds about whether a revision had applied.)

The envelope is public — GET /health is an unguarded probe and the push reaches an anonymous tab — which is why the fleet and the alarm table appear on it as counts (nodes, alarms) and never as their rows. Which node holds what is the operator-only fleet answer, and what each alarm measured is work_retention’s.

{
"status": "ok",
"node": "core-1",
"configured": true,
"version": "v0.4.0",
"started_at": "2026-04-01T11:58:03Z",
"queue": "jetstream-embedded",
"clients": 3,
"event_history_seconds": 2592000,
"spend_history_seconds": 15638400,
"in_flight": 2,
"shutting_down": false,
"posture": "serve",
"applied_epoch": 41,
"seats": ["ceo", "cto"],
"unproven_seconds": {"eng": 312.5},
"nodes": 3,
"alarms": {"count": 1, "worst": "backup_age"},
"seeded_from": {
"nodes": [
{"id": "core-1", "answered": true, "error": ""},
{"id": "core-2", "answered": true, "error": ""}
],
"complete": true
}
}
FieldMeaning
statusshutting_down, unconfigured, a diverged posture (shed, stuck or isolated), or ok, in that order of precedence. A draining engine is draining first, whatever else is true of it, and a node with no active revision is that before it is anything else. The two ordinary postures, serve and wait, read as ok. A node applying its first company reads unconfigured until that apply completes and ok from then on; a later apply reads ok throughout.
nodeThe name this node’s engine runs under — node.id, else CREWLET_NODE_ID, else node-0 — which is what its presence lease carries and the only way a caller can tell which node a load balancer sent it to.
configuredWhether a company revision is active. Read off the engine’s live epoch on every call, so an apply that brings this node its first revision flips it. When false the node refuses every inbound webhook with 503, so an operator watching empty screens needs to be told this rather than left to infer it.
versionThe crewlet version this process is running.
started_atWhen this node’s engine started, which is when the node started: the API is served inside the engine’s process. The fleet view reports the same instant for this node.
queueThe event queue’s backend — jetstream-embedded (a NATS server inside this process), jetstream (an external NATS cluster this node dialled), or memory. Read off the EventQueue contract’s own Backend(), never sniffed from a type name. Display only; nothing may branch on it.
clientsDashboards currently connected to this node.
event_history_secondsHow far back the event log can be read — the hard bottom of paging: once a cursor crosses it every page is empty forever, so a client that cannot name the floor draws the store’s own horizon as “the org went quiet”. The store’s constant, not a number this API picked, so a change to the retention reaches every screen without an edit. Seconds rather than days, because the retention is a duration and a client re-deriving the unit is a second place the number can be wrong.
spend_history_secondsHow far back a named spend window can reach: the replicated usage domain’s own history (181 days, ADR-0020), which is not the event log’s. A spend chart states this floor and never event_history_seconds — the two answer “can I still chart that month” and “can I still open that turn”.
in_flightTurns running on this node. Always present, and a 0 is a real zero: every process that serves the API runs the engine beside it.
shutting_downtrue from the first moment of a drain, so a dashboard shows the drain while it happens: the listener keeps serving until the drain has completed. See During a drain.
postureThe node’s config posture: serve, wait, shed, isolated or stuck. The only place an operator can see why a node left rotation, since /ready answers a bare 503 either way. A node applying a revision reads wait until the apply ends — serve once it succeeds — so a client polling across a PUT /config sees no diverged posture from a write that applies, a corrected revision written after a failed one included; shed, isolated and stuck follow only an attempt at that epoch that concluded in failure, or (shed) a lag on that epoch past 45 s while a peer has it.
applied_epochThe activation epoch this node last applied.
seatsThe handles of the seats this node holds, [] on a node holding none.
stall_lag_secondsPresent only when the node’s watched duty is behind: how far, in seconds. It climbs towards the seat lease TTL, at which the watchdog ends the process.
nodesHow many nodes hold a presence lease — the fleet this node’s fan-outs (search, fleet history) divide their work by. Absent when the presence read failed or did not finish inside the probe’s coordination budget (an eighth of the 15-second reconcile interval, under two seconds) (it runs beside the posture read, so a wedged broker slows /health by that budget rather than hanging it); never 0, since the node answering is itself one. A screen says “node count unavailable” for an absence rather than guessing.
alarms{count, worst}: how many of this node’s alarms are firing, and worst, the one that has been firing longest (absent when count is 0) — the table asserts no severity of its own, and the condition that has gone unanswered longest is the one a health card names. From the same evaluation the crewlet.alarm.active gauge and the alarm_raised / alarm_cleared log lines come from, which runs every ten seconds on every node. Absent before that evaluation first runs and on a node running no state log: neither has looked, and {count: 0} would read as healthy.
seeded_fromWhich nodes this node’s live projection was seeded from at boot — the activity feed and each seat’s last turn, which the pushed screens start from, and the live spend window behind the tokens push — in the fleet coverage shape. Absent until the seed has run. A seed that missed a node started those surfaces a node short, and this is where that stays visible after the log line has scrolled away.
unproven_secondsEach seat whose teardown this node could not prove, mapped to how long it has been stranded, present only when one is. Such a seat is still leased by this node, so no peer can claim it, and this node will not run it: it is absent from seats for exactly that reason. Alert on the duration rather than on the field’s presence: a release that fails once and succeeds on the next heartbeat is a working system. See Seat ownership.

Per-socket facts, such as how many envelopes this connection dropped or how deep its queue is, are deliberately not here. The tick encodes one JSON string and hands the same string to every client, so a per-client field would force one encode per client per tick. A connection that lost envelopes to backpressure is logged as stream_client_left_behind, with the count, when it disconnects.

GET /health always returns 200, including when status is unconfigured or a diverged posture: the status code is liveness, and an engine waiting for a configuration is alive. Steer traffic with GET /ready instead, which answers 503 and names the reason.

A node whose node.roles leave out ingress — a satellite, or the seats and workers half of a split deployment — serves no API, no dashboard and no webhook. But an orchestrator runs every node, and one it cannot probe it can neither restart when it wedges nor wait on during a rollout. So such a node binds api.port for exactly two routes, and on a seats node with CREWLET_MCP_BRIDGE_URL set the tool bridge beside them — or, with api.public set, on that public listener alone, where it is the only route; every other path answers 404 no_route, or 401 for a write and for an always-guarded prefix, as an unknown path does on an ingress node. The node logs api_probes_listening with tool_bridge saying whether the bridge is mounted and public_addr naming the public listener when the bridge has one. api.port: 0 binds nothing, on this node as on any other.

GET /health is liveness, 200 while the process is alive — through a drain and through a broker link that is down. Its body is the node’s half of the health envelope, under the same field names, plus the node’s roles; it leaves out what describes a dashboard this node does not serve (clients, event_history_seconds, spend_history_seconds, seeded_from):

{
"status": "ok",
"node": "sat-eu-1",
"roles": ["seats"],
"configured": true,
"version": "v0.4.0",
"started_at": "2026-04-01T11:58:03Z",
"queue": "jetstream-embedded",
"in_flight": 1,
"shutting_down": false,
"posture": "serve",
"applied_epoch": 41,
"seats": ["eu-support"],
"nodes": 4
}

GET /ready answers whether the node is doing its work, because it takes no traffic for the probe to steer: what waits on it is a rollout that must not replace the next node until this one has joined. 200 once every condition below holds; otherwise 503, with the first that does not in reason, in this order of precedence:

reasonThe node is not ready because
drainingIt is stopping.
broker_unlinkedIt cannot reach the fleet’s broker: a leaf whose link to every member is down, or a connection to an external cluster that is reconnecting. detail carries the cause. Ahead of what follows, because each of those is what a lost broker looks like.
unconfiguredNo company revision is active here.
shed, stuckIts config posture is diverged. wait and isolated stay ready, as they do on an ingress node.
no_presenceIt does not hold its presence lease — its last renew is older than the lease’s TTL — so no peer counts it and nothing routes to it.
admission_withheldIt runs seats, and its latest placement pass was not admitted to claim: no data node has yet answered that a copy of the estate admits a seat (a stateless node), its own copy is behind (a data node), or it cannot serve the seats it holds. A node that runs no seats, or one started in a maintenance mode, is not judged on this.
{
"ready": false,
"node": "sat-eu-1",
"configured": true,
"draining": false,
"posture": "serve",
"reason": "broker_unlinked",
"detail": "jetstream: this leaf has no link to any member of the fleet, so every stream and bucket it uses is out of reach"
}

An ingress node’s /ready is unchanged by any of this: it answers whether traffic should come to it, on draining, unconfigured, shed and stuck alone. The two share one judgement (judgeReadiness, internal/api/health.go) and one precedence, so a reason both can give means the same thing on both.

GET /events and the events query return rows ordered by (timestamp, id) descending, and accept an exclusive keyset cursor:

GET /events?limit=100&before_time=2026-04-01T12:00:00Z&before_id=<event_id>

Pass the oldest row you already hold to get the page beneath it. The id half is not optional — burst writes routinely share a timestamp at microsecond resolution, and a cursor over a non-unique key silently skips or repeats whatever collided with it.

The end of the history is exhausted: true, next: null — and only that. A page SHORTER than limit is not the end: the page is merged from every node’s store (see Reading the fleet’s history), and it stops at the newest point any node’s page stopped at, so a node whose reply was cut to fit the transport shortens the page without ending it. The agent filter adds another reason — it also pulls in every event sharing a trace with a direct match, from every node, so a caller must dedupe by id.

The persistent store retains 30 days, and event_history_seconds on the health envelope is that floor on the wire — read it rather than restating the number, which is the store’s own constant and not a promise this page makes. Once a cursor crosses that floor every page is empty — which is why a client must distinguish it from quiet, rather than drawing the gap as silence.

category is a filter for the same reason paging exists at all — filtering a paged list client-side silently excludes, because a 100-row page holding 2 matches reads as “only 2 exist”. Its vocabulary is a closed set of eight values, and which event type lands under which is in Deployment § What gets stored.

Every node’s event store holds what that node published and nothing else, so the history questions — events, event, event_series, trace, turn, turns, phases, a seat’s llm_history on agent, the delivery and outcome counts on integrations, and the live projection’s boot seed — are answered by every live node at query time: the node you asked reads its own store and scatters the same question to its peers, merges the answers, and says which nodes it heard from. Each of those answers carries one shape:

"coverage": {
"nodes": [
{"id": "node-a", "answered": true, "error": ""},
{"id": "node-b", "answered": false, "error": "no answer within the 2s fleet read budget"}
],
"complete": false
}
  • nodes is every node asked or heard from, sorted by id; the node you asked is always one and always answered — a failure of its own store is an error, not a gap.
  • complete is true only when the node roster could be read and every node on it answered. A node that did not is named with why: no answer inside the two-second fleet read budget, a build speaking another version of the scatter’s protocol, a reply that could not be read, or its own read failing.
  • Every node answers in one protocol version, its own, and refuses a question in any other — a node on another build is named, with that reason, on every answer, and nothing of what it holds is merged. No node is asked to answer around a filter it would drop, a count it would never send or the instant it would ignore, so every part that is merged was read exactly as asked: the axis’s window is cut alike on every node, a window the history clips down to the bucket the horizon falls in, and the serving node drops that partial first bar after the sum (see the time axis).
  • It sits at the top of each answer — beside the record’s own fields on event and on event_series — and is null on agent and integrations when the history could not be read at all.
  • A node that has left the fleet is not asked, because it is not live: its turn-level detail left with it. The spend and turn counts it recorded are answered by the replicated usage domain instead.
  • A stateless node’s row is listed and counted once, though for a moment two data nodes can hold it — a custody batch the second keeper wrote before the first had settled it. A listing holds it once by its (timestamp, id); every count — event_series’s bars, total, failed and by_category, trace’s and turn’s total, a turns row’s tokens, phases and duration_ms, and the outcome counts on integrations — counts it once, because each node counts the rows it keeps and names the ones it has not settled, and the serving node asks which node keeps each named row before it adds it.
  • event not found answers not_found naming any node that did not answer, because a link whose node was merely silent is a different fact from a dead one. An event older than the 30-day history is not found either: it is past the horizon every read of the log stops at, and a row a node’s sweep has not deleted yet is not history it serves — every node floors its lookup by id at the serving node’s horizon.
  • Every question is asked at one instant — the serving node’s clock when the question arrived — and that instant travels with it, so every node floors the 30-day history at the serving node’s horizon rather than at its own, and one node’s part of an answer (a turn’s rows, its count, its ending and its traces) is read at that one instant throughout. A node answering a second late, or with a clock a little ahead, therefore holds back nothing the others include, and a question carrying no instant is refused rather than answered as of a node’s own clock. turns merges in two passes — every node’s page, then every node’s share of exactly the turns any page listed — so a turn resumed on another node after a restart is one row folded from both halves. The window is pinned to the asker’s clock first — and the instant travels with the question, so every node floors the window and its share of each turn at the asker’s 30-day horizon rather than at its own — and the merged turn is held to it whole: a node’s half of a resumed turn can start inside the window while the turn began before it elsewhere.

A turn is paged where a node lists it: at the earliest start at which any node’s own page lists it, which each node says of its share in the second pass. That is not always its started_at, which is where the turn began: a turn resumed across nodes is selected by each node on its own half, so under a filter the whole turn passes, the half where it began may not — the clean half of a turn that failed later, under failed=true; the half that names no item of one charged to an item only by its later records, under work_item=. Paged at its started_at, such a turn could fall below where a page was cut and behind the cursor that page returned, and so on no page of the walk. Walking next lists every turn exactly once, two turns that start at one microsecond included, whichever side of a page’s cut they fall: the cursor is that position and the turn’s id, and every node resumes after exactly the turn the last page ended on. The rows are newest first by that position, the higher turn_id first within one microsecond, so a row whose started_at is earlier than where it is listed can sit above rows that began after it. next is the fleet’s cursor: it can be present on an empty page, where a node’s page stopped before any turn above it could be shown. With sort=-tokens the page ranks each node’s top turns by their MERGED totals and carries no cursor (before= is refused bad_params): a ranking has no position to resume from, and it can miss a turn split across nodes whose every half fell below every node’s cut.

Every change a person makes through a running node leaves one event, whether or not it went through:

Event typeWritten for
operator_actedEvery operator tool call that is not a proven read, on both transports — a button on the dashboard (/operator/act) and a person’s own assistant (/operator/mcp)
backup_requestedEvery POST /backup that began copying, whether the copy finished or not

Each carries source: "operator" on the envelope, so GET /events?source=operator is the runtime audit on its own; its actor is the token’s own name (operator_id), with actor_seat naming the person the token is bound to by contact.crewlet_operator_id (absent for a token nobody bound) — the seat the call was made as, resolved once for the call, so a config apply landing while it ran never makes the record name a different person. Both are filed under lifecycle.

{"type": "operator_acted", "source": "operator",
"operator_id": "founder", "actor_seat": "jane-founder",
"transport": "act", "tool": "update_work_item",
"request_id": "0192f1a4-9b2d-7e51-8c3a-6d7e8f9a0b1c",
"outcome": "applied", "position": "CREWLET_TRACKER_LOG@1:4711"}

outcome is one of five: applied, pending and unknown are the write’s own answer, exactly as the call returned it — and a call interrupted before it answered is unknown, because whether it landed is precisely what nobody knows; refused names the tool’s refusal class in refusal; failed is a failure the tool did not classify. The last two also carry failed: true, so the row is tagged failed and the log’s failure filter finds it. request_id is the act transport’s own (MCP sends none), so every retry of one gesture reads as the same request. A backup’s record names the dir and the number of streams the manifest covers; the node whose disk it was written to is the envelope’s node, which the queue stamps on every event with the node that published it — and the backup route publishes from the node that took the copy.

The arguments are never recorded. A page body or a comment is the company’s content and already lives in the history of the object it changed; the audit says who called what, and what became of it.

A listing carries what a row needs without its payload. GET /events never returns payloads, so the store promotes the runtime audit’s dimensions into each row’s tags: actor_seat (the bound person), tool (an operator_acted call’s tool) and dir (a backup_requested copy’s directory) — beside node, the envelope’s publishing node, which every event row carries. That is what the dashboard’s Audit log and backup history draw from. A row stored before these tags existed reads back without them.

What is not here. A request refused before any tool ran — a read sent to the write transport, a token that is nobody, a malformed body, a backup with no destination or with one it refused (the 400s below) — changed nothing and is not recorded. A proven read over MCP is not recorded either: an assistant asks many questions, and a row per question would bury the writes. Configuration and credentials keep their own records (a revision names who created it, a credential who stored it), which is why /config and /secrets publish nothing here. And like every event, the row is written by the node the call reached, so a fleet’s audit is the union of its nodes’ logs.

A record that could not be published is logged as operator_audit_not_published at error on that node; the call itself has already been answered, and its tracker or page history is unaffected.

GET /events/series counts the matching rows per bucket over a window. It takes every filter GET /events takes bar the cursor and agent (see below), plus:

NameDefaultDescription
bucket(required)minute, hour or day. A closed set rather than a duration, for the reason the spend series gives for its two: an axis with an arbitrary bucket width is one nobody can label. The log has minute and the spend series does not, because “what just happened” is the commonest question asked of a log and an hour is the whole of that answer’s window. An unknown value is refused naming what is accepted, never defaulted.
since / until(the 30-day history, to now)RFC 3339 instants, half-open. Both edges are snapped outward to whole buckets, so the first and last bars are whole ones — a partial bar has a height that means something different from its neighbours’ and a reader has no way to know. The one edge no snap crosses is the 30-day history (event_history_seconds): a window it clips — the default one included — begins at the first whole bucket inside it, since a bar reaching below it could count only its upper part, and a window lying wholly below that bucket covers nothing (since equals until, no bars). The answer is labelled with the window it actually covers.

The answer is {bucket, since, until, bars, total, failed, by_category}. bars is every bucket in the window including the empty ones, so a quiet hour is a gap of full width rather than a bar the chart squeezed out.

The whole answer is cut against one instant — the serving node’s clock when the question arrived — and every count in it, bars, total, failed and by_category alike, keeps rows from 30 days before that same instant. Across a fleet that instant is sent with the question, so every node cuts the same bars — a window the history clips is cut down to the bucket the horizon falls in, its first bar counting only what lies above the horizon — and floors what it counts at the serving node’s horizon rather than its own. The serving node sums every node’s bars and only then drops that partial bar, so no bar in the answer starts below the horizon (see since above), every bar lies wholly inside the history, and a row any node holds inside a bar — the first one included — is counted in it however long that node took to answer.

Each bar is {at, count, failed}. failed is how many of the bar’s rows reported a failure — by the rule a turn’s own failed mark uses: the event said so, or its type is one (llm_unavailable, budget_exhausted, turn.guard_breach, sandbox_run_failed) — and it is a split of count, never an addition to it. The answer’s own failed is the window’s, the sum of the bars’. Counted by the engine for the reason the bars are: a screen holds a page of rows and never the window.

by_category is how many rows each category would give, with the category filter lifted and every other one applied. That is the only meaning a facet count can have: counted through its own filter, every value but the selected one reads zero. A category with no rows is absent from the map, so a caller rendering the closed set reads a missing key as the zero it is.

It is counted over the window that was asked for, not the snapped one since and until report: a chip says how many rows choosing it would show, and the rows come from GET /events, which takes the caller’s own edges. So the chips need not sum to total — total describes the bars, which are whole buckets. On a window the history clips, the chips reach down to the horizon itself and so also count the rows between it and the first whole bucket, which no bar draws.

Two refusals, both 400 rather than a smaller answer:

  • related_agent is not accepted. That filter over-fetches and post-filters (see above), so a count over the predicate alone is a smaller set than the listing shows — and a bar that disagrees with its own rows is the one thing this axis exists not to be.
  • A window of more than 1,500 buckets is refused naming the bucket and the span. Truncating would put a month’s heading over a day of bars and coarsening would answer a different question from the one the axis is labelled with; the caller’s fix is a coarser bucket or a shorter window. The buckets counted are the ones every node cuts, so on a window the history clips they include the one the horizon falls in, which the answer then drops: the number every build refuses at is one count, not two.

budget carries the fleet’s shared token counters as the budget gate enforces them: for the company and for every seat whose token_budget caps a window, one entry per capped calendar window — the day, the ISO week and the month on the company’s clock — with that window’s spend, its ceiling in the active revision, the gate’s refusal stamp and the engine’s judgement of it. It is the only figure that can honestly be divided into a configured ceiling, because both cover the same span. The dashboard’s other token figures are spend rollups over a window of time the reader chose; dividing one of those into a ceiling produces a percentage that is wrong by however much was spent outside the window.

{
"meter_id": "node-a:7f3c…", "seq": 42, "timezone": "Europe/Berlin",
"org": {
"windows": [
{
"period": "day", "window": "2026-09-23",
"starts_at": "2026-09-22T22:00:00Z", "resets_at": "2026-09-23T22:00:00Z",
"used": 2710450, "limit": 3000000, "state": "near"
}
]
}
}

A seat’s meter rides on its row of the agents push as budget: {windows: […]} in the same shape.

  • period is day, week or month, and window is its label on the company’s clock — the identity every node computes alike. starts_at and resets_at are the window’s half-open span in UTC; resets_at is when its allowance comes back without a ceiling being raised. timezone names the clock the windows were cut on.
  • Only capped windows are listed, each with its limit; a scope that caps none carries windows: [], and a seat that caps none carries no meter.
  • state is the engine’s own judgement, computed once beside the counter so no screen holds a threshold of its own: refusing when the window has no room left for a single token — which every window that refused a round has, since the refused round is counted, and the condition a seat is parked on — near once used reaches nine tenths of limit (engine.BudgetNearFraction, served as near_fraction on GET /budgets), and ok otherwise.
  • refused_at is when the window last turned a call away, in UTC, and absent while it has not: a round whose charge it refused, or work turned away before its first call because the window was already full — a turn’s next call, a seat’s delivery parked, a person’s question, a reflection pass — each recorded as the gate’s refusal on the scope a charge would be refused by, the company’s before the seat’s. A round it turned away is counted in used like any other — the vendor billed it — so a window that refused one reads past its limit by that round, and every later charge is refused against that figure. The stamp is kept in the shared counter beside the spend, so every node reports the same one, and it clears on the scope’s next admitted charge or when the window turns over. A window whose ceiling was raised after a refusal keeps its refused_at until then, under a state of ok or near: it has room again, and the stamp is history.

Every node publishes a budget_meters snapshot of the counters as soon as its seat host is running and every 15 seconds (engine.BudgetReportInterval) after that — and at once when a budget window first refuses a charge, so the moment a company or a seat stops reaches every open screen with the refusal rather than up to a tick later; a repeat refusal in the same window waits for the tick — and the projection folds each one in as it arrives. Until the first one lands, budget is null — nobody has read the counter — which is a different fact from a report whose org.windows is [], “nothing is capped”. A company with no ceiling anywhere publishes exactly that: an empty list and no seats, without reading the counter. A node whose read of the counter fails publishes nothing that interval, so the meter keeps the last reading it had rather than drawing zeroes.

  • meter_id identifies the node incarnation whose report is held. Every node reads the same counter, so reports under different ids describe the same figures read at different moments. A report is a complete snapshot, so a consumer replaces what it holds rather than merging or taking a maximum: a window turning over has to be able to lower the figure, and a ceiling removed has to take its window’s bar with it.
  • seq is monotonic within a meter_id. The feed it arrives on is best-effort: an ephemeral broadcast subscription that takes no acks, starts at the stream’s tail on every (re)connect, and lets a slow consumer miss frames rather than hold them. So a report at or below the held seq from the same meter is dropped, a report from another meter that was read earlier than the held one is dropped, and a gap is closed by the next report rather than replayed.
  • {} means no report has arrived yet. Per-agent, budget: null means the same, or that the seat has no per-agent ceiling at all: the engine meters a seat only when its token_budget caps a window.

It is deliberately never persisted: a report is a reading of a counter that moves every round, so a copy replayed from history would show figures the counter has since left behind as the current ones.

Each agent’s live_call is null between turns, or { turn_id, phase, iteration, model, prompt, prompt_messages, response, tool_executions, round_narration, partial_round, round_num, rounds_used, rounds, max_rounds, round_ceiling, round_started_at, running_call, steers, cache_read_tokens, cache_write_tokens, work_item, node, in_progress, versions } while an LLM call is under way. The fields are these:

  • prompt_messages is [{role, content, sections}], the conversation the phase opened with. sections is that message’s outline as the prompt’s builder recorded it — [{key, title, bytes, headed?}], tiling content exactly, with headed: true on a part that opens on its own heading line whose text is title (absent on a part the builder named, which may still open on a heading of its content’s) — and is absent on a message no builder outlined or whose text is not valid UTF-8; a finished phase record carries the same maps as system_sections and user_sections. See a prompt carries its outline for the keys and the rules a reader checks before slicing by one.
  • round_narration is [{round, reasoning, content, declined}], what the model said in each round so far; declined: true marks a round that answered in prose where the phase had to end in a call, and is absent otherwise.
  • rounds_used is the round the phase is on, counted from one: the rounds that have come back, and the one in flight from the frame the engine publishes as that round’s provider call is made. A finished phase record carries the same count under the same name.
  • rounds is each round’s own timing and tokens, {round, started_at, duration_ms, model, input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, tool_calls}. This used to be the count, under the name the phase record uses for the list.
  • max_rounds is the round cap currently granted, which an extension can raise mid-phase.
  • round_ceiling is the highest value any extension may raise max_rounds to.
  • round_started_at is when the latest round made its provider call. It is later than every entry in rounds while that call is out, and equal to the last entry’s start while the round’s tools run — only the first is a model call in flight.
  • running_call is {round, name, arguments, started_at}, the tool call running right now. It is absent between calls, and it is never carried forward from an earlier frame. A frame that stops naming a call means the call returned.
  • steers is every person’s note the phase has read so far, {round, note_id} — the round whose provider call first saw it. What the note said and who sent it are on the turn’s agent_turn_steered rows. See steering a running turn.
  • node is the node running the call.
  • versions is {prompt, response, narration, executions, rounds}: the version of the copy the call holds of each of its heavy fields — prompt numbers prompt and prompt_messages, response the response, narration the round_narration, executions the tool_executions and rounds the rounds — a number the node’s projection never hands out twice, taken whenever the field is written: every field of a call when the call begins, and after that each field a frame changes. So on the node a tab’s socket reads, a version only grows, for every call under every key — and a call built again under its own key (turn_id, phase, iteration), as a suspended Execute phase is when its resumed rounds stream after the checkpoint cleared it, is newer in every field than the copy a tab held from before. A tab that reconnects is sent a snapshot, which replaces every copy it holds. Every surface that carries a call carries it.

Beside live_call, each seat carries live_call_seq: the node projection’s own sequence — the one versions are taken from — read when the seat’s call last changed: began, took a frame, froze as failed, or was cleared to null. It is on every surface that carries the call, a null call included, and it orders the call slot itself where a field’s version cannot: across a clear, which leaves no version behind, and across two calls, whose fields are never compared (see What the projection carries, and what the wire sends). Like a version it is a fact about one node’s projection, and a snapshot resets what a tab holds.

Beside live_call, each seat carries turn, last_turn and paused: the turn the seat is on, with its stage of context, phase or parked, the newest turn it ended, and who paused the seat, when and why (null while nobody has). See Agent States. A call whose phase failed keeps in_progress: false plus failed: true and an error object, so the dashboard renders the failure instead of an answer that never arrives — and no partial_round, because the phase is over and nothing is still arriving.

What the projection carries, and what the wire sends

Section titled “What the projection carries, and what the wire sends”

A running phase moves its seat’s row on every progress frame — five times a second while a round streams — so both the frames into the projection and the pushes out of it are trimmed, and each is reassembled on the other side:

  • Into the projection, the prompt and prompt_messages are sent once, on the phase’s opening frame, and carried forward from there (a 30 KB system prompt does not change mid-phase, and past queue.MaxPayloadBytes the publish is refused outright and the live row simply stops); a tool result and the joined response are sent tail-bounded on a live frame, because they are what a reader is watching the end of. The durable agent_phase_completed record keeps every one of them verbatim, which is what a reader opens the finished card for.
  • Out to a tab, an agents push carries a call’s HEAVY fields — prompt and prompt_messages, response, round_narration, tool_executions and rounds — only when their version (versions) moved since the last push for the same call, and leaves them out otherwise: the key is ABSENT, never empty, and the tab keeps the copy it holds. They move once a round or once a tool call; between those moments a push is the round being written (partial_round), the call in flight and the light fields. Measured over one long executor phase (thirty rounds, 1 311 frames), each open tab was sent 185 MiB with every push whole, and is sent 5.9 MiB.
  • A dropped push is noticed by the next one. The socket drops a slow tab’s oldest envelope rather than stalling every other tab, which is why the pushes used to be whole: a frame that carried only what changed, dropped, would leave a tab wrong with nothing to say so. Every push names every version, so the push after a dropped one names a version past the copy the tab holds, and the tab asks for the call whole (live_call) rather than drawing the older copy as current. A field a push carries at a version older than the one a tab holds — a push overtaken by a snapshot or an answer — is not taken. The handshake snapshot, GET /agents, the agent answer and the live_call answer always carry every field.
  • An answer is ordered against the pushes that overtake it. The live_call answer is computed on its own and can reach the tab after a push the node generated later, so it carries the seat’s live_call_seq, and the tab compares it with the newest it has applied for the seat. An answer at or past it is merged like a push carrying everything. One BEHIND it describes the slot as it was before those pushes: when they cleared the call or began another, it is dropped — so a stopped seat never shows the running call read a moment before its completion — and when they are the same call, it still brings each heavy field it holds at a newer version than the tab’s, while the light fields stay the pushes’. And if a push found the call behind while the answer was out, the tab asks again once that answer lands behind it, rather than waiting for a next push a frozen call never sends. An agents push whose live_call_seq is behind an answer already applied leaves the call as the answer had it, and its other fields still apply.
  • A version is never handed out twice, so “moved” and “newer” are one question. Counted per call, a call cleared and built again under the same key — a suspended Execute phase, whose completion checkpoint clears its call and whose resumed rounds stream under the same turn, phase and iteration — numbered its fields from one again, and a tab that had missed the clearing push and the first one after it read every later push as older than the copy it held from before the suspension, and kept that copy.

Upgrades to a WebSocket. All frames are JSON envelopes of the form {"kind": "...", "data": ..., "ts": "<iso8601>"}.

The socket is the dashboard’s channel for state, not for everything. The projection arrives here as pushes, and every question the query registry answers is asked and answered here too. What the socket does not carry is REST: every write — through /operator/act as the person the token is bound to, or through the credential-scoped /config, /secrets, /setup and /backup, because a write has to be able to say whether it happened and a frame into a dropped socket has no answer — and the few guarded reads no query answers (GET /secrets, GET /config/references, GET /setup/integrations and its passes), which the dashboard reads through one loader that re-reads on a new token and honours a Retry-After. The dashboard survives losing the socket by polling /stream/snapshot every five seconds, which is exactly the kind of failure that is easy to miss: nothing looks broken, the page is simply always a few seconds stale. internal/e2e closes that gap by replaying the frames a real server produced through the dashboard’s own protocol module (static/dashboard/protocol.js, the same source its bundle contains), so both halves of the protocol are checked against each other rather than each against its own idea of the other.

A kind a client does not know is ignored, never an error. A fleet part way through an upgrade has a node pushing kinds an older bundle was built before, so the dashboard drops such a frame — and counts it, because the same fall-through is what a kind this build’s engine sends and its own client forgot looks like, and the e2e replay fails on a non-zero count.

Server → client kinds:

kindWhendata
snapshotFirst envelope after the upgrade succeeds, and again on reconnect.Same payload as GET /stream/snapshot — agents carry their in-flight live_call and its live_call_seq, so a reconnect re-renders the live row and starts the ordering of the call slot afresh.
eventEvery engine event published to crewlet.events.>.{ id, type, timestamp, source, actor, summary, category, trace_id, span_id, parent_span_id, topic, agent_id?, channel_id?, payload } — agent_id the seat the event concerns and channel_id the agent-to-agent channel it belongs to, each read by the rule that fills the store’s own column, so the event log narrows its live rows to a seat or a channel exactly as the store narrows its pages — the same shape as a /events row, plus the full event payload (from which the snapshot feed’s failed flag is derived). agent_phase_completed events carry the system prompt, response, and tool calls, so LLM invocations stream live; agent_turn_progress events (per tool-call round, tagged with turn_id / phase / iteration) stream the in-flight call before its phase record exists.
agentsAfter an event moved one or more agents — or a read moved their state: a run record reconcile, or the seat-lease read the five-second tick makes.The changed agents’ overlays, each with its role, its activity and its stopped_reason — the result of applying the change, so a client merges them rather than running its own state machine over the raw stream. A live_call carries its heavy fields only when their versions moved since the last push for the same call, and a client keeps the copy it holds of one left out; every row carries the seat’s live_call_seq, a null call included, which orders the call slot against a live_call answer (see What the projection carries, and what the wire sends).
seatsAfter a config revision changed the roster.The COMPLETE seat list, replacing what the client holds. Distinct from agents on purpose: that one is a per-role merge, and a merge cannot express the deletion of a role a revision removed.
sandboxesAfter a detached sandbox run started, asked a question, finished or was lost, and after a reconcile against the durable run record changed the set.The full in-flight sandbox list.
tokensOn the shared 5-second tick, when a spend record arrived or one aged out of the live window since the last one. The fold runs on the tick rather than on the publish, so a busy company costs one aggregation every five seconds rather than one per phase.The spend rollup, same shape as GET /tokens/breakdown.
budgetAfter a node’s token meter report is applied (every node reports at start, every 15 seconds and at once when a budget window first refuses a charge, a company that caps nothing included).{ meter_id, seq, timezone, org: { windows: [...] } }, the org-wide half: one entry per capped calendar window, each with its span, spend, ceiling, refusal stamp and state. Per-seat figures ride on each agent’s overlay in the agents push. See the live token meter.
org / tools / schedulesAfter a config revision is activated.The new org tree / tool surface / schedule list, so open tabs stop showing seats that no longer exist.
healthPulsed every 5s by a single shared tick (one timer for all clients, not one per connection).The whole health envelope, exactly what GET /health answers. There is no query for it.
resultReply to a client query that succeeded.{ id, what, data } — id echoes the request’s.
errorReply to a client query that could not be answered.{ id, what, error, retry_after_seconds?, detail? } where error is a code: unknown_query, unauthorized, not_found, bad_params, unavailable, or query_failed for every other failure (the reason goes to the log, never to the socket). unknown_query covers a surface this process does not have: a question whose source is not wired here is never registered, so it is unknown rather than empty and never carries a Retry-After, because waiting cannot give this node a store it was not configured with. Its REST twin is 404. unavailable is not query_failed: it says this node understood the question and cannot answer it yet — a projection still catching up after a restart or a fresh join, a coordination store it could not reach, or no copy of the estate to answer from (the tracker and the knowledge base are read through the estate router, so a copy out of service or a data node restarting is a holder to ask again rather than a fault) — so a client says “ask again in a moment” rather than reporting a fault. It is the one code that carries retry_after_seconds: how long to wait, from the same helper as its REST twin’s 503 Retry-After header, so the two transports never disagree — the refusal’s own derived hint where it has one (how far behind this node is, over how fast it is draining), rounded and never below a second, and the health tick’s five seconds otherwise. The dashboard asks again after exactly that wait and keeps no wait of its own, so this hint — and the Retry-After on a REST refusal — is the only retry clock it has. bad_params is not query_failed either, in the opposite direction: the node understood the question and refused it — a parameter missing, malformed, or outside the set the field accepts — so the fault is the caller’s and retrying sends the same bad request again. Its REST twin is 400. It is the one code that carries detail: the refusal’s own sentence, which names the parameter to change and what it accepts (days is 91, and a spend window is 1 to 90 company days — ask for at most 90), written for the person who will read it: no class name, the engine’s or a finer one (tokens.ErrWindowLength, which a Go caller tests with errors.Is), and no echo of the query’s name — the REST 400 body carries the same detail beside its error. Every other code’s text stays in the node’s log, since a failure’s own text can carry a path.
pongReply to a client ping.null

Client → server kinds:

kindPurpose
pingKeepalive; server replies with pong.
queryRequest one thing, answered with exactly one result or error frame. { kind, id, what, params, token? } — id is any client-chosen value echoed back on the reply, and token carries the operator bearer token that the config-family queries require (validated with the same constant-time comparison the /config middleware performs). Queries run concurrently with each other and with the push stream, so one database read cannot stall a tab’s live rows — at most four at a time per socket, which is the size of the node’s reader pool: one tab may use every reader connection and no more, and a fifth query waits on its own socket rather than in the pool the engine’s own reads share.

Each query (what) is answered by the same function the matching REST route calls, so the two surfaces cannot diverge:

whatparamsAnswers with
agent{id}GET /agents/{id} — config + live state + llm_history
live_call{role} (or {id}, the seat’s handle)GET /query/live_call?role=…. {role, live_call, live_call_seq}: the seat’s call in flight with every field and its versions, or live_call: null while it has none (a seat the projection has never seen included, at live_call_seq 0), and the seat’s live_call_seq beside it either way, which the tab orders the answer by against the pushes that may overtake it. From the projection alone, so it costs what a push does — the REPAIR a tab makes when an agents push names a version of a heavy field newer than the copy it holds: the push that carried the field was dropped (see What the projection carries, and what the wire sends)
agent_memory{id, limit}GET /agents/{id}/memory. ANSWERED BY THE NODE HOLDING THE SEAT, which it names (held_by, or none with an empty answer for a seat no node holds; unavailable while the holder is silent or still taking the seat) — every node keeps a copy of a seat’s memory and only the holder keeps it current. Four collections, each a page (limit, at most 50) with its counted total beside it: the diary (diary_total), the episodes (episodes_total), the synthesized skills (skills_total) and the COUNTERPARTY PROFILES (counterparties_total) — what this seat has learned about the colleagues it works with, both instants carried because last_updated_at moves on every interaction and last_corroborated_at only when the traits changed. Plus latest_reflection (the newest live diary entry, whatever the page) and onboarded_at. An episode’s ask and plan_summary are their OPENINGS, at most 600 bytes each — the figure a seat is shown of a past turn’s words before they are condensed (learning.EpisodeAccountBytes) — with each whole text’s size in ask_bytes and plan_summary_bytes. See the route
agent_episode{id, episode}GET /agents/{id}/memory/episodes/{episode}. ONE EPISODE WHOLE — what the turn was asked and what it did, complete — answered by the seat’s holder as agent_memory is (held_by); episode: null for one the seat no longer holds. See the route
memory_overview{}EVERY AGENT SEAT’S memory totals — diary_total, episodes_total, skills_total, last_reflection_at and the latest_reflection itself — each counted by the node holding the seat, gathered in ONE scatter rather than a read per seat, with held_by per row (none for a seat no node holds, nothing counted), an unavailable reason on a row whose holder did not answer, and the fleet coverage. Every agent in the chart, handle order, no cap. See the section
conversations{handle, conversation, limit}GET /agents/{id}/conversations. The seat’s own thread ledger — the engine’s only account of what a seat said on a surface it does not own, and what stops it replying twice in one thread. TWO SHAPES IN ONE ANSWER, because a screen asks two questions with one navigation: conversations is every thread this seat holds entries in, and naming one in conversation adds that thread’s turns as entries. Each turn’s reply and unsent carry the same artifact and WHICH ONE HOLDS IT is the whole record of whether anybody received it — a turn can end with real work done and no way to say so. The listing is a page (default 50, at most 200) with conversations_total beside it. ANSWERED BY THE SEAT’S HOLDER, as agent_memory is and for its reason: the ledger travels with a seat’s memory and only the holder’s copy is current (held_by). Same scope rule as work_my_work
event{id}GET /events/{id} — one event with its full payload
events{limit, type, source, category, trace_id, channel_id, seat, actor, agent, turn_id, work_key, work_item, suspended, failed, feed_only, since, until, before_id, before_time}GET /events. feed_only=true selects the rows the activity feed carries — every stored type but the ones it keeps out (auxiliary_spend), a trace’s agent siblings included — and is what every read merged with the live feed asks; true, false or absent, any other word a 400. failed is THREE-VALUED the same way — true, false or absent, any other word a 400 — and selects by the rule every row’s own failed is stamped by: a stored failed tag, or a type that is itself a failure (llm_unavailable, budget_exhausted, turn.guard_breach, sandbox_run_failed). It is the event log’s “Failures only”, applied by the engine so the axis and every page are one set. suspended is THREE-VALUED — true, false or absent for every row, any other word a 400 — and selects by whether a completion record PARKED its turn on a detached coding run: a turn that parks and resumes writes two agent_turn_completed records, the first marked suspended, so type=agent_turn_completed&suspended=false is the turns that ENDED, one record each, and is what a turns axis counts over. trace_id selects one trace as a FILTER — paged, windowed and combinable with every other filter, where GET /events/trace/{trace_id} is the whole trace oldest first. channel_id selects one agent-to-agent conversation’s events, by the channel id every A2A event carries (an index seek, migration 0034). seat is a seat’s handle and selects the events that seat published, resolved on the server to the id every node derives for it — so a seat since removed from the chart still names its history, and a role name, which a rename changes, is never the key; anything that is not a handle is a 400, and before a company configuration is applied it is 503 unavailable, since the id is derived from the company’s name. agent is the broader question — every event that names the seat as its actor, role, target, recipient or sender, plus every event sharing a trace with one — and takes the name those fields hold. turn_id selects ONE RUN of a turn; work_key selects every run of one unit of work — the attempts at a trigger that was redelivered; work_item selects every event on one work item — each turn’s start, its phases and completions, a coding run it launched — named by the item’s identity across trackers, <backend>:<id> (native:<task id>, jira:<issue id>), never by its key, which a move rewrites. A value that is not that shape (a key such as ENG-4, or either half missing) is a 400 rather than an empty page, because an empty answer reads as “nothing happened on this item”. The filter reads the work_item COLUMN (migration 0033), filled from the same payload field as the row’s work_item tag. Every row answers with its own work_key read off the work_key COLUMN (migration 0029) rather than out of its tags — the column is what the work_key filter reads, so a row and the filter that matched it never disagree about its unit of work
event_series{bucket, since, until, type, source, category, trace_id, channel_id, seat, actor, turn_id, work_key, work_item, suspended, failed, feed_only}GET /events/series. Every bar carries failed beside count — how many of its rows reported a failure — and the answer a failed total beside total (see the event log’s time axis). THE SAME ROWS WITH A TIME AXIS, which a page of rows has no dimension for: a burst at four in the morning and a steady trickle across a week are the same hundred rows in the same column. A second question rather than a flag on the first, because the two answers have different shapes and one route returning either would make every caller branch on what came back — the same split tokens and token_series carry. Both halves compile their filters through ONE predicate in the store, so a bar can never claim rows the listing beside it would not show
trace{trace_id}GET /events/trace/{trace_id}. Answers {trace_id, events, truncated}; truncated is true when the read stopped at the store’s per-trace cap (500) rather than at the end of the trace, which the caller must say — a trace shown short with no note reads as a complete causal chain that simply ends. It is counted, not inferred from the row count: a trace of exactly the cap holds every row it has, and len(rows) == cap would put a truncation warning on a complete one. Across a fleet the rows are every node’s oldest, merged, and they stop at the last row a node sent when that node holds more than it sent — its own read stopped at the cap, or its reply was cut to fit the transport — rather than show another node’s newer rows past the ones it did not send: such a view can hold fewer than 500 rows, and truncated says it is cut
turns{days, since, until, seat, model, work_key, work_item, failed, sort, before, limit}GET /turns. The window is on the turn’s START: days back from now (default 7, at most 30), OR since (inclusive) and until (exclusive) as RFC 3339 instants — either alone is a one-sided window — which is what a bar in the past needs, since days counts back from now; naming both forms, or a since not before until, is 400. A turn is in the window when it STARTED there: one that began before since and ran on into the window is not listed, rather than listed from the window’s edge with half its tokens. seat is a seat’s handle and narrows to that seat’s turns, resolved on the server exactly as on events (and refused the same way); there is no role-name filter, because a role name is changed by a rename while the history keeps the old one. ONE ROW PER RUN of a turn — a wake, a decision, its rounds and its reply — which is the view of a working company that did not exist anywhere. A turn that broke before reaching outside the engine is redelivered, so one TRIGGER is legitimately several rows; each carries the work_key they share and work_key= narrows to every attempt at one (see a turn’s two identities). A turn is what this engine DOES and every other surface is a projection of one: the spend rollup groups them, the seat page shows one seat’s, an item’s history links to the ones that touched it, and none of them is a list of them. The dashboard faked one by paging the raw event feed sixty-one times and folding in the browser — slow, capped at whatever the caller gave up on, and wrong at the page boundary, where a turn straddling two pages appeared twice. The aggregates are over PROMOTED COLUMNS (migration 0015) rather than payloads; only the duration, the summary and the work item come from the completion record’s own payload, read from the one row per turn that carries it. work_item is {backend, id, key, project} — the one item the turn was charged to (see which work a turn is on) — and ABSENT for a turn on nothing, including one still running, since a sole write names its item only at the end. work_item= narrows to the turns on one item, by the same <backend>:<id> identity /events takes (and the same 400 for anything else); it selects TURNS rather than rows, so a turn is listed whole — every phase and every segment folded — when any of its records names the item. There is no task_id on a row: the key it read was declared on the completion and never set, and task_id elsewhere means a delegated worker’s task or a schedule fire’s run, never a tracker item. A turn that launched a detached coding run PARKS: the segment that launched it publishes a completion with suspended: true, and the same turn completes again when the run is collected — so one turn can hold several completion records and “a completion exists” is not “the turn ended”. The NEWEST completion decides: complete is true when it is not a suspension, and parked when it is, so the two are never both true; a turn with neither is running or died mid-flight. A sandbox_run_failed also counts as an end. It is a coding run that was LOST, and nothing resumes the turn that was parked on it, so a list reading completions alone called that turn parked for good. The newest end still decides: a run lost while it was still launching is followed by its turn’s own completion. duration_ms is the turn’s OWN measurement — the SUM of every segment’s, so the wait for the coding run between them is not counted — which is not the span of its events either: the span covers the reflection pass that publishes afterwards. A turn’s tokens are its COST: its phases’ and its in-turn auxiliary_spend records’ (its turn-start context, its ledgers’ rewrites, its card), never the reflection after it — whose records carry the turn’s id, so the Turn screen can draw them, and which is the seat’s learning rather than what the work cost; sort=-tokens ranks by the same sum, and model= selects over the same rows: a turn is listed for a model one of its phases or in-turn calls ran on — exactly what its models names — and never for one only its reflection used. cache_read_tokens and cache_write_tokens are that cost’s prompt-cache counts (migrations 0032, 0040), a breakdown of input_tokens rather than an addition to total_tokens. failed is THREE-VALUED and absent means every turn, because folding it into false would hide every failing turn from an unparameterised list — and a turn counts as failed when ANY of its events carried a failure OR was a failure BY ITS TYPE (llm_unavailable, budget_exhausted, turn.guard_breach, sandbox_run_failed), which is the same rule /events applies to a row. The second half is what a turn the engine killed BETWEEN phases leaves behind — a refused charge, an exhausted chain, a breached guard — so reading the failure flag alone reported those as clean turns with no completion record, which is indistinguishable from a turn still running. The rows are newest first by where each turn is listed, and the higher turn_id first within one microsecond, so the order is total. The cursor is on the TURN — a keyset on any one event pages a turn twice — and it is OPAQUE: next is a token, and before takes exactly the next a page answered (anything else is 400 bad_params, an RFC 3339 instant included). It is opaque because the position is not a row’s own fields: it is where the page’s last turn is LISTED, which is not always its started_at (see Reading the fleet’s history), and that turn’s id, without which two turns starting at one microsecond either side of a page’s cut were one position and the one past the cut was on no page; and it can stand on an empty page, which has no row to build one from. Where /events and phases echo before_time/before_id, the position IS the last row’s key. next is present only while more turns lie past the page — null on the last one, which is how a reader paging a seat’s record knows the walk has ended rather than offering “older” onto an empty page. sort is -started (the default, newest first) or -tokens (the most total tokens first, the spend screen’s drill-down per turn) — anything else is 400 naming both. Read from EVERY node and merged, with a coverage — see Reading the fleet’s history
turn{turn_id}Every event of ONE RUN of a turn, oldest first, payloads included — each phase, the turn’s own completion, and the fallbacks and guard breaches that happened inside it. Not a slice of the trace: one trace can span several turns and one turn several traces. Answers {turn_id, events, truncated}; truncated is true when the read stopped at the store’s per-turn cap (500) rather than at the end of the turn. A cut answer is the turn’s opening and its ending, not its opening alone: a turn is read oldest first, so a head-only read would drop agent_turn_completed and turn_completed — the two records a reader takes the outcome, the duration and the plan summary from — and a turn cut at the cap would be indistinguishable from one that never finished. The last rows are recovered beside the first (up to 32 more, merged on the store’s own identity, (event_time, event_id), so the two reads cannot overlap into duplicates — the id alone is not unique, and a narrower key would drop a row the two reads legitimately both carry and then report a gap over a page holding the whole turn), so what truncated names is a gap in the middle — and it is counted, not inferred from the row count, because a turn between the cap and the cap plus thirty-two ends up whole on the page and must not carry a truncation warning. Across a fleet the opening stops, as trace’s rows do, at the last opening row of a node holding rows it sent in neither half, so it never shows a row past a gap it does not report. It also answers work_key and attempts: the unit of work this run was an attempt at, and every run of it the store holds, OLDEST FIRST — over the SAME thirty-day horizon the events above come from, not the turns list’s own default week, and read at the same instant they were, so a turn between eight and thirty days old, or one whose rows sit at the very edge of the horizon, names its attempts rather than reporting none while displaying one — so the screen a deep link lands on can say “attempt 2 of 2” and link the other, rather than leaving a reader to conclude the company did the work twice. One element is the ordinary case; an empty work_key means the trigger had none to collapse on, and attempts is then empty too. And nodes: the nodes whose own store holds any of this turn’s events — where it RAN, since each node’s store holds only what that node published — sorted, and an empty list rather than an absent key where no node this read could reach holds any
phases{seat, limit, before_time, before_id}The company’s agent_phase_completed records, newest first, payloads included, keyset-paged. seat is a seat’s handle and narrows to that one seat’s phases, resolved on the server exactly as on events (and refused the same way) — never a role name, which two unit seats stamped from one template share, so a role filter answered one seat’s phases with every such seat’s. events?type=agent_phase_completed is not a substitute: the event listing deliberately never selects the payload, and a phase record without one has no prompts, no response, no tool calls and no decision. next is present only while more records lie past the page — each node’s read asks one row past it rather than guessing from a page that filled, so a history exactly a page long ends without offering “older” onto nothing
tokens{days, since, until, previous, seat}GET /tokens/breakdown. With no parameters, the live 24-hour window from the projection; with any, whole company days from the replicated usage domain — every node’s, up to 90 days, within the 181-day horizon (see Token Spend Breakdown)
token_series{days, since, until, previous, seat, group, bucket, groups}GET /tokens/series. THE SAME SPEND WITH A TIME AXIS, which the breakdown has no dimension for: every one of its rows is a sum over the whole window, so a runaway loop, a spike and a quiet weekend are the same number. A second question rather than a flag on the first, because the two answers have different shapes and one route returning either would make every caller branch on what came back. Bucketed by the ENGINE, by company day or ISO week from the usage domain. An unknown group or bucket is refused naming what is accepted, never defaulted: a chart legended by one dimension over another’s bands is worse than an error
seat_activity{seat?, days?, previous?}GET /agents/activity. Every seat’s TURNS over the days (1 to 90, default 7) company days ending today, summed across every node from the replicated usage domain — so the answer is the same on whichever node is asked and still counts a node that has left. {since, until, days, previous_since?, previous_until?, seats, quantile_resolution}; each seat is {handle, role, agent_id, in_chart, turns, failed, reviewed, first_pass, first_pass_pct?, sent_back, p50_ms?, p90_ms?, tokens, per_day, last_turn_at?, previous?}. first_pass_pct is first_pass over REVIEWED turns (0–100) and is ABSENT when none was reviewed — a 0% for a seat nobody reviewed would be a verdict nobody gave. sent_back counts reviews that sent work back. p50_ms and p90_ms are read from the merged turn-duration histogram and are within quantile_resolution (0.06) of the true value; absent when no turn ended. per_day is every day of the window, oldest first, a quiet day included as zeros. previous (with previous=true) is the seat’s {turns, failed, reviewed, first_pass, sent_back, tokens, per_day} over the same number of days before, per_day being every one of those days, oldest first — so a profile draws the fortnight its week-on-week figure is made over from this one answer. Every AGENT seat of the current chart has a row, a quiet one with zeros (“took no turns” is a measurement); a seat that has left the chart appears with in_chart: false while its days are in the window; a human seat has none. seat= narrows to one handle, and a handle with no rows answers one row of zeros rather than none; a human seat’s handle is refused (bad_params, naming the person), because the engine runs no turns for a person and a zero row would say one took none. Ordered by handle
schedule_runs{scope_type, scope_id, name, limit}GET /schedules/{scope_type}/{scope_id}/{name}/runs. ONE schedule’s dispatch history, newest first, fifty to a page. schedules.recent_runs is the COMPANY’s fifty most recent fires across every schedule, so twenty hourly ones fill it in two and a half hours — “did the standup fire this week” was unanswerable while every row of the answer sat in the table. The identity is all THREE parts and each is required: two units may each declare a standup, and a role and a unit may both, so a name alone merges two teams’ histories. truncated says the page filled, because a full page is otherwise indistinguishable from a schedule that has fired exactly that many times
schedules{}GET /schedules
fleet_broker{}GET /fleet/broker: what every live node advertises about its broker, the metadata group as a member reports it, and where the two disagree. Operator-only
access{}GET /access: the token labels, the people and the posture (auth.company_writers included) the Settings › People & access screen draws. Operator-only. See below
credential_pool{}GET /credential-pool: every model’s keys and their cooldowns, the Settings › Models & keys screen. Operator-only. See below
backups{}GET /backups: each owner’s newest backup and the backup history, the Settings › Backups & retention screen. Operator-only. See below
mcp_servers_status{}GET /mcp-servers: each MCP server’s condition and its per-node counts, the Settings › Tools & MCP screen’s Servers section. Operator-only. See below
fleet{}GET /fleet: leases move with no event to push, so Settings › Nodes polls this rather than waiting for one. Operator-only, like the rest of Settings. A lease table that could not be read answers unavailable, which is a blip to ask again about rather than a fault (the REST twin answers 503 with a Retry-After)
sandbox_runs{audience?}GET /sandbox-runs: unknown_query on a company with no sandbox configured, and unavailable when the fleet’s run record could not be read. Each run carries work_item ({backend, id, key, project}, or null), the item the launching turn was charged to, and launch_id, the job the row holds now — what sandbox_tail is asked by
sandbox_tail{turn_id, launch_id, epoch?, after?, digest?}GET /sandbox-runs/{turn_id}/tail?launch_id=…[&epoch=…&after=…&digest=…]. What ONE running coding job has said so far, read from its box by the node that owns the run (see Watching a run live). Both ids are required (bad_params otherwise): a turn can launch more than one job, and the launch id is the one sandbox_run_started and the run’s phase record carry. Answers {outcome, turn_id, launch_id, node?, status?, output?}: outcome is tail with output; launching for a job whose box is still being made (not stopped — ask again); not_running with the record’s status (awaiting_clarification, resumed, reseed, replaced for a job a later launch replaced, or absent where no record is left); owner_silent naming the owning node that did not answer inside the 2 s fleet read budget — a draining owner, or one whose heartbeat lapsed, included (no node for a run nobody holds right now); or box_paused for a running record whose box is paused where its backend cannot read it without waking it (a peek never wakes a box). The asker says what it holds — epoch, after (the byte offset it holds through: a whole number of zero or more, bad_params otherwise — a fraction is refused rather than rounded) and digest, all three absent for one holding nothing yet — and output is {text, source: transcript|stderr|none, cut, as_of, finished, epoch, start, end, digest, reset, front, held}, redacted — epoch, start, end and digest on every answer, 0 included (a run that has settled nothing yet is answered at end: 0 and followed from start: 0): text is what follows after (start = after), or on reset: true the last 256 KiB in whole lines, which replaces what the asker held (start says where it begins; cut that earlier output exists). Send end and digest back as the next after and digest. front says the owner’s reading began after the job’s own start; held counts bytes written but not shown until their redaction is settled. A record that could not be read, or a box the owner could not read, is an error carrying the reason. There is no event and no row: the dashboard asks it every 3 s for each running job it is showing — an open span on a turn’s timeline, a parked turn’s coding-run card, and the run’s own page — and nothing outside the dashboard asks
budgets{}GET /budgets
a2a_channels{}The fleet’s agent-to-agent authorization record: who asked whom, how many messages crossed, and when. available: false when this node could not reach the coordination store — which is not the same as no channels having been opened
knowledge{q, mode?}The company’s own knowledge search, run live through the same knowledge.Searcher seam a seat’s own search_knowledge tool uses. Searched as the ORG with no seat, so it applies the engine’s own account and nothing more — searching as a named seat would let a dashboard reader read, through that seat’s credential, material their own account may not have. Registered whenever a company is active, NOT only when a searcher exists — “this company has no knowledge backend” is a fact the company establishes on its own, and it is a far more useful answer than an unknown query. available: false covers all four of no company, no backend, a backend wired with no org-wide read scope, and a node whose index is still building. reason (no_company / no_backend / no_scope / building, empty when the search ran) is the value to branch on and note is the prose for a person — a screen picking which remedy to offer must not string-match the note, nor infer the state from an empty backend, which means “no backend” and “no company” alike. The no_scope note names knowledge.scope, because an operator whose integration is correct must not be sent to re-check it. A search that RAN and fell short says so in its outcome rather than its note (which is empty whenever available is true): served_mode empty with coverage.complete false is a backend that did not answer, and a partial coverage is part of the corpus unsearched. An EMPTY q is the seam’s PROBE: nothing runs and nothing is read, no hits are returned, but modes and — for the mode asked — the degraded its CONFIGURATION decides are answered, so a screen offers the modes honestly before anybody types (the dashboard asks the probe for semantic, the mode whose reason names what is missing). A transient reason (embedding_failed) is a property of a search and is never predicted. mode is hybrid (the default), keyword or semantic — the screen’s label for the last is “Meaning”, and any other value is refused rather than run as the default. A q past 400 bytes (trimmed) is refused bad_params, naming its size and the limit, rather than cut — the bound every search surface shares (Search). Every answer carries what the search actually did: mode (resolved), served_mode (the ranking the hits came from; "" when nothing ran), modes (what this backend can serve as asked — [keyword] with no embeddings provider and on Confluence), degraded (no_embeddings / embedding_failed / semantic_partial / unsupported, empty when it served what was asked) and coverage{nodes:[{id,answered,error}], complete, buckets_missing} — the fleet the scan was divided across, which is how a partial answer is told from a short corpus. A hybrid search with nothing to rank by meaning serves its keyword half; a semantic one serves nothing and says why. See Search
colleague{q}A NAME TO A SEAT, through the same four tiers an agent’s lookup_colleague and a2a_ask resolve through (internal/agent/colleague): exact handle, chat id and role name, then case and separators folded, then part of a name, then a close spelling — each tier answering only when every tier above it found nothing, so a name that is exactly somebody’s handle is never diluted by near misses. Answers {match, candidates[{handle, why}]}: match is the one seat the text names when EXACTLY ONE does and null otherwise — including when several do, which is a list for a person to choose from and never a pick — and candidates is every seat it could name, best tier first (the match included; empty when nothing matched, which reads differently from ambiguity). why is the tier in a person’s words (“handle matches exactly”, “part of the name matches”). The chat-id tier needs a credential: a seat’s contact identities are operator-gated configuration, so an anonymous caller resolves over handles and names alone. q is at most 200 bytes — a name, not the sentence around it. The command palette’s assign and ask pickers read it. Registered whenever a company is active
integrations{}GET /integrations
work_items{container, status, status_group, assignee, reporter, watcher, collaborator, tag, type, priority, parent, root, q, key, references, linked_page, removed, blocked, blocking, has_dependencies, has_open_asks, flag, asked_of, asked_by, subtasks, show_closed, closed_since, f.<slug>, view, preset, viewer, group_by, group_by2, group, subgroup, group_limit, totals, sort, fields, around, cursor, limit, …}GET /work. container is the scope — workspace, or project:ENG (a bare ENG works too, and the key is upper-cased because the column is); any other <kind>: — unit:eng, person:ana — is REFUSED, because a task lives in the workspace or a project and a team’s or a person’s work is unit= or assignee= (read as a project key, unit:eng answered an empty board) — and an ABSENT container is neither: the engine refuses to default it, because an omitted key would otherwise be the most expensive query in the system. Every list key is comma-separated, because a socket frame’s JSON object cannot carry a repeated key and a filter only one transport can express is exactly the divergence this channel exists to prevent; status also takes ! negation. There is no open flag — open and closed are STATUS GROUPS (not_started, active, done, closed), which is the level every rule in the tracker is written at. f.<slug>=<value> filters on a custom field — resolved against the company’s catalogue by slug, id or label, and compared on the column its DECLARED TYPE says, so f.effort=gt:9 is a numeric comparison and not a lexical one; the seventeen operators are eq, ne, lt, lte, gt, gte, contains, startswith, in, range, any, all, not_any, not_all, me, null and not_null — and which of them a field admits is a property of its TYPE, so eq on a labels field is REFUSED naming any, all, not_any and not_all rather than compiling to a clause that matches nothing and reads as “no task has this label”. null and not_null are on every type, because “is this set” is a question about the ROW. A bare value is the type’s NATURAL comparison — any on a set, because naming a value is not claiming the set IS it, and eq everywhere else. A set operator takes a comma-separated list (any:api,ui, at most 16) and range takes both ends (range:3..8), because a range with one end is gte or lte. A value whose text begins <scheme>:// is a VALUE rather than an operator call, so a url field can be filtered by what it holds — anything else before a colon is carried through as an operator, so a typo is refused naming the set rather than silently answered. f.<slug>=me is resolved to the reader by the SURFACE before the query is parsed, which is what makes one saved view mean whoever opens it. A ref nothing resolves is REFUSED naming it. unit= and routing_unit= name a team by its id or by its NAME, in any case, and each matches the work filed under either spelling: a unit’s durable key is its id where it has one and its name where it does not, so a company that adds an id holds both across its own history and a filter comparing against one string would answer with half the team’s work. A team the chart does not have matches nothing rather than refusing, because a task’s filed unit is a record of what was true and may name a team since dissolved. q= is a FIND rather than a search — a substring of a key (from the front) or a title (anywhere), which is what finds the item somebody half remembers; there is no mode, because this grammar has no ranker and ranked search over what the work says is work_search’s. A q longer than the longest title (256 bytes) is refused naming the bound, because it can match nothing and an empty board reads as “there is no such item”. key=ENG-1,ENG-7 narrows to keys a caller already holds — upper-cased, like container= and references=, because a key is what somebody pasted and the column it is compared against is minted upper-case — and removed=true is the TRASH — the only way to list what a removal hid, which is what a restore is a gesture about. references=ENG-12 lists the tasks whose DESCRIPTION mentions that key (a task’s former key resolves like its current one; a comment’s mentions are not references), and linked_page=<page id> the tasks that name that knowledge-base page through linked_pages. asked_of= is whose answer an open ask is waiting on and asked_by= whose question it is — the work a person is WAITING on — and asked_by matches every name the asker writes under: a founder’s asks put through their own assistant are authored by the token and the ones they put from the dashboard by their seat, so for the viewer’s own handle it also matches the token bound to them, and “asked by me” is the whole list rather than whichever half they typed. A parameter this grammar does not read is REFUSED naming it, never ignored: a filter nobody parsed is a board showing more than the person asked for, silently. An unknown status or group is refused naming the closed set rather than matching nothing. A custom field VALUE is checked against its own declaration at the write and refused naming the rule — never rounded or coerced to fit; see the coercion table in the work tracker guide. flag= is the ATTENTION queue and its values OR: cycle, too_deep, inconsistent_project and key_collision are facts about a task’s own row, and one_sided and one_sided_final are about a DEPENDENCY of it — an authored waiting_on whose blocker does not list it, and one whose mirror was refused permanently (the blocker is gone, was removed, or is full). The first is what the tracker duty repairs 30 seconds on; the second is what a person resolves. They OR because an attention queue asks “is anything wrong with this”, and a conjunction over six flags answers nothing on every company. totals=<column>:<op> adds aggregates over the WHOLE matched set rather than the page — a number that changed as somebody scrolled would be the one thing a header must not do. The seven ops are sum, avg, min, max, count, median and p90. median and p90 are ORDER STATISTICS by nearest rank — the value at position ⌈p·n⌉ of the set’s values in ascending order — so each answers a value some task actually holds: the median of four values is the second, never an average of two. The columns are the summable ones (points, estimate_min, the spend_* family — spend_workers and spend_sent_back included — reassignments, reopens, depth), the date columns for the four order statistics min/max/median/p90 only (a sum of dates is a number of microseconds nobody meant), tasks:count, and f.<slug> for a declared number or date field. A total with nothing to add up is ABSENT rather than zero: “nothing is estimated” and “everything is estimated at nothing” are different facts. subtasks= is how a tree is filtered: collapsed (the default) and expanded filter ROOT tasks and let their subtrees ride along unfiltered — so a todo root brings its done subtask — while separate filters every task on its own. The first two answer the same SET and differ only in how a caller renders it. Asking for a subtree with parent= or root= turns the mode off, because those are questions about subtasks and filtering their roots would answer the parent’s siblings. any=[{…},{…}] is one level of disjunction, ANDed with the top-level keys: a branch is a PREDICATE, so it may not carry the keys that decide the answer’s own shape (removed, archived, show_closed, closed_since, subtasks, fields, around) or how fresh it must be (read_level, max_lag_seconds, max_lag_seq, min_position) — those are the same decision at every branch or they are incoherent, and a branch that carried one would narrow what was asked for at the top level rather than widening it. An empty branch is refused, because it matches every task and makes the others decoration. view=<id> and preset=<name> are loaded FIRST and every explicit key overrides them — a saved view is a set of defaults rather than a lock, so somebody who opens a board and picks another assignee gets the view with that one key changed. A view beats a preset (somebody saved it) and what was typed beats both. The five presets are my_queue, priorities, triage, blocked and overdue. my_queue is what can I pick up: a DISJUNCTION of the work the viewer holds and the work in their OWN project nobody holds, open and unblocked, most important first — both arms matter, because written as “assigned to me” alone a seat with an empty queue reads the company as having nothing for it while its project’s unclaimed backlog sits there, and the second arm is scoped to their project because unscoped it offers every unassigned task in the company. priorities is the viewer’s own ordered list, open tasks only, IN THE ORDER somebody arranged it — that order is the answer, so nothing sorts over it, and a finished task drops out of the answer without the list being rewritten. triage is the unassigned open work, which with one fixed status set is the honest definition of “needs somebody to decide”. my_queue and priorities both need viewer= and are refused without one, because a list with nobody’s name on it is everybody’s. A view= nothing resolves is REFUSED, never answered as the whole board. group_by= turns the answer into a BOARD: groups replaces rows — returning both would be the same rows twice — and each column carries its own count over the whole set beside a bounded slice of its rows (group_limit, default 20, max 100). The axes are status, status_group, assignee, priority, type, tag, project, unit, routing_unit, parent, due:day, due:week, start:week, due:bucket and f.<slug> for a custom field; anything else is REFUSED naming the key rather than answered ungrouped. unit and routing_unit group on the TEAM rather than on the stored string, for the reason unit= matches both spellings: a company that gives a team an id after work is already filed into it holds that team’s name on the older rows and its id on the newer ones, and grouping on the column drew one team as two columns — both headed with its name — with its counts split between them. The expression folds every spelling onto the unit’s key, and group= is folded the same way, so group=eng and group=Engineering load the one column. A stored unit the chart no longer has keeps its own column under the literal its rows hold, since folding it into anything would invent a home for work whose team is gone. A grouped answer mints no cursor, because across a set of columns there is no single order to be after; group=<value> is how a board loads one column further, and it narrows the WHOLE query, so the hint and the totals describe that column too. group_by2= adds swimlanes inside each column and subgroup= names one — a swimlane board is bounded by its CELLS rather than by either axis alone, because the work it costs is the PRODUCT of the two, so asking for lanes lowers the column cap and subgroups_dropped says how many lanes a column has beyond it. A group_by= over the WHOLE COMPANY is refused when the query’s own narrowed predicate still matches more than 20 000 tasks: a board is drawn by sorting every one of them, and the refusal names the ceiling and what narrows it. It is a bounded COUNT rather than a check for the presence of a filter key, deliberately — status_group=not_started,active is a filter and narrows nothing, so a gate spelled “needs a narrowing filter” is one a caller clears in a single attempt without making the query any cheaper. Scoping to one project with container=project:<key> lifts it, because there the input is an index range whose width is one project’s own size. An absent value is its own labelled column — “nobody is assigned” is a question a board answers rather than a row it hides — and group= with no value is how that column is loaded one further, because a key named and left empty asks for the rows with no value where an absent key asks for all of them. group_by=due:bucket is the one axis that is not a stored value: it is WHEN the work is due, read against the query’s own day — overdue (still open and past it), earlier (finished, and past it — work delivered late is not overdue and calling it so would be a false claim, so it is its own band and is empty unless show_closed brings finished work into the answer), today, this_week (through the end of the current Monday-anchored week, which is the week due=range:sow..eow means), later, and the empty key for a task with no due date. Its six headings read Overdue, Earlier, Today, This week, Later and No due date. The day it cuts on is the COMPANY’s midnight in the company’s own zone — the same instant the row’s overdue flag is derived from and the same one every due= filter compiles against — so the bands, the flag and the filters cannot disagree about a task. A band cut in the reader’s own browser could and did: for anybody whose local day differs from the company’s, a task landed under Earlier on a row the same answer flagged as due today and not overdue. A CLOSED axis — status, status_group, priority and the due:bucket bands — carries every column the query itself admits, the empty ones at count: 0 with rows: [], in the declared order: a board is the shape of the process rather than of this week’s rows, so an open-work board draws To do, In progress and In review whether or not anything is in them — and never Done, which the query excluded, because “nothing is done” said about a set that was never asked is a claim rather than an absence. The admission is the predicate’s own (status, status!, status_group, show_closed, the overdue alias, due= for the bands, and group= down to one column — where a key NAMED AND LEFT EMPTY admits the undated band on due:bucket and nothing at all on the three whose values are never empty). An OPEN axis — assignee, tag, type, a field — carries only the values present, since every seat as an empty column is a roster rather than a board, and the second axis is never filled. group_by=tag is the one axis where a task is on several columns at once; the answer sets groups_overlap so a reader knows the counts do not sum to total_hint, and groups_dropped says how many columns did not fit. sort= takes rank, updated, due, start, priority, created, title, estimate, points, spend_tokens, reopens, status_entered and removed, each reversible with a leading - — spend_tokens is the column’s own name, as a total’s is, since spend on a row is an object of five. The numeric filters estimate=, points= and spend_tokens= take lt:, lte:, gt:, gte:, range:a..b, null or not_null. closed_since=<date token> is the open work PLUS what finished (done or cancelled) at or after that date, resolved on the COMPANY’s clock through the same calendar as due= — closed_since=sow begins at the company’s Monday midnight, so a Done lane on it is this week’s work and empties itself when the week ends; it is the Board’s Recent scope, and naming it beside show_closed is refused because both say which finished work is in the answer. fields=tags,dependents_count,open_asks,spend puts a board card’s facts on each row, and only when asked: tags (sorted, [] when none), dependents_count (live tasks waiting on it through the authored waiting_on edges blocking= reads), open_asks (unanswered questions, by the same test has_open_asks= makes) and spend: {tokens, turns, workers, sent_back, reopens} — each present, 0 included, when asked, and absent otherwise, because the same row is what a seat reads through list_work_items, which never asks for them even through a saved view that does; an unknown name is refused. around=<task> (a key, a former key or an id) adds around: {position, prev, next, total_hint, total_capped?} — where that task sits in this answer’s DRAWING order over the whole answer rather than the page: a flat answer’s own sort, or a board read column by column (and lane by lane within a column) in the answer’s column order with each column’s rows in the row order, so the card after the last one in a column is the first in the next. prev/next are keys, null at the ends; on a label board the order counts cards and a task is placed at its first column; position is null past the 10 000 a count stops at. A task the answer does not hold answers around: null — never a refusal — and a saved view cannot carry around. A priorities= answer is paged in the list’s own order: the whole list (at most 32) is read each page and its cursor is a place in the list. An absent value sorts LAST in both directions: “soonest first” and “latest first” are both questions about values, and a task with no due date is the answer to neither — so sort=due puts the undated at the end rather than ahead of the one due tomorrow, and a cursor resumes in the same place the order put it. sort=f.<slug> orders by a custom field, LEFT-joined so a task that set no value still appears — a sort that also filtered would be two things the caller asked for once, and such a task sorts last by the same rule. The answer carries total_hint (capped — an exact total over an unbounded set turns a poll into a scan), next_cursor, totals, groups, and an echo of the view/preset it was expanded from — these answers travel detached from their requests, so a board restored from a URL can still say which saved view it is showing — plus the coverage half below
work_item{id}GET /work/{id} — key or id. Answers {task, comments, comments_cursor?, comment_seats?, reporter_seat?, history, links, parent?, fields, units, blocked, due_standing?, keys, reassignment_budget} plus the same coverage half. due_standing is where the due date stands on the COMPANY’s calendar — {days, overdue?}: whole calendar days from the company’s today to the due day (negative once it has passed, counted on the day labels so a DST day is still one day), and overdue for open work whose due instant is before the company’s midnight, the same predicate as a board row’s overdue — and is absent for a task with no due date, so a task page never re-derives either from the reader’s own clock. task.spend is the item’s running totals as its turns added them — read from the counters the turns wrote, never from the create’s copy in the document. parent names the item this one is filed under — {id, key, title, status}, read in the same transaction — and is absent for a top-level item and for a parent this node holds no row for. comments is the newest page of the thread and comments_cursor the page before it, which work_comments reads; comment_seats names the person behind each comment a token wrote, as there. reporter_seat is the same for the task itself: the seat the filing token was bound to when it wrote the create, read off the create’s own history row so it survives the create leaving the history page, and absent for a task a seat filed — task.reporter stays the credential, which is the record. reassignment_budget is how many times agents may hand the item on before the engine refuses the next hand-off — the limit task.reassignments counts against, served so no screen carries a figure of its own — and each history row carries reassignments, the item’s hand-off count AS THAT CHANGE LEFT IT (absent on a row the answering node holds no count for); an assignment made with a reason shows the reason as that row’s excerpt. units is the task’s two unit references RESOLVED against the org chart — {filed: {key, name, resolved}, routing: {…}}, the same shape a project row’s unit carries — because what the document holds is the unit’s KEY, which is its id on a company that gave its units one: a word chosen so that a rename moves nothing, and therefore a word nobody reads. The document’s own filed_unit / routing_unit are untouched beside it — they are the record, and key repeats exactly what they hold, so a filter link built from it reaches the same rows. resolved: false is the finding “this names a team the chart no longer has”, and it is ABSENT for a task filed into no team at all, because that is what its two empty strings already say. blocked is on the ANSWER rather than on task because it is DERIVED — an open dependency edge, computed in the same transaction as the task, so the badge here and the badge on the board row cannot disagree; links say what the relations are, not whether any blocker is still open. fields are the task’s CUSTOM fields resolved against the company’s catalogue — each carrying its declared name, type and whether its declaration was archived — because a stored choice is an option’s UUID and a panel rendering the raw value would print it under a heading. keys names the tasks this answer’s history POINTS AT, id to item key — the same map work_activity carries, resolved by the same walk, and present only when history was asked for. A delta names the other end of a relation by its ID because a key is a fact about another task’s row, so without this map the item’s own History tab rendered a re-parent as Parent: 1d573f85-… → 50a01576-… while the company-wide log rendered the same commit as Parent: — → ENG-1. An id this node holds no row for is simply ABSENT, so a renderer falls back to the id
work_catalogue{archived}GET /work/catalogue. Answers {types, fields, policy_version, types_version, fields_version} plus the coverage half. policy_version moves on every fields edit and is what a task’s policy stamp records having validated against; a types edit does not move it, because the two are separate objects on separate subjects so an unrelated edit never invalidates every task’s stamp
work_projects{q, unit, archived, sort, limit}GET /work/projects. task_counts ({todo, active, done, closed}) is read from tracker_projects.open_count/active_count/done_count/closed_count — todo being the unfinished work less the started part — MAINTAINED by the task apply whenever a status group changes or a task enters, leaves or is removed — never aggregated per poll, which over every task in every project is what a sixty-second refresh used to cost. last_change ({at, actor, actor_kind}) is maintained on the same row and by the same apply, from the commit that writes the project’s own history row: it is the HEAD OF THAT PROJECT’S ACTIVITY FEED, so the two never disagree — a turn’s spend, a board re-order and an edit to the project’s own settings write no such row and do not move it. It is OMITTED for a project no work has ever been filed into, because “nothing yet” is a different answer from any instant. unit.resolved is a FIELD rather than an absence: “this project names a unit the chart no longer has” is a finding, and an absent unit would be indistinguishable from a project that names none. archived and sort are both CLOSED SETS the engine owns, and both act on the whole company rather than on the page: the answer is capped at 200 rows, so a caller that widened and then narrowed found no archived row at all once the live projects filled the page, and a caller that re-sorted the page ranked the first two hundred keys rather than the company. total counts the SELECTED set, which is what makes “N of M” readable on every one of them — and it IS census[archived], computed from the one aggregate rather than a second COUNT(*), so the number printed beside the rows cannot disagree with the counts printed on the control that chose them. The census is what a SEGMENTED screen needs and cannot derive: on the live segment the archived count has no row on screen to be derived from, so without it an empty answer could not tell a company with no projects from one that has archived every one of them
work_project{key, for_type}GET /work/projects/{key}. A SET read, not a point read: its shape is dominated by aggregates over task rows — the counts — so it carries complete/incomplete under the same contract every set answer takes, and its closure is the project’s container plus both catalogues. It carries the listing’s task_counts, target_date and last_change too, read from the same row so the directory and the project’s own page cannot disagree. shadowed names the workspace field ids this project redeclares, which is what a field in the middle of a move between scopes looks like
work_workload{unit, read_level, …}GET /work/workload. Who is carrying how much, for everybody at once. It counts OPEN work — every task assigned to somebody, whatever its dates — because “who is carrying the most” is a question about a whole queue. Both measures a company may size in are carried rather than one chosen, since which is used differs by team and an answer that picked one would be wrong for everybody sizing in the other. Beside them it carries blocked, overdue and unscheduled, because a person whose whole queue is blocked has a different problem from one who is simply busy; overdue is cut on the company’s own midnight, the day work_items and work_my_work cut on. Ordered heaviest first and capped, with truncated when the cap was reached. unit= narrows to the people whose open work sits in projects that unit owns, named by the unit’s id or by its name, in any case
work_flow{bucket, points, project, read_level, …}GET /work/flow. The company’s work as a SERIES: points windows (1–90, default 14) of a bucket (day, the default, or week) on the company’s clock, oldest first, the last the window now falls in. Each point is {window, start, end, not_started, active, done, closed, completed} — the census at the window’s END (at now for the current one) and how many changes in it took a task from not delivered to delivered (a cancellation is not one; done → closed is not a second). Replayed BACKWARD from today’s census over tracker_history (status, create, remove, restore, purge and project moves, through the partial index replicated migration 0029 adds), so the cost is the window’s changes, not the company’s history. now is {not_started, active, blocked, overdue} — overdue cut on the company’s midnight — and blocked_history is always false: nothing records when a task became blocked, so no blocked series is drawn from a guess. project= narrows to one project, following a task to where it was at each instant
work_item_turns{id, cursor?, limit?, read_level, …}GET /work/turns. One page of the agent TURNS charged to an item, newest first: {item, key, turns, next_cursor?} plus the coverage half. Each turn is {turn_id, ordinal, seat, trigger?, segments, tokens, cache_read, rounds, wall_ms, workers?, sent_back?, outcome, phases, failed_in?, summary?, review?, tools?, at} — a turn’s SEGMENTS folded into one (a turn that parked on a coding run is charged once per segment): tokens, rounds and wall time summed, phases in the order each first ran, and outcome, summary (what it did, in the agent’s words) and review (what the newest review that sent the work back asked for) off its newest segment; tools is each tool its executor called with its calls. failed_in names the phase that failed when outcome is failed because one did — phases lists every phase that ran, the failed one included — and is absent for a failure outside every phase and on a turn that did not fail. ordinal is “Turn n” — the position of the turn’s counted segment among every counted segment on the item, so the newest turn’s ordinal is the item’s spend.turns; a turn with no counted segment on this item (more of a turn charged to another) has 0. at is when the newest segment landed. Read from the tracker’s own turn rows — the account spend sums — so it answers on any node and outlives the thirty days a node keeps its event history, where turns?work_item= reads the event history. next_cursor is the position of the oldest turn’s first segment; the next page is the turns whose first segment is below it, so a turn that gains a segment while you page never moves between pages. limit defaults to 20 and is held to 50; a cursor this endpoint did not hand out is bad_params
work_comments{item, cursor?, limit?, read_level, …}GET /work/comments. One page of an item’s THREAD, newest page first and each page in the order it was written: {item, key, title, comments, next_cursor?, comment_seats?} plus the coverage half. next_cursor reads the page before and is absent when this page reaches the first comment; it continues from work_item’s own comments_cursor exactly, since both are the same read. limit defaults to the detail’s 20 and is held to 50. comment_seats names, per comment id, the PERSON behind a comment an operator token wrote — the seat the token was bound to — since author is the credential, which is the audit trail and not a name; a comment a seat wrote has no entry. Its own question beside work_item because nothing followed that cursor, so every comment older than the twentieth was unreachable, and because a pane that draws a conversation wants the thread without the history, links and fields a detail assembles
decisions{handle?, read_level, …}GET /work/decisions. What waits on ONE PERSON’s decision: the open asks put to any of their identities (the seat and the token bound to it) and the coding runs parked on a question put to them (awaiting_clarification or reseed), newest first, at most 50. {handle, items: [{kind: "ask"|"run", at, ask?, run?}], total, capped, oldest_at?} — total counts every one, not the page; capped says the ask count stopped at the engine’s ceiling, so total is a floor; oldest_at is when the longest-waiting began. An ask is work_my_work’s asked_of_me row — with asked_by_seat, the person behind a token that asked — and a run is sandbox_runs’ row. The same scope rule as work_my_work: a caller reads the person their token is bound to, and naming anybody else needs an operator credential
company_feed{kinds, actor, limit, cursor, read_level, …}GET /feed. What the company DID, newest first: completed, created and handoff from the tracker, page (created or saved) from the knowledge base, and schedule fires from the usage domain — kinds a comma list, empty for every kind this node keeps; asking for one it keeps no source for is bad_params. limit is 1–50 (default 20). {rows: [{kind, at, work?, page?, schedule?}], next_cursor?, complete, read_level}. A tracker row carries its task’s key and title, and — by kind — the task’s running spend ({tokens, turns}, tokens only) and first_pass (no reviewer sent it back), the create’s origin ({surface, conversation}: the chat surface the filing turn was woken on), or the hand-off’s from, to, reassignments and reassignment_budget. A schedule row is a RUN — a fire the scheduler dispatched, never a tick it skipped (those stay in schedule_runs) — and consecutive runs of one schedule for one runner that no other row falls between are ONE row: schedule.runs counts them (at least 1) and schedule.since is the oldest one’s instant (set when runs is above 1). A page that folds many runs reads its source again from where it stopped, at most four reads per source; past that the page is answered short, never out of order. MERGED BY INSTANT, CURSORED PER SOURCE: next_cursor resumes each source exactly where this page stopped consuming it, so a scroll never repeats or skips a row however the three interleave, and is absent once every source is read to its end. A read_level=session floor names the TRACKER’s log (this is a tracker session question); the pages half is read at the dashboard’s own level
work_activity{task, container, kinds, actor, actor_kinds, assignee, q, notified, since, from, to, limit, cursor}GET /work/activity. Every commit writes a history row, so this is an account of what HAPPENED rather than of what was announced — notified is how a reader tells the two apart, and kinds is what the change WAS whether or not anybody heard about it. The two used to be one: a record’s kind was read off its notification, so the same change filed under one word with an audience and another without, and a quiet removal reached the feed as tombstone while the filter spells it removed. kinds accepts the twenty-seven change kinds and nothing else. Each record carries BOTH instants: at is the authored one a card renders, and effective_at is the fleet-agreed one every duration is measured on. actor is who made the change; assignee is whose work it is, and the two are routinely different people. actor_kinds is a CSV of agent, human, operator and system and narrows to WHO WAS WRITING rather than to which handle — which is not the same question and cannot be asked as a set of handles, since an operator commit carries a token’s own label where an agent one carries a seat handle, and the set of people is the roster, which changes. It is REFUSED when it names a kind this build does not have, unlike kinds: a filter silently ignored answers a wider question than the caller asked, and on an audit feed that reads as a company where everybody is an operator. The q gate is on what the query would SCAN, never on which keys were named — container=workspace&q= and a five-year since both name a key and narrow nothing. actor_seat is the person an operator token was bound to when it wrote, beside actor, which stays the token; ask is the question a commit asked or answered, in work_inbox’s shape and read the same way. fields is what MOVED, in the same {from, to} shape for every kind, the full list of names a task’s own row draws on is there too, and keys names the tasks those deltas point at — see GET /work/activity
work_my_work{handle}GET /work/my-work. Seven lists, each bounded at 20 so no block crowds out another — the whole answer is read as one page. totals counts each list IN FULL — {priorities, assigned, asked_of_me, checklist_items, collaborating, watching_recent, unblocked_recent}, each {total, capped} — by the predicate that drew its page and in the same read, so a count is drawn from there and never from a list’s length; capped means the count stopped at 10,000. priorities is NOT re-sorted: the order is what somebody decided. A finished or removed task is filtered out of it rather than rewritten out, because a read must not write to somebody’s own object. Every row’s overdue mark is cut on the company’s own midnight — the timezone a board’s due bands are cut on — so a task due today is not overdue here while a board says today. handle DEFAULTS to the caller’s own seat and naming anybody else’s needs an operator credential — see below
work_inbox{handle, unread, primary_only, snoozed, reasons, limit, cursor, since}GET /work/inbox. One person’s notices, newest first, 50 to a page. Each names the ONE reason of eighteen it reached them under, addressed (it asks something of them rather than informing them), fallback (nobody better was found), and their own read and snooze marks. actor is who made the change — for an operator, the TOKEN, which is the audit trail — and actor_seat the person that token is bound to, which is who a row draws; comment_id and turn_id are the comment the change wrote and the turn that made it; ask is the question the notice is about — the ask itself for asked, the ask it answered for answered — as {comment, asked_of, open, answered_by, answered_at, resolved, choice, decision}, read NOW, so an ask somebody has since answered reads open: false. snoozed is exclude (the default), include (snoozed notices kept and marked) or only (just what was put off); a snooze whose time has come is back under every scope and is not only. snoozed, unread, primary_only and reasons all narrow the SCAN rather than the page, by the same rules the marks are computed with, so a page holds limit notices whenever the scope does — they used to be applied to the page after it was read while the cursor was taken before, so a person with their newest fifty notices snoozed or read opened an empty page with a cursor behind it. For a person with two identities a reason filter keeps a change only under the reason it is HEARD under — the strongest of its rows — so one change never answers two filters under two reasons. primary_reasons is the split that was APPLIED, defaulted, so a caller renders you are seeing these because without repeating the rule; unread and primary are counts over the PAGE and say so, because a total over the table is a second scan of rows this answer did not return. reasons FILTERS rather than classifies — the primary split classifies the same rows — and an unknown one is refused naming the eighteen. since is a log POSITION (<stream>@<generation>:<sequence>, what seen_through renders), never a bare sequence. Same scope rule as work_my_work
work_search{q, limit, mode?}GET /work/search. The company’s work RANKED against a phrase — hybrid by default (BM25 over the engine’s own inverted list fused with the semantic scan over the replicated vectors), or keyword / semantic alone, which is the same fan-out a seat’s search_work_items runs. Not a filter: work_activity’s q is an escaped LIKE over an excerpt, gated to a span of days, and answers a different question. Registered only where this node HOLDS an index, which is separate from holding the board: a node that joined recently has every row and no index, and answers available: false with reason: "building" rather than an error or an empty result — nothing is wrong, and a reader told nothing matched files the duplicate. It carries the same outcome fields as knowledge — mode, served_mode, modes, degraded and coverage{nodes, complete, buckets_missing} — and refuses a mode it does not know, and a q past 400 bytes as bad_params naming its size and the limit, rather than cutting it. A hit carries the item’s id, key, title, project, type, status, assignee and priority, its index snippet, and its 1-based rank and no score, because what ordered it is a fusion across rankers and slices, and a fused number means nothing beside a BM25 one
work_routing{record_id}GET /work/routing/{record_id}. Who ONE change woke, and under which reason — the fact no other tracker records. tracker_notifications has always been readable by RECIPIENT (work_inbox); this is the same rows by RECORD, which is a primary-key prefix scan and needs no index of its own. Each recipient names the ONE reason of eighteen that found them, addressed (it asks something of them), and fallback/fallback_rank (nobody better was found). notified is the history row’s own flag and means the commit CARRIED a notification — never that somebody was woken, since the applier deliberately does not hold the roster that would need. So an empty recipient list is THREE facts and delivery tells them apart: nobody (announced, inside the retention window, and every candidate was the actor or has left), swept (older than tracker.native.inbox_retention_days, so their absence is not evidence), unknown (no horizon stated) and quiet (the commit announced nothing, which is most of them). retained_from is the instant that decision was made against
viewer{}GET /viewer. {operator_id, operator, handle, name, kind, acts, project, config_writer, config_managed_by}; config_writer is whether this credential may change the company document — false for an anonymous caller and for every credential a managed document’s api.auth.company_writers does not list — and config_managed_by is that list, named to an operator only (which token can rewrite the company is the line of access most worth stealing) and [] when the document is not managed, so a screen draws every editing control disabled with who manages it rather than offering a save the engine refuses; acts is what /operator/act serves this caller, empty unless the token is bound to a seat. project is where this person’s create_work_item lands when it names none — the engine’s own default for their seat (the seat’s project, else its unit’s, else the nearest ancestor’s), so a screen offering “Create task” says where rather than working out a second answer; "" when there is none, and a create must name one. Registered on EVERY build with no seam of its own: who is asking is a property of the request rather than of anything this node stores. Answers three states apart — anonymous (operator_id empty), bound (handle set), and presented-but-unbound (an id with no handle), which is an ordinary state rather than a refusal
work_person{handle}GET /work/people/{handle}. Scoped like work_my_work: absent is the caller’s own seat, somebody else’s needs an operator credential. due is the snoozes whose time has come, REPORTED rather than promoted: putting one back in the unread list is a write, and a read that performed one would change fleet state from a path with no operation id and no record. priorities_set_by is who last set the queue when it was not this person — the PERSON, a bound token’s seat rather than the token’s id, which nobody the stamp is shown to can resolve — which is how a lead’s authority is made visible beside the prioritised wake the write sends them — a wake that names the task now at the top of the list, because a notification here is task-shaped and “your list changed” names nothing to act on. max_snooze_ahead is how far ahead a snooze may be set, in SECONDS — the engine’s bound, so a screen offers only the presets mark_inbox will accept
work_views{container, viewer, counts}GET /work/views. container is the strip’s own — workspace, project:ENG, unit:engineering, person:ana — and it is REQUIRED, because a strip belongs to exactly one. viewer is whose personal views appear and whose pins come first, and it takes the personal scope rule: your own seat, or an operator credential for anybody else’s. Absent is the shared strip — no pins and no personal views but the shared ones — which is what a screen asks for before it knows who is looking, and it needs no credential. Every row carries builtin, which is what tells the six nobody saved from the ones somebody did: a builtin row has no id, so there is nothing to rename, protect, rank or pin. params is the saved query in work_items’ own parameter names — this channel’s, not the list_work_items TOOL’s, which renames four of them for a model — so a caller either hands them straight back or, simpler, passes the view’s id as view= and lets the engine expand it. counts=true adds count to every row PINNED for viewer — the total_hint work_items{view, container} answers for it, run as that viewer on the company’s clock, with count_capped when it stopped at 10,000 — or count_refused naming why a view that no longer compiles could not be counted. At most 32 counts, in the strip’s one read; it needs a viewer, since only a viewer has pins, and is refused bad_params without one
work_saved_views{viewer, counts}GET /work/views/saved. Every SAVED view — never a builtin — that viewer can see across every container: the shared ones, and viewer’s own personal ones, pinned-for-them first and then by container (the workspace, projects, units, people) and the strip’s own rank. Each row is work_views’ row shape, so container says where it lives and params is its saved query. A sibling of work_views rather than a container= it takes, because a strip is ONE container’s tabs and this is one PERSON’s views — the inventory of what somebody saved and the pins a sidebar draws — which the workspace strip could not answer: a view saved on a project board appeared in neither. viewer takes the same personal scope rule, and counts=true the same rule as work_views: every pinned row’s count is its view run IN ITS OWN CONTAINER, which is what opening it runs
work_files{project, folder, removed, limit, after}GET /work/files. Answers {project, files, next} plus the coverage half. Each file carries its path, content_type, size, hash (the SHA-256 of its whole content), version, who created and last wrote it and when, and — only with removed — who removed it and when. next is absent on the last page. The listing is rows only: the bytes are read through the byte route, never through the socket
pages{container, parent, roots, status, label, watcher, title, skills, onboarding, limit, after}GET /pages. skills is three-stated: only the tool-skill pages, everything but them, or everything. roots=true is the TOP of a container — the pages with no parent — and is refused beside parent, because an empty parent already means “under any parent” and the two ask opposite questions. The answer carries total (every page the filter matches, counted in the same transaction as the rows) and, while there is more, after — pass it back as after for the next window; a cursor this listing did not mint is bad_params. Each page carries children: how many pages sit directly under it that the SAME listing would show (its status and skills narrowing applied to them), absent when none — what a tree draws its expander off. With skills=true on a node that reads the replicated usage domain the answer also carries skill_loaded_by: page id → the page answer’s skill_loaded_by list for that page, answered for the whole window in ONE read rather than one per row
page{id}GET /pages/{id} — id or CONTAINER/Title. children is the first 50 by title and children_total all of them. skill says the page is a TOOL SKILL (admitted to the skills registry) and onboarding that it is an onboarding page. On a tool-skill page, skill_loaded_by is every seat it reached as a skill over the last 30 company days, most recent first — {handle, last_at, count, loaded, offered}, where loaded counts the seat asking for the body (load_tool_skill) and offered a phase’s catalogue putting its summary in front of the seat; absent on a node that does not read the usage domain, and [] when nobody was. linked_from is who LINKS here, from this node’s index: pages ({id, container, title}, published pages whose body carries this page’s address) and tasks ({id, key, title, status, via}, where via is linked_page for a page relation and description for an address in the description), each capped at 50 with pages_total / tasks_total beside it. A link is a page id in either address the engine reads — /pages/<id> or the dashboard’s #/knowledge/pages/<id> — outside code; a title is not an address, since a rename moves it. Absent on a node with no index. A node that HAS an index and cannot answer sends linked_from_status instead of an empty list: building while its index is on its first lap over pages and tasks (a fresh or joined node’s first minutes — an empty list then would claim nothing links here before every body was read), unavailable when the read failed; the page itself is served either way
page_reads{page, days}Who READ a page: one row per (seat, way of reading) over days company days (1–30, default 30, anything else bad_params) from the replicated usage domain, so every node’s reads are counted — a departed node’s included — and every node answers alike. via is get_page, search, prefetch or skill_loaded; a skill’s catalogue OFFER is deliberately not a read (it would make every skill “read by every agent today”) and is what page’s skill_loaded_by counts. Each row carries count, last_at, and the newest read’s last_turn_id, last_work_key, last_query and, where the run was charged to a task in the engine’s tracker, last_work_item{id, key, title, ordinal} — “turn 2 on ENG-412”. At most 100 rows, newest first, with readers_total; distinct_seats_today is the seats that read it since the company’s midnight, and elided how many (page, way) entries the per-seat-day cap dropped across the window, COMPANY-WIDE rather than for this page — some may have been this page’s and none may have been, so a non-zero value says the list MAY be short (0: it is certainly complete). since, until and days name the window. Registered only where this node reads the usage domain
containers{}GET /containers. A separate question from pages rather than a facet of it: a browser draws the container list once and the page list on every navigation
page_activity{page, container, kinds, actor_kinds, since, cursor, limit}What happened to a page, or to everything in a container — the wiki’s own change log, mirroring work_activity. kinds is a CSV of the ten change kinds and actor_kinds of the three author kinds (agent, human, operator), the latter refused when it names one this build does not have — work_activity’s note says why. since bounds the window and cursor pages it: the same unit, two parameters, because the cursor moves with every page and the bound does not
page_revision{page, version}One revision’s own body, message and author. Revision N is the body AT version N — including the newest — so a reader comparing two versions asks for both rather than for one and the head
config{}GET /config (operator token required)
config_audit{limit}The revision history — no REST twin; GET /config/revisions serves the same records (operator token required)
config_diff{revision_id}GET /config/revisions/{id}/diff — the listing is cut at 500 and changes_total is how many there are (operator token required)
config_entities{kind, id}One addressable collection of the active revision: its ids, or one entity out of it. The read half of Settings › Configuration, whose write half is PUT /config/{kind}/{id} (operator token required)

Every tracker answer carries how far this node had got, and both halves matter. read_level is the level the read was ACTUALLY served at, never the one asked for — a level never silently downgrades, so the two can differ only by a refusal you can see. log_seq is this node’s committed position and applied_through the prefix whose consequences it actually holds; they are two numbers because a node applying nothing while its position advances looks identical to a caught-up one from either alone. log_lag is absent rather than zero when the broker could not be reached, because a read asks how far behind an answer may be and an unreachable broker answering “not at all” is the confident wrong answer. A caller that will not accept an answer of any age says so with max_lag_seconds and/or max_lag_seq beside read_level=stale, and the read refuses too_stale past whichever is reached first — the two are readings of ONE distance rather than two, the record count being what the broker actually answers and the duration being derived from it through this node’s own drain rate. Both are refused at every other level, because they are a staleness bound and nothing else is. min_position is the read-your-writes half: a log position — <stream>@<generation>:<sequence>, the form every write’s answer carries as position and every wake carries — and the answer is served from no earlier than it, at whatever level. At stale it turns “whatever this node holds” into “whatever this node holds, from here on”, still labelled with its lag; a node that has not reached it within the read budget refuses behind rather than serving rows from before the write, and a position on another domain’s log is refused wrong_stream rather than waited for. It is also what makes read_level=session an honest ask here: a session read waits for the caller’s own last write, so it is accepted only beside a min_position and refused without one, the refusal naming the key and the two other asks a caller who typed session might have meant — linearizable, or stale with max_lag_seq. Absent, the level resolves to this surface’s default, which is stale; a seat’s own tools resolve theirs to linearizable and cannot be asked for anything else. See Read consistency.

read_level, max_lag_seconds, max_lag_seq and min_position work on every question that carries a read level, not only on /work/items — /work/items/{id}, /work/views, /work/catalogue, /work/people/{handle}, /work/projects, /work/projects/{key}, /work/activity, /work/my-work, /pages, /pages/{id} and /containers all resolve them the same way. They did not until recently: each of those wrote a hardcoded stale into its read and never looked at the keys, so a caller asking for a stronger answer was served this node’s rows and told, in the answer’s own read_level, that it came back at the level asked for — and the two single-object reads, /work/items/{id} and /pages/{id}, then took the level and dropped the bounds beside it, on the claim that a bound is enforced against a set’s coverage and one row has no set. It is not: a bound is about this node’s lag, checked before any row is read, and one row is exactly as far behind as a board on the same node. And complete: false — with incomplete naming the count, the lowest position among the records (from, as {stream, generation, seq} — the object form every position in an answer takes, seen_through and a listing’s position included; a position in a parameter is the token <stream>@<generation>:<sequence>) and the record version — says rows may be missing, rows that should have gone may still be present, and the totals were computed over the incomplete set. That is a different fact from staleness, and a client that renders read_level and swallows complete looks confidently right.

A read no data node could answer is unavailable, never a list. The tracker and the knowledge base are answered by a data node’s copy of the estate — this node’s own where it holds one, another data node’s otherwise — and a read nobody answered is refused unavailable (REST: 503 with a Retry-After) rather than answered as an empty list, which would say the company has nothing in it. knowledge is the one exception, because a search is best effort: a search no data node ran answers no hits with served_mode empty and coverage.complete false, which a screen renders as “could not be searched” rather than “nothing matched”.

Whose record a personal question answers for

Section titled “Whose record a personal question answers for”

Four questions answer about one person rather than about the company: work_my_work, work_inbox, work_person and conversations. Every one of them takes a handle, and the rule for whose is the same, in one place:

  1. No handle answers for the seat the caller’s own credential is bound to.
  2. Your own handle is the same thing said explicitly.
  3. Anybody else’s handle needs an operator credential.

The binding is a chain of two, and neither link is new. Tier A’s api.auth.tokens maps a presented credential to an operator id; a human seat claims one with contact.crewlet_operator_id. viewer is the question that walks it, and it lives on the seat rather than in Tier A deliberately — Tier A is the root of trust and holds the keys to the secret store, so it resolves with EnvOnly and may never read a value out of Tier B.

A caller with no seat behind their credential is refused bad_params, not unauthorized. Nobody was denied anything: there is no person to answer about, and the remedy is a line of company configuration rather than a different token. A surface that reported it as an authorization failure would send somebody looking for a credential that does not exist.

Three of the four read rows, and those three carry BOTH of that person’s names. A write made through somebody’s own credential is attributed to the token, with author kind operator — never to a seat handle, because a tracker whose author field is chosen by the writer is not an audit trail — so one person’s rows carry two names. work_my_work, work_inbox and work_person therefore match the seat handle or the crewlet_operator_id bound to it, and report the answer under the seat. A change that named both is one notice, under the stronger of the two reasons. conversations takes the handle alone: its rows are written by the seat’s own turns and no credential appears in them.

work_views?viewer= takes the same pair for the same reason — a saved view is owned by whoever wrote it and a pin lives on their own record, and both verbs exist only on the operator surface. Its absent case stays the shared strip rather than the caller’s own seat: a strip is about a container, and the sidebar polls for it before anybody is known.

The alias belongs to the person asked about, resolved from the chart — so an operator reading a report’s day gets that report’s own token id, never the one in the caller’s hand.

These were registered operator-only until recently, which made the one screen a human teammate lives on unreachable to them — and, for an operator, a stranger’s day: the dashboard had no viewer at all, so “my work” fell back to the alphabetically first seat in the chart.

Every node that serves the API feeds its projection through one entry point, stream.Service.Ingest, from an ephemeral broadcast subscription to the event stream (observe.Projector). It receives every event from every node, because a dashboard served by one node must show turns that ran on another, and that is why it is not a publish listener: a listener sees only what its own node published. The event store takes the other route, a publish listener inline on the publishing node, so no two nodes can write one row. Each event updates the live-state projection and fans out to connected dashboards. Backpressure is per-WebSocket: a stalled tab drops the oldest queued envelope so it cannot stall the publish path or other tabs.

The dashboard itself is a React + TypeScript application, built by Vite from crewlet/dashboard/ into crewlet/static/dashboard/, which the binary embeds. Its wire half is src/protocol/: a store that mirrors the projection and derives nothing (protocol/store.ts), a reconnecting WebSocket client with heartbeat, query channel and REST-snapshot fallback (protocol/socket.ts), the one REST transport (protocol/rest.ts) and the one write client (protocol/act.ts). Every list the engine owns and the dashboard must repeat — the event categories, the push kinds, the act refusals, the tools a button may call — is declared once in src/contract/, and a Go gate holds each against the engine’s own value. Around that sit a hash router that keeps every screen, section and filter in the URL, and one file per screen. /dashboard serves the shell; /static/{path} serves its assets. The build output is COMMITTED, so go build ./... needs no Node.

The shell and its assets are served under two caching classes, decided by where the build put the file:

FilesCache-ControlWhy
Everything under /static/dashboard/assets/ — the entry module, every chunk, the stylesheetpublic, max-age=31536000, immutableEach name carries a content hash, so different bytes are a different URL. A browser that has the file never asks again, reload included
Everything else — the shell (/dashboard), /favicon.ico, /static/dashboard/crewlet-icon.svg, the fonts, the notices and protocol.jsno-cacheThe name does not change with the bytes. The browser keeps its copy and revalidates it on every load, so a redeploy is picked up on the next one; the shell is what names the new hashed files

Every file answers with a strong ETag, and If-None-Match is read as a list under weak comparison, so "a", "b", W/"a" and * each earn a 304. HEAD and Range are answered too.

Text is gzipped for a client that asks for it: HTML, JavaScript, CSS, SVG, JSON, plain text and the .ico favicon are compressed once per file per process, at gzip’s best level, and served with Content-Encoding: gzip when the request’s Accept-Encoding admits gzip (or x-gzip, or *) with a weight above zero and the result is smaller than the file. A member that names gzip outranks the wildcard, so *, gzip;q=0 gets the file as it is, and so does a request with no Accept-Encoding at all — the clients that send none are scripts and probes, which would print the compressed bytes. Fonts and images are never recompressed: woff2 and PNG already are. The gzip representation has its own ETag (the identity tag with -gz before the closing quote), and every response for a file that has one carries Vary: Accept-Encoding, its 304 included. Measured on the committed build, the four files a first load fetches go from 1.47 MB to 401 KB. A reverse proxy in front of the engine needs no compression or caching rule of its own for the dashboard; one that compresses leaves an already-encoded response as it is.

/static/dashboard/THIRD_PARTY_NOTICES.txt (served as text/plain) is the license text of every npm package the bundle contains, the design system’s three among them, written by Vite’s build.license, followed by the SIL Open Font License of the embedded Geist and Geist Mono faces and the ISC License of the Lucide drawings every glyph is one of (with Feather’s MIT text for the glyphs Lucide derives from it). The release archives and the container image carry the same file, beside the notices for the Go modules the binary links.

The product’s mark is /static/dashboard/crewlet-icon.svg, emitted by the build from @crewlethq/icons beside the raster favicon.ico; the tab icon, the dashboard’s lockup and the GitHub App landing page all draw that one file.

A second build target, /static/dashboard/protocol.js, is the wire protocol alone as plain ESM — src/protocol/, the store (protocol/store.ts) included, with nothing of React: internal/e2e replays a real company’s captured frames through it under node, so the client’s understanding of this contract is checked against a real server rather than against a fixture.

Its visual system — the token layer, the measured palette, and the rules a change has to keep — is documented in Dashboard Design System.


The company’s charter and its organization tree, as an anonymous reader may see it. The same object is the org section of the handshake snapshot and the body of every org push, so all three surfaces carry exactly one shape.

{
"name": "Nimbus",
"mission": "...",
"vision": "...",
"policies": ["..."],
"timezone": "Europe/Berlin",
"token_budget": {"month": 40000000},
"roles": [
{"name": "Founder", "kind": "human", "manages": ["CTO"], "availability": "CET business hours"}
],
"units": [
{
"name": "Engineering",
"type": "department",
"purpose": "...",
"lead": "CTO",
"goals": ["..."],
"channel": "engineering",
"knowledge": ["..."],
"roles": [
{
"name": "CTO",
"handle": "cto",
"goal": "...",
"backstory": "...",
"responsibilities": ["..."],
"behavioral_guidelines": ["..."],
"manages": ["Platform"],
"token_budget": {"day": 2000000},
"llm": {
"execute": ["fast", "backup"], "review": ["big", "fast"],
"subagent": ["fast", "backup"], "auxiliary": ["cheap"], "judge": ["cheap"],
"sandbox": ["fast", "backup"], "onboarding": ["fast", "backup"]
},
"tool_sources": ["builtin", "mcp:search", "mcp:github"]
}
],
"children": [
{
"name": "Platform",
"purpose": "...",
"roles": [{"name": "Platform Engineer", "goal": "..."}]
}
]
}
],
"derived": {
"seats": [
{
"handle": "cto", "name": "CTO", "kind": "agent", "placed_by_ref": false,
"manager": "founder", "managers": ["founder"],
"reports": ["platform-engineer"], "auto_reports": ["platform-engineer"],
"onboarding_chain": ["Engineering"]
}
],
"units": [
{
"name": "Platform", "type": "team", "lead": "cto", "lead_inherited": true,
"channel": "engineering", "channel_inherited": true,
"seats": ["platform-engineer"]
}
]
}
}

derived is the hierarchy the engine derives from that document, so a client draws a chart rather than deriving one. Each rule in it is one a second implementation gets wrong: a handle is a slug with Go’s own case mapping, a root seat carrying unit: moves into that unit, a lead and a channel cascade to child units that set none, a manages entry naming a unit stands for the seats in its subtree, a unit’s lead manages the members nobody else manages, and the primary manager is the first seat in the engine’s own order that manages a seat. The dashboard derived these in TypeScript and had already diverged on three of them. It is on every answer for a company, derived from the same document in the same call as the tree beside it, and absent only from {}.

The fields above it stay as WRITTEN, so a reader can still tell a declared lead from an inherited one. Every list here may arrive as null (Go marshals a nil slice that way); a reader treats null as empty. The authored path and unit_path of each entry are omitted, because an anonymous reader is given no document to point into, and membership is each unit’s seats; the same block with paths comes back from a configuration write or dry run.

What it carries, and nothing else. The company’s name, mission, vision, policies, timezone, token_budget and derived; for each seat its name, kind, handle, goal, backstory, responsibilities, behavioral_guidelines, manages, availability, token_budget, llm and tool_sources; for each unit its name, type, purpose, lead, goals, channel, knowledge, roles and children. Every value is the one the company document holds, as written: a seat with no declared handle has none here (the engine derives it from the name), and a unit that inherits its lead has no lead of its own. An empty field is omitted, and a node with no active company answers {}.

timezone is the one value the engine resolves, and it is not a mixture: it is the company’s one clock, and an unwritten clock IS UTC, so a running company always carries a zone name here — UTC where the document writes none — rather than an empty string a client would default to its own browser’s zone. Every day the engine cuts is cut on it (“today”, a due band, an overdue mark, a person’s own day), so a screen deciding which day something falls on cuts on this.

token_budget is the ceilings as written, on the company and on each seat that names its own: one number per calendar window it caps (day, week, month, on the company clock), and nothing for a window it leaves open — so an absent key is “no ceiling”, never zero. How much of each window is spent, and when it resets, is GET /budgets; this is the rule those meters count against.

A seat’s llm and tool_sources are RESOLVED, for the reason derived is: the rule is one a client would get wrong. Both are absent on a human seat, which runs neither.

  • llm is every phase’s provider chain exactly as a turn resolves it — keyed by phase (execute, review, subagent, auxiliary, judge, sandbox, onboarding), the first key the model that phase runs on and every later one a fallback in the order it is tried. A flat llm_<phase> field wins over the same phase inside the llm mapping, a phase naming nothing takes the seat’s llm, and a seat naming nothing lands on the company’s default provider or, without one, the first provider declared. The values are provider KEYS, the labels providers.llm gives its entries; the model, endpoint and credentials behind each stay guarded. Absent when the company configures no provider at all.
  • tool_sources is where the seat’s tools come from, in the tool registry’s own origin grammar: builtin first, then mcp:<server> for each server the seat is granted, in the order mcp_servers declares them. A shared server is granted to every agent seat; a shared: false template only to a seat that declares credentials for it under mcp_env, its own or its unit’s — the rule the engine starts a seat’s own server instances by. It is the GRANT, not what is running: a server that failed to start is still listed, and the node heartbeat’s MCP report is what says it failed.

What it never carries. A seat’s contact identities, email, unit reference, workers, learning_enabled, mcp_env, sandbox, placement, integrations and schedules, and its authored llm / llm_* fields (their effect is the resolved llm above); a unit’s mcp_env, integrations and schedules; and every company block outside the charter and its budget (providers, MCP servers, integrations, knowledge, the tracker, notification and learning settings, worker templates). Those are read through the operator-gated config query or GET /config, which masks credentials. Schedules also have a read surface of their own, under the same posture as /org, and the tree does not repeat them: GET /schedules answers every configured schedule with its task and next run.

Why an explicit shape. /org is readable without a token under the default api.auth.allow_anonymous_read: true. Serialising the config’s own seat and unit types would make every field added to a seat public the day it landed, whatever it held. The shape is declared field by field in internal/api instead, and a test fails when the config gains a company, seat or unit field nobody has classified as public, resolved or guarded.

Founder prose is served as written. Nothing in the public fields is resolved as a ${VAR}, so a reference typed into a goal is shown as the text it is; keep credentials in the fields built for them.


The premise of running Crewlet is that an AI manages your company. The person doing that management is very often working through an AI of their own — a coding agent, an assistant, whatever they already have open — and this is the endpoint that lets it read the board, file the thing you just decided, and read what the company has written down. Without it they are reduced to describing the dashboard to it.

Point any MCP client at it:

{
"mcpServers": {
"crewlet": {
"type": "http",
"url": "https://crewlet.example.com/operator/mcp",
"headers": { "Authorization": "Bearer ${CREWLET_API_TOKEN}" }
}
}
}

The same tools a seat holds, not a parallel implementation: list_work_items, get_work_item, create_work_item, update_work_item, comment_on_work_item, merge_work_item, move_work_item, search_work_items, get_work_catalogue, list_projects, describe_project, write_project, task_activity, my_work, list_pages, get_page, write_page, save_page, comment_on_page, and search_knowledge. A schema, a default, a trimmed field and the wording of a refusal are each written once — two copies of “file an item” drift on exactly the parts nobody looks at, and only one of the two is ever tested.

Every write’s answer carries its three-valued outcome, the object’s new version, and position — where the record landed, as <stream>@<generation>:<sequence>, or null on an unknown outcome, where the broker never said. That is the value a client hands back as min_position on any read above, so an assistant that files an item here and redraws the board over GET /work/items is served an answer that includes it rather than whatever this node happened to hold. The tools’ own reads need none of it — they read linearizable — but the answer travels to a client that does not.

A write’s answer here also carries op_id: the operation that call was. A seat’s writes derive their operation ids from its turn, so a seat repeating a call is the same operation; a call here has no turn, so each one is minted an operation of its own and told it. To finish a write that came back unknown, or a gesture that stopped part of the way through, send the call again with exactly the same arguments and that op_id — every write it makes is then the same operation again, a created item or a new comment or saved view included, answered from the ledger where it landed and finished where it did not. A call without one is a new operation: repeated, a create files a second item. The argument is offered by create_work_item, update_work_item, comment_on_work_item, merge_work_item, move_work_item, place_work_item, remove_work_item, restore_work_item, set_priorities, set_pins, mark_inbox and save_work_view, and held to the rule the purge and gate routes hold theirs to: an id this engine minted, at most 128 bytes of visible ASCII — anything else is refused naming op_id. An op_id belongs to the one call it was answered for: the id names that call’s tool and carries a digest of its arguments, so brought back with the same tool and exactly the same arguments it is that operation again, and with any other argument — another item, another project, another title, another blocker — or with another tool, it is refused naming op_id before anything is written. Nothing is ever made twice under one, and nothing is half-made: a looser rule would answer the steps the two calls share from the first call and write the ones they do not. To make a different write, leave op_id out. One case is refused later, by the ledger, and says so: an operation whose step now meets another object — a move whose item somebody else moved in between, so the key it aliases is another — is refused op_reused, and the answer says to leave op_id out.

The one field that differs is who the call acts as. There is no turn and no seat here, so this surface supplies its own identity, and every tool resolves the caller through that rather than through the turn — a read that asked the turn first refused every call on this endpoint while the writes beside it worked. A comment’s @handle is resolved here too, against the company chart current when the comment is written, so mentioning somebody from your own assistant wakes them exactly as it does from a seat.

That identity is two facts, not one. A write is attributed to the token, with author kind operator, always — there is deliberately no way to ask this surface to act as a seat. But if the chart binds that token to a human seat with contact.crewlet_operator_id, the surface also knows who the caller is, and that is what every tool keys on that asks who the caller is rather than who wrote it: mark_inbox, set_pins and set_priorities write that person’s record and sign it with the token; my_work, get_person and work_inbox answer for both of that person’s names, and so does the viewer list_work_items expands preset=my_queue and preset=priorities against; a watch: true, a comment and a create all record that person as the watcher; a create with no project files into their team’s; and the lead relation routing_unit and write_project are gated on is resolved for them. It is also what keeps a wake from coming back to you: the change you just made is not announced to the person who made it, and that exclusion reads the seat as well as the token — so filing work through your own assistant does not wake you about it. An unbound token is an operator outside the chart and writes its own record under its own id, which is ordinary. See Humans in the org.

A person’s own credential also carries an authority of its own, bound or not, and it is the same one a human seat writing from the dashboard has: a re-route, a project’s field declarations and somebody else’s queue are open to a person where they are a lead’s alone for a seat. A lead relation is between two people in the org chart, and a token nobody bound is in no chart — so without it an operator could not re-route the work they own, and with it the company’s own credential is never locked out of its own tracker.

The three person writes take moves, not lists, and each is resolved against the person’s record as the write finds it — so two screens writing at once both land, and nothing a call does not name changes:

ToolArguments
mark_inboxread, unread, unsnooze — lists of the record_ids work_inbox returns; snooze — [{record_id, until}], until RFC3339, in the future and at most a year away; read_through — a log position, <stream>@<generation>:<sequence>, that only ever moves forward; primary_reasons — omitted leaves the choice, [] takes the default back. A notice named twice in one call is refused invalid; a list the call would leave past 256 entries is refused inbox_full.
set_pinsviews and favorites, each {add, remove} or {set} — never both, and a bare list is refused naming the shape. The caps are held against the list the change would leave.
set_prioritieshandle, items — the whole order, most important first — and if_match, the version get_person answered: given, a reorder against an older record is refused stale_version.

Plus sixteen no seat is given: list_work_views, save_work_view, write_work_catalogue, get_person, work_inbox, mark_inbox, set_pins, set_priorities, remove_work_item, restore_work_item, place_work_item, answer_run, pause_seat and resume_seat, steer_turn, and answer_knowledge. A view is furniture — a name, a shape and a filter, arranged so a person finds the same question tomorrow — and a seat’s job is the work rather than the furniture around it. And the catalogue is the company’s own vocabulary: a seat adding a type so its own create succeeds is a seat editing the rules it is judged by, and the refusal it was working around is the signal a person needs to see — which is why reading the catalogue is a seat’s and writing it is not. And a person’s record is a HUMAN’s: a seat has a mailbox — the durable subscription the engine attaches when it acquires the seat — and nothing on a person’s record describes one. The trash is another: a removal takes an item off every board in the company, and a seat that could hide work it did not want to do would be marking its own homework in the one way that leaves no trace. Neither destroys anything — a removal is reversible at any age, and crewlet work purge is the one that is not. A board’s manual order is furniture too — where a card sits says what a person wants looked at first — so dragging one is a person’s; a seat moves work between lanes with update_work_item. And a coding run’s question is a person’s to answer: a seat that could answer its own run would be guessing on its own behalf. Whether a seat works at all is a person’s decision about it too: a seat that could pause a colleague, or resume itself, would be overruling the people who run the company. And a note to a running turn is a person redirecting the work: a seat that could steer a colleague’s turn would be directing it past the person who asked for the work.

place_work_item is a board drag: one item dropped beside another of the same project, in its own lane or into the next one, in one call. It is never a move to another project — that re-keys the item and everything under it, and is move_work_item, which a seat holds too.

Argument
itemThe item being moved, by key or id.
before / afterThe item it now sits directly above, or directly below — one of them, never both. It names a neighbour, never a position: the engine mints the new place inside its own write, between that item and the one beside it as the board stands when the move lands, so two people dragging in one project at once both land where they dropped.
statusThe lane it was dropped into — send it even when it is the item’s own lane, which writes nothing on the item. Omitted keeps its own. Given with neither neighbour, it is a drop into an empty lane: the status changes and its place in the order does not.
if_matchRequired: the item’s version as the board read it. A move of an item somebody changed since is refused stale_version, and nothing lands. A move never changes the version itself, so dragging the same card twice needs no re-read.

It answers {key, status, placed, rank, version, outcome, position}. A move across lanes is two records — the status change on the item, which is history and wakes the people on it exactly as update_work_item does, then the place in the project’s order, which wakes nobody. If the item changes between the two, the lane change stands and the answer says placed: false with the reason in unplaced, rather than failing a call whose status write landed. Dropping a card where it already sits writes nothing and answers applied.

A retry is the same call. The two records are steps of the call’s one operation — its op_id over MCP, its request_id over /operator/act — so a retry is answered from the ledger step by step, the lane change the first attempt made included. That is why the call sends the lane the card was dropped into, never the lane it reads now: compared with the card and left out after the first attempt’s lane change had landed, the placement would be conditioned on the version read before that change and refused as stale by nothing but the retry itself.

A coding run that stops to ask a person something parks until it is answered. A reply on the conversation it was asked in answers it — but a run started by a schedule, a task assignment or a colleague’s ask has no conversation, so answer_run answers any parked run by naming it:

Argument
turn_idThe parked run’s turn_id, as GET /sandbox-runs lists it.
answerWhat the coding agent should be told, at most 32 KiB — it is spliced into the run as one tool reply.

It answers {"turn_id", "agent_handle", "question", "outcome": "pending"}: the answer is on the inbox of the seat holding the run, and that node resumes the run with it — so an answer given while the seat is paused waits for the resume, and one given while the seat is busy with another coding run waits until that run settles or parks and is then the first thing the seat takes — exactly as a chat reply would. The answer is to the question the run was waiting on when you gave it, and it resumes the run only while the run still waits on that one: if somebody else’s answer resumed the run meanwhile and it has asked something new since, yours is announced not_awaiting rather than taken as the answer to a question you never saw. What it became is announced on the event stream as sandbox_run_answered (resumed, not_awaiting, gone, or declined when every attempt to resume the run with your answer failed and it was let go of — the run waits on its question again, and you answer it again), naming the credential and the person. The node holding the seat records your answer on the run before it resumes with it, so a node that stops on the way to the resume leaves it for the seat’s next holder rather than losing it. An unknown outcome is a delivery the broker never confirmed; answering again is harmless, because whichever copy arrives second finds that question no longer waiting.

It refuses not_running for a run that is not waiting for an answer, or has no record at all (a run that ended has none). It is served on every company, native backends or not: the run record is the fleet’s, and a company on Jira runs coding agents too.

pause_seat stops an agent seat taking work until somebody resumes it: it starts no new turn, its incoming mail waits on its inbox in order, and its scheduled runs are recorded skipped_paused rather than sent. The turn it is on finishes first, unless the pause asks to stop it. resume_seat lifts the pause, and what waited is delivered first. See Agent Runtime § Pausing a seat.

ToolArguments
pause_seathandle — the agent seat; reason — one line, at most 500 characters, optional; stop_running — also end the turn the seat is on at its next round. A stopped turn is not run again.
resume_seathandle

Both answer {"handle", "outcome", "paused", "changed", …} — a pause adds paused_by, paused_by_seat, paused_at, reason and stop_running. applied means the pause is the fleet’s record; the node holding the seat carries it out from its own copy, typically within a second. changed is false for a pause of a paused seat and a resume of a free one: the seat is already in the state asked for, and nothing is announced. The one exception is a pause that adds stop_running to a pause without it — the record is amended, names whoever asked for the stop, and is announced again. Each real change is announced once, as seat_paused or seat_resumed, by the caller whose compare-and-set won. unknown is a write the store may or may not have taken, and a retry is safe.

They refuse not_found for a handle that names no agent seat (a person’s seat takes no work a pause could hold), invalid for a reason past its bound, and conflict after losing four compare-and-sets in a row to other changes to the same pause.

steer_turn sends a short note to a turn while it runs. The turn reads it at its next round — after the tool call in flight returns — as a correction or addition to the work in hand, and keeps to it for the rest of the turn: the reviewer that judges the work and every later executor iteration open with it too. See Turn Engine § Steering a running turn.

Argument
turn_idThe running turn’s turn_id, as the agents push names it on each seat’s live_call.
noteWhat the turn should take into account, at most 2,000 characters. Longer is a brief, and belongs on the work item.

It answers {"turn_id", "note_id", "agent_handle", "outcome": "pending"}: the node running the turn took the note, and the turn reads it at its next round. What became of it is recorded there, as agent_turn_steered — delivered naming the phase and round that read it, or expired if the turn ended or parked first. note_id is the request’s own id, so a retry of one request is one note however often it is sent.

unknown means no node answered inside two seconds. A reply lost on its way back is indistinguishable from none, so the note may have been taken; sending it again is safe for that reason.

It refuses not_running for a turn that has ended or parked, conflict for one already holding five notes it has not read yet (once it reads them, the note may be sent again), steer_unsupported for a turn whose executor runs as a coding CLI’s own agentic loop — its rounds are the CLI’s, and the engine has no round boundary to hand a note to — and invalid for an empty or oversized note.

Answering a question from the company’s knowledge

Section titled “Answering a question from the company’s knowledge”

answer_knowledge answers a person’s question — the dashboard’s ⌘K answer — from what the company has written down: it searches the knowledge base (hybrid, auto-drafts hidden) for five pages and the work tracker for three items, reads each WHOLE where this node holds it (a native page’s body, an item’s description; an external wiki’s search snippet, said to be one), condenses any source past 4 KiB for the question with the same auxiliary model — never cut — and asks one model to answer from those sources alone, citing each claim as [n]. A source that cannot be condensed is dropped from the answer and from its sources alike, so nothing is cited that the model was not shown. See Knowledge System § Answering a question.

Argument
qThe question, in plain words; at most 400 bytes — a longer one is refused, never cut.

It answers:

{
"answer_md": "Run `make deploy` from `main` [1]; it is being automated [3].",
"sources": [
{"kind": "page", "ref": "0f7c…", "title": "Deploy runbook"},
{"kind": "page", "ref": "5a1d…", "title": "Rollback"},
{"kind": "task", "ref": "ENG-7", "title": "Automate the deploy"}
],
"tokens": {"input": 7120, "output": 184},
"model": "claude-haiku-…",
"cached": false
}

Source [n] is the n-th entry of sources: a page by its id (and its url, for a page on an external wiki), a work item by its key. tokens is what this call spent, and nothing converts it to money. A question nothing matches is answered in one sentence with no sources, no model and no tokens.

Only a person may ask. It spends the company’s tokens on somebody’s behalf, so the act transport admits only a bound token (as for every act), and over /operator/mcp a token bound to no seat is refused forbidden. No seat is given it: a seat has search_knowledge and a model of its own.

It is charged to the company’s windows. The model is the asker’s own seat’s auxiliary one (llm_auxiliary, falling back as every auxiliary pass does). A person has no seat budget, so before any model call it reads the company’s day, week and month and refuses budget_exhausted — naming the window that ends last and when it resets — if one has no room; after the call it records exactly what the reply spent on the company’s counter alone, past a ceiling included. It is a write rather than a read for this reason: every answer that misses the cache is a model call, and the dashboard’s reads are refetched on focus and on reconnect.

A repeated question spends nothing. Answers are cached on each node, 256 of them, keyed on the question (case and spacing folded) and the node’s corpus position — where its tracker, pages and vector logs are applied through — so any write that could change the answer retires every older one. A cache hit answers "cached": true and "tokens": {"input": 0, "output": 0}, and is served even while the budget is spent. A company whose knowledge base is not native has no position to key on, so its answers are never cached.

It refuses invalid for an empty question, forbidden for an unbound token or a seat the chart no longer has, budget_exhausted as above, and unavailable for a counter it cannot read (nothing is spent), a knowledge base and tracker that could not be searched at all, no model configured, or a model that failed or wrote nothing — a reply that spent tokens and wrote nothing is still charged.

One catalogue, and a call is a fresh write

Section titled “One catalogue, and a call is a fresh write”

The tool set is built once per company and every operator transport serves that one value — so a verb, its schema, its hints and the wording of its refusals cannot differ between the ways a person reaches it. The hints each tool is listed with (read-only, destructive, idempotent, open-world) are the catalogue’s own, and a verb this company is not served is not listed at all.

An MCP call names no request of its own, so every call here is a fresh write: an assistant that files the same item twice gets two items, because that is what it asked for twice, and there is nothing to deduplicate against. /operator/act carries the caller’s own request identity, and derives every operation it writes from it, so a retry of that one request — sent again after an unknown — is the first attempt’s operations rather than new ones; a create, a comment, an update, a page write, a project or catalogue change, a saved view and a person’s own marks, pins and queue all follow the one rule.

Each tool appears only where its half of the company is native: a company on tracker.backend: jira gets the page tools and not the work tools, and one on neither gets only answer_run, which is about the fleet’s own run record. search_knowledge is the exception and is offered against any knowledge backend, Confluence included — a ranked search over the company’s own wiki is exactly as useful to an assistant there.

The turn-only tools are deliberately absent: the memory tools (a diary belongs to a seat), the skill tools (a skill is loaded into a phase), a2a_ask (a colleague’s answer comes back by waking a seat, and there is nobody here for it to reach) and run_sandbox (a detached run resumes a suspended phase that does not exist — answering a run a seat started is answer_run).

On the native backend this endpoint is also how a directory of version- controlled markdown gets published — there is no import CLI for it and there does not need to be one:

Publish everything under examples/nimbus-docs/ — one container per directory, the page title from each file’s first # H1.

The assistant calls write_page per file. That handles what a flag-driven CLI handles badly: the parent chain, a title that already exists (save_page with the version it read), and a file that turns out to be a tool skill rather than prose. The reserved containers are refused to it exactly as they are to a seat.

The token’s own name — the key in api.auth.tokens — with author kind operator. Not a seat: a token is not a colleague, and a name in the handle field would render as one in every thread it appeared in. So an audit can tell an operator’s edit from an agent’s, and a person and the credential they used stay two separate facts on the record.

There is deliberately no way for the caller to name a seat to act as. That would let anybody holding the token write as anybody, and a tracker whose author field is chosen by the writer is not an audit trail.

Every call that is not a proven read also leaves one operator_acted event — the same record a dashboard press leaves, with transport: "mcp" — in the node’s event store: see the runtime audit.

/mcp/ is exempt from authentication wholesale, because the sandbox tool bridge lives there and a box running generated code holds no API token — a signed per-run token in its own path is what authenticates it instead. Mounting a writable company surface under the same prefix would have put it behind no credential at all. /operator is its own always-guarded prefix, alongside /config and /secrets.

/operator/act — the dashboard’s write surface

Section titled “/operator/act — the dashboard’s write surface”

The dashboard changes the company as you: every button that writes posts one tool of the operator catalogue here, and the write is made by the person your token is bound to. It is the same catalogue, the same tools and the same attribution as /operator/mcp — the only rule this transport adds is who may use it.

POST /operator/act/update_work_item
Authorization: Bearer ${CREWLET_API_TOKEN}
Content-Type: application/json
{"request_id": "0192f1a4-9b2d-7e51-8c3a-6d7e8f9a0b1c",
"args": {"item": "ENG-8", "status": "in_progress"}}

Only a person acts. The token must be bound to a human seat by contact.crewlet_operator_id. A token no seat binds — a CI token, an ops bot — is refused unbound (403), and so is every caller while api.auth.disabled is set: a caller the guard never checked is nobody, and a change here is made by somebody. Both keep /operator/mcp, where a credential acting as itself is ordinary. GET /viewer’s acts names what this route would serve the caller, and is empty for exactly the callers it refuses, so a screen can disable a control with the reason rather than offer a press that fails.

Who the write is attributed to does not change. The author is the token, the kind operator, and the bound seat rides beside it as the person whose own state it is — so a write from the dashboard and one from the same person’s assistant read identically in the audit and in every thread. The seat is resolved once per call, when the call is admitted, and the write and its audit record both name that one answer: a config apply that rebinds or unbinds the token while the call is running applies from the next call on, and never admits a call as one person and makes it as another.

The body is JSON and nothing else: {request_id, args}, declared Content-Type: application/json (UTF-8). A form post, text/plain or no content type is 415 before anything is read — a cross-site form can send those without a preflight — and a key the envelope does not take is refused rather than ignored. args is the tool’s own arguments, exactly as the tool’s schema names them. The body is capped at 1 114 112 bytes: twice the largest legal page, because a page escapes to up to twice its length as a JSON string, plus 64 KiB for the rest.

request_id is what names a retry. A UUIDv7 the client mints once per gesture and sends again, unchanged, if it never heard the answer. Every id the write derives is derived from it and from what the call sent — each record’s operation, a created item’s, page’s or view’s own id, a comment’s — so a retry is the first attempt’s operations rather than new ones, and two different calls under one id are still two. It must be version 7 because its own instant is when the gesture began: every retry reproduces it, and it is what a node reads to decide whether its operation ledger can still vouch for the operation. An id carrying no instant — a version 4, say — would read as minted before every row the ledger has ever lost, so on any node whose ledger was ever swept every write under it would answer unknown without being written; it is refused instead, naming what to send. For the same reason the arguments may not carry an op_id: the request already names the operation, and a second retry identity beside it is refused invalid before anything is written. What that buys, precisely:

  • a retry whose append reaches the log while the first attempt’s is still in flight is collapsed into that record by the log’s two-minute duplicate window, and answered at its position. One that arrives after the first attempt’s record waits until this node has applied it (or answers unavailable if it does not catch up), and is then decided as below;
  • a create retried after the first attempt landed files under the same id and the same address, so it is refused exists rather than filed twice, and a page save retried after it landed is refused stale_version — each the sign the first attempt went through: re-read rather than retry again;
  • any other retry is decided again against what the first attempt produced. A comment edit to the text it already holds, or a rename to the title the page already has, is applied with no record; anything else is written again under the same operation — a comment under the same comment id.

It is compared in canonical form, scoped to the token that sent it — one id from two people is two requests, so nobody can make their write the first attempt of somebody else’s — and the nil UUID is refused.

A write that went through answers 200:

{"tool": "update_work_item", "outcome": "applied",
"position": "CREWLET_TRACKER_LOG@1:4711",
"receipt": {"key": "ENG-8", "labels_created": null, "outcome": "applied",
"version": 4711, "position": "CREWLET_TRACKER_LOG@1:4711"}}

outcome is the write’s three-valued answer and position where its record landed — the value a read hands back as min_position so the answer after the write includes it. pending is durable and not yet applied here; unknown means the broker never said, and the only safe retry is the same request under the same request_id. A write that appended nothing answers applied at a null position: the state asked for already holds. receipt is the tool’s own answer, verbatim.

A tool that appends more than one record — an item filed or edited together with a dependency, a page saved and renamed in one call, a project’s tags and settings, the catalogue’s types and fields — answers for all of them: the least certain outcome (unknown over pending over applied) at the latest position, and for a work item the version it is at after the last of them. So min_position never floors a read below a record the call made, and a write with one unconfirmed record is never reported applied. A later record refused after an earlier one landed answers the later record’s refusal, and its detail says what did land — with that record’s outcome and position — because the earlier change is on every node and “not made” would send a person to redo it.

A refusal is {error, tool, detail}, and the transport’s own unbound adds hint. There is no separate field key: a tool’s refusal is its own sentence, which names the argument it refused — for invalid and forbidden, the classes a person is shown as they stand, without the Go error’s package prefix — and no refusal in this build carries a hint except unbound. The transport’s own:

errorStatusWhen
invalid_token401No valid bearer token (the guard’s own refusal)
unbound403The token is bound to no seat, or the guard is disabled. hint names the line of configuration that fixes it
unknown_tool404The company’s catalogue serves no such tool
read_only_tool400The tool is a read; ask it over the socket or its REST route
unsupported_media_type415The body is not declared application/json
invalid_request_id400request_id is absent, not a UUID, not a version 7 UUID, or the nil UUID
invalid_body / body_too_large / unreadable_body400 / 413 / 400The envelope is not one JSON object of {request_id, args}, is over the cap, or did not arrive
draining503This node is draining
internal_error500A tool failed without classifying its failure; the detail is in the node’s log

And the tool’s own refusal, carrying the tool’s sentence as detail:

errorStatus
invalid422
not_found404
forbidden403
stale_version, conflict, exists, already_answered, reassignment_budget, inbox_full, not_running, steer_unsupported, budget_exhausted409
unavailable503

A call interrupted before the tool answered is 503 unavailable, never a refusal: whether it landed is unknown, so the answer carries outcome: "unknown" beside the class and says to send it again with the same request_id. A tool that made its write and could not confirm it — the broker’s acknowledgement was lost, or this node’s operation ledger cannot vouch for it — is answered exactly the same way, with the tool’s own sentence as detail saying what to read and that the same request is the safe retry, and its audit record says unknown rather than refused. That key is what tells either from a tool’s own unavailable refusal (a node in maintenance, a sealed log), which wrote nothing and carries no outcome — the class alone would read both as “nothing happened”. No transport code is also a refusal class, so a client otherwise branches on error alone. The dashboard reads any other 5xx — a gateway that gave up waiting, a success whose body was cut short — as unknown too: it says nothing about an engine that may have written.

Every act is logged as operator_act at info with the tool, the operator id, the seat, the request id and the outcome or refusal — never the arguments — and every act that reached its tool leaves an operator_acted event with the same facts in this node’s event store, whatever became of it: see the runtime audit.

GET /work, GET /pages and their neighbours read this node’s own projection of the company’s tracker and knowledge base — the same copy a seat’s tools read, so an operator and an agent looking at one item see one item.

Three properties are worth stating because each is a decision rather than an implementation detail.

They are registered only where the company runs that backend. A company on tracker.backend: jira has no work_items question at all, and asking gets unknown_query / 404. That is deliberate: an operator who wired Jira and then found a blank Crewlet board would reasonably conclude their integration was broken. There is no native record for this node to have a copy of, and an empty board would claim otherwise.

A node that has not caught up refuses. Every one of these answers unavailable (503, with a Retry-After) while this node is behind the log, rather than returning an empty list. “This company has no work” is an answer a person acts on — they file the duplicate, they conclude the migration failed — so a node that has not applied what the fleet holds must not be able to say it.

The hint is derived, not fixed: how far behind this node is over how fast it is actually draining, so a node grinding through a bulk apply asks for longer than one that caught up in milliseconds. A refusal that waiting cannot clear — a node holding a record its build cannot decode — is not a 503, because a client told to come back would go round a loop that cannot terminate; those are ordinary failures and the log names them. See Read Consistency.

These routes read; a write is a tool. There is no POST /work. An item is filed and moved by a seat’s own tools, by an operator’s assistant through the MCP surface, or by a person at the dashboard through /operator/act — the same tools every time, each write attributed to somebody who can be asked why. There is no write as “the dashboard”, which is not a person and not a seat. The three routes under /work/retention/ below are the exception, and they are not about items: they are operator gestures against the log’s own history, attributed to the token that made them.

Both listings take limit (default 50, max 500) and page by an opaque CURSOR, never an offset: both are read while seats write, and an offset over a set a create or a rename just moved skips one row and repeats another with nothing to say it did. The work listing takes cursor and answers next_cursor; the pages listing takes after and answers after, beside its total. (The pages listing took an offset on the argument that a title order is an arrangement nobody moves; a page created ahead of a reader’s place moves every row after it, which is exactly what a tree’s “Load more” noticed.)

Multi-valued filters are comma-separated — ?status=todo,in_progress — because the socket’s query channel carries a JSON object, which cannot express a repeated key. A filter only one transport could send is exactly the divergence the shared answer function exists to prevent.

Openness is not a boolean on the work listing. The four status groups are what every rule in the tracker is written at, so open work is ?status_group=not_started,active — and show_closed is the separate three-stated question of whether finished work belongs in an answer at all: absent or false leaves it out, true puts it in, and recent:<dur> puts in only what finished inside that window. A single open=false could not express the third.

The pages listing’s skills is three-stated for the reason a boolean would lose:

FilterAbsenttruefalse
skills (pages)every pageonly tool-skill pageseverything but them

An unknown enum value is refused naming the closed set — ?status=finished answers 400 saying which statuses exist — rather than matching nothing. A listing that answered empty for a typo would send somebody looking for items that were never missing.

A project holds files beside its work: a report a seat wrote, a spreadsheet an operator uploaded, the output of a run somebody wants to keep. Each is a row in the tracker — a path, a type, a size and SHA-256, a version and who wrote it — and its content is one object in the object store, the one store the whole fleet shares, rather than a copy in every node’s database. See A project’s files for what a seat does with them.

The listing is an ordinary read (GET /work/files, the work_files query). What the three byte routes add is what a question cannot carry: a download streams the content back and an upload streams it in, neither holding the file in memory.

Terminal window
# upload, replacing version 3 and nothing newer
curl -X PUT -H "Authorization: Bearer $CREWLET_API_TOKEN" \
-H "Content-Type: text/csv" -H 'If-Match: "3"' \
--data-binary @q3.csv "http://localhost:8080/work/files/ENG/reports/q3.csv"
# download
curl -OJ -H "Authorization: Bearer $CREWLET_API_TOKEN" \
"http://localhost:8080/work/files/ENG/reports/q3.csv"
{
"project": "ENG", "path": "reports/q3.csv", "size": 48213,
"hash": "9f2c…", "content_type": "text/csv",
"outcome": "applied", "op_id": "…", "position": "…", "version": 4
}

The content goes first. An upload streams the whole body into one new object in the store before it records the file, so a file that is listed is a file whose content the fleet holds. A write cut off halfway, or refused once its object is stored, leaves an object nothing names, which the collector removes a day later; it never leaves a row pointing at content that is not there. A body whose Content-Length is already over 1 GiB, and an upload into a project that does not exist (404 no_project) or is archived (400 invalid), are refused before any of it is read.

A write names the version it replaces. If-Match carries the ETag the download answered. Absent, the write creates the path or overwrites whatever is there; present, a file that moved since is 412 version_moved and nothing is written. Unlike a work item, a file has no per-field merge — two people editing one spreadsheet would lose one of them.

A write is attributed to the operator on the request. A token with no operator identity is refused 403 operator_required: an upload names the person who made it in the file’s history, never the process that relayed it.

The answer is three-valued. outcome is applied, pending or unknown, as every tracker write’s is (see Read Consistency), and position and version are present only where the write is known to have landed. An unknown is retried by sending the same request again with the op_id it answered.

Each mebibyte has 30 seconds to cross, in either direction: an upload’s mebibyte has 30 seconds of the time the server spends waiting on the client to arrive, never counting the time the server spends storing what already arrived, and a download’s next to be taken by the client once it has been fetched — a floor of about 35 KB/s, far under any real link, and the same floor the object store holds its own transfers to. A whole-body deadline cannot bound a stream of up to a gibibyte without capping real uploads, and none at all let a client trickle one a byte at a time, or open a download and never read it, holding the handler and its connection for as long as it liked. An upload that stops arriving — a connection closed mid-body included — answers 400 unreadable_body and records nothing; a download the client stops taking, or whose content stops being readable part way, is cut, which the client sees as a body shorter than its Content-Length.

A download is served as an attachment, with the file’s own type and the same default-src 'none' policy every non-dashboard response carries, so a file somebody uploaded as HTML downloads rather than running on the dashboard’s origin.

StatuserrorWhen
400bad_pathThe path is empty, longer than 1024 bytes, not UTF-8, ends in /, has an empty folder or a . or .. in it, or carries a control character or a backslash. A leading / is dropped rather than refused
400bad_if_matchIf-Match is not a version
400invalidThe write can never land as sent: the project is archived (unarchive it first), or the content type is over 255 bytes or spans lines
400unreadable_bodyThe upload’s body stopped arriving: the connection closed, or a mebibyte of it took more than 30 seconds (see below)
403operator_requiredA write whose credential names no operator
404no_project / no_fileNo such project, or no file at that path
412version_movedIf-Match names a version the file has moved past
409upload_expiredThe upload took longer than a write may name its object after (twelve hours from when its key was minted); send it again
413body_too_largeThe content is over 1 GiB — refused before reading where Content-Length says so, and otherwise once the body passes it
500content_corruptThe store answered with content that is not what the file records; trying again reads the same bytes, so restore the file from a backup or upload it again
503unavailable / content_unavailableThe tracker or the object store could not be reached, or the file’s content could not be read right now

The byte routes are mounted only where the tracker is native and the node runs the object store. A company on Jira has no project files to serve.

GET /work/retention — what the log is holding

Section titled “GET /work/retention — what the log is holding”

Operator-only, reads included, on the same rule /config and /secrets follow: the answer names every node in the fleet, its position, its disk and its snapshot repository, which is a map of which machine to take out to lose the company’s history. It is never eligible for allow_anonymous_read.

The document is assembled by the node you ask, and says so: half its fields are facts only that node can state — its own applier’s lag, its snapshot, its disk — and half are fleet-wide, read from coordination. node_id is on the document rather than beside it, because a report pasted into a ticket without its author is three per-node facts attributed to a fleet.

{
"node_id": "node-1",
"at": "2031-04-02T03:14:00Z",
"backup_owner": "platform-oncall",
"register_readable": true,
"domains": [
{
"domain": "tracker",
"stream": "CREWLET_TRACKER_LOG",
"generation": 0,
"first_seq": 918100000,
"last_seq": 918280001,
"bytes": 67108864,
"max_bytes": 4294967296,
"reserve_bytes": 268435456,
"headroom_fraction": 0.983,
"trim_floor": 918100000,
"trim_to": 0,
"blocked_by": "backup_floor",
"blocked_since": "2031-03-30T02:00:00Z",
"prose": "Nothing is being trimmed on tracker: ...",
"terms": [{"name": "applied", "state": "ok", "seq": 918279004, "detail": "..."}]
}
],
"nodes": [...],
"snapshots": [...],
"replica": {"store_bytes": 10415140864, "projected_join_seconds": 308,
"rejoin_window_seconds": 1800},
"alarms": [{"kind": "backup_age", "detail": "...", "remedy": "..."}]
}

trim_floor and trim_to are two different numbers, and a blocked domain is where they part. trim_floor is the floor the fleet has published: everything below it may already have been deleted, it is written before the delete it licenses, and it never moves down within a generation — so a blocked domain keeps the floor its last advance reached. trim_to is what the last tick itself concluded: zero while blocked, and below the floor whenever the lowest counted node is. The first is what a node must hold to replay; the second says whether the trim is moving. Both, with terms, blocked_by and blocked_since, come from the row the trim published at the domain’s OWN generation only: just after a reanchor that row is about the stream the domain left. trim_floor_state says which it is — published (the floor, the conclusion, the terms and the blocking term are that tick’s), none_at_generation (the trim has concluded nothing about this generation yet: a fresh fleet before its first tick, or any fleet just after a reanchor, until the trim’s first tick on the adopted stream) or unreadable (the floor register could not be read). Only published makes trim_floor and trim_to a number worth reading; terms is always a list, empty rather than null where nothing is concluded. See Retention.

not_ready is present when the answering node refuses every read of the domain right now — the same refusal its readiness probe reads — with code (wrong_stream, stalled, below_floor, behind, floor_unknown, evicted, or broker_unreachable where its health could not be read), causes naming each identity finding behind a wrong_stream (recreated, ahead_of_log, log_diverged, generation_passed) and detail, the sentence the refusal carries everywhere else. writes_refused is present while the node serves the domain’s reads and refuses its writes: code log_truncated, with detail naming the peer whose rows hold records the log lost. Both are the answering node’s own facts, like the replica block.

Each node’s per-domain row carries generation_state — current, left (a generation the log has since left), ahead (one above the answering node’s) or unknown (a domain the answering node does not run) — and lag only where it is current, because a sequence from another generation is a number in another space and the difference is not a distance. It also carries the node’s own log_diverged, and stream_created_at and checkpoint_stored_at — the stream its rows are keyed to and the record its checkpoint stands on — where the node published them.

A term’s state is one of four, and seq is an answer in exactly one of them: ok carries the sequence the term permits, unknown is a term that could not be evaluated (which blocks), n/a is one this domain does not have, and unbounded is one that was read and binds nothing — no hold pins the log, or a solo fleet takes no snapshots. seq is 0 in the other three rather than absent, so a term permitting nothing yet — the state that holds a young fleet’s trim, and the one worth reading — is distinguishable from a term with no sequence to give.

register_readable is the field that keeps an empty nodes block honest: “this fleet has no nodes” cannot happen, and “coordination could not be listed” happens during exactly the outage somebody is running this in. Without the flag a renderer prints the impossible one.

A node row’s evicted is its tombstone once every log holds one, dated from the latest and naming in by the operator whose eviction reached the last of them.

evictions_unreadable is true on the tracker’s or the pages log’s row when the answering node could not read that log’s evictions as it assembled the report — a failed store read, or its replicated estate closed for a snapshot adoption’s rename or a shutdown — and absent otherwise. It is the same kind of honesty for the node block’s evicted: an unread log contributes no tombstone, so every node reads as not evicted there and stays counted (the conservative side, which is the trim’s own), and “not evicted” is then not an answer. Read it before concluding a node was readmitted; both crewlet retention status and Settings › Backups & retention say so above the node block, and that screen keeps an eviction it just made on the row until a report whose evictions were read.

headroom_fraction is a pointer and is absent when the broker could not be asked. A fraction of an unknown ceiling is not zero headroom, and zero is what the one alarm an operator cannot ignore fires on. It is a fraction of the ceiling ordinary writes are held to: max_bytes less reserve_bytes, which on the tracker and pages logs is the top sixteenth kept for the records that install or lift a gate, so that an eviction still lands on a log full for everything else (the gate reserve). reserve_bytes is absent on a log that keeps none — the vector changelog — where the headroom is of the whole ceiling.

crewlet retention status renders exactly these bytes.

All three are POSTs, so the anonymous-read posture never reaches them: moving the floor the trim deletes against, stopping a machine writing and letting it write again are not reads, whatever a laptop deployment allows.

RouteWhat it does
POST /work/retention/ack?stream=NAME&position=NPublishes an operator backup floor. Refused 400 naming both when either is missing, and 404 when the stream is not one this node runs, which on a node running no state log is every stream. The point is stamped with that stream’s own generation, read from the running log: a bare sequence at another log’s generation names a number space the copy does not cover.
POST /work/retention/evict/{node}?confirm={node}Installs the eviction gate on every identity-claiming log — the tracker’s and the pages log; the vector log counts no node and gets none. Refused 409 eviction_refused, with nothing written to either log, while the node still holds a live presence lease: it is still reaching the fleet, and an eviction would drop everything it writes. The body carries detail, hint, op_id and actions — ["wait", "force"]: stop the node and let its lease lapse, or force it. force=true overrides that refusal for a node wedged in a way that still renews its lease. A node that cannot read the presence leases at all answers 503 eviction_unjudged — a judgement nobody could make is not one that came back clear — with a hint and actions ["retry_same_op", "force"]; force=true takes the eviction past that too, since the leases are the judgement’s only input, and the node logs retention_eviction_forced_unjudged. A node in a capacity window answers 409 not_publishing, before anything is judged, with detail, hint and actions ["wait"]: a gate record is an append, and the window’s mode appends nothing; readmission answers the same.
POST /work/retention/readmit/{node}?confirm={node}The inverse commit, on every one of those logs. Refused 409 readmission_refused, with nothing written to either log, when in the tracker’s or the pages log the node has not applied every record up to the one just before the higher of that log’s published floor (at the log’s current generation) and its first surviving sequence — its last published position is more than one below that bound — or its last published position is from a generation the log has since left. The body carries the sentence (detail), what to do (hint, and actions ["wait"]), and the numbers: domain, position, generation, floor, first_seq, floor_generation, and published — false for a node that has never published a position and is judged as holding nothing. A judgement that could not be made at all refuses too, with nothing written, as 503 readmission_unjudged: the register or a floor unreadable, a log’s first surviving sequence unreadable, or this node behind a reanchor the fleet has made. The body carries detail, hint, actions ["retry_same_op", "other_node"] and, where one log could not be judged, that log. force has no effect here and nothing overrides this refusal: the floor is a fact about what the node holds, not a lease it might be wedged into renewing.

confirm echoes the node id, and a mismatch is 400. A node id no node could run under — the node.id rule, since the id becomes a subject token on every log — is 400 invalid_gate, with nothing judged or written. A request carrying no operator identity is 403 operator_required, because every log’s record names the operator who made the gesture and crewlet retention status prints it beside the eviction.

The gesture is judged once, before any log is written, and then its record goes to every identity-claiming log at once, each written by the node the request reached through its own write authority — every data node holds the whole estate (the retention guide). The trim counts nodes per log, so a record on one log lifts that log’s pin and no other. A 200 answers per log:

{
"node": "node-4",
"evicted": true,
"op_id": "01a0cd85-735a-7294-9d3e-38998abd698c.evict-node-4",
"complete": false,
"domains": [
{"domain": "tracker", "stream": "CREWLET_TRACKER_LOG",
"op_id": "01a0cd85-735a-7294-9d3e-38998abd698c.evict-node-4.evict.tracker:node-4",
"outcome": "applied",
"position": {"stream": "CREWLET_TRACKER_LOG", "generation": 1, "seq": 918280002}},
{"domain": "pages", "stream": "CREWLET_PAGES_LOG",
"op_id": "01a0cd85-735a-7294-9d3e-38998abd698c.evict-node-4.evict.pages:node-4",
"outcome": "unknown",
"actions": ["retry_same_op"],
"hint": "its outcome is unknown: the same gesture under the same operation id answers from this log's own ledger if the record landed, and writes it if it did not"}
]
}

Each entry is one log’s own answer. outcome is the write’s three-valued outcome — applied, pending or unknown — with the position the record holds, absent for unknown, which has none: a gate the caller believes has landed and which is only pending is the difference between a node that has stopped writing and one that is about to. An entry with error in place of an outcome was not written, and reason names why in the vocabulary every write refusal uses (log_full, evicted, …). A log that answered holds its record whatever the other did.

An unknown entry may also carry "unvouched": true: this node’s operation ledger may have lost the row the operation needs — it was minted before the ledger’s sweep reached it — so the node published nothing and cannot tell whether the record landed, and the same request here answers the same way every time. Such an entry offers other_node rather than retry_same_op: send the same request, with the same op_id, through a node whose ledger reaches back that far.

A log the gesture did not finish also carries actions — what to do, in order — and hint, the sentence saying why. Both are absent on a log that holds its record. The actions are a closed set a client switches on, and hint names no client’s controls: crewlet retention evict renders an action as its own flags and the dashboard’s evict dialog as its own buttons, so a sentence spelling -force is never shown beside a screen that has no such flag. Every refusal answer above carries actions beside its hint too, and the op_id the request was sent under — so retry_same_op on a refusal (503 eviction_unjudged) is the same request with that op_id, read off the answer like any other. Nothing was written under it, so sending it again with a fresh one is equally safe.

ActionWhat the operator doesWhere the gate answers it
retry_same_opSend the same request again with the answer’s op_idAn unknown outcome (unless it is unvouched), a lost race, a failure before the write answered, and a refusal that clears on its own (behind, deferred, floor_unknown, below_floor); a gate refusal — evicted, overtaken, abandoned — of another node’s copy of the operation, which this node’s append was collapsed onto (hint names that node): the reason is that node’s standing, not this one’s, so this node finishes it once the log’s duplicate window (two minutes from when the copy landed) has passed; 503 eviction_unjudged; and 503 readmission_unjudged (beside other_node)
new_gestureStart a new gesture, without op_idsuperseded — the operation’s record landed and a later gate record on the same node has undone it since (an eviction retried after a readmission) — and op_reused, an operation id that already names a record on another object
forceSend the eviction again with force=true409 eviction_refused, 503 eviction_unjudged
other_nodeSend it, with the same op_id, through another node the fleet still countsOf this node’s own standing or its own record: evicted (this node is evicted itself), overtaken and abandoned (this node wrote the record from rows a reanchor left behind — send it through a node on the log’s current generation), an unvouched unknown (this node’s ledger cannot say whether the record landed), and beside retry_same_op on deferred, below_floor and 503 readmission_unjudged. An overtaken or abandoned refusal, and an evicted one that names a position, is a record that landed and applies nowhere: it holds the op_id on that log for the log’s duplicate window (two minutes) from when it landed, and the same op_id sent sooner — through any node — is answered the same way again, so hint says to wait that out first
reanchorRe-anchor the log first, then send the same request with the same op_idwrong_stream
set_capacityRaise the log’s ceiling, then send the same request with the same op_idlog_full — a gate record is admitted into the log’s gate reserve, so this is a log full to its broker ceiling past even that
waitWait for what hint names to clear on its own, then run it again — what that is depends on the refusal, so a client reads the answer’s error beside the action409 eviction_refused (the lease to lapse), 409 readmission_refused (the node to catch up), 409 not_publishing (the fleet to leave its capacity window)
restoreRestore the store and the stream from one backupskew

A log with a hint and no actions is one no gesture clears — a record larger than the broker’s max_payload, or a refusal reason this build has no word for — and the hint says what does. Only retry_same_op makes the same request, sent again now, the remedy. But four actions — retry_same_op, other_node, reanchor and set_capacity — keep the gesture’s own op_id as how it is finished once the operator has acted: a gesture sent afresh after raising a ceiling is a second one, which writes every log that already holds the first one’s record again, re-dates that eviction, restarts its fence window and turns the first op_id into superseded. Only new_gesture ends the operation for good.

complete is true only when every log answered applied or pending. When it is false and a log offers retry_same_op, send the same request again with op_id set to the operation id the answer carried. Each log’s record is published under an id derived from it — the entry’s own op_id, which carries the gesture’s sign, the log and the node — and each log’s snapshot reads that id’s ledger row before anything is decided: a log whose record is already the gate in force answers at the position it has and is not written twice, and only a missing log is written — through a node whose applier had not reached the first record yet as well, and however long after the first request the retry comes, since the broker’s two-minute duplicate window plays no part in it. Without op_id the route mints a fresh one, which is a second gesture rather than this one finished; an id carried from an eviction to the readmission after it is a different operation on every log, and one carried to another node is that node’s own operation.

An op_id is held to one rule, here and on the purge route, and anything else is 400 op_id_invalid with nothing judged or written: an id in the engine’s grammar — a UUIDv7 whose leading 48 bits are the Unix millisecond it was minted at, optionally followed by . and a name — of at most 128 bytes of visible ASCII with no space. A client may mint its own, and both of the engine’s own clients do, before the first request: crewlet retention evict on the workstation’s clock and the dashboard’s evict dialog on the browser’s. The mint instant is read by the state log to decide whether its ledger can vouch for a retry, so a client clock far ahead of the fleet’s is the one case this trusts a caller with — the same one it trusts the engine’s own nodes with. An id with no instant carries nothing to read, so no node could tell whether it already ran. The rest of the rule is the broker’s: only the id’s leading uuid is minted and the rest is free text, while the whole id travels as the broker’s message-id header, which trims its ends and turns a line break into a space — so such an id would be deduplicated at the broker as a different id from the one every log’s ledger answers for. It is refused rather than cleaned, because a retry has to send back the id it holds, byte for byte; every id an answer carries already fits.

The gesture does not stop when its caller does. The judgement runs under the request, within twenty seconds, so a request abandoned before it wrote nothing; once the first record is about to be written the node finishes the gesture under its own budget — thirty seconds for the logs — so a dropped connection or a client timeout never leaves a node evicted on one log and counted on the other. A node answers one gesture within fifty seconds at most — the judgement and the logs together — so a client waiting a minute, as both of the engine’s own do, always reads the node’s own answer, and so does a reverse proxy at its usual sixty-second read timeout. A caller that sends its own op_id — crewlet retention evict and the dashboard both do — can ask again with it and read every log’s answer; one that let the route mint it has lost the id with the answer that never arrived.

Every node running the state log serves both routes, whichever backends the company uses: a company on an external tracker still runs both logs. A node running no state log has no log to write a gate to and answers 503 no_state_log rather than 404 — the routes exist on this build, and telling an operator they do not sends them looking for a version mismatch that is not there.

A log’s Tier A ceiling is only the value its stream is created with, and these routes are the window in which a running log’s ceiling changes. The procedure is documented once; what follows is the wire surface.

RouteWhat it does
POST /work/retention/capacity?stream=NAME&bytes=N&confirm=NOpens or resumes the operation and drives it as far as this node’s mode allows. confirm repeats bytes and a mismatch is 400: the target is chosen once for the life of an operation. assert_excluded=true is required only on an external broker.
GET /work/retention/maintenance?stream=NAMEThe operation, every acknowledgement, every admission, and — computed here rather than by each client — whether the seal holds, what is blocking it, and which admissions block activation.
POST /work/retention/maintenance/abandon?stream=NAMEFrom opened clears the operation outright; from anywhere else enters the seal.
POST /work/retention/maintenance/exclude?stream=NAME&node=ID&confirm=IDWaives one participant’s acknowledgement and withdraws its admission.

A node in normal mode refuses the write routes, naming the restart: the usage a resize is decided against has to be a quantity nothing can move. The three-mode boot is crewlet run -mode; see the CLI reference.

The sealed / blocking / admissions_blocking fields are computed server-side, from the same predicate the coordinator itself runs. A client that re-derived them would be a second opinion about one barrier, and the two would drift.

RouteWhat it does
GET /work/retention/reanchor?stream=NAMEThe LIVE stream’s own created_at, read from the broker on the call — the instant a wrong_stream refusal names — the generation that stream’s domain stands at, and what a reanchor run now would do: case — recreated (another stream than the rows are keyed to, followed from its first surviving record), restored (the same stream brought back from an older copy, ending below the checkpoint or holding another record at it since, followed from its end — one below the node’s own generation record where an earlier attempt already appended it) or abandoned (the same stream, continuing in a generation only an evicted peer held, followed from this node’s own checkpoint with that generation’s records void) — with cursor, the sequence the new checkpoint would sit at. With nothing to re-anchor there is no case, and nothing_to_reanchor says why. A restored log holding records this node’s rows do not hold, written after the restore, also answers discards (the newest of them: seq, kind, subject, writer, op_id, stored_at, and ledger_lost_before when its operation is older than the instant this node’s operation ledger may have lost rows from — the rows may hold it after all) and discarding (why a reanchor refuses without discard=true). 404 unknown_stream for a stream this node does not run; 503 stream_unreadable when the broker did not answer the read, which is worth retrying.
POST /work/retention/reanchor?stream=NAME&confirm=<created_at>[&force=true][&discard=true]Runs that ONE domain’s generation transition, answering with the new generation, the case it answered, the cursor its checkpoint went to (for a restored log, one below the generation record the transition appended; once that record is appended the transition finishes on its own time, so a client that disconnects does not stop it halfway) and, when it discarded records written after a restore, the newest as discarded. No other domain’s checkpoint moves, and the domain’s applier resumes with no restart.

confirm is the value the GET returns, supplied by the caller: the confirmation means I looked at the thing I am re-anchoring, so the two are deliberately separate round trips rather than one route that reads and acts. It is compared as an instant at microsecond precision, so the RFC 3339 value the GET answers is accepted as it came. force=true overrides the rule that only the most caught-up node may re-anchor, for a fleet whose register cannot say; it never overrides a peer that has already re-anchored the stream, whose rows are the fleet’s history in the new generation and which the refusal names. discard=true accepts that a restored log’s records this node’s rows do not hold — written after the restore, and named by the GET as discards — are applied on no node; without it that reanchor is refused 409 reanchor_refused, and with it the answer carries the newest one as discarded. The two flags are separate because neither answers the other’s question. A node in a capacity window refuses every reanchor 409 reanchor_refused, naming its mode: the reanchor appends a generation record, and the window’s mode appends nothing.

What one seat has learned, in one round trip — answered by the node holding the seat. Also served as the agent_memory query, which takes {id, limit}.

{id} is the seat’s handle — the canonical identifier everywhere in the system; a caller holding the seat’s derived agent id instead is resolved to the handle through the chart. The halves are keyed differently in the store (the diary and the onboarding marker by the derived agent id, the rest by the handle), and the answer resolves that itself rather than making a caller know which.

{
"id": "<handle>",
"diary": [
{ "id", "content", "retention", "source", "turn_id",
"created_at", "ttl_until", "retrievals" }
],
"diary_total": 142,
"episodes": [
{ "id", "turn_id", "agent_handle", "task_summary", "ask", "ask_bytes",
"plan_summary", "plan_summary_bytes",
"review_outcome", "tool_sequence", "skills_used",
"conversation_key", "work_key", "created_at", "ended_at",
"duration_ms", "compacted", "count",
"compaction": { "common_task_pattern", "done", "notable_patterns" } }
],
"episodes_total": 38,
"skills": [
{ "id", "key", "title", "summary", "version", "updated_at", "uses" }
],
"skills_total": 4,
"counterparties": [
{ "subject": { "handle" | "external_id" + "platform", "name" },
"resolved", "traits", "interactions",
"first_seen_at", "last_updated_at", "last_corroborated_at" }
],
"counterparties_total": 11,
"latest_reflection": { "id", "content", "…": "a diary row" },
"onboarded_at": "2026-09-01T08:02:11Z",
"held_by": "node-2"
}

Who answers. A seat’s memory is written to the store of the node running it and carried to every other node on a compacted changelog, so every node that ever held a seat keeps a copy and only the holder keeps it CURRENT. The node serving the request reads the seat’s lease and:

The lease namesThe answer
nobodyempty, with held_by: "none" — the copies on disk are of unknown age, and none is shown as the seat’s memory
this node’s own incarnationread here, once the seat is attached; while it is still arriving (hydrating before its mailbox attaches) the read is unavailable rather than short
a peerasked of that incarnation on an ephemeral scatter (crewlet.held.read), with a 2 s budget; silence is unavailable naming the node, never an empty memory — a holder that is draining, or whose heartbeat lapsed, included

held_by is the node that answered. A node with no broker is the whole fleet and answers every read from its own store. unavailable is a 503 with a Retry-After on REST and the socket’s unavailable code with retry_after_seconds — a moment’s wait, not a fault.

Every collection is a page with its total beside it. limit (1–50, default 50) pages all four; diary_total, episodes_total, skills_total and counterparties_total are COUNTED in the store over the seat’s whole set, so a screen renders the total rather than the length of a page that was cut. latest_reflection is the newest live diary entry whatever the page — a profile’s summary asks for limit=1 and reads the totals and this. It is null when the diary holds none.

Every key is present on every answer, as an empty list or a zero rather than an absent one: a caller cannot tell “this seat has learned nothing” from “this answer does not carry that half” if the key is simply not there.

An episode says what it is. task_summary is the label of the event that woke the turn, ask what it was asked ("" for a turn that was told nothing) and plan_summary what it did. A compacted row stands for count turns and carries compaction instead — what they had in common, how many of them ended done, counted from the members, and what varied — which is null on a raw row.

A listed episode carries the openings of its two long texts. An ask is bounded only by the event that delivered it and an account by nothing, so a page of fifty whole could be more than the transport carries — and the holder would refuse the whole read. So ask and plan_summary are each at most their first 600 bytes, cut between characters — the figure a seat is shown of a past turn’s words before they are condensed (learning.EpisodeAccountBytes), so a text listed whole is one a seat was shown whole — and ask_bytes and plan_summary_bytes are the whole texts’ sizes: a text is whole exactly when its size is its length. The whole row is GET /agents/{id}/memory/episodes/{episode}.

GET /agents/{id}/memory/episodes/{episode}

Section titled “GET /agents/{id}/memory/episodes/{episode}”

One episode whole: the same row as the listing’s, with ask and plan_summary complete. Also served as the agent_episode query, which takes {id, episode}. Answered by the node holding the seat under exactly the rules of GET /agents/{id}/memory above, and naming it.

{
"handle": "<handle>",
"episode": { "id", "task_summary", "ask", "ask_bytes", "plan_summary",
"plan_summary_bytes", "…": "an episode row" },
"held_by": "node-2"
}

episode is null for one the seat does not hold — the lifecycle dropped it or folded it into a compacted row since it was listed, or it is another seat’s — which is an ordinary absence. An episode’s ask and account came in one completed turn’s event, which the transport held to 8 MiB, so one read whole fits as that event did; should one ever not, it is unavailable, with the holder saying how large the answer was, rather than cut.

Every agent seat’s memory at a glance — the list Knowledge › Agent diaries draws. A socket query with no parameters (there is no REST route: it is a screen’s list, and GET /agents/{id}/memory is the one seat’s record).

{
"seats": [
{ "handle": "swe", "diary_total": 142, "episodes_total": 38,
"skills_total": 4, "last_reflection_at": "2026-09-28T16:02:11Z",
"latest_reflection": { "id", "content", "…": "a diary row" },
"held_by": "node-2", "unavailable": "" }
],
"coverage": { "nodes": [{ "id": "node-1", "answered": true, "error": "" }],
"complete": true }
}

Every agent seat in the chart, in handle order, and no cap — a person keeps no memory the engine writes and is not listed. Each row is counted by the node HOLDING the seat, under exactly the rules of agent_memory above, but gathered in ONE round: the serving node lists every seat lease once, groups the seats by the incarnation holding them, reads its own from its store and puts ONE request on crewlet.held.read naming each holder’s seats; every holder answers for its own in one reply, inside the same 2 s budget. So a row is one of three things:

RowMeaning
held_by a node, unavailable emptycounted by that node — the totals are the ones agent_memory carries
held_by: "none"no node holds the seat; nothing is counted, because no copy anywhere is current
held_by a node, unavailable setthat node did not answer (or its answer could not be read), or is still taking the seat — the reason is here and the zeros beside it are not a count

coverage is the shape every fleet answer carries: this node and every holder that was asked, each answered or with its error, and complete only when all of them answered. A lease table that cannot be read fails the whole answer as unavailable, since without it every row would be a guess at who holds what.

The rows are projected by internal/learning/memread rather than being the learning package’s own structs marshalled directly: those are domain types whose fields exist for the recall path, they carry no json tags, and marshalling them shipped Go field names plus every row’s raw embedding vector to a screen with no use for one.

Sources:

  • diary — the seat’s private observation log, written by reflect_and_persist and the reflection pass, live entries only, newest first. retention is diary_long or diary_short; a short entry carries the ttl_until it lapses at. retrievals is how often it has actually been recalled, which is the difference between a memory that keeps proving useful and one written once and never read.
  • episodes — one row per completed turn (or per compacted cluster), newest first. duration_ms is milliseconds: a Go time.Duration marshals as an integer count of NANOSECONDS, which renders as a plausible and wildly wrong number.
  • skills — the seat’s own synthesized skills, drafted from its repeated work and loadable mid-turn via use_skill. Archived rows are hidden and stale ones shown, because a stale skill still works and still revives on use.
  • counterparties — what the seat learned about the people it works with, most recently updated first. Both instants are carried and they measure different cadences: last_updated_at moves on every interaction and last_corroborated_at only when the traits changed. traits is a bag whose keys the model invents. subject carries a handle for a seat of this company and an external_id with its platform for anybody else.
  • onboarded_at — when the seat first finished onboarding, or ""; a pass claimed and never finished is not an onboarding.

A store that cannot be read fails the read rather than answering an empty section: “this seat remembers nothing” and “the store did not answer” are opposite facts. The table is strictly per-agent; cross-agent procedural artefacts are promoted as draft pages in the shared knowledge backend, reachable by all members via query-time search.


Backs Settings › People & access: who can reach the company through this engine’s own surface, and as whom. It needs a token, reads included — which labels the guard accepts and whom each one is, is a map of which credential to take — and api.allow_anonymous_read does not open it.

Labels, never values. The answer is built from the auth guard the API mounts, through a type with no member a token’s value could travel in, so no edit to the answer can put one on the wire. It is read off the GUARD rather than off Tier A because the guard is what decides: with api.auth.disabled it accepts no listed token at all, so tokens is empty and every binding reads no_token.

It is one join walked from both ends. Each token names the human seat whose contact.crewlet_operator_id binds it — through the same lookup the viewer and /operator/act make — and a bound token’s scope is person (it acts from the dashboard as that seat); every other token’s is operator (every guarded surface under its own label, never the act transport). Each person names the state of their own binding:

bindingMeansThe remedy
boundThe id resolves to a label the guard accepts—
unboundThe seat names no id — an ordinary state: agents reach them on their other surfacesAdd crewlet_operator_id to act as themself
unresolvedThe id is a ${VAR} this engine’s environment does not set (or it resolves to the reserved anonymous)Set the variable, or write the label
no_tokenThe id resolves to a label api.auth.tokens does not carry — every binding under a disabled guardCorrect the label on either side

operator_id is the binding as WRITTEN, and each of contacts — the seat’s other identity fields, one per config key — is its value as written with reference and resolves beside it: a ${VAR} is its name, never the variable’s value. The binding is not among the contacts, because it is an attribution rather than an address.

auth.company_writers is api.auth.company_writers in Tier A’s order: the token ids that alone may change the company document when another system manages it. It is always a list, and [] means every token may. The screen names the writers beside the tokens, so an operator reading the deployment’s api.auth sees whether the document is managed and by which credential.

{
"auth": {"disabled": false, "anonymous_read": true, "allowed_origins": [], "company_writers": []},
"tokens": [
{"id": "ci", "scope": "operator", "seat": null, "yours": false},
{"id": "founder", "scope": "person", "seat": {"handle": "ana", "name": "Ana Diaz"}, "yours": true}
],
"people": [
{"handle": "ana", "name": "Ana Diaz", "email": "[email protected]", "availability": "",
"operator_id": "founder", "binding": "bound",
"contacts": [{"key": "slack_user_id", "value": "U0FOUNDER", "reference": false, "resolves": true}]}
]
}

Backs Settings › Models & keys: every model the company configures, in config order (the order a seat that names no model falls back through), the keys each rotates through and which of them is benched. It needs a token, reads included — which variable holds each model’s key and when each is refused is a map of which credential to take — and api.allow_anonymous_read does not open it.

Names, never values. A key is ref, the variable a whole ${VAR} names (or the vendor’s conventional variable, source: "default", for a model that names no api_keys), or its position alone for a value written into the document (source: "inline", ref: ""). hint is the 12-character, non-reversible identifier the engine’s credential_cooled log lines carry.

The pool and the fleet’s ledger, the later deadline winning. A bench is published to the fleet when a node takes it and pulled by every other node every 15 s; this answer reads the ledger directly, so a key a peer benched a second ago is cooling here before the answering node has pulled it, and a bench whose publish failed still reads cooling on the node that took it. A ledger that cannot be read does not fail the answer: fleet is false, fleet_error says why, and every deadline is the answering node’s own. uses and in_flight are the answering node’s leases of the key since it applied its configuration, and unresolved is what ITS environment and the company’s secrets resolve.

key stateMeans
readyA call can lease it now
coolingBenched after a rate-limit or auth refusal until cooling_until
unresolvedIt resolved to nothing on the answering node and is not in the pool
duplicateThe same value as the key at same_as (1-based), held once
model stateMeans
readyEvery key resolves and none is cooling
degradedSome keys can be leased and some cannot
exhaustedEvery key that resolves is cooling: each call falls through to the seat’s next model
no_keyNo key resolves: every call is refused as unauthorised
loginA cli-agent entry — one login held by the CLI, no key bag
{
"node": "node-1",
"fleet": true,
"fleet_error": "",
"providers": [
{"key": "smart", "type": "anthropic", "model": "claude-sonnet-5", "state": "degraded",
"ready": 1, "rate_limit_seconds": 3600, "auth_seconds": 300,
"keys": [
{"ref": "ANTHROPIC_KEY_A", "source": "reference", "hint": "3f9a1c0b7e2d", "state": "ready",
"cooling_until": null, "same_as": 0, "uses": 12, "in_flight": 1},
{"ref": "ANTHROPIC_KEY_B", "source": "reference", "hint": "8c41d2e9a0f7", "state": "cooling",
"cooling_until": "2026-09-29T12:40:00Z", "same_as": 0, "uses": 4, "in_flight": 0}
]}
]
}

To change a model’s keys, PUT /config/llm-providers/{id} — see Per-entity read and write.

Backs the Servers section of Settings › Tools & MCP: every MCP server, what the configuration declares for it and what each live node did with it. It needs a token, reads included, for the reason /fleet does — it names the nodes, the launch commands and the first line of each failure — and api.allow_anonymous_read does not open it.

Off the heartbeats, not a fan-out. Each node re-publishes what its MCP starts concluded on its presence lease, one row per server with its instances counted (a per-seat template has one instance per seat that node holds), so one read of the lease table is every node’s answer at once. A node whose last heartbeat carried no status (its status hook overran the beat) is reported: false and its cells are UNKNOWN — never a row of zeros, which would read as “started nothing”. A node that reports and started nothing is reported: true with no rows.

servers lists every server this node’s active configuration declares, then any a node reports that the configuration does not carry (configured: false — a node still on an older revision mid-rollout). The launch is the parts that are not credentials — transport, command, args, url; env and headers are never here. started and failed are summed over the nodes, tools is the most one started instance serves, and state is decided here so no screen re-derives it:

stateMeans
runningEvery instance any node launched started and listed its tools
partialSome started and some did not — a node’s environment or one seat’s credentials rather than the server
failingInstances were launched and none started — the server, its command or address, or credentials every seat shares
not_startedEvery reporting node started nothing for it — a per-seat template no seat on a live node declares credentials for
unreportedNo live node’s latest heartbeat carried the report

error is one failed instance’s reason, cut to 240 bytes on the heartbeat, and error_seat the seat it was launched for; the whole text is the node’s mcp_server_failed log line.

{
"nodes": [{"id": "node-1", "reported": true}, {"id": "node-2", "reported": true}],
"servers": [
{"name": "github", "configured": true, "shared": false, "transport": "stdio",
"command": "npx", "args": ["-y", "@modelcontextprotocol/server-github"], "url": "",
"state": "partial", "started": 3, "failed": 1, "tools": 26,
"nodes": [
{"node": "node-1", "reported": true, "started": 2, "failed": 0, "tools": 26, "error": "", "error_seat": ""},
{"node": "node-2", "reported": true, "started": 1, "failed": 1, "tools": 26,
"error": "401 Bad credentials", "error_seat": "backend-dev"}
]}
]
}

To add a server, PUT /config/mcp-servers/{name} with If-None-Match: * — see Per-entity read and write.

Backs the dashboard’s Settings › Nodes screen — the questions /health cannot answer, because it answers about the node that served it and a load balancer sends the next refresh somewhere else.

Read from the lease table, so every node gives the same answer: node presence carries each node’s node.roles and node.labels, seat and worker leases name their holder, and the per-node config epoch comes from the control plane’s apply status.

It needs a token, reads included, like every other answer the dashboard’s Settings draws. What it describes is the DEPLOYMENT rather than the company’s work — the node ids, which node holds which seat, the lease epochs, how far a rollout has reached — so it is scoped the way /integrations beside it always has been, and api.allow_anonymous_read does not open it.

Presence also carries what each node is doing — in_flight, draining, posture and started_at — because only the node running a seat knows those, and /health answers about whichever node served the request. They ride on the heartbeat that already re-sends roles and labels on every beat, rather than over a request/reply to the owning node: every answer would then be partial, it opens a new trust edge, and it duplicates the mechanism the lease table already is.

Absent is not zero. A node whose status read overran its share of the heartbeat (seat.StatusBudgetRatio) publishes no status on that beat and omits those fields entirely, and the dashboard draws an em dash. A confident 0 would render an idle row for a process that is simply not saying.

Two fields report the failures that are otherwise invisible, because their only symptom is an absence: unmanned_roles lists roles no live node performs, and unplaceable lists seats whose role.placement matches no live node. A lease table that could not be read answers 503 with a Retry-After rather than an empty fleet: “no node is live” is a claim, and a store blip is not evidence for it.

Each seat row carries acquired_at — since when its node has held it, as an RFC 3339 UTC time on the coordination store’s clock. It is the tenure’s start, stamped when the lease’s epoch was minted and carried unchanged through every renewal, so it moves exactly when epoch does: on a takeover, and on the same node re-claiming after its own lease lapsed.

{
"nodes": [
{
"id": "core-1", "roles": ["ingress", "seats", "workers"], "broker": "member", "labels": {},
"owner": "core-1:8f2a", "protocol": 4, "seats": 4, "expires_in": 41.2,
"config_epoch": 7, "config_status": "ok", "config_error": ""
}
],
"seats": [
{"handle": "ceo", "node": "core-1", "owner": "core-1:8f2a",
"epoch": 4, "expires_in": 41.2, "acquired_at": "2026-09-23T08:02:11.482Z"}
],
"duties": [{"duty": "maintenance", "node": "core-1", "expires_in": 41.2}],
"unplaceable": [{"handle": "gpu-eng", "placement": "labels=gpu=true"}],
"unmanned_roles": [],
"this_node": "core-1",
"objects": {
"state": "reported",
"backend": "s3:https://s3.example.com/files/acme/",
"node": "data-a",
"collect": {
"at": "2026-09-01T12:00:00Z",
"completed": true,
"listed": 1840,
"aged": 1702,
"deleted": 12,
"referenced": 1690,
"abandoned": 1
},
"audit": {
"at": "2026-09-01T09:00:00Z",
"completed": true,
"referenced": 1828,
"missing": 0,
"damaged": 0,
"found": {
"at": "2026-09-01T09:00:00Z",
"completed": true,
"referenced": 1828,
"missing": 0,
"damaged": 0
}
}
}
}

objects is where the company’s files are kept, and what the object store’s collector last found — read from the coordination store like the rest of this answer, so every node reports the same record whichever node ran the passes. It is absent on a node that reads no record.

state names three things apart rather than drawing any of them as an empty report: unavailable (the coordination store did not answer), not_yet (no collector has finished a pass — a new fleet, or no data node running the duty) and reported. Only a report carries the other fields. The example is the engine’s own rendering: a test holds its objects block to what the renderer writes.

FieldWhat it is
backendThe store every node agreed on at boot: nats (the data nodes’ replicated bucket) or s3:<endpoint>/<bucket>/<prefix>
nodeThe data node holding the collector’s duty when it ran the passes below
collectThe last collection: when it ended (at), whether it listed the whole store and judged every object past the day’s grace (completed), and its counts — listed (the objects under the engine’s own namespace), aged (past the grace by both the store’s clock and the key’s own, so judged), deleted (no row named them), referenced (a row still did) and abandoned (uploads begun more than a day ago and never finished, which no listing shows). skipped says why it stopped judging — an estate this node could not fully read — and the counts are what it did before it stopped; sweep_error what kept it from abandoning unfinished uploads, which does not fail the collection (on S3, an identity without s3:ListBucketMultipartUploads or s3:AbortMultipartUpload); error what stopped it. Absent before the first one ends
auditThe last audit attempt, which asks the store about every object a row names: when it ended (at), referenced, missing (objects the store does not hold), damaged (objects it holds at another size, or under another digest where it keeps one), completed — false over an estate that was not complete, when the counts are floors — and error, what stopped it. Absent before the first one ends
audit.foundWhat the last audit to run to its end found — the attempt above, or the one before it when that one failed, so a failed attempt never hides what was found: at, completed, referenced, missing, damaged and missing_files, the first hundred files that cannot be read, each {"object", "named_by", "damaged"} — the object’s key, the file as PROJECT/path, and true where the store holds it wrong rather than not at all (absent when none). A non-zero missing plus damaged here raises objects_missing. Absent before any audit has run to its end

A collection runs hourly and an audit daily, on one data node at a time; a pass that fails is tried again ten minutes later.

Each node row carries broker: how the node’s broker takes part in the fleet’s — member (embedded, a voter in the JetStream cluster), leaf (embedded, joining the members through stream.leaf.urls, holding nothing), client (stream.type: nats) — or unknown for a node advertising a kind this build does not know (a newer build’s). It is derived from the node’s own stream block, never from its roles; GET /fleet/broker holds it against what the broker itself counts.

GET /fleet/broker
POST /fleet/broker/remove/{node}?confirm={node}[&force=true]
POST /fleet/broker/remove-peer/{peer}?confirm={peer}[&force=true]

Two records say who the fleet broker’s members are, and they can disagree. Every node advertises its broker kind on its presence lease; the broker’s metadata group — the raft group that places every stream and consumer — counts its voters for itself. A member that is gone for good is the disagreement that matters: its presence lapses with its process, while the metadata group goes on counting it in every election and every create until it is removed, so a three-member fleet that loses two for good has no quorum left to create anything with.

GET /fleet/broker answers both and where they disagree. Any node answers it: only a member holds the metadata group, so a node that is not one asks every live member at once, and group_from names the member whose view group is. A fleet on an external cluster answers external: true and lists no group — that cluster’s membership is its operator’s. A group no member could report within five seconds is absent with group_error saying why, never an empty one.

The group counts each voter by its raft peer id, which nats-server derives from the server’s name — the node id. A member learns another’s name only from that server itself, so a voter whose survivors have restarted since it died is listed with no name at all, by its peer alone; each node row carries the peer it is counted by, and the two records are held against each other by peer id, so such a voter is still recognised as the live node it is where one is.

{
"node": "node-a", "kind": "member", "external": false,
"nodes": [
{"node": "node-a", "peer": "ePFsSWs4", "kind": "member", "roles": ["data", "ingress", "seats", "workers"]},
{"node": "old-1", "peer": "mTNLbbt3", "kind": "unknown", "roles": ["data", "ingress", "seats", "workers"]},
{"node": "sat-eu-1", "peer": "zYqdsRpE", "kind": "leaf", "roles": ["seats"]}
],
"group": {"cluster": "acme", "leader": "node-a", "peers": [
{"name": "", "peer": "X57jblDH", "current": false, "offline": true, "active": 7200000000000},
{"name": "node-a", "peer": "ePFsSWs4", "self": true, "leader": true, "current": true, "active": 0},
{"name": "node-c", "peer": "9iReXzcw", "current": false, "offline": true, "active": 5400000000000}
]},
"group_from": "node-a",
"findings": [
{"kind": "dead_member", "node": "", "peer": "X57jblDH", "detail": "the metadata group counts it as a voter, no live node is it, and the member that answered has not heard its name since it started"},
{"kind": "dead_member", "node": "node-c", "peer": "9iReXzcw", "detail": "the metadata group counts it as a voter and no live node is it"},
{"kind": "unknown_kind", "node": "old-1", "detail": "its presence does not say what its broker is"}
]
}

active is nanoseconds since the answering member last heard from the peer. The finding kinds:

kindMeans
dead_memberA voter no live node is — its process is gone, or its node came back as a leaf or a client. It is counted in every election until it returns or is removed
not_in_groupA live node advertising a member that the group does not count: still joining, or removed while it ran, in which case it rejoins as a voter at its next restart
unknown_kindA live node whose presence advertises no broker kind this build knows — a newer build’s. It is counted as a member wherever that is the safe reading, a capacity seal included

It is the fleet_broker question, so the socket’s query channel answers it too. A lease table that could not be reached answers 503 unavailable with a Retry-After rather than a fleet with no nodes.

POST /fleet/broker/remove/{node} removes a member from the metadata group, by the peer id its node id hashes to — so it reaches a member whose name no survivor remembers, as long as you know which node it was. nats-server answers that request only on the broker’s system account, which a member reaches inside its own process and nothing else can — so the node that receives the request asks a live member, never the one being removed, and that member’s system account carries it. The answer comes once the group has committed the change, and names the voter, the member that carried it and the group as it reads afterwards:

{"node": "node-c", "peer": "9iReXzcw", "by": "node-a", "group": {"cluster": "acme", "leader": "node-a", "peers": [...]}}

POST /fleet/broker/remove-peer/{peer} is the same removal naming the voter by the peer the listing shows — the form for a voter listed with no name, which has no node id to give. Its answer’s node is empty unless a live node is that voter.

A removal is refused while the voter’s node holds a live presence lease as a member, or without saying what its broker is: the node is running, and a running member removed from the group rejoins it as a voter at its next restart. force=true removes it anyway — for a member wedged in a way that still renews its lease. A node alive as a leaf or a client under the voter’s name is removed without force: the member it was is gone for good, and its broker never rejoins. The whole gesture is bounded by one wait — the carrying member’s own commit budget and one round trip; a member that does not answer within it ends the gesture as outcome_unknown rather than another member proposing it again. crewlet fleet broker remove is a client of both routes. Operator-only, and refused during a drain. Every refusal carries detail and hint:

StatuserrorWhen
400confirm_required?confirm= does not repeat the node id or the peer id
400voter_invalidThe name is not a node id (or, on remove-peer, not a peer id)
404not_a_memberThe metadata group counts no voter by that peer id — a typo, or a removal already made
409member_liveThe voter’s node holds a live presence lease as a member (or does not say); stop it first, or force=true
409membership_changingAnother membership change is still being committed; ask again once it has
409external_brokerThe fleet’s broker is an external cluster, whose membership is its operator’s
503no_leaderNobody answered as the group’s leader for the whole wait. The request is asked again every second, so an election alone does not end here: the group has lost its quorum and can change nothing about itself. A removal whose answer was lost may still have been committed — read the group first
503no_memberNo live member could carry the removal — including when the only live member is the one being removed, which never carries its own
504outcome_unknownThe member carrying it did not answer; it may have committed the removal first — read the group before asking again
500broker_remove_failedAnything else; read the group before asking again

Every detached coding run the engine still holds, oldest first — launching, running, awaiting_clarification, reseed, answered and resumed run records. An answered run has a person’s reply recorded as the answer to its question and is owed the resume that answer drives; it is no longer waiting on anybody, so audience= — what is waiting on one person — lists only runs still waiting (awaiting_clarification or reseed).

A run that has settled, whether its turn finished or it was lost, is not listed because it has no record: its record is deleted once its box is reclaimed. How it ended is on the event stream, in the resumed turn’s own events or a sandbox_run_failed event naming the reason.

A launching run is one whose coding job has started while the turn that started it is still unwinding, so the suspended conversation a resume re-enters is not on the row yet; it is listed but never polled, because a row nobody lists is a box nobody reclaims.

This query reads the durable row directly. The live projection’s panel is reconciled against the same record every 30 seconds, but it carries only what a running-runs panel draws. This board needs the row’s own facts: the branch, the placement, the pause TTL, whether a box still exists, and the bridge’s call log. It also lists resumed runs, which the panel drops because their turn has already taken back the result. A reseed run (pause expired, box reclaimed, work preserved on a pushed branch) is listed on both.

{
"runs": [
{
"turn_id": "<uuid>", "launch_id": "<uuid>", "agent_handle": "eng", "role": "Engineer",
"status": "awaiting_clarification", "coding_agent": "claude-code",
"task_description": "Add retry to the webhook client",
"question": "Which backoff ceiling should I use?", "audience": "manager",
"audience_handles": ["founder"], "audience_fallback": false,
"branch": "crewlet/eng/retry", "trace_id": "<hex>", "owner": "core-1:8f2a",
"box_exists": true, "paused_at": "2026-06-08T07:30:02+00:00",
"pause_ttl_seconds": 3600,
"started_at": "2026-06-08T07:12:44+00:00",
"updated_at": "2026-06-08T07:30:02+00:00",
"answerable_in_chat": true
}
]
}

launch_id names the job the row holds now. A turn can launch more than one — a resumed executor that calls run_sandbox again reuses the row — and sandbox_tail is asked by it, so a run’s page polls the live output of the job it shows rather than of whichever replaced it.

box_exists and paused_at stand in for the sandbox id: a board wants to know that a box exists and that it is currently paused (and being billed for as a snapshot), not which box it is. paused_at is the answer the pause reaper acts on rather than the raw stamp, so a run parked on a question whose pause instant never reached its row still reads as held: that box is being paid for, and a board drawing the stamp alone showed it as a live one. answerable_in_chat is false for a run whose turn was triggered by something other than an inbound message — a schedule tick, a task assignment, an A2A wake — because the resume path matches an inbound message’s conversation identity against the one the run was parked with, and those runs stored no conversation at all: their trigger names neither key, so neither is stamped and neither reaches the row. Telling somebody to “reply in the thread” would send them to a thread that does not exist. Such a run is still answerable: answer_run names it by its turn_id instead.

audience is the coding agent’s own label for who should answer — requester, manager, team, or a name it typed. audience_handles is that label resolved against the org chart when the run parked: the seat whose message or ask woke the turn, the seat’s manager, its unit’s lead and the people in that unit, or the one colleague an exact match names. audience_fallback is true when the label named nobody the chart has and the question was put to the seat’s lead chain instead (its managers, or the leads of the units above it). Both are empty on a run that is not parked. See who is asked.

?audience=<handle> narrows the board to the runs whose question is put to that person — every identity their rows may carry, the seat and the operator credential bound to it. It is a filter over the board and not a personal read, so it has no scope rule of its own: the unfiltered board already names every run’s audience. A run whose question resolved to nobody — the label named no one the chart has and the seat has no lead chain, or the node held no company when it parked — matches no one.

execute_state — the serialised Execute-loop conversation — is deliberately not returned: it is the largest column in the row and every prompt in it is already reachable through the event store.

The run record lives in the fleet’s coordination store, which every node opens, so every node answers with the fleet’s runs. A company with no sandbox configured does not register the question at all, so the route answers 404 with unknown_query rather than an empty board; a record that could not be read answers 503 with a Retry-After, because “no run is parked” is a claim and a store blip is not evidence for it.

Backs the dashboard’s Budgets screen and crewlet budgets show. Every scope — the company, and each agent seat — states all three calendar windows, the day, the ISO week and the month on the company’s clock, cut at the moment of the answer:

  • used is the fleet’s shared counter for that window, in the coordination store, written by every node running the company and surviving restarts. It is what the engine actually enforces against, and it is the same counter the live token meter pushes. A window no ceiling caps is still counted, because what a seat spent this week is a fact whether or not a ceiling is written for the week;
  • limit is configuration, from the active company revision, and absent where no ceiling caps the window — never 0, which would state a range of nothing that is already full;
  • refused_at is when a capped window last turned a call away — a refused charge, or work turned away unsent because the window was already full: a turn’s next call, a parked delivery, a person’s question, a reflection pass — kept in the same counter and cleared by the scope’s next admitted charge or by the window turning over, and absent while it has not;
  • state is the engine’s judgement — refusing, near or ok, exactly as on the live meter — and near_fraction beside it is the one threshold behind near (0.9), for a screen that draws it as a mark.

Where the counter is already on a later window than the moment of the answer — a peer’s clock a few seconds ahead across a boundary, or the company’s timezone moved west — the row states the later window’s spend, because that is what the gate refuses against. Each window’s allowance comes back when it turns over; there is no route that resets a counter, and room before then is made by raising the ceiling.

What a seat spent over a window you choose is not here: that is the per-agent row of the spend breakdown, a different span that must not be divided into a ceiling. The ceiling and the durable counter are the pair that can be, which is how this screen can say “this seat has burned 94% of today’s ceiling across two restarts”.

{
"timezone": "Europe/Berlin",
"durable": true,
"near_fraction": 0.9,
"org": {
"windows": [
{"period": "day", "window": "2026-06-08", "starts_at": "2026-06-07T22:00:00Z",
"resets_at": "2026-06-08T22:00:00Z", "used": 1284410, "limit": 5000000, "state": "ok"},
{"period": "week", "window": "2026-W24", "starts_at": "2026-06-07T22:00:00Z",
"resets_at": "2026-06-14T22:00:00Z", "used": 4015220, "state": "ok"},
{"period": "month", "window": "2026-06", "starts_at": "2026-05-31T22:00:00Z",
"resets_at": "2026-06-30T22:00:00Z", "used": 9120045, "state": "ok"}
]
},
"seats": [
{
"agent_id": "<uuid>", "role": "Engineer", "handle": "eng",
"windows": [
{"period": "day", "window": "2026-06-08", "starts_at": "2026-06-07T22:00:00Z",
"resets_at": "2026-06-08T22:00:00Z", "used": 102120, "limit": 100000,
"refused_at": "2026-06-08T07:29:51Z", "state": "refusing"},
{"period": "week", "window": "2026-W24", "starts_at": "2026-06-07T22:00:00Z",
"resets_at": "2026-06-14T22:00:00Z", "used": 301877, "state": "ok"},
{"period": "month", "window": "2026-06", "starts_at": "2026-05-31T22:00:00Z",
"resets_at": "2026-06-30T22:00:00Z", "used": 702311, "state": "ok"}
]
}
]
}

durable carries the honesty. It is false when the shared counter could not be read, and every window list is then empty: a counter that cannot be read is not a counter that reads zero, and without the flag a coordination blip renders every seat at the bottom of its ceiling, which is the most reassuring possible picture drawn at the moment nothing is known. Human seats have no row, because they spend nothing.

Exhaustion is the engine’s refusing, never a ratio a client computes, so every surface and the budget park agree on which windows can take another round. A refused round is counted like any other — the vendor billed it — so a window that refused one reads past its ceiling by that round: the Engineer above had 99 120 of 100 000 when a 3 000-token round was refused, and reads 102 120.

Copies this node’s durable state — both of its store files, every JetStream stream and coordination bucket, and every object holding a company file that the store copy names — into ?dir=, a directory on the engine’s host.

Terminal window
curl -X POST -H "Authorization: Bearer $CREWLET_API_TOKEN" \
"http://localhost:8080/backup?dir=/var/backups/crewlet/2026-08-30T18-00"
{
"taken_at": "2026-08-30T18:00:00Z",
"finished_at": "2026-08-30T18:00:01.412Z",
"node_id": "node-0",
"engine_version": "v0.1.0",
"stores": [
{"estate": "node", "file": "store.db", "source": "/data/company.db",
"bytes": 258048, "sha256": "…", "migrations": ["0001_events.sql", "…"]},
{"estate": "replicated", "file": "store-replicated.db",
"source": "/data/crewlet-replicated.db", "bytes": 131072, "sha256": "…",
"migrations": ["0001_tracker.sql", "…"]}
],
"objects": {"dir": "objects", "objects": 40, "bytes": 41943040, "reused": 36,
"reused_from": "/var/backups/crewlet/2026-08-29T18-00/objects"},
"streams": [
{"name": "CREWLET_AGENT", "file": "streams/CREWLET_AGENT.snapshot",
"bytes": 1087, "messages": 5, "config": {…}, "state": {…}}
]
}

The answer is the manifest, which is also written into the directory as manifest.json — and its presence there is what marks the backup complete. A failure anywhere leaves the directory without one, because a backup missing an estate is unrestorable rather than partial.

This route exists because the state it copies is reachable only from inside the engine, twice over. The store is locked to the engine’s process and the driver refuses a second process on a database file, so nothing outside can read it; the embedded broker binds no socket, so nothing outside can reach the stream estate either. crewlet backup is a client of this route.

It is synchronous and can take a while — the duration is a property of the data, not of this handler. That is deliberate: a job outliving its request would need somewhere durable to record itself, and the only place is the store being copied. The work is safe to be cut off, since the store copy is renamed into place only after it verifies and the manifest is written last, so a client that gives up leaves an unfinished directory rather than a false one.

Four refusals, each pointing somewhere different:

  • 401 without a token. allow_anonymous_read is on by default and opens the read surface; this writes every credential the company holds to a path the caller chooses, so it is never eligible.
  • 400 for a destination this node cannot use — relative, already occupied, one this host cannot create, read or make private (a path through a regular file, a parent that does not exist or is not writable, a read-only mount), or a path the database engine mishandles. The reason is returned in detail rather than only logged, unlike every other route here, because it is the caller’s own command to fix. A disk that fails or fills while the directory is prepared is the node’s failure, not the path’s, and answers 500.
  • 503 objects_unreachable — the store copy names objects holding the company’s files that the object store did not answer for. Nothing is wrong with this node or the command: check that it reaches the store and take the backup again. A backup without those objects is refused rather than written, because they may well be intact there. An object the store answered it does not hold, or holds as bytes other than the ones the file records, is different — it is lost whatever the backup does — and is listed in the manifest’s objects.lost (each with the file that named it, as {"object", "named_by"}) rather than refused, so one lost file never stops every later backup and, with them, the trim (see Backup). On an S3 bucket each object is copied into objects/ under the bucket’s own layout and checked against its row as it is written. On the default nats backend the objects are in a stream the backup snapshots with every other (objects.stream names it, and no object is copied on its own); once the snapshot is taken the store is asked about each object the copy names, so the same refusal and the same lost list apply there too.
  • A copy without the stream estate. A node that dialled an external NATS cluster has no connection to snapshot the streams over, so its manifest carries the store copies alone and crewlet backup says where the rest lives. Back that half up at the cluster, from the same moment.

Every backup that began copying leaves a backup_requested event naming the caller, the node, the directory and whether it finished — a failed one included, since it may have left files there. A destination refused with a 400 wrote nothing and leaves no event — a 400 is only ever answered before a byte is copied; a refusal that arrives with part of the copy already in the directory is a failed backup and is recorded as one: see the runtime audit.

The request has no deadline of its own on the engine’s side, and a client should give it a long one: crewlet backup and the dashboard both wait up to 30 minutes for the answer, because the copy is bounded by the size of the store and the stream estate and a client that gave up early would report a failure while the engine finishes a good backup. Taking one from the dashboard is Settings › Backups & retention › Take a backup.

Backs Settings › Backups & retention: what the fleet has backed up. It needs a token, reads included — every row names a directory on a named host that holds the company’s sealed credentials — and api.allow_anonymous_read does not open it.

Two records, because they answer two questions:

  • points — each owner’s NEWEST backup, from the fleet’s backup register: what each node announced when its manifest was written, plus the operator’s acknowledgement (crewlet retention ack, kind: "operator"). counted says the trim may count it — the backup_floor policy (policy) takes this owner’s word and the copy was verified — and exactly one counted point is newest: the one the trim’s backup term reads. bytes is the whole artefact (every database copy and stream snapshot) and is absent on an acknowledgement, which asserts a copy the engine never saw. covers is how far the copy reaches in each state-log stream.
  • history — every POST /backup a person made, newest first, from the runtime audit every node keeps for the event log’s 30 days: when it finished, the node whose disk holds it, who asked (operator, and the bound person in actor_seat), the dir, and outcome — applied (the manifest was written) or failed (it was not: the directory holds debris, not a backup). At most 100 rows; more says the page filled. coverage names the nodes the history was read from, since a node that did not answer takes its backups’ rows with it.
{
"policy": "engine",
"points": [
{"owner": "node-a", "kind": "node", "taken_at": "2026-09-30T02:00:00Z",
"dir": "/var/backups/crewlet-20260930-0200", "verified": true, "bytes": 83886080,
"covers": [{"stream": "CREWLET_TRACKER_LOG", "generation": 2, "seq": 9001}],
"counted": true, "newest": true}
],
"history": [
{"id": "6ac15845-97c8-4ac2-9ac8-bfa7729a3572", "at": "2026-09-30T02:00:21Z",
"node": "node-a", "operator": "founder", "actor_seat": "jane-founder",
"dir": "/var/backups/crewlet-20260930-0200", "outcome": "applied",
"summary": "founder (jane-founder) backed up to /var/backups/crewlet-20260930-0200 (25 streams)"}
],
"more": false,
"coverage": {"nodes": [{"id": "node-a", "answered": true, "error": ""}], "complete": true}
}

Backs the dashboard’s Integrations screen: how each external surface is wired, and what has come through it.

Every count is over one window, from traffic_since (inclusive) up to the instant the answer was read (exclusive), and traffic_since is what makes the counts a measurement: “42 inbound” alone could be an hour or a year; “42 since Tuesday” is not. The window is the widest one whose every delivery the answer holds. The deliveries are page-capped, not time-bounded — the most recent page of the delivery log, at most 400 deliveries across the fleet — and which window that is depends on whether any delivery may lie past the page, never on how long the page is:

  • Nothing lies past the page when no node holds 400 or more deliveries in the 30-day history, no node’s reply had to be cut to fit the transport, and the fleet holds no more than 400 between them. The page is then every delivery the history keeps — it can hold exactly 400, two nodes holding 200 each — and the window is the whole history: traffic_since is 30 days before the instant the answer was read, so that instant is traffic_since plus 30 days. An empty page is this case too: no delivery in 30 days is a count of 0 over those 30 days, beside the drops and merges of the same 30 days.
  • Deliveries may lie past the page otherwise — the fleet holds more than 400, or one node holds 400 or more (its own page filled, which it reads as more behind it even at exactly 400), or a node’s reply was cut to fit the transport (a page that can then be shorter than 400). The window reaches no further back than the page does: traffic_since is one microsecond after the oldest delivery the page reached. Every later delivery is on the page, but one sharing that instant may not be, so the deliveries at that instant are left out of the window rather than counted short.

traffic_since is null only when no event log could be read, which traffic_known: false says. It is written to the store’s own resolution (RFC 3339 with fractional seconds, at most microseconds), because an edge cut to the second would name a window up to a second wider than the one counted; last_at is written the same way.

Integrations had close to no surface at all before this. The dashboard branded an event once it had already been accepted and routed, so every failure mode an operator actually hits was invisible — a Mattermost SiteURL that blinds every browser while agents keep working, a revoked bot token, a mis-pasted webhook secret. Rejected deliveries are deliberately never written to the event store (verification runs before the row is logged, which is correct), so a signature mismatch left no trace anywhere except the provider’s own delivery UI.

{
"traffic_known": true,
"traffic_since": "2026-06-07T09:12:00.418207Z",
"integrations": [
{
"key": "gitlab", "configured": true, "enabled": true,
"url": "https://gitlab.example.com",
"inbound_kind": "webhook", "inbound_path": "/webhooks/gitlab",
"routes": true,
"secret_present": true, "secret_usable": true,
"seats": ["eng", "pm"],
"inbound": 128,
"skipped": 30,
"coalesced": 2,
"last_at": "2026-06-08T07:31:10.052914Z"
}
],
"tools": [
{
"key": "gitlab", "surfaces": ["gitlab"],
"state": "attention", "label": "Credential expiring",
"reason": "the group Owner token this integration runs on expires on 2026-06-20, …",
"surface": "gitlab"
}
]
}

tools is one roll-up per tool this build serves — slack, mattermost, atlassian (the organization, Confluence, Jira and the Forge relay), github, gitlab, datadog — whether or not the company configured it, so a reader never invents a state for a missing one. state is attention (a person has to act), not_connected (configured and not working yet, with nobody owing anything — or this node could not read the status), connected or not_in_use (no block, or every block switched off; label says which). label is the state in a reader’s words, reason one sentence on why, and surface the surface it was taken from. The rules are One state per tool.

A reconcile finding of kind credential_expiring carries expires_at, the instant the credential stops working.

inbound, skipped and coalesced answer one question together and are misleading apart. inbound counts deliveries — one delivery presented to one seat, counted once across the fleet — that arrived at this row’s own ingress (inbound_path, or the websocket for Mattermost): a message two agents’ Slack apps both receive is two, a Mattermost post two bots’ sockets both read is two, and the same post read by every node holding the bot’s socket is still one, because the fleet claims it before it is published. Counted by ingress, not by integration: the Forge relay hands Jira and Confluence events on under its own token and they belong to those products, so a relayed event counts once, under forge, and the Atlassian card, which sums its surfaces, counts it once — while what became of it (skipped, coalesced) counts under the product it was parsed as. skipped counts those the routing gate dropped without waking anybody; coalesced counts merges, where N same-conversation notifications became one turn. “128 arrived” on its own cannot tell a working integration from one whose every delivery reaches nobody — “128 arrived, 30 dropped, 2 merges” can, and a seat draining a thread’s backlog as one turn stops looking like a seat that ignored twelve messages. All three cover the window above, and so does last_at, the newest delivery in it (null when it holds none), never earlier than traffic_since.

The two outcome counts are three-valued like the secret fields: null means this node could not read its event log, and reporting that as 0 would claim every delivery woke a seat on a node that cannot tell. They come from the engine’s own notification_skipped and notifications_coalesced events rather than from the inbound rows, and they are bounded by the same event-log window traffic_since names: every node counts its own outcome events from traffic_since up to the instant the inbound rows were read — the same instant every node floors the 30-day history at — and the node you asked sums them, counting once a row two data nodes hold while a stateless node’s custody batch is still being settled (Reading the fleet’s history). They are counted, not paged, so every outcome in the window is in the count however many more of them there are than deliveries, and one before traffic_since or written after the answer was read is not. A row counts under the third-party app its notification_source tag names; one whose event names no app carries none and is not counted. The window is one for the whole answer, so a surface whose drops and merges outnumber its deliveries has its outcomes counted over the window every surface’s deliveries leave: the whole 30 days whenever nothing lies past the page (a company receiving nothing at all included), and only as far back as the page reaches when the busiest surfaces’ traffic fills it.

secret_present and secret_usable are two different facts, and the gap between them is a silent outage.

secret_present is a claim about the document: an operator wrote a secret down. It is three-valued because the cases mean opposite things: null — this surface does not use one (Mattermost authenticates its websocket with the bot’s own token, and Atlassian receives no delivery to verify); false — it does, and none is configured, which means the webhook route answers 503 to every delivery.

secret_usable is a claim about what this process resolved. A secret lives in the config as a ${VAR}, so secret_present: true, secret_usable: false is a route refusing every delivery while the config shows a secret and the third-party app’s settings page shows a healthy hook, with nothing anywhere naming the variable. For GitLab the bar is higher than non-empty: the value must be whsec_ over standard base64 of a 32-byte key, the only shape the third-party app signs with. For Slack, whose material is one signing secret per seat, it is lower: one seat whose secret resolved makes the surface usable, because a delivery addressed to that seat’s path would be accepted, and a seat whose own secret is unresolved is reported by that seat’s identity finding rather than by the whole surface. null means this node cannot say (nothing has resolved yet), or the surface has no secret to resolve.

Only the booleans are ever returned; no secret value leaves the process.

A row exists for a surface the company’s document turns on, and for no other. That sounds like a restatement of “the block is present”, and for six of the eight it is. Slack and GitHub are the two where a seat carries its own app — its own credential, its own inbound path — and both used to be reported on those per-seat values as well as on the company block, so that a company holding nothing but per-agent apps still had a row.

Neither half of that survives contact with what the engine does. The parser that turns a verified delivery into work for a seat is registered only where the company block is present and enabled: drop integrations.slack and the transport is retired (slack_retired); drop or disable integrations.github and the same happens (github_retired). A seat app without it delivers to a route that verifies the signature and then has nowhere to send it. So a company in that state is not a surface missing a row — it is a surface that is off, and a row for it is a row with enabled: false, which the dashboard draws as Paused.

That matters because enabled: false is the one field on this row that claims somebody’s intent. A disconnect produced the other reading every time: the block went, the seats kept their sealed credentials, and the card an operator had just disconnected settled on Paused — a word for a state nobody had chosen — and stayed there. enabled: false now means a block that says enabled: false, on every surface, and a company whose block is gone gets no row and a Connect button. What each seat is still holding is on the setup screen’s seat roster, which is where a seat is acted on.

seats lists the agents carrying their own identity on that surface: a Slack app, a Mattermost bot, a per-seat project or space, wherever they sit in the hierarchy. A seat in a unit is a seat: the list walks the whole tree, not just the top-level roles: block, which is by definition the seats belonging to no unit.

routes is the third of the same family: whether a verified delivery would wake a seat. The three fail independently, and an operator staring at a silent integration needs to know which half broke.

It is null for a surface nothing ever arrives from, which is a different answer from false and the only honest one. Atlassian is that surface: an organization is where an agent’s account is created, and the products that account then works in — Jira, Confluence — are separate surfaces with their own webhooks and their own parsers. Reporting routes: false there described a real fault (“deliveries are verified and stored and no parser turns them into work”) about a surface that is not asked the question, beside a secret_usable: false for a secret it does not have. Both are null now, and so are inbound_kind and inbound_path — the row used to name /webhooks/atlassian, a route this engine does not serve, as the address to check a settings page against.

Mattermost is false rather than null when it does not route: it has no inbound address because the engine dials out, but everything said in its team arrives, so a missing parser there is the outage the field is for.

A surface is asked about the source its deliveries are published as, which is not always its own name. The Forge relay is the case: a Cloud event it relays is republished as the product it belongs to — jira or confluence — and parsed by that product’s parser, so nothing is ever registered under forge. Asked about itself the relay answered false on every Cloud deployment for ever, and because the dashboard groups it under the Atlassian row, a tenant whose relay was feeding both products correctly carried a permanent Forge relay — routes nowhere beside the two rows saying they routed fine. It now answers true when either product’s parser is registered; the finer answer is on those two rows, immediately below it.

Health is deliberately not inferred. An idle Slack and a 401-ing Slack are indistinguishable in the event store, so silence is reported as “no traffic seen” — never as healthy, never as down. traffic_known is false when this node could not read the event log at all, so the zeros below it are absence of measurement rather than measurement of absence.

Backs the dashboard’s Schedules view. Returns every configured role/unit schedule with its cron, effective timezone, target → resolved runner handles, and a per-request next_run (computed from the cron), plus the most recent rows from the scheduled_runs dispatch ledger.

{
"schedules": [
{
"scope_type": "unit", "scope_id": "Backend", "name": "daily-standup",
"cron": "30 9 * * 1-5", "timezone": "Europe/Amsterdam",
"task": "Post your standup…", "target": "each",
"enabled": true, "timeout_seconds": 180, "catchup": true,
"runners": ["backend-lead", "backend-dev"],
"next_run": "2026-06-09T07:30:00+00:00"
}
],
"recent_runs": [
{
"scope_type": "unit", "scope_id": "Backend",
"schedule_name": "daily-standup", "target_handle": "backend-dev",
"scheduled_at": "2026-06-08T07:30:00+00:00",
"fired_at": "2026-06-08T07:30:02+00:00", "outcome": "fired"
}
]
}

A run’s outcome is fired, skipped_catchup (a missed tick outside the catchup window) or skipped_paused (the runner seat was paused when the fire came due). recent_runs is empty when the dispatch ledger cannot be read (the configured list and next_run still render). A schedule with no next fire — disabled, an unparseable cron or timezone (problem says which), or a date the calendar never reaches — carries no next_run key, never a zero instant.


Two sources, one aggregation, and which one answers is decided by the parameters:

  • The live window — a request naming no days, no dates, no seat and no previous — is the projection’s: the spend records of the last 24 hours (livestate.LiveSpendWindow, rolling), held in memory and pushed as the tokens push, for a client that wants the last day as it happens. No screen of the bundled dashboard draws it: the Spend screen reads named windows only, so its figures are the company’s. It is the only answer with a per-turn tail (by_turn) and a watermark (aggregated_through).
  • Every named window is whole company days read from the replicated usage domain (ADR-0020): every node’s day, applied on every node. So the answer is the same whichever node is asked, reaches back 181 days (the domain’s history, not the event log’s 30), and still counts a node that has left the fleet. It replaced a scan of the answering node’s own event log, which reported a third of a three-node fleet’s spend under the company’s name and drew a ninety-day chart over thirty days of rows.

Both are folded by internal/tokens, so a reader moving between the two compares like with like. The guide Budgets and spend explains the windows, the counter a budget enforces and how it differs from this rollup.

The window parameters, shared by both routes:

NameDefaultDescription
days1 on a named windowThe company days ending today, on the company’s clock: 7 is today and the six before it. 1 to 90 (tokens.MaxSpendRangeDays); anything else is 400 (tokens.ErrWindowLength) naming days.
since / until—Instead of days: two company dates, 2026-06-01, both inclusive — since=2026-06-01&until=2026-06-08 is eight days. A pair or neither; at most 90 days — a longer pair is 400 (tokens.ErrWindowLength: 2026-05-01 to 2026-09-29 is 152 days, and a spend window is at most 90 — bring since and until closer together) wherever in the history it lies; never together with days.
previousfalseThe same number of company days ending the day before the window begins — compare-to-previous, cut on the company’s calendar rather than a browser’s, so the two windows are never different weeks.
seat(every seat)One seat, by its handle. Matched on the agent id every node derives from the org name and the handle, so a seat since removed from the chart still answers for the days it left behind.

A window whose first day — or whose previous window’s first day — is older than the history’s floor is 400 (tokens.ErrOutOfRange, a different class from a window that is merely too long) naming the parameter to change, never answered short: the rows before the floor are gone on every node, and a heading over fewer days than it names is a lie about the numbers under it. At 90 days the previous window begins 179 days back, inside the 181.

The rollup: the window’s spend by phase, model, provider entry, worker and seat.

Response (a named window)

{
"since": "2026-06-08T15:00:00Z",
"until": "2026-06-15T15:00:00Z",
"from": "2026-06-09", "to": "2026-06-15", "days": 7,
"horizon": { "days": 181, "floor": "2025-12-16" },
"totals": {
"input_tokens": 17700, "output_tokens": 2750,
"total_tokens": 20450, "calls": 6,
"cache_read_tokens": 12100, "cache_write_tokens": 900
},
"by_phase": [
{ "phase": "execute", "input_tokens": 14000, "output_tokens": 2000,
"total_tokens": 16000, "calls": 2 },
...
],
"by_model": [
{ "model": "claude-sonnet-5", "total_tokens": 19300, "calls": 4, ... },
...
],
"by_provider": [
{ "provider_key": "anthropic", "models": ["claude-sonnet-5", "claude-haiku-4-5"],
"seats": ["pm", "coder", "reviewer"], "seats_total": 5,
"total_tokens": 20100, "calls": 5, ... }
],
"by_worker": [
{ "worker": "persist_decider", "total_tokens": 900, "calls": 1, ... }
],
"by_agent": [
{ "role": "PM", "handle": "pm", "agent_id": "<derived uuid>",
"turns": 4, "failed": 1,
"total_tokens": 20200, "calls": 5, ...,
"by_phase": {
"execute": { "total_tokens": 16000, "calls": 2, ... },
"review": { "total_tokens": 1400, "calls": 1, ... },
"auxiliary": { "total_tokens": 900, "calls": 1, ... }
}
},
...
]
}

Notes:

  • since/until are the window as instants — the first instant of its first company day and the first instant after its last, until exclusive — and from/to/days name the same window by its days. The live window carries only the instants.
  • horizon states how far back a named window can reach: the history in days and floor, the oldest company day still answerable. It is named horizon rather than coverage because nothing here was asked of a node — the rows are replicated whole, and what bounds them is time, not presence.
  • by_phase covers every phase the Turn Engine emits (onboarding, execute, review, subagent, judge, sandbox), as recorded, and auxiliary — what the seats’ auxiliary model spent, from the auxiliary_spend records: the turn-start context, every compaction, the reflection workers, the background learning passes and a person’s answered question. The series folds these into four bands; the rollup does not.
  • by_worker breaks the auxiliary phase down by what it was FOR — its purpose: memory_filter, knowledge_query, episode_summary, persist_decider, counterparty_profiler, skill_synthesizer, skill_refiner, episode_compaction, skill_clustering, skill_promotion, answer_knowledge, and condense_<kind> for each kind of text a compaction rewrites. Its rows sum to that phase.
  • calls is PROVIDER CALLS on every bucket: a phase’s model rounds, an auxiliary record’s coalesced calls. It counted records once, which made a forty-round executor one call.
  • by_provider answers “which configured entry (providers.llm.<key>) do we pay for”, which by_model cannot: a fallback chain serves several models under one key. models is every model the entry answered with, biggest first; seats the three handles that spent the most through it and seats_total how many did at all. A call recorded before the key was promoted (node migration 0032) is under unknown.
  • by_agent[].turns and failed are how many of the seat’s turns ENDED in the window, and how many of those failed — a named window only. The live window holds phase records, not endings, so it carries neither rather than a count of “turns that spent”, which is a different number. A seat that ended a turn without spending is still listed.
  • by_turn — one row per RUN, newest first, capped at 50, each the turn’s COST: its phases and its in-turn auxiliary records, never the reflection after it, which is the seat’s learning rather than the work’s cost — and aggregated_through, the newest record counted, are the live window’s only. A company day holds no turn and no per-call instant, so a named window has neither. Per-turn spend over any window is GET /turns?sort=-tokens. Both carry each record’s stamp as it was published, so a record from a node whose clock runs fast puts them past until for the day the window holds it — which is how that node is found.
  • On a named window a seat is one row per derived agent id, named by the newest day’s record: a role renamed mid-window is one row under its current name. The live window keys a seat on its role, as the live projection keys every seat’s state, so a role renamed inside the last day is two rows until the old name’s records age out; a row’s agent_id is the one its newest record carries.
  • A person is a row of their own with person: true, on both windows: what the auxiliary model spent for a human seat — a question answered with answer_knowledge, a background pass on a unit a person leads — named by that seat’s handle and the role its newest record or day names (a record naming none leaves it), with no agent_id and, on a named window, no turns or failed, since a person takes no turns. seat=<handle> narrows a window to one person as it does to one seat. A person’s spend reaches the named windows as its own usage record, published with the rest of the node’s day.
  • All lists are sorted by total_tokens descending, ties on the name.
  • Every bucket also carries cache_read_tokens and cache_write_tokens: the share of input_tokens the providers’ prompt caches served and stored. A breakdown of the input, never an addition to it — input_tokens already counts the cached prefix on every backend, so the cache’s share of a bucket is cache_read_tokens / input_tokens, and total_tokens stays input plus output. Both sources carry them, so a window reads the same share whichever answered it. A phase recorded by a build that did not count the cache reads 0.
  • Every live-window bucket also carries cost_usd and priced_calls. Two numbers, because zero dollars is two different facts: only a subscription coding CLI reports a price, so a cost_usd of 0 over priced_calls: 0 means nobody said what this cost, while 0 over 3 means three runs were billed nothing. The usage domain does not carry a price, so a named window’s are zero over zero. The dashboard reads neither field: it measures spend in tokens and never in money (rule 19).

The same spend with a time axis: one bucket per company day or ISO week over a named window, each split into bands on one dimension. There is no live path — its buckets are company days, which only the usage domain holds.

Query parameters — the window parameters above, plus:

NameDefaultDescription
groupphaseThe dimension the bands are: phase, model, provider, seat, unit or worker. Anything else — turn included — is 400 naming the set. phase is the four bands below; seat is keyed and labelled by handle; unit resolves through the org chart’s DIRECT unit for each seat’s handle — not the chain, because a band per nesting level would count the same spend for the team and again for the department above it. A usage row carries no project and no work item; that attribution is the tracker’s own per-item counters.
bucketdayday (a company day) or week (the company’s ISO week, from Monday midnight on its clock). Nothing finer: a day is the smallest thing every node’s usage agrees on, and hour is 400.
groups4How many bands before the rest fold into the residual. Capped at 20. Four is how many data hues the design system has, and exactly the phase breakdown’s band count — a phase grouping never folds.

The four phase bands, folded once in tokens.PhaseBand:

BandPhases
executeexecute, sandbox (a detached coding run is the executor’s own work done elsewhere)
reviewreview
workerssubagent — the workers an executor delegated to
auxiliaryauxiliary, judge, onboarding, and any phase this build does not know

Response

{
"group": "phase",
"bucket": "week",
"since": "2026-06-09T15:00:00Z", "until": "2026-06-23T15:00:00Z",
"from": "2026-06-10", "to": "2026-06-23", "days": 14,
"horizon": { "days": 181, "floor": "2025-12-24" },
"series": [
{ "at": "2026-06-07T15:00:00Z", "window": "2026-W24", "days": 5,
"total_tokens": 80, "calls": 1, ...,
"groups": { "execute": { "total_tokens": 80, "calls": 1, ... } },
"other": { "total_tokens": 0, "calls": 0, ... } },
{ "at": "2026-06-14T15:00:00Z", "window": "2026-W25", "days": 7, ... },
...
],
"by_group": [
{ "group": "execute", "other": false, "folded": 0, "total_tokens": 16000, "calls": 2, ... },
{ "group": "review", "other": false, "folded": 0, "total_tokens": 300, "calls": 9, ... }
],
"totals": { "total_tokens": 20450, "calls": 6, ... },
"grouped": { "total_tokens": 20450, "calls": 6, ... }
}

Notes:

  • Every bucket in the window is present, including the empty ones. A series with holes is a chart the client has to repair. A quiet day is a gap of full height, not a column the chart squeezed out.
  • at is the bucket’s start and window its label on the company calendar (2026-06-14, 2026-W25). A week’s at can be before the window’s since: the first and last weeks can be partial, and days says how many of the window’s days each bucket holds.
  • by_group is the legend and the grid: each band’s total over the whole window, with the residual last. By phase ALL FOUR bands are listed, in the stacking order above whatever their size — a band nothing spent in is there at zero, so the legend is the same four every time; every other grouping is biggest first and lists only what spent. Which bands survive the cap is decided over the WHOLE window, never per bucket — a per-bucket decision would put a band in the chart for the days it happened to lead and in the residual for the rest.
  • A unit band carries seats, how many seats spent in it; a seat band carries its handle.
  • The residual carries an empty group and other: true, rather than a reserved name: a model or seat genuinely called other must not be mistaken for the fold. folded is how many distinct groups it stands for, and it sums every field of the bands it stands for, the cache counts included.
  • totals is every cell in the window, including the ones this grouping places in no band at all — grouping by worker leaves out every phase that is not a worker’s. grouped is what the bands do cover, so the gap is a number rather than an inference a reader has to make by subtracting.

Every delivery — one delivery presented to one seat, counted once across the fleet — writes one row to the event log under category: "webhook", so the listing answers “what has been arriving” without reading a payload. Both inbound edges write it the same way: a webhook route for each delivery it verified, and the Mattermost socket for each post it claimed. Each publishes an inbound_delivery event, which the node’s own event log files — or, on a node without data, a data node’s, through custody — so a delivery read on any node is a row:

FieldWhat it carries
typeThe delivery’s own label: webhook:<event> from a webhook route, forge:<event> for an Atlassian Cloud relay, socket:posted for a Mattermost post
sourceThe integration the payload belongs to — the route for six of the seven webhook routes, the relayed product for Forge, and mattermost for a post
summaryThe delivery in one sentence
tags.recipientThe seat a per-seat delivery was addressed to — a Slack or GitHub app’s seat, or the bot whose socket read a post — absent for a company-wide one. It is one of the four keys the log indexes as a party, so GET /events?agent=<handle> also returns what reached that seat from outside
tags.delivery_keyThe provider’s own delivery id — a delivery header, or a Mattermost post id — absent for the providers that send none (the Forge relay, Datadog, and an Atlassian build without the identifier header, which this edge deduplicates on a hash of the body instead): what an operator has in front of them in the provider’s console
tags.routeThe ingress that authenticated it, which is not always its source: a Forge-relayed Jira event is source: jira, route: forge. The integrations answer counts inbound by it
payloadThe raw body the provider sent (a Mattermost post as its socket carried it), on GET /events/{event_id} only. A listing never carries a payload, so a deliveries screen is one request rather than one per row

Note that a row exists only for a delivery that was verified, claimed and queued. A refusal — a bad signature, an unset secret, a body too large — is answered at the edge and appears in the engine’s log rather than here; a post a peer node already claimed is that node’s row; and a post whose publish failed is read again rather than recorded.


Carried by GET /tools, by the tools push, and inside the socket snapshot. One row per registered tool:

{
"name": "post_message",
"description": "Post a message to a channel",
"source": "slack",
"title": "Post message",
"annotations": {
"read_only": "no",
"destructive": "no",
"idempotent": "unknown",
"open_world": "yes"
},
"delivers": "slack",
"input_schema": { "type": "object", "properties": { "…": {} }, "required": ["…"] }
}
  • Every hint is a WORD, never a bool, and the third word is the point: unknown means the server did not advertise the hint, which is a different fact from no. A bool cannot hold the difference — an absent hint would arrive as false and read as a positive denial, so a fresh MCP server’s unannotated tools would look like proven reads on the one screen an operator audits them on. The engine’s own delivery fence has always read them this way (see Registry.KnownReads): an unannotated tool is not a known read.
  • delivers names where calling this tool puts something in front of somebody outside the turn, and is empty for a tool that reaches nobody. It is the registry’s own predicate, not “was this served by MCP”: a proven read-only MCP tool delivers nowhere, and the native tracker’s comment tool delivers although it is a builtin.
  • title is the human-readable name a server advertised, omitted when it advertised none.
  • input_schema is the JSON Schema the model is offered, verbatim. It is absent when the tool takes no arguments, which is not the same as {}: only the first means this build did not send one.

Receives Data Center Jira webhook payloads (issue created, updated, commented, assigned). Verifies HMAC-SHA256 over the raw body against X-Hub-Signature, keyed on integrations.jira.webhook_secret; a route with no resolved secret answers 503 rather than accepting the delivery. Deduped on X-Atlassian-Webhook-Identifier, which is stable across Jira’s own retries. Jira Cloud does not use this route — a Cloud webhook belongs to an app, so those events arrive through /webhooks/forge with their own JWT, and webhook_secret is unused there. Publishes to crewlet.notifications.inbound. See Jira Integration — Webhooks.

Receives Slack Events API payloads for a specific agent (identified by handle). Verifies the signing secret for that agent’s own app — Slack gives each seat its own, so the handle in the path is what selects the key. Publishes to crewlet.notifications.inbound. Slack’s url_verification challenge is answered unconditionally (no engine or company config needed), so a freshly provisioned app’s Request URL verifies even before the engine is configured — it has to, because during provisioning that app’s signing secret does not exist yet. See Slack Integration.

The OAuth install landing page for crewlet slack provision. Every provisioned Slack app has this as its OAuth redirect URL. After the operator approves an install, Slack redirects here with a temporary code (and state carrying the agent handle); the page displays the code for pasting back into the waiting CLI prompt. Unauthenticated by design: the code expires after 10 minutes and is useless without the app’s client secret, which only the provisioning CLI holds. Every value on the page comes from the query string, so it is served under a policy that allows its one inline style by hash and no script at all (see Security headers on every response).

Receives GitHub webhook payloads. Verifies HMAC-SHA256 over the raw body against the x-hub-signature-256 header, keyed on the required webhook_secret from the github config block; invalid or missing signatures are rejected with 401, and a route with no resolved secret answers 503 with a Retry-After so the delivery is held for retry rather than blamed on the sender. Deliveries are deduped on X-GitHub-Delivery, which is stable across GitHub’s own retries and an operator’s manual redelivery. The event name is in the X-GitHub-Event header, not the body — the payload carries only the action — so the header is carried onto the envelope and read by the parser. Publishes to crewlet.notifications.inbound. The same handler serves POST /webhooks/github/{handle}, which is the address a seat’s own GitHub App is created with: the handle travels onto the published event so five agents’ apps reporting one comment are five wakes rather than four duplicates, and both forms verify against the same company webhook_secret, so a seat in the path is not a way past the signature check. See GitHub Integration — Webhooks.

Where GitHub returns an operator’s browser during the per-agent GitHub App flow, and one of the two /webhooks/* routes that render a page rather than accept a delivery (the other is /webhooks/slack-oauth). Two arrivals, one route: after the app is created, with a one-time code to convert, and after it is installed, with nothing but ?installed=<handle>. Unauthenticated, because a redirect from GitHub carries no engine credential; the state minted by POST /setup/integrations/github/app stands in its place and is a signed token naming the seat, checked before the code is converted — and spent there, so a state that reached a browser history or an ingress access log cannot be presented a second time within the fifteen minutes it stays valid. The spend goes through the fleet’s claim registry, so it holds when the two halves of the flow land on different nodes, and it fails CLOSED: a registry that cannot answer has not said the link is unused. The install arrival carries no such proof and so changes nothing — it renders a page and no more; the reconcile loop is what records the installation, from GitHub’s own list rather than from the query. On a conversion the app’s private key and webhook secret are sealed before anything else can fail, because GitHub returns both exactly once and reissues neither. Answers 200 for a completion or an install, 400 for a refusal from GitHub, a missing code, or a state or code the engine will not accept, and 503 when this process has no setup surface. Error text is always the engine’s own wording: GitHub’s response body here carries the private key. The page runs only its own inline style and install countdown script, allowed by hash (see Security headers on every response).

Receives GitLab webhook payloads. The signature is the only credential: webhook-signature is verified as a Standard-Webhooks HMAC-SHA256 over {webhook-id}.{webhook-timestamp}.{body}, keyed on the signing_secret’s base64 payload, constant-time against any of the header’s space-separated v1,… entries, with a ±5-minute timestamp tolerance. A missing or wrong signature is rejected with 401 — the plaintext X-Gitlab-Token is not accepted, so omitting the signature header is not a downgrade path. Answers 503 with a Retry-After when no signing_secret is configured, or when its value is not a usable whsec_ key, so the delivery is held for retry rather than blamed on the sender. GitLab signs whenever the hook has a signing_token (GitLab 19.1+); see GitLab § Verification. Publishes to crewlet.notifications.inbound. See GitLab Integration — Webhooks.

Receives Data Center Confluence webhook payloads (page created/updated, comments). Verifies HMAC-SHA256 over the raw body against X-Hub-Signature, keyed on integrations.confluence.webhook_secret; a route with no resolved secret answers 503. Confluence Cloud does not use this route — those events arrive on /webhooks/confluence/{event} or through /webhooks/forge, which is why webhook_secret is required on Data Center and unused on Cloud. Publishes to crewlet.notifications.inbound. See Confluence Integration.

Receives one Confluence Cloud event, named by the path because a Cloud payload does not say which event fired — the registered URL is the only thing that knows. Cloud signs nothing and honours no registration field for a header, so the authentication is a shared token, compared constant-time: X-Crewlet-Token is read first and ?token= in the query is the fallback, which is where crewlet confluence provision puts it. A route whose integrations.confluence.webhook_token is unset answers 503, and so does one whose token is shorter than 26 characters — the token is the entire check, so its length is the entire strength. The engine never logs the query string on this route. Deduped on a hash of the raw body, because Cloud sends no per-delivery identifier. See Confluence Integration — Webhooks.

Receives a Datadog monitor alert. Datadog’s Webhooks integration attaches custom headers with fixed values only, so there is nothing varying with the payload to sign and the authentication is a shared token, compared constant-time against X-Crewlet-Token and keyed on integrations.datadog.webhook_token. An unset token answers 503, and so does one shorter than 26 characters. A mismatch is 401. Deduped on the payload’s own id, which Datadog repeats across its retries. The alert routes by the monitor’s TAGS — crewlet:<handle> by default — falling back to integrations.datadog.route_to, because a monitor is addressed to nobody. Publishes to crewlet.notifications.inbound. See Datadog Integration.

Receives events from the Atlassian Forge app. Every request must carry a Forge Invocation Token (FIT) as an Authorization: Bearer JWT; the token is verified against Atlassian’s JWKS endpoint and its aud claim must match the configured forge_app_id (401 on failure, 500 when no app id is configured). The request body is drained before FIT verification — verification can block on a JWKS fetch, and the body must be off the socket before the sender’s delivery deadline aborts the request. Maps avi:jira:* / avi:confluence:* events onto the native Jira/Confluence pipeline and publishes to crewlet.notifications.inbound. Self-generated events (an agent’s own actions echoed back by Forge) are acknowledged and dropped. Jira Cloud and Confluence Cloud both ride this route and are served end to end — see the integration pages.

Webhook senders enforce delivery deadlines and abort requests that respond too slowly. When a sender hangs up before the request body is fully read, the read fails part way: there is nothing to verify and nobody left to tell, so the receiver logs webhook_body_unreadable (component=api.webhooks, keyed by path and error) and still writes a 400 — a handler that returns without writing one answers 200, telling the sender a delivery it abandoned was accepted. The aborted delivery is dropped, and whether it is redelivered is up to the sender’s retry policy, so recurring webhook_body_unreadable warnings on a webhook path mean events are being lost because the API is answering too slowly.

The body is read whole even when the request will be refused, and bounded at 25 MiB (body_too_large, then 413 — every JSON surface answers a 413 with that one code) and at 30 s (see Request timeouts). Answering without draining leaves unread bytes in the socket and the sender sees a connection reset instead of the status — which for a 401 reads as “retry forever” rather than “your signature is wrong”.


Terminal window
crewlet run -config crewlet.yaml -roles data,ingress -api-host 0.0.0.0 -api-port 8000

The API is served inside the engine’s own process, over the engine’s own store, broker and coordination plane. Its reads answer from that node’s store and projection. Its writes are few and each is named above: the operator surface (/operator/mcp, /operator/act) writes the tracker and knowledge base through the tools a seat holds, /config, /secrets and /setup write the company’s configuration and credentials, /backup and the /work/retention gestures act on the node and the log, and the webhook edge publishes inbound deliveries onto crewlet.notifications.inbound. It runs no agent itself — a seat’s turn is the engine’s.

See Deployment for how the API and engine run together, and the integration docs (Slack, Jira) for webhook setup.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.