Skip to content
You are reading documentation for unreleased main. Read the 0.1 version.

Agent Runtime

The agent runtime (internal/agent, wired together by internal/engine) runs a seat’s turns: what wakes a seat, what a turn is handed, and how the process stops without abandoning one.

Per-turn execution: every agent turn runs through the two-stage Executor and Reviewer Turn Engine. A seat’s inbox partition reaches engine.Dispatcher.Dispatch, which screens it (ownership, posture, duplicates, the completion ledger), merges a coalesced conversation, restores the trigger’s trace and hands one request to the turn. The turn pins the epoch, builds the seat’s runner, runs the onboarding pass when one is due, and then calls turn.Run, which owns the rounds: the delegation-depth check, the wall-clock cap, the engine’s own delivery check, the stall guard and the iteration cap. The sections below describe the surrounding lifecycle; the turn-engine doc describes what happens inside a turn.


Each agent seat (kind: agent, the default) is one roles: entry, and there is no long-lived agent object behind it. The authored config.Role becomes an org.Role in the epoch’s Organization, and every turn builds a fresh runner for that seat from the epoch it pinned. Human seats are never run: they exist in the Organization and resolve through the party registry (notify.Registry).

Identity is deterministic. A seat’s agent id is org.DeriveAgentID(company name, handle): a UUIDv5 over "<company name>:<handle>" in a fixed namespace (org.Organization.AgentIDFor applies it to an agent seat). The same seat in the same company lands on the same UUID across processes, machines and restarts, which is what lets any node address a seat another node is running. The seat’s memory is keyed by that id or by the handle itself: agent_diary and agent_onboarding_markers rows by the agent id, episodes and synthesized_skills by the handle, and counterparty_profiles by the observing seat’s handle. All of it survives engine restarts.

Rename caveat. Both inputs are part of the derived id: changing a seat’s handle or the company’s name creates a new derived id and orphans the prior per-agent rows (diary, onboarding markers, counterparty profiles). The seat keeps working; it has simply lost its memory. A company rename does this to every seat at once, so settle name and each handle before the company runs. (An explicit handle on each role pins half of it; nothing pins the company name.)

Role.Seat()

Company.RunnerFor

config.Role (authored)name, handle, emailgoal, backstory, managesresponsibilitiesbehavioral_guidelinesllm, integrations, mcp_env

org.Role (in the epoch)per-phase provider chainschat identitiesnormalized manages and mcp_env

runner.Runner (one per turn)per-phase promptsthe seat's tool registryprovider chains, budget meterprefetched context blocks

For a seat with direct reports, the executor prompt includes a team roster: each report’s name and handle with a compact profile (background, goal, responsibilities), rendered from the in-memory Organization so the lead can reason about who to assign work to.


The engine keeps no per-seat state machine in the turn path. What a seat is doing is computed by the live projection (internal/api/livestate) that every node serving the API holds, and it is served as one word per seat — activity — on every seat row: the handshake snapshot’s, GET /agents and every agents push. The dashboard maps the word to a colour and a label and derives nothing itself, so no two screens can disagree about a seat.

activityWhenstopped_reason
workingA turn is running on the seat (from agent_turn_started, so the prefetch counts as work), or a coding run it launched is launching or running.null
needsA coding run the seat launched is waiting on a person: awaiting_clarification, or reseed (the box was reclaimed and only the question survives).null
stoppedThe seat cannot take work.paused, unplaced, budget or provider
idleNone of the above: the seat is held somewhere in the fleet and waiting for work.null

The rows are checked in that order, and the order is the rule. A seat whose turn is still running is working even while a person’s pause waits for that turn to end, or while an older run of its waits on a question: what it is doing now is the truer word, and paused and the running-runs panel say the rest.

The four reasons a seat is stopped, in the order they are stated when more than one holds:

  • paused — a person paused the seat. Its mail waits on its inbox until somebody resumes it.
  • unplaced — no node in the fleet holds the seat’s lease, so nothing will run it however much mail it has. This is read from the seat leases (the same table fleet reads), never from which seats the node serving the dashboard runs: a seat a peer holds is placed and reads idle or working like any other. The leases are read on every snapshot and on the dashboard’s five-second tick, because a seat moves between nodes on a lease and publishes nothing. A read that fails changes nothing, and before the first read no seat is called unplaced — “no node holds it” is a fact about a lease table somebody read.
  • budget — a capped token window the seat is charged against, its own or the company’s, is refusing: the seat is parked until the window resets or the ceiling is raised. Read from the live meter’s per-window state, and only while that window lasts.
  • provider — the seat’s model provider chain was exhausted (llm_unavailable), and the seat has done no work since. It clears the moment the seat starts a turn or a phase again, or a new instance of it is spawned.

How much a seat has done is a different question, with a different source. The state is the projection’s, and describes now; a seat’s turns over days — how many ended, how many failed, what share a reviewer approved on the first pass, how long they take — are summed from every node’s replicated usage days by seat_activity, so they are the same on whichever node is asked and still count a node that has left.

A failed last turn is not a stop. A turn guard that fired (stall, max iterations, the delegation-depth cap, a scheduled turn’s wall-clock cap, an unhandled exception) fails that turn, and the seat takes its next wake like any other: it reads idle, and why its last turn failed is last_error and last_turn.outcome, which are separate fields.

The runs are read from the durable record, not from the turn’s stage. A turn that launched a detached coding run is parked, and a park says only that a run was launched; the run record — reconciled into the projection at boot and every thirty seconds, and never aged out — says whether that run is still running, has stopped to ask somebody, or is gone. A run parked on a question for thirteen hours is exactly the seat a person most needs to see, and it reads needs for as long as it waits.

Beside the state, the projection holds the turn the seat is on and the last turn it ended. Both are on every seat row of the agents push:

  • turn is {turn_id, work_item, work_item_basis, started_at, stage, node}. stage is one of three values. context means the turn has started and is assembling what it knows. phase means a phase is running. parked means the turn launched a detached coding run and is suspended until the run is collected. The run is collected later, possibly on another node or after a restart, and the same turn resumes then. A parked turn therefore stays the seat’s turn, and it is not reported as an end. The suspension publishes a turn completion with suspended: true, and the projection used to read that completion as the end of the turn, so the seat said it was idle while its work ran on in a box. When a parked turn resumes, its started_at is still the turn’s first start and not the segment’s. node is the node that published the turn’s newest event. A turn whose coding run is lost (sandbox_run_failed) has nothing left to resume it, so the loss ends that turn.
  • last_turn is {turn_id, ended_at, outcome}. outcome is completed or failed, and a turn is failed when any of its events was a failure. That is the same rule the turn list applies, so a seat seeded from the store and a seat watched live report the same outcome. ended_at is the turn’s newest event: its completion, or the reflection pass that runs after the completion.

A third key says whether a person has paused the seat: paused is {by, at, reason, stop_running} — who paused it (the person their token is bound to, or the token itself), when, why, and whether the pause also ended the turn the seat was on — and null while nobody has. It is an input to the seat’s state: a paused seat reads stopped with stopped_reason: "paused" once it is not working. See Pausing a seat.

activity, stopped_reason, turn, last_turn and paused are always present, and each of the last four is null when it has no value. The client merges each pushed row over the row it holds, so an omitted key would leave a finished turn, or a lifted stop, on the card. After a restart, the turns are seeded from the fleet’s turn list: each seat’s newest turns are read from every live node, so the “last turn 24m ago” line survives a restart, and so does a turn that is still parked.

How a seat comes to be held, and what happens when it is released, is Seat Ownership.


Each agent, when triggered (by event or task assignment), executes a turn through the two-stage Turn Engine:

1. Collect context (task, knowledge, trigger event, delegation chain),
and derive from the trigger WHO IS WAITING for this turn and WHICH
WORK ITEM it is on (the trigger's, a colleague's ask's, or a parked
run's) — announced on agent_turn_started before the prefetch
2. Executor phase
├── Tool surface = every first-party tool except mark_onboarded,
│ plus submit_work, activate_tool, list_mcp_server_tools
├── The system prompt carries a slim catalogue: builtin tool names
│ and MCP SERVER names, never the 50-150 MCP tool schemas
├── To use an MCP tool: list_mcp_server_tools(server) to discover
│ names, then activate_tool(name) to promote it onto the surface
│ so its schema arrives on the next round. Nothing is named in
│ advance, so nothing has to be reconciled afterwards.
└── Ends by calling submit_work: outcome, summary, deliveries,
checked against the engine's own record of the turn — every
round of it, so a delivery an earlier round made is citable.
A round that ends in prose instead is not an end: the loop
asks again, naming submit_work, at most twice in a row, before
the phase is rescued as incomplete. A round that did not FINISH
never gets that far: a refusal, a response cut off at the output
cap, a full context window or a paused turn ends the phase by
name, with nothing rescued and nothing re-asked
3. Engine check (no model call), over the whole turn's record
├── no_action nobody asked for and nothing acted on -> the turn ends
└── a claim the record refutes -> loop back with a correction
naming the tools that deliver where the asker is waiting
4. Reviewer phase
├── submit_review emits a decision
└── done | self_iterate (loop back, carrying the prior-work ledger
so the next round does only the gap) | failed
5. Publish agent_turn_completed and turn_completed; reflection consumes the latter.
A turn nothing named an item for is charged here to the one task its
writes committed to, if there was exactly one

Every turn is charged to one work item or to none — see Which work a turn is on.

The executor and the reviewer can run on different LLM models — see the Turn Engine doc.


The LLM is an external HTTP service — it cannot access local code, MCP servers, or engine internals directly. The shared tool-call loop (internal/agent/toolloop, driven by each phase of the Turn Engine) acts as a proxy that translates between the LLM’s text-based tool calls and local execution:

LLM responds without tool_calls

stop reason: refusal, max_tokens,context_exceeded or paused

1. Build messages + tool definitions (JSON schemas)

2. Request

3. Response: content + tool_calls [name, arguments]

4. Execute LOCALLY

5. Append tool results to message history

6. Loop back to step 2 (up to max_tool_rounds)

Per-role MCP tool?forward to MCP server(role-specific credentials, checked first)

Global tool?builtin function or global MCP

LLM API (external)Claude, GPT, …

phase ends

phase fails by name(its calls are not run)

YOUR MACHINE — the tool loop (executor / review / sub-agent)

Step 3 reads why the model stopped before anything else reads the response. Every backend normalises its own stop reason (end_turn, length, content_filter, …) onto one vocabulary, and a round the output cap cut off, one that filled the context window, a turn the provider paused, or a model that declined the request ends the phase with that reason as its error_kind — the tools that round asked for are not run, and no corrective is sent. A refusal is never handed to the next model in the fallback chain and its turn is not redelivered. A tool call whose arguments did not parse is answered with a failed result saying why rather than run with none. See A round that did not finish.

Both builtin and MCP tools produce identical tool definition schemas. From the LLM’s perspective, lookup_colleague (builtin) and an MCP server’s issue-creation tool look the same: a function it can ask the engine to call.


Under the two-stage Turn Engine, each phase builds its own narrow system prompt — there is no single monolithic prompt for a turn. Each builder lives in internal/agent/prompts/; the detail layer is in internal/agent/prompts/sections.go. Founder-defined role/org context (mission, vision, policies, backstory, responsibilities, behavioral guidelines, team roster) renders directly from the in-memory Organization model into the executor’s prompt via the section builders — no DB seed step, no reconcile pass.

PhaseWhat’s in the prompt
ExecutorIdentity (role, unit, goal, manager, direct reports, team channel), company mission and vision, full policy text, role profile (backstory, responsibilities, behavioral guidelines), unit context (purpose, goals), team roster with per-member profile (leads only), the ## Human colleagues note (only in a company with human seats), the executor’s contract, ## When a decision is not yours to make (the escalation guidance below; only when the seat holds comment_on_work_item), Tool Skills catalogue (one-line summary per triggered skill), slim tool catalogue (builtin tool names + MCP server names; MCP tool names hidden behind list_mcp_server_tools). Plus ## The thread so far — the chat thread the turn was woken in, read at turn start and handed over rather than left for the agent to fetch (see the thread block below); it is not a learning prefetch but the trigger’s own context, which is why it leads. Then the six learning prefetches, in the order they render: ## First-turn onboarding (until mark_onboarded fires), ## Personal memory (diary), ## Synthesized skills you've learned, ## Relevant knowledge (a knowledge-base search built from the trigger), ## Similar prior work (episodes), and ## Known counterparty. On rounds after the first, the user message also carries the prior-work ledger as ## Already done earlier in this turn.
ReviewOne-line identity, the round’s own account of what it set out to do, the outcome word (and who wrote it), the verbatim tool log, the text it produced, the decision-enum contract, and the Tool Skills catalogue for MCP-server-keyed skills (operator-scoped to the review phase). On rounds after the first, a ## Earlier rounds (already delivered) section carries the prior-work ledger so the duplicate-delivery rule holds turn-wide. No tool catalogue, no policies, no roster, no prefetch.
Worker (delegate)The worker’s persona (a workers: template or the parent’s inline prompt), the Tool Skills catalogue scoped to the tools the worker was granted, the slim tool catalogue, then the mandated runtime preamble (no further delegation, no colleague contact, read-only discovery only, and end by calling submit_result).

Why the split: the executor is the frame making every ownership / delegation / policy-sensitive decision AND acting on it, so it gets the whole picture — the two-prompt engine’s real cost was never the tokens saved by splitting them, it was sending the identity scaffold twice and throwing away everything the planner had read. The reviewer’s question is narrower: is this round’s work right, given what the record says it did. Standing memory, the team’s docs and the requester’s traits are what the executor needed to DO the work; in front of a reviewer they compete with the evidence it is meant to judge.

Every builder returns its prompt with an outline beside it — the parts it appended, each with a stable key, a title and its length — and the phase record publishes both, so a screen shows a prompt as the parts the engine assembled rather than guessing them from ## lines that embedded content brings with it. The outline never changes a byte of the prompt. See a prompt carries its outline.

Engine guardrails (event triage, tool usage, knowledge-system usage) are carried by tool descriptions (search_knowledge, colleague-surface tools) and by the executor and review contracts themselves, not by dedicated prompt prose. Each tool’s one-line description tells the LLM when to use it; the per-phase contract tells the LLM what output shape is expected.

Escalation is the one piece of guidance with a section of its own, ## When a decision is not yours to make, rendered right after the executor’s contract whenever the seat holds comment_on_work_item. It exists because the shape of an escalation is the engine’s rather than any role’s: a question that needs somebody to choose is a structured ask — the question, two to four options, the one the seat recommends and why, the evidence it looked at, and the role the person is asked in (approver or contributor) — put on the work item with ask and decision, or filed as the item itself with create_work_item. The person answers by choosing an option, and the asker is woken with the choice, by its label. What the section insists on is the half a model gets wrong: after asking, the seat ends the turn blocked on that branch — it finishes whatever does not depend on the answer and stops, rather than choosing for the person or waiting inside a turn that cannot receive the reply.

There is still no special escalation mechanism: an ask is an ordinary comment, a colleague reached any other way (a Slack mention, a2a_ask) re-triggers the agent the same way, and the reviewer routes a turn that has not yet reached anybody back through self_iterate so the next round makes that outreach, and ends one that already has as done, because the reply is what re-triggers the agent (no escalate tool, no ask_colleague decision, and no waiting state).

Tool- and MCP-server-specific guidance (when to call reflect_and_persist, how to mention teammates on Jira vs Slack, when to author code via the code sandbox and what the GitHub tools are for) lives in the Tool Skills registry — modular knowledge-base-sourced fragments (Confluence pages) where each skill carries a short summary (always inline in the per-phase catalogue) and a rich body that loads on demand via the always-on load_tool_skill builtin. The engine ships no skill prose; operators seed the skills container with crewlet confluence import and edit pages in the backend’s editor thereafter.

There is no single monolithic system prompt to read: internal/agent/prompts builds one per phase (BuildOnboarding, BuildExecutor, BuildReview, BuildSubagent) from the same identity sections, and each phase sees only the guidance and the tool catalogue that phase is meant to act on.

A chat thread reply is usually thin — “yes”, “+1”, “what about the other one” — and the thread is the context. The engine used to say so in the prompt and tell the agent to go and read the thread with its chat tools. On a company whose chat tools come from a per-role MCP server that costs three rounds before a word is read: list_mcp_server_tools, then activate_tool, and only then the call, with the tool’s schema arriving on the next message. An agent that skipped the trip answered the eleven words of trigger text with no idea what the thread was about.

So the engine reads it instead, at turn start, and renders it as ## The thread so far. Everything it needs is already on the node: the channel and the thread root are stamped on the trigger’s own metadata by the chat parser, and each seat’s authenticated client is held by its transport — so the read is made as that agent, on that bot’s own token and its own channel membership. No new credential, and no new scope: Mattermost reads GET /api/v4/posts/{root}/thread on the bot token it already holds, and Slack reads conversations.replies on the *:history scopes the app manifest already requests.

What lands in the prompt:

  • Oldest first, with the thread’s root always kept — it is what the thread is about — and then the newest messages. The root keeps its place even when there is nothing to render in it: an alert app posts its payload in blocks or attachments and no text at all, a system line is channel bookkeeping, and Mattermost leaves a deleted root out of its answer while every reply to it stays. The block then carries one line saying the first message could not be shown, which is the honest alternative to the silent one — dropping it promotes the oldest surviving reply into the root’s slot, where every bound protects it and the preamble calls it what the thread is about.
  • Bounded by whole messages, never a cut inside one: 8000 bytes. A thread within it is handed over whole, however many messages it has. Past it, the root (or the line standing in for it) and the newest messages that fit take three quarters of the room verbatim, and the messages between them are condensed into one account by the seat’s auxiliary model in the last quarter, marked as a rewrite and with the number of messages it stands for — so the decision made in the middle of a long thread is still in front of the seat. Only where no rewrite can be had are those messages left out, and the block then says how many it dropped. The root and the newest message survive however long they are. 8000 is one third of the conversation ledger’s 24000-byte budget, because both blocks are frozen into the same system prompt and re-sent on every round of every phase, and the thread must not crowd out the seat’s own cross-turn history. It is not configurable. A second bound of 30 messages used to drop everything older whatever its size; with the middle condensed rather than dropped it bounded nothing the bytes do not.
  • Senders resolved through the party registry, so a colleague reads as Tech Lead (lead) rather than an opaque platform id; a stranger renders as whatever the backend volunteered and then as the raw id, and never as a blank.
  • The seat’s own earlier replies marked **you**, resolved by the transport from the identity it learned at connect. On Slack that takes both the bot user id and the app id, because a bot_message echo of the seat’s own post carries the app id and no user id at all.
  • A thread too long to read says so, in place of the claim it would otherwise make. Slack pages conversations.replies from the oldest end, 100 messages at a time, up to 10 pages — so a thread past ~1000 messages cannot be reached at its newest end at all. The walk keeps the root and the newest of what it did reach rather than the oldest of the thread, and the block then drops its ordinary “the newest messages are what woke you” framing for one that says it stops short and tells the seat to read the rest with its chat tools: on that path the newest message is exactly what is missing, so the ordinary sentence would be guaranteed false. Everything either end dropped is in the count. The page size is 100 rather than the 200 Slack recommends because the client reads at most 1 MiB of a response before decoding it, and 200 messages carrying blocks, attachments or unfurls exceed that — which fails the read outright rather than shortening it. Mattermost has no such bound: GET /api/v4/posts/{root}/thread answers the whole thread in one response.

It says which waiting messages it showed. A message in the thread that was still waiting for this seat — somebody else’s, after the seat’s own last reply, before the message that woke the turn — is one the turn answers whether it was woken for it or not, because it answers the thread as the block shows it. That matters when the earlier message’s own turn failed: a failed delivery comes back behind its conversation’s newer mail, so the newer message’s turn runs first. The block reports those messages (by the chat backend’s own ids — never one a bound dropped, and none at all when the read stopped short or did not reach the trigger), and when the turn completes they are recorded in the completion ledger as worked through, so the earlier message is dropped when it comes round instead of being answered a second time, out of order.

It is best effort, always. A thread that could not be read — a node in maintenance mode runs no chat transport, a chat instance unreachable at boot leaves the company running without its chat surface, a seat whose token was refused has no client, a channel the bot is not in — renders a different sentence from a thread that was read and had nothing in it. “There is nothing earlier” says answer the trigger as it stands; “it could not be read from this node” tells the seat to go and read the thread itself, which is the one case where the old instruction was right. Neither ever fails a turn.

It does not make the trigger thick. RequiresRecon still gates the three relevance prefetches on a chat thread reply, and deliberately so: that flag describes the trigger body, which this block does not change, and those filters judge relevance against the trigger text — “+1” is exactly as useless a search query with the thread in the prompt as it was without. The flag is also stored on every past event and read by the dashboard, so flipping its meaning would rewrite what every historical turn claims about itself. The block is reported separately on prefetch_summary (thread_context_hit / _bytes / _posts / _read / _stopped_short). _hit and _bytes cannot separate the block’s states on their own — every path renders non-empty prose — so _read says whether a backend answered at all, which is what tells a thread that was read and empty from one the seat was told to go and find, _posts says how much was handed over, and _stopped_short says the read could not reach the thread’s newest message.


The engine ships these tools (internal/agent/builtin, registered in the epoch’s tools.Registry with the origin builtin). A tool whose dependency is absent is omitted rather than registered and broken, so a company without a store, a knowledge backend or a sandbox gets exactly the tools it can serve, and the node logs the list it registered (builtin_tools_registered).

ToolRegistered whenPurpose
lookup_colleaguealwaysResolve any colleague identifier (handle, role name, a human’s contact ID) to one seat, case-insensitively, with partial and fuzzy fallbacks; returns the candidate list when more than one seat matches
a2a_askthe node has a stream and a coordination storeAsk one AI colleague one question. The colleague is woken on its own inbox and answers in its own turn, so the call returns as soon as the question is sent
use_skilla learning storeLoad one of the seat’s own synthesized skills on demand
refine_skilla learning storeReplace a synthesized skill’s body with a corrected procedure; the previous version is kept
query_episodesa learning storeRecall the seat’s own past turns: by meaning (query), by conversation, or most recent first
refresh_memorya learning storeRe-run the personal-memory filter mid-turn with a context hint
reflect_and_persista learning storeKeep a durable fact in the seat’s private diary (kind: long, the default, or short)
mark_onboardeda learning storeStamp the seat’s onboarding marker after reading the relevant knowledge-base pages (offered to the onboarding pass, not to the executor)
run_sandboxproviders.sandbox is configuredHand a code task to a coding agent in a sandbox; the executor suspends until the run reports
load_tool_skillthe company publishes Tool SkillsLoad the full body of a Tool Skill by exact key (the catalogue carries only the summary). Required skills (the default; required: false opts out) must be loaded this way before the tools they cover can be called, and the engine rejects earlier calls with a “load this skill first” error
search_knowledgea knowledge backendSearch the company knowledge base on a query the executor writes itself, over the same seam as the turn-start ## Relevant knowledge prefetch. It is what a seat woken by a bare pointer uses once it knows what the task actually needs: the prefetch’s own search is gated off on such a trigger, because a query built from “PR #42 got a comment” matches the wrong pages or none
delegateper turn, executor onlyHand narrowly-scoped work to one or more short-lived workers, optionally as a dependency graph. Built per turn rather than registered once: it carries that turn’s grant (the parent’s own live tool set, minus the control tools and anything that writes to a shared surface) and that seat’s visible worker templates, so it cannot be a shared registry entry. Absent when the seat’s remaining token allowance cannot be read, because delegating with no readable ceiling is delegating with no ceiling

The phase tools are not in the registry: the runner adds submit_work, activate_tool and list_mcp_server_tools to the executor’s surface, submit_review to the reviewer’s, and submit_result plus a discovery pair of its own to a worker’s.

Colleague outreach happens through the upstream MCP tools directly (on the common stack: a chat server’s post-message tool, the tracker’s comment and update tools, the wiki’s comment tool, the code host’s review tools; these are examples, not engine-known names). The engine ships no chat or tracker wrappers of its own; a2a_ask is the one colleague tool it registers, and it is narrowly scoped to tight-loop, mechanical sync between agents. Use whichever chat, issue-tracker, wiki or code-host tools your MCP servers expose for any collaboration a human teammate would reasonably want to see. The engine prompts name none of these: they describe the capability and the LLM picks the tool from its catalogue (see Tool Capabilities). See Turn Engine: Colleague-surface tools for when to use each.

A decision a seat needs from somebody is a structured ask on a work item — comment_on_work_item or create_work_item with ask and decision — answered with a choice that wakes the asker; how a company decides stays behavioural guidance, see Decision Framework.

MCP tools (Jira, Slack, GitHub, and so on) are discovered from the configured MCP servers and registered alongside builtins: a shared server’s tools when an epoch is applied, and a shared: false server’s tools into the seat’s own registry when the node acquires that seat’s lease. The executor does not see every MCP tool name in its system prompt (a role with 50–150 MCP tools would push 15–25 KB of catalogue into every prompt); instead the prompt lists MCP server names and the LLM walks the discover-then-activate flow:

  1. list_mcp_server_tools(server) — returns the name: description listing for one server.
  2. activate_tool(name): promotes a tool from the catalogue onto the phase’s active tool list so the LLM can call it on the next round.

Both meta-tools are available to the executor and to the onboarding pass. A worker cannot use the parent’s pair (activate_tool and list_mcp_server_tools are on the worker denylist, because they would activate tools onto the parent’s surface); it gets a pair of its own, bound to its filtered grant, so it can discover and activate only read-only tools the parent could already reach.

Roles with GitHub credentials in mcp_env.github get a per-role instance of the remote GitHub MCP server (declared as a shared: false http entry in mcp_servers), giving them the full GitHub toolset for reading/reviewing/tracking code (issues, PRs, repos, code search, actions); code authoring goes through the code sandbox. See GitHub Integration.


There is no pool of agent instances. Three structures answer the questions a pool would:

  • Which seats exist: the epoch’s Organization. Company.Seats() lists its agent seats for placement; human seats are left out.
  • Which seats this node runs: the seat host (internal/seat), from the leases it holds. A node claims its fair share, attaches each seat’s mailbox last, and releases a seat whose role an apply removed. See Seat Ownership.
  • Who an event is for: the party registry (notify.Registry), rebuilt on every apply, which resolves a handle, a role name, a derived agent id, an email or an external ID on any connected surface. Resolution is derived from the organization, so the node that consumes a delivery can route to a seat it is not running.

A failure is scoped to a turn, not to an instance: a phase that breaks fails that turn (see Turn Engine), and a seat whose acquire hook fails is released and not re-attempted on that node for one lease TTL, which gives a peer a clear run at it.

Since each agent is a unique individual, there is no load-balancing or role-based routing. Task assignment is a team lead decision, not an engine algorithm.


Agents are callback-driven. When a node acquires a seat it attaches a handler to the seat’s durable subscription (agent-{handle}) on its inbox topic (crewlet.agent.{handle}.inbox), and the queue invokes that handler as messages arrive. There is no per-agent loop to run.

Inbox delivery is batched per conversation (see Event System — Inbox batching): events that queued up while the agent was busy — or within the configured linger window — are drained together and partitioned by partition key, so ten comments on one Jira issue or Slack thread reach the handler as ONE batch and trigger ONE digest turn instead of ten. Every partition — one event or several — reaches the same dispatcher and runs one Turn Engine turn; what differs is only the ask it is handed. A single-event partition is handed its own event, and a multi-event one is merged into a single coalesced notification first, so the third-party app’s scaffolding renders once and the seat is told to treat the thread as one piece of work.

The engine runs genuinely parallel work within a single process:

  • Each delivery is handled in its own goroutine, so seats make progress independently rather than taking turns
  • Multiple agents can be in the Working state simultaneously
  • A per-node concurrency gate limits how many agent turns run at once: Tier A’s node.max_concurrent (default 32)
  • A turn takes a slot after the ownership check and before its first model round-trip, and releases it when the turn ends

What the gate is for, and what already bounds itself. A seat’s mailbox is a durable subscription whose handler runs one batch at a time, so a seat is already serial — it never runs two turns at once. What is not bounded is how many seats run at once: that is placement’s arithmetic over the company’s seat count and the live node count, not a statement about the machine this process is on. A node handed forty seats would open forty simultaneous model round-trips and their tool loops. max_concurrent is the knob that says how many the host can actually take.

A turn past the ceiling waits, in this process. It is not handed back for the broker to redeliver on the broker’s own schedule — for a chat message someone is waiting on, that turns a busy moment into a visible stall. The waiting turn starts the instant a slot frees. The one exception is a drain, below.

What it does not gate, and why that is not an oversight. Post-turn reflection does not take a slot. It does not need one: it consumes turn_completed through a single durable subscription whose handler runs one delivery at a time, so a node runs at most one reflection pass at a time however many turns finish at once. Making it compete for turn slots instead would let a backlog of completed turns starve live seats — a company under load would stop answering people in order to finish learning from what it already answered — and the reverse, learning starved indefinitely by traffic, is what the separate consumer group exists to prevent. Auxiliary spend is bounded where it belongs, by the token budget, and every learning worker resolves its model through the engine’s one auxiliary seam, so that counter sees it and every spend figure records it (Budgets and spend § Auxiliary spend) — a pass the model refused included: a refusal returns no answer, but the response the vendor billed travels with it and is recorded like any other (llm.Billed is the one reading every meter uses).

Sizing it. It is per node, so a fleet’s ceiling is N × the value — see Scaling Out. The default of 32 is above the seat count of a single-node company (the example company runs a handful of seats; a large one runs tens) so it changes nothing for a company running today, while still bounding a node that has been handed far more seats than its host can serve. Raise it on a bigger host; lower it on a satellite running one agent. There is no “unbounded” — 0 means “take the default”, and effectively-no-limit is a large number you can see in your config. Note this is a different knob from a cli-agent provider’s own max_concurrent, which caps that provider’s subprocesses; see Subscription LLM backends.

That is real parallelism rather than one cooperative loop, which is the single biggest behavioural difference from the engine’s first implementation: anything shared between turns is guarded rather than safe-by-construction, and the whole suite runs under the race detector for exactly that reason.

A token ceiling is per calendar window — the day, the ISO week and the month on the company’s clock — so a seat that has run out of room has not run out for good: it has room again when the window turns over, or as soon as somebody raises the ceiling. The dispatcher acts on that before a delivery is claimed: before the completion ledger reads it, before it is offered to a parked coding run as an answer (resuming one charges tokens exactly as a turn does), and before any model is asked anything.

delivery reaches the seat

every capped window has room

a capped window is refusing

a round is refused, nothing written outside yet

a round is refused after an outside write

the window turns over

an apply changes the ceilings or the clock

Asked

Runs

Parked

Recorded

  • Asked. The node reads the seat’s counters and the company’s, in the current windows. A window is refusing when it has no room left for a single token — and a window that refused a round always has none, because the refused round is counted past its ceiling. The refusal stamp is not asked: after a ceiling is raised it stays until an admitted charge clears it, and a park taken on it would hold the seat back from the very charge that could. A counter that cannot be read parks nothing: the turn runs and its own meter, which fails closed, is the gate.
  • Parked. The seat’s inbox takes a pause hold (budget_window), the delivery is deferred — handed back unacked, for one of its deliveries, with a reason naming the window: budget: day window 2026-09-23 resets 2026-09-24T07:00:00Z — and an alarm is set for the end of the refusing window, the one that ends last where several refuse, since nothing can run before it. seat_budget_parked is logged with the scope, the window, its figures and the reset. And the park is recorded as the gate’s refusal: the delivery it defers is the turn whose first charge would have been refused, so the window’s refused_at — “last refused” on every screen and in crewlet budgets show — is stamped on the scope a charge would be refused by, the company’s before the seat’s, exactly as that charge would have stamped it. A window a coding run or a background pass filled used to park every delivery its seat was sent while reading as one that had refused nothing.
  • Released. At the reset, or at once when an apply changes the ceilings of either scope or the company’s clock (which moves every window’s end), the hold is lifted, seat_budget_park_released is logged, and the held mail is delivered again in order and asked again. A revision that lowers a ceiling parks it again at the cost of one delivery.
  • Refused mid-flight. A window that had room when the delivery was claimed can run out during the turn. That turn stops with budget_exhausted and makes no model call after the refusal — its meter holds it, so the next phase, the round-cap judge and the turn’s other workers are refused before they are sent, and the first call it holds in a window is recorded on the counter as the gate’s refusal, the stamp a refused charge would have written (Turn Engine, invariant 4) — and the seat is parked exactly as above — unless the turn had already written outside the engine (an MCP write, a colleague ask, a coding run), in which case the trigger is recorded and acked, because running it again after the reset would repeat those writes.

Why a park and not a retry. Before the park, a wake reaching a spent seat ran a turn that was refused on its first charge, and a refusal proves nothing left the engine, so the delivery was NAKed and redelivered — and refused again, twenty-five times over about ten minutes, until the broker dead-lettered a perfectly healthy message. A seat that ran out at ten in the morning lost every message it was sent for the rest of the day.

Two stops with two owners. The deferral quiesces the attachment, and that quiesce is the seat host’s: it is resumed on the next lease renew, as every deferral is. The pause hold is the park’s, and it is what keeps the resumed attachment from being handed anything until the release. So a park costs one delivery per message, not one per renew, and the release never resumes a consumer the seat host stopped because it could not prove it owns the seat.

This node alone agrees on it. A park is a hold in this process’s queue client and an alarm in its memory, derived from the fleet’s shared counters on every delivery. A restart, or the seat moving to a peer, simply asks the counters again on the first delivery. It is not an alarm: nothing is wrong with the node when a ceiling does its job.

A person can pause a seat — from its profile, from their own assistant through the operator catalogue, or with crewlet seats pause — and resume it later. While it is paused the seat starts no new turn, its incoming mail waits on its inbox in order, and its scheduled runs are skipped. The turn it is on finishes first, unless the pause also asked to stop it (stop_running), in which case that turn ends at its next round.

seat not paused

pause_seat

pause_seat with stop_running, a turn running

the turn reaches its next round and ends

pause_seat adds stop_running, a turn running

resume_seat — what waited is delivered first, in order

Taking

Paused

Stopping

  • The pause is one record for the whole company. It lives in the coordination store (seat_pause), with no age: it ends when somebody resumes the seat, or when the seat leaves the company (the apply that removes a seat clears its pause, so a seat added later under the same handle is not born paused). It is written by compare-and-set, so two people pausing one seat at once make one change: the second is told it was already paused. The writer whose change won publishes seat_paused or seat_resumed.
  • Every node carries it out from its own copy. Each node watches the records and, the moment a pause lands, takes a pause hold (seat_paused) on the seat’s inbox; the resume lifts it and the mail that waited is delivered. A node that acquires a paused seat — placement moved it, or its holder restarted — takes the hold before attaching the seat’s mailbox, so the mail it has been holding is never the first thing the new holder runs. A delivery that races the hold is held and parked by the inbox screening. A node that has not yet read the pauses at all (the seconds after a boot) defers its deliveries rather than guessing. A store that cannot be reached is not a resume: a node keeps the pauses it last read, and a paused seat stays paused.
  • Stop now. The per-round fence every turn runs under — the same one that stops a turn whose seat moved to another node — also closes on a pause that asked to stop, so the turn ends at a round boundary, never between a tool call and its result. What it had not yet done is lost. The trigger is recorded as worked and acked, not retried: a person who stopped a turn did not ask for it to run again the moment they resume the seat. The turn’s completion reads stopped rather than failed, and agent_turn_stopped names who stopped it. A detached coding run is not fenced — it outlives its turn by design — but the turn that resumes it is, and ends there with the run.
  • Answers wait too. An answer to a parked coding run, by chat or by turn, is held behind the pause like the seat’s other mail: resuming a run is work. A chat reply already recorded as a run’s answer before the pause waits the same way — the coordinator does not retry its resume while the seat is paused, and charges the wait to none of the answer’s attempts.
  • Scheduled runs are skipped, not queued. A fire that comes due on a paused seat is recorded skipped_paused in the dispatch ledger and not sent — a standup held behind a week’s pause would otherwise run once for every day of it. See Scheduling.

Short of stopping a turn, a person can steer it: send a note that the turn reads at its next round boundary and keeps to for the rest of the turn — from the live view, or from their own assistant with steer_turn. The note travels on an ephemeral scatter to every node; the node running the turn is the one that answers, and the turn’s own record says what became of it (agent_turn_steered: delivered at the round that read it, or expired when the turn ended first). A turn whose executor runs as a coding CLI’s own loop cannot take one. See Turn Engine § Steering a running turn.

SIGINT / SIGTERM trigger a drain with the probes up, designed so a restart picks up cleanly without a half-finished turn: the node stops taking new work, lets the turns already running finish, hands its seats back, and only then closes its HTTP listener and its backends. The engine owns the process signals exclusively. Nothing else in the process may install a handler, the embedded API server included.

The listener stays up for the whole drain, and the door is a refusal rather than a closed port. An orchestrator watches a node precisely while it drains, so GET /health keeps answering 200 with status: "shutting_down", and GET /ready answers 503 with reason: "draining". Traffic moves elsewhere, and nothing kills the node in the middle of the turns the drain exists to finish. What the drain must not do is keep making work for itself, so from its first moment every route that would start new work answers 503 with {"error": "draining"} and a Retry-After: the webhook edge, the /config, /secrets and /setup writes, the operator’s writes, POST /operator/mcp and backups. Reads keep being served, the dashboard included, and so do the sandbox bridge (/mcp/{token}) and the telemetry edge (/otlp/{token}), because those carry the tool calls and spans of coding runs that started before the drain and would only be broken by a refusal. See During a drain for the exact rule.

2nd signal:immediate exit

Signal arrives (1st)signals handed back to the OS

1. The drain begins/health 200 · /ready 503 · work routes 503

2. Stop claiming, give up presence

3. Close the concurrency gate

4. Quiesce every held seat

5. Wait for in-flight handlers

6. Release every seat

7. Close the HTTP listenerdashboard · REST · webhooks · probes

8. Stop the dutiessandbox waiter · notifications · maintenanceintegrations · memory sync · learning · scheduler

9. Reap shared MCP servers; last auxiliary-spend flush;custody flush; close stream + store

Process dies

  1. The drain begins. The engine reports that it is shutting down from this moment, before anything that can block: the HTTP surface starts refusing new work and both probes report the drain. The watchdog is disarmed, since the drain and the teardown legitimately block for longer than it would tolerate; no further config revision is applied, because everything one would build is about to be torn down; and the stop is announced in the audit log (org_stopped).
  2. Stop claiming and give up presence, so peers stop reserving a share of the company for a node that will never claim again.
  3. Close the concurrency gate. Turns still parked at it are released at once and their deliveries deferred (left unacked, so a peer picks them up rather than waiting out a redelivery timer) instead of starting fresh LLM rounds mid-drain. The gate closes before the mailboxes quiesce: quiescing stops new deliveries, but a turn already delivered and parked behind a slot is past that point.
  4. Quiesce every held seat. The node stops taking new work while staying attached. This is what makes the wait below terminate: without it the mailbox keeps feeding this node work for as long as its peers keep publishing, and “wait until nothing is running” never comes true.
  5. Wait for in-flight handlers, indefinitely: running turns finish their rounds until the count hits 0, with drain_in_progress logging the in-flight count every 10 s. A turn parked on a detached coding run is not one of them. It suspended when the run detached and its trigger is already recorded as worked, which is exactly why coding work is detached: a drain that waited for a real coding job would wait for its whole runtime. The run’s record lives in the fleet’s coordination store rather than in this node, so whichever node holds the seat next picks it up rather than it being lost with this process.
  6. Release every seat, each lease given back with its mailbox intact, so peers can claim them at once rather than waiting out the lease TTL, and each seat’s last lifecycle event published. The drain then logs drain_complete.
  7. Close the HTTP listener, and not before: until now the probes are what the orchestrator reads, and they read the stream and the coordination store the next steps close. Requests still running get a five-second grace and are then cut, the live feed stops, and every dashboard socket is closed (api_stopped).
  8. Stop the duties: the sandbox waiter first (its keepalive is what stops a running box being reaped while turns are still finishing), then the notification transports, the maintenance duties, the integration reconcile loop, memory sync, the learning passes, the cron scheduler and the credential cooldown refresh.
  9. Close the backends: the shared MCP servers are reaped; the auxiliary-spend ledger publishes its last records — what the seats’ auxiliary model spent since its last 15-second flush — on a budget of five seconds, once every producer of them has stopped and before a node without data flushes its custody batches, which carry those records too; then the custody flush, and the stream connection and the store file are closed (engine_stopped).

The stop’s coordination shares one allowance. Every lease the stop gives back — the presence and seat leases of steps 2 and 6, and in step 8 its fleet duties and the integration lease a reconcile pass still in flight gives back as its loop is stopped — draws on ONE allowance of one heartbeat interval: the lease TTL over three, 15 s at the shipped 45 s. A third of it, 5 s there, is reserved for withdrawing the node’s admission to a ceiling change in step 8, and no step before it can spend that share: every lease falls back to lapsing on its TTL, but an admission has no TTL, so one that is not withdrawn stays until the node restarts under the same id or an operator excludes it (admission_not_withdrawn says which). The leases’ two thirds are charged only while one of their round trips is in flight, so the wait of step 5 spends none of it, and seats released together are charged once for the time they overlap. On a healthy fleet each step takes milliseconds and the allowance never binds. On a member that has lost its coordination store — every step failing, each lease falling back to lapsing on its TTL — the stop’s give-back costs that one allowance in total, after which every remaining step fails at once with its own …_unavailable / …_not_released warning. The integration lease is charged only from step 8, where its loop is stopped and the stop waits for the pass in flight — from that moment on, a give-back that began before it included. The loop runs through the drain, so a pass that ends on its own while step 5 waits is an ordinary pass, and its give-back ends on its own client’s request timeout rather than spending what the seats are owed when the wait is over. Each step used to wait out a bound of its own, so a member whose coordination store had gone spent the sum of them on deadlines that could not succeed; one that has lost that store alone now stops in about one allowance, 15 s at the shipped TTL.

The stop’s events are not on the allowance. The org_stopped announcement of step 1 and each seat’s last agent_terminated in step 6 travel on the event stream rather than through the coordination store, and the two are separately replicated streams, so one can lose its quorum alone. Each event is therefore bounded on its own, by the five seconds the stream’s client gives a request with no deadline (past it, lifecycle_event_not_published or seat_lifecycle_not_published, and a live screen ages the seat out), and published beside the give-backs rather than in front of them: the announcement while the drain runs, each seat’s event while that seat is torn down — and finished before its lease is given back, so a peer that takes the seat over announces it only after this node has let it go. A stream that will not acknowledge costs the stop those bounds where nothing else is running, and never the time the leases and the admission are owed. The auxiliary-spend and custody flushes of step 9 are not on the allowance either: they carry records rather than give a lease back, and what they cannot publish is lost rather than lapsed, so each keeps the budget of its own named there.

A member that has lost both stores stops in about 25 s. A lone member of an embedded fleet whose peers have gone loses the event stream and the coordination buckets together, so it spends the whole allowance and, wherever nothing overlaps them, every one of those bounds besides: the leases’ share of the allowance (10 s at the shipped TTL, spent by the presence give-back while the announcement waits out its own five seconds inside it), each seat’s last event beside its last memory publish (5 s), the admission’s share (5 s) and the auxiliary-spend flush (5 s). That is about 25 s at the shipped TTL, as such a stop was measured before the allowance existed, and it fits Kubernetes’ 30 s default grace by five seconds rather than by the fifteen a member that lost only its coordination store has to spare.

Let LLMs finish their rounds — but only the running ones. The drain distinguishes two kinds of in-flight turn. Turns already past the concurrency gate (model rounds under way) run to completion: they may have fired side effects, and abandoning that work buys a faster deploy by throwing away what was nearly done. Turns delivered before the quiesce but still waiting for a slot abort immediately — they have called no model and fired nothing, so their trigger is simply deferred. Without this split, a backlog parked behind max_concurrent would run full multi-minute executor → reviewer turns one after another during a shutdown that waits for them indefinitely.

No engine-level timeout on the drain. Step 5 waits as long as in-flight turns need, and the listener stays up for all of it. We don’t try to second-guess “too long”, because the host already provides that cutoff:

  • Interactive: a second Ctrl+C tells us you’re done waiting.
  • Kubernetes: terminationGracePeriodSeconds (default 30 s) — after which the kubelet sends SIGKILL.
  • systemd: TimeoutStopSec (default 90 s) — same SIGKILL fallback.

Embedding our own grace window would duplicate that decision in two places and inevitably disagree. Size the orchestrator’s grace period to cover your expected turn length (a multi-tool executor → reviewer turn can comfortably take 2–5 minutes).

Force stop (second signal). The first signal starts the drain and hands the signals back to the operating system, so a second SIGINT or SIGTERM does what it always does: the process dies immediately, with no cleanup. That handover is what makes the unbounded drain above safe to offer — without it the engine would still be the installed handler, every further press would be swallowed, and an operator watching a drain from the terminal they started it in would have no way to abort it short of SIGKILL from somewhere else.

A turn killed that way leaves its trigger unacknowledged rather than NAK’d, because nothing gets to run. The broker redelivers it once its ack window elapses, so the work is not lost — it is just slower to come back than after a graceful drain, where each finished turn acks normally and each turn still parked at the concurrency gate is deferred, so a peer takes it straight over. A redelivered turn runs from scratch, and side effects the killed turn already fired (a chat post, a work-item comment) may duplicate — the completion ledger covers a turn that finished, and this one did not. That is the trade-off you opted into by sending the second signal.

Watching the drain. On the dashboard and over the API, for as long as it lasts: the listener closes only once the drain has completed, so the dashboard shows the node as draining with its in-flight count, GET /health reports it, and GET /ready names the reason. The log says the same and outlives the process: engine_draining on the first signal, with what is being waited for and how to stop waiting, then drain_in_progress with the in-flight count every 10 seconds, then drain_complete, api_stopped and engine_stopped. Set logging.file if you want that record to survive the terminal it was watched in: the file is closed last of everything, after the drain and after the trace flush, so engine_stopped is in it.

A node’s drain is reported by that node: its own probes, its own dashboard and its own log. It gives up its presence at step 2, so a peer’s Settings › Nodes screen stops listing it rather than showing it draining. On a split deployment a -roles data,ingress node drains the same way; it holds no seats and runs no turns, so its drain is short.

The drain is available programmatically up to the moment the listener closes:

  • Engine.ShuttingDown: true from the first moment of the drain and never false again. The HTTP surface refuses work on it and both probes report it.
  • GET /health: 200 throughout, with status: "shutting_down", shutting_down: true and the in-flight count as in_flight.
  • GET /ready: 503 throughout, with draining: true and reason: "draining".

Per-agent visibility is finer-grained: each working agent’s row carries current_phase (onboarding, execute, review, or subagent for a worker) plus the round number, derived from the agent_phase_started events the runner publishes at the top of each phase.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.