Skip to content
You are reading documentation for unreleased main. This page is not in 0.1 yet.

Architecture

Crewlet is one Go binary. Start it and a company is running: an HTTP surface, an event stream, a database, a runtime that holds seats and executes turns — with no broker to operate, no database to provision, and nothing to install beside it.

This page draws that binary at six zoom levels, from what sits outside its boundary down to the path a single trigger takes through it. It is a map: every box belongs to a page that goes deeper, and the last section says which. The Overview carries the one-diagram version; everything past that is here.


Everything Crewlet needs to run is inside the process. Everything it needs to be useful is outside it — the surfaces your company already works on, the models that think, and the tool servers that act.

Two things about that outside half are easy to miss and are the whole design, so it takes two pictures. The first is the round trip — work arrives on those surfaces, and it is delivered back to the same ones.

YAML seed · REST · dashboard

webhook — or, for Mattermost,an outbound websocket

tool calls, made as each agent

the agent's own credentials

Founder / operatorwrites the company documentreads the dashboard

Teammatescolleagues, some of themholding seats in the org chart

A nodeingress · seats · workersembedded event stream · local store file

MCP serversstdio children this process supervises,or remote http endpoints

Where the work happensall optional, all independentChat — Slack · MattermostTracker + knowledge base — Jira · ConfluenceCode host — GitHub · GitLab

crewlet run — one process, one binary

People

The engine never calls a third-party app’s API on its own account. It calls MCP servers, and each seat’s server carries that seat’s credentials (role.mcp_env), so a comment on an issue is written by the agent, not by a service account fronting for it. The engine’s own integration packages exist for the inbound half — verifying a delivery, parsing it, deciding whose it is — plus provisioning and the two calls a chat transport makes on the seat’s own bot token: the working indicator it raises while the seat thinks, and the thread it reads back at the start of a turn woken in one. See Tool capabilities for why no engine prompt names a vendor tool.

The second picture is the supply — what a node reaches out to while a turn runs. Its MCP servers are the ones above; the two dashed edges are the process-level connections a deployment may never make at all.

per phase, per role

tool calls, made as each agent

detached coding runs

traces

stream + coordination

A nodecrewlet run

LLM APIAnthropic · OpenAI · anyOpenAI-compatible endpointor a coding CLI on your own subscription

Embeddings APIoptional — without it, search iskeyword only and nothing isrecalled by similarity

MCP serversthe agents' hands

Code sandboxan E2B VM, or this host

OTLP collectoroptional — Jaeger, Tempo, …

External NATS clusteroptional — a fleet may dial oneinstead of embedding the stream

What a turn consumes

Nothing outside the boundary is a hard dependency except a model. No embeddings means no similarity search — the diary’s candidate pool falls back to recency and the “similar prior work” block renders nothing, both first-class states rather than failures. No sandbox means no run_sandbox tool. No chat, tracker or code host means an agent has fewer surfaces to be reached on and to act through, and every other one still routes end to end.

Outside the boundaryRequired?What it is forWithout it
An LLM providerYesEvery phase of every turnNothing runs
MCP serversIn practiceThe agents’ hands — chat, tracker, wiki, code hostAgents have only the builtins
Chat · tracker · code hostNo, eachWhere triggers arrive and work is deliveredThat surface simply routes nothing
Embeddings providerNoVector recall over agent_diary and episodesRecall degrades to recency; episodes render nothing
Code sandboxNoCode authoring as a detached, suspended executor phaseNo run_sandbox; agents still read and review code over MCP
OTLP collectorNoExported tracesTrace ids are still minted and still stored on every event row
External NATSNoThe stream for a fleet that will not embed itThe process embeds its own — the default

crewlet run is the node, and what it does is a config value: node.roles picks from data, ingress, seats and workers. Declaring none means all four, which is the single-process company — one process holding the company’s records, serving the API and the dashboard, running every agent, and performing the company-wide duties. A node without data holds no copy of the replicated estate and answers its seats’ tools through a node that does; see Scaling Out.

events

live push

ingress: terminate inbound trafficwebhooks · REST · /ws/stream · OTLP · probes

Always on, whatever the rolesnotifications · reconciler · presence ·observability edge

seats: run agentsmailbox → batching → turn engine · MCP bridge ·post-turn reflection

workers — company-wide singletonseach on a worker:DUTY lease

Event streamembedded NATS JetStream, an embeddedcluster, or an external one

Coordination KVrides the stream's own connection

Storeone local file thisprocess owns exclusively

That is one crewlet run process. node.roles picks from data, ingress, seats and workers — declare none and you get all four — and the always-on set runs on every node whatever it says. The three backends are opened together and closed together; on a node without data the store is scratch and holds the node estate alone.

Three of those four groups are inventories, not pipelines: no box inside ingress, workers or the always-on set hands work to another box in the same group. They are named here rather than drawn.

Two slots and a file. The stream and coordination are the two chosen backends, validated together — a multi-node fleet cannot coordinate locally, and a two-member fleet has no quorum. The store is not a third choice: it is this node’s own file, opened with them because everything that writes to it is driven by them. A node holding one without the others could hear work it may not do, hold seats it cannot serve, or run turns it cannot record.

Coordination rides the stream’s connection. Not a second dial that could fail on its own, and never the store file — see Coordination for the line between the two estates and section 5 for which fact lives where. Only the lease half follows coordination.type: on a single node it falls back to an in-process store, while the shared buckets — the activation pointer, the ledgers, the counters, the company’s secrets — are opened on every topology, because a lone node still has to read back what it wrote.

The routes ingress terminates.

RouteWhat it is
/webhooks/slack/HANDLE · /webhooks/github · /webhooks/github/HANDLE · /webhooks/gitlab · /webhooks/jira · /webhooks/confluence · /webhooks/confluence/EVENT · /webhooks/datadog · /webhooks/forgeThe webhook routes. A delivery is verified, then claimed once per fleet, then handed to the notification service. Slack’s OAuth landing and the GitHub App return live beside them.
/config · /secrets · /setup · /agents · /org · /tools · /query · /backupThe REST and config plane. It reads and writes the coordination KV and the store directly.
/ws/streamThe dashboard’s only read channel: live pushes plus a query channel. The observability edge’s projector is what pushes onto it. It carries no write.
/operator/mcp · /operator/actThe operator catalogue — the tracker and knowledge tools a seat holds — served to a person’s own assistant over MCP, and to the dashboard as the person the token is bound to. Always guarded.
/otlp/{token}/v1/{signal}Signed-token trace ingest.
/mcp/{token}Signed-token tool bridge: one running seat’s own tool surface, served to a coding agent in a box. Per-run, expires with the run. The exception on this list: a session lives in the process that opened it, so this route belongs to the node that runs the seat, and a seats node without ingress binds its listener for this route alone.
/health · /readyThe two probes — section 6 says why they answer different questions.

workers is five company-wide singletons, each held on its own worker:DUTY lease.

LeaseDuty
worker:schedulerRole- and unit-scoped cron; a fire is published to the stream.
worker:sandbox-waiterPolls detached runs and resumes the turns waiting on them, over the stream. The same tick is the box keepalive.
worker:maintenanceThe retention sweep over the records that answer “recently” rather than “ever”, and the retirement of a removed seat’s mailbox and coding runs.
worker:integration-reconcileThe integration reconcile loop: every connected third-party app’s pass, on a cadence set by who has to act.
worker:learningEvery learning background pass: skill ageing, episode compaction, clustering and cross-agent promotion.

Five more services run on every node, whatever the roles say.

Always onWhat it does
Notification serviceOne fleet-wide group, notify-inbound: parse → resolve → valve → wake, publishing the wake to the seat’s inbox on the stream.
Config reconcilerPolls the activation pointer, applies an epoch, reports status — all in the KV.
Node presenceThe node:ID lease in the KV, plus the posture heartbeat.
Observability edgeTwo routes off one published event: a publish listener writes the store row, a projector pushes it live onto /ws/stream.

seats is the exception — the one group whose boxes hand work to each other. This is what a wake becomes after the stream delivers it:

events

Seat hostclaims seat leases, attaches the mailboxacquire → equip → THEN attach

Per-seat mailboxdurable consumer oncrewlet.agent.HANDLE.inbox+ .control for sandbox resumes+ .reflect for post-turn reflection

Inbox batchingdrain · partition by conversation· one digest turn per partition

Turn engineexecutor → reviewerone per running turn, gated bynode.max_concurrent

Per-seat tool registry + bridgethe shared catalogue, CLONED,plus this role's own MCP children

Provider chainfallback chain overa credential pool

Event stream

Coordination KV

seats — run agents

Two MCP lifetimes, and the difference is a security boundary. A shared: true server is one company-wide child bound to the config epoch. A shared: false server is a template, and each seat gets its own child bound to that seat’s lease — spawned before its mailbox opens, killed when the lease goes. The credentials in one of those children are that seat’s identity, and two children of one template publish identical tool names: in a single registry one would shadow the other and every seat would call whichever won, acting in the tracker or the chat backend as somebody else. So a claimed seat gets its own cloned registry and its own bridge.

One model call has three layers of failure handling, each owning exactly one decision. The backend classifies and does nothing else — it never retries. The credential pool decides whether the key is at fault: a rate-limit or auth failure benches that key for a cooldown and the next call leases another, least-in-flight. The fallback chain decides whether the model is worth abandoning and moves to the next one in the role’s chain, publishing a provider_fallback event for each hand-off — addressed to the turn, the phase and the iteration it happened in, so a chain that only flaps under load is visible on the turn that paid for it rather than only in aggregate. Exhausting the chain publishes llm_unavailable and fails the turn. That is why the same 429 produces three different behaviours at three different altitudes, and why cooldowns are fleet state rather than per-process: a limit belongs to the key at the vendor, so four nodes should not each pay their own 429 to learn it. What a classified failure says — in a log line, llm_unavailable, a turn’s error — is the backend’s own line: the status, the provider’s id for the request and what the provider said about it, redacted, which is also what an exhausted pool reports its last key was told. It is never the vendor SDK’s own error text, which for Anthropic is the request URL (a password in base_url with it) and the raw response body.

The observability edge is two routes, not one, and the split is deliberate. A published event forks. It is written to this node’s crewlet_events inline, in the publishing goroutine, through a publish listener with no consumer group — so no two nodes can ever write one row. (A node without the data role has no log to write, so its listener batches its events onto the custody topic instead, and exactly one data node keeps each batch — see custody.) And it is read back off the broker by an ephemeral broadcast subscription on crewlet.events.> that feeds the live projection — so a dashboard tab attached to node B shows turns that ran on node A. Swap either mechanism for the other and you lose the guarantee the other one was providing.

The two sets are not identical, and each exception carries its reason in internal/events: a few types are live-only — a per-round progress signal, a snapshot of the shared token counter, because persisting them would fill the log with intermediate states of rows it already holds finished, or replay a reading of a counter that has moved on as though it were current. What the projection shows and what the store keeps are two questions with two answers.

The API is served inside the engine’s process, and only there. crewlet run builds it over the engine’s own store, broker and coordination plane; a node without the ingress role serves only its two probes on api.port, and a node with api.port: 0 serves nothing at all. node.roles changes what the engine does, never whether there is one, so every API answers its node’s in-flight count, seats, posture and applied epoch rather than reporting them as unknowable.

The HTTP surface binds before the engine starts. A seat is not claimed until its per-role MCP children are up — one subprocess per server per seat, each a spawn, a handshake and a tools/list — and on the example company, whose one shared: false server is Mattermost, that is 7 children; a company with a tracker, a wiki and a code host wired in runs three times that. Binding first means the dashboard, the REST API and every webhook route answer during that window, and /ready says honestly that this node holds no seats yet. Webhooks arriving in the window are retained rather than dropped, because a seat’s mailbox is created before any claiming.


A person mentions an agent in a Slack thread. Here is every hop between that message and the agent’s reply — across a fleet, where the node that receives the delivery is rarely the node that runs the seat.

It is drawn in three legs, one per process that touches the delivery: on a fleet those are usually three different processes, and on one node they are three parts of the same one. Each leg ends by publishing to a subject and the next begins by consuming from it — the seam between the pictures is the seam in the architecture, and no hop is hiding in it. The coordination store and the event stream appear in more than one leg; there is one of each, company-wide, not one per picture. The step numbers run 1–17 straight through.

Leg 1 — the node that took the delivery. Verify it, claim it, publish it, and answer the third-party app only once it is on the stream.

Event streamCoordination KVAny node(ingress)SlackEvent streamCoordination KVAny node(ingress)SlackA retry that lands on another nodefinds the claim taken and is answered"duplicate" — one wake, however manycopies the third-party app sendsPOST /webhooks/slack/HANDLE1verify the signature(per-seat signing secret)2claim the delivery id3publish RawWebhook →crewlet.notifications.inbound42005

Leg 2 — any node in the inbound group. It takes the delivery off crewlet.notifications.inbound, works out whose it is, and publishes a wake onto that seat’s own subject.

Coordination KVAny node(notify-inbound)Event streamCoordination KVAny node(notify-inbound)Event streamone fleet-wide group: notify-inbound6the third-party app's parser: who is this for?mention · assignee · watcher ·thread follow · project lead7resolve to a seat through theorg-derived party registry8notification valve — is this seatover its rate for the window?9publish ExternalNotification →crewlet.agent.HANDLE.inbox10

Leg 3 — the node holding that seat. Its mailbox is a durable consumer on crewlet.agent.HANDLE.inbox. Everything from there is the turn, and the reply leaves as the agent rather than back down the path it arrived on.

LLM + MCPThe seat's owner(seats)Event streamLLM + MCPThe seat's owner(seats)Event streamthe seat's durable consumer,group agent-HANDLE11drain the backlog, partition bypartition key, one digest per partition12take a slot at node.max_concurrent13executor → reviewer14the reply is posted by the agent'sown Slack tool, as itself15turn events → crewlet.events.*16ack the delivery17

Five properties of that path are worth stating on their own, because each one is why a step exists at all.

The delivery is claimed before it is published. Two concurrent retries must not both wake a seat, and a third-party app retrying reaches whichever node a load balancer picks — so the claim lives in the fleet’s coordination store, not in the receiving node’s memory. Publishing then comes before the store row and the live push, because the publish is the only step that has to happen: a delivery that reached the stream will be worked even if the receiving process dies in the next instruction. A publish that fails releases the claim and answers 503, so the third-party app’s retry finds the delivery unclaimed.

Routing is a publish, never a call. The inbound consumer group is fleet-wide, so the node that wins a delivery is usually not the node running the recipient. Every resolution goes through the org-derived registry — seat ids are a UUIDv5 over (org name, handle), so every node computes the same answer with no database and no running instance — and every wake is a publish to the seat’s inbox subject. A service that resolved against local state would drop most of a fleet’s mail.

The mailbox exists whether or not the seat is running. A seat’s inbox is a durable consumer created with nothing attached, and it retains what is published while no node holds the seat. That is what turns three otherwise alarming moments — a rolling upgrade, a seat moving between nodes, a node that has not finished booting — into non-events rather than lost messages. It is also why every node creates a mailbox for every seat in the company rather than for its own share: the mailbox streams use interest retention, so a publish to a subject no durable consumer covers is dropped in silence.

One trace covers the whole path. The webhook edge starts the span, the wake event carries trace_id and parent_span_id forward, and every event the turn publishes hangs beneath it — so a delivery and the turn it woke are one story at the collector, and the same ids are columns on the event rows whether or not a collector exists.

There are two inbound edges, not one. Six third-party apps plus Atlassian’s Forge relay arrive as verified HTTP on /webhooks/* and take every step above — five of them verified by an HMAC over the body, and Datadog by a constant-time comparison of a shared token, because its provider attaches only fixed-value headers and so has nothing varying with the payload to sign. Mattermost does not: every node holds one websocket per seat, outbound, so it needs no public URL and no signing secret — and it joins the picture only at the republish onto crewlet.notifications.inbound, after the same kind of fleet-wide claim the webhook edge takes, because every node reads every seat’s socket and only one may deliver each post (Running on a fleet). Both edges record each delivery as an inbound_delivery event. Everything from that subject onward is identical for both.

A schedule firing, an a2a_ask from a colleague and a sandbox run completing enter further down still: they publish straight to crewlet.agent.HANDLE.inbox (or .control), and everything from the mailbox onward is identical again.


Every trigger that reaches a seat runs the same two-stage turn — a batch of triggers for one conversation is one turn. What varies is how many rounds it takes, and whether the turn survives its own process.

run_sandbox

minutes or days later,possibly on another node

no_action, and nothing acted

a claim the record refutes

done

delivered

nothing reached anybody

self_iterate — carrying theprior-work ledger

failed

Who is waiting?derived from the trigger's own type,before any model runs

Executorone agentic loop: decide, discover,act, then account for it

Suspendeda detached coding run: the loop is serializedinto the pending-run record and left

Engine checkdoes the record supportwhat it says it did?

skippednobody was askingthis seat to do anything

Revieweris the work any good?

Overridea done that answered in texton a turn somebody is waiting for

done

failed

What the turn leaves behindevents · episode · diary entries ·conversation-ledger entry · token spend

Two nested loops, not one. The outer loop is the turn: executor → reviewer, re-entered on self_iterate up to turn_engine.max_iterations. The inner loop is the model↔tool round trip inside each phase, bounded by its own cap (max_tool_rounds for the executor; a constant for the reviewer, which holds one submission tool). A turn is therefore up to max_iterations × (executor rounds + review rounds) priced model calls, which is the number to reason about when sizing a budget — not “one call per turn”.

One loop where there were two, and that is the point. The engine used to plan in one conversation and act in another; the actor lost everything the planner had read, and the planner had to name its tools in advance against a catalogue it was never shown. The executor decides and acts in one place, so it carries the whole picture — identity, policies, the team roster, and the turn-start prefetch. That prefetch is seven blocks rendered concurrently before the turn starts: the chat thread the turn was woken in, personal memory, relevant knowledge, similar prior work, known counterparty, synthesized skills, first-turn onboarding. The reviewer’s question is narrower and its prompt is smaller: is this round’s work right, given the record. A frontier model can do the work while a cheap one reviews; see Turn Engine.

Tools are discovered, not enumerated. A role with 50–150 MCP tools would push 15–25 KB of catalogue into every prompt, so the executor sees server names and walks list_mcp_server_tools → activate_tool to promote what it actually needs. From the model’s side a builtin and an MCP tool are the same thing: a function it can ask the engine to call.

A suspended turn is the reason there is no parked goroutine anywhere. A detached coding run stops the executor’s loop mid-round with its run_sandbox call unanswered, and the conversation is serialized into the pending-run record rather than held in memory. The run outlives the process: it may be resumed after a restart, on a different node, days later once a person answers the agent’s question. The record carries an explicit version, a permanent reader for the previous one, and a build that understands neither refuses loudly and leaves the row untouched.

Whether anything was DELIVERED is the engine’s judgement, not a model’s. Who is waiting comes from the trigger’s own type before the turn starts; what actually ran is the tool loop’s own record. The two are checked three times, in increasing cost: a delivery claim is refused at decode time where one bounced tool call fixes it; a claim the record refutes loops the round back without spending a review call; and a reviewer’s done on a turn that answered in text where a tool was owed is overturned. The two failure modes this exists for are a seat that composed a reply and never posted it, and a seat that posted it twice.

Nothing about a turn’s shape is ambient. The configuration a turn reads is taken by value at the top of the turn from one immutable epoch, so a config apply landing mid-turn cannot change the round cap between the executor and the reviewer. A turn holds the company it started under until it ends.


Four estates, and which one a fact belongs to is decided by a single question: who has to agree on it? — with the fourth answering a second question the first three cannot: and does everybody have to reach the same answer by the same route? Beside them sits the one thing that is named rather than held by everybody — the bytes of a company’s files, which a replicated row names and one store the whole fleet shares keeps (below).

nobody — it is this node'sown record of what it did

everybody, and each derives itfrom the same ordered log

every node, or the answeris wrong on all of them

it is a message, or arow that has to travel

Who has to agreeon this fact?

This node alone — the node storeone file, one process, exclusively owned

Every node, identically — the replicated storea second file, written by a state log's applier

The whole company — coordination KVa bucket per lifetime, on the stream's own connection

In flight, or keyed — the streams6 message streams + one ordered log per domain

Why the store is two files and not one. A snapshot is a copy of ONE estate: a node too far behind to replay fetches a peer’s replicated file and installs it wholesale, and that file must not carry the donor’s audit log, its learning rows or the bootstrap half of its secret store. Taken from a single file the artefact would be a copy of everything followed by a delete — and with no in-place VACUUM, the deleted pages ride along in the artefact, the transfer, the checksum and the integrity check anyway. No transaction spans the two and no read joins across them.

What each of the four holds, in full:

This node alone — the store.

TablesWhat they hold
crewlet_events · crewlet_event_partiesThe audit log and its party index — this node’s own events, and on a data node the batches of a stateless node’s events it keeps (custody)
custody_unsettledThe custody batches this node has written and not yet learned whether it keeps: the fleet decides which data node keeps each batch, create-only in coordination, and a node that crashed between writing a batch and claiming it asks at its next pass and deletes the rows if another node keeps them. Empty but for the batches of the last few moments
agent_diary · episodesA seat’s notes and turns, each with the vector of the model it came from — recall is a per-seat scan the database ranks, with no vector index
synthesized_skills · synthesized_skill_versions · counterparty_profiles · agent_onboarding_markersThe rest of the learning subsystem — skill induction and its versions, counterparty profiles, first-turn onboarding markers
memory_change_sequenceOne counter, which every new diary note, episode and skill version — and every vector set on a note or an episode — is stamped from, so the memory changelog carries what changed since its last cycle. It only ever increases: a row’s rowid is handed out again once the newest row is deleted, and a watermark over it skipped the next row written
conversation_sessionsWhat this seat already said in that thread
company_config · scheduled_runs · secret_valuesRevisions, cron bookkeeping, and the secret store’s bootstrap half
kb_docs · kb_postingsThe lexical half of the knowledge search index over those rows, built asynchronously behind them and droppable wholesale when the analyzer changes. The semantic half is not here — an embedding costs a provider call, so it is derived once by the fleet and lives in the estate below
page_linksWhich pages and tasks link to which page — the backlinks, derived by the same indexer from the same bodies it tokenises (pages.Links), so a page answers “linked from” with no scan of anybody’s text. Cascades from kb_docs, and rebuilt by the same local walk
statelog_adoptionThis node’s own history of the peer snapshots it has adopted — which donor’s artefact, when the join began and whether it completed. Nothing on the write path reads it: the operation ledger travels inside the snapshot with its own watermark
statelog_divergedThis node’s finding that a log diverged from its rows — the broker was restored from an older copy and another record now sits at the node’s checkpoint — keyed to the stream, the checkpoint and both records’ broker instants. Kept here rather than only in memory because the log can lose the record it was found by (an evicted node’s checkpoint stops pinning the trim), and a restart must not then apply the other history on top of these rows. It holds while the checkpoint names the same record, so a reanchor or an adoption is what ends it
stream_identityWhat this node last saw of each stream’s identity, which is how it notices one that was recreated underneath it. A per-node observation rather than shared state: two nodes can legitimately have seen different generations, so one agreed value would destroy the comparison it exists to make

Every node, identically — the replicated store.

One file, the replicated estate, which every data node holds whole — written by a state log’s applier: records arrive in one order from the log, every node applies the same ones, and the rows plus this node’s position on the log commit in a single transaction. There is no leader and no node whose copy is the real one.

Two writes in the engine are not records, and each is a column or a row no record could own: the clear of a duplicate-rank probe flag each node’s own applier sets, and the inbox sweep over rows whose class lets two nodes legitimately hold different ones. They are named with their reasons in one place, and a third fails the build — the rule and its exceptions are adr/0002, held by internal/store.TestOnlyTheApplierWritesTheReplicatedEstate. A log reanchor is not among them: it writes the framework’s own checkpoint, and the audit row it leaves is applied from the generation record it appends to the adopted log.

TablesWhat they hold
tracker_tasks · tracker_comments · tracker_history · …The company’s work — the tracker’s whole state, derived from CREWLET_TRACKER_LOG
pages_heads · pages_revisions · pages_titles · …The company’s knowledge base, derived from CREWLET_PAGES_LOG: a page’s current body, the immutable revisions behind it, and the title claim that is what makes a name an address
kb_vectors · kb_vectors_bin · kb_ivf · kb_ivf_centroids · kb_ivf_rollout · kb_ivf_listsPage and task embeddings, their 1-bit codes, and the semantic index filing those codes in lists — its head, its centroids and the re-filing its training cut are a record on the same log, and the per-list counts are kept beside the rows they count (ADR-0028) — derived from CREWLET_TRACKER_VECTORS. The fleet pays the provider bill once and every node holds the answer, which is precisely why these are not in the node’s own file
usage_tokens · usage_turns · usage_reads · usage_schedule_runs · usage_person_tokensWhat each node’s seats, schedules and people did each company day — spend by phase, worker, model and provider slot; ended turns and how they ended; the pages seats read; every fire; and what the auxiliary model spent for a person (a question answered on the operator surface), who has no agent id to file it under a seat — derived from CREWLET_USAGE_LOG. Every node publishes its own days and applies everyone’s, so spend history is answered fleet-wide and outlives the node that spent it (ADR-0020). Kept 181 days, aged out by the applier rather than a sweep
statelog_cursor · each domain’s operation ledger and deferred records · statelog_ops_lostWhere this node is on each log, which operations have already been applied — a record that travels inside a snapshot — and how far back that record may have lost rows, to the sweep, so a retry older than that is answered unknown rather than applied twice; and any record a newer build wrote that this one cannot decode

The whole company — coordination KV.

BucketWhat it holds
crewlet_leasesnode: · seat: ownership, and nothing else. The bucket’s age is the lease TTL
crewlet_dutiesworker: ownership, and every other lease — the work tracker’s move: · merge: · bulk: claims among them. Each record is judged by its own deadline; the bucket’s age only has to outlive the longest duty
crewlet_epochsThe monotonic fencing counter. No age at all — see below
crewlet_configThe activation pointer and its payload — the pointer’s own revision is the epoch
crewlet_statusOne key per node: which revision it applied
crewlet_ledger · crewlet_claims · crewlet_firesTurn completions, webhook delivery claims, scheduled-fire claims
crewlet_rebasesThe instant a unit of work’s writes are minted at when its own start is older than the operation ledger remembers, so a crash re-run, a retried resume or the next half of the turn mints where the attempt before it did. The bucket’s age, 30 days, is the ledger’s retention
crewlet_token_windows · crewlet_rate · crewlet_cooldownsThe token counters — one record per scope, a slot for each calendar window, aged 32 days past its last charge — the notification valve, benched credentials
crewlet_secrets · crewlet_channels · crewlet_sandbox_runsThe company’s sealed credentials, open A2A channels, detached coding runs
crewlet_followsThe chat threads each seat follows, so the next reply wakes it whichever node claims that delivery. The bucket’s age, 90 days, is the last-activity horizon
crewlet_objectsThe object store’s backend record — nats or s3:<endpoint>/<bucket>/<prefix>, written create-only by the first node to boot and compared by every node after it, which refuses to boot configured with another store — and the object-collector duty’s last report, replaced by each pass so every node’s /fleet shows it whoever ran it. No age — an expired record would let the next node record a different store and split the company’s files between two
crewlet_custodyWhich data node keeps each batch of a stateless node’s events: claimed create-only by the data node that wrote the batch, after it wrote it, so exactly one node’s log keeps it (custody). The bucket’s age, 32 days, outlasts the event log’s retention, so a node settling a batch it wrote before a crash always finds the answer
crewlet_integrations · crewlet_mailboxesEach surface’s reconcile status, and the seat mailboxes that may exist so a removed seat’s can be retired
crewlet_statelog_positionsFour key classes, all answering what the log may delete: each node’s position per domain; the trim holds a backup or a join takes; what each owner’s newest backup covers, which is the only input the backup term has; and the floor the trim published, with the term holding it and how long it has been holding — the last is the one nothing can re-derive, because a duty that moves on a lease carries no memory across the move. No age at all, and this is the one where an age would be worst — an expired position reads as a node that has applied nothing, which either pins the trim for ever or, read the other way, deletes records that node still needs

In flight, or keyed — the event stream.

StreamSubjects
CREWLET_AGENTcrewlet.agent.>
CREWLET_NOTIFICATIONScrewlet.notifications.>
CREWLET_EVENTScrewlet.events.>
CREWLET_CONFIGcrewlet.config.>
CREWLET_MEMORYcrewlet.memory.> — one message per subject: a keyed table, not a log
CREWLET_DLQdlq.> — deliberately outside crewlet.*
CREWLET_CUSTODYcrewlet.custody.> — the events of nodes without data, in batches, until one data node keeps each; a mailbox (interest retention) aged at the event log’s horizon. Outside crewlet.events.*, whose broadcast every dashboard streams
CREWLET_TRACKER_LOGcrewlet.tracker.log.> — the write-ahead log the replicated estate’s tracker tables are derived from. One subject per object, which is what makes the subject the unit two writers contend on; retention is bounded by durability rather than by age. Two of its subjects carry no object at all: …log.barrier, which every linearizable read appends one record to and then waits for — the acknowledgement is what proves a quorum agrees on a position, where a field read can be served by an isolated former leader; and …log.rankorder.<PROJECT>, which is where a board drag is arbitrated, so two people reordering one project’s board contend and two reordering different ones never do
CREWLET_TRACKER_VECTORScrewlet.tracker.vectors.> — the same shape for embeddings, compacted: one message retained per subject, because the current embedding of a source is the only one anybody wants and a history of superseded vectors is a bill nobody asked for
CREWLET_PAGES_LOGcrewlet.pages.log.>, the ordered log the replicated estate’s knowledge base is derived from, and the state log’s third domain. The same shape as the tracker’s: one subject per object, so two writers saving one page contend at the broker and two saving different pages never do. Retention is bounded by what every node has already applied rather than by age, because a page is a fact for the life of the deployment and removing one is a decision somebody takes rather than a horizon that reaps it while a person is still reading it
CREWLET_USAGE_LOGcrewlet.usage.log.> — the state log’s fourth domain, compacted: one message per (kind, node, company day, seat or schedule), each the object’s whole cumulative value, republished by the node that owns the day whenever it moves and aged out after 181 days. The node is part of every subject, so each object has exactly one writer and nothing is arbitrated
OBJ_crewlet_files$O.crewlet_files.> — the object store’s bucket, crewlet_files, on the default nats backend only: one object per upload, named files/<key> by a key minted for that upload and never reused, and kept as a run of 128 KiB messages in NATS’s own object store format. An ordinary stream at stream.replicas copies, so the company’s files survive what its logs survive and every backup snapshots it with the rest. Absent when store.objects.backend is s3

Named, not replicated — the object store. One kind of state answers “who has to agree on it?” with the row does, and the bytes do not: the content of a company’s files. The row naming a file — its path, its version, its size and SHA-256, and the key of the one object holding its bytes — is in the replicated store like every other tracker row. The object itself, stored once per upload under a key minted for it, is kept in one store the whole fleet shares, named by every node’s Tier A store.objects: by default a JetStream object store bucket on the fleet’s own broker, crewlet_files, backed by the stream OBJ_crewlet_files at stream.replicas copies like every other stream, or an S3-compatible bucket. Every node reaches it directly, a node without data included. Nothing here is derived by replay, and the engine places nothing: the store keeps its own copies, and the one thing the engine runs beside it is the object-collector duty, which deletes the objects no row names, abandons uploads that never finished, and audits that every object a row names is there and whole. See Object Store.

Mailboxes and event history are different kinds of stream. The two mailbox streams use interest retention — a message lives until its durable consumer acks it, which is what makes a seat’s inbox a mailbox rather than a log. The event, config, memory and dead-letter streams use limits retention: events and dead letters age out, config nudges age out within the hour because losing one costs a poll interval and never a revision, and memory has no age bound at all — what a seat should still remember is the learning subsystem’s decision, not the broker’s.

A seat’s memory follows the seat. Memory is written to the node’s store, and placement moves seats — so every memory row also rides crewlet.memory.HANDLE.TABLE.DIGEST, one subject per row, on a stream that retains exactly one message per subject. A node acquiring a seat replays that seat’s rows in a single pass and hydrates them before the mailbox is attached. Deletes deliberately do not travel: the lifecycle re-converges, and a tombstone protocol would be a second thing to keep correct forever.

The line was learned, not designed. Migrations 0010–0013 are what breaking it cost: the delivery dedupe, the completion ledger, the config activation pointer, per-node apply status, the token counter, the A2A channel ledger and the detached-run record all started as tables in one node’s file, where each of them answered a company-wide question with one node’s opinion. They were moved, and the rule is now the one above. See Coordination.

Retention here is a bucket’s age, never a per-write TTL. On the embedded broker a per-key TTL is create-only — an update clears it, leaving the key immortal — so a horizon has to be fixed when its bucket is created, and that is why there are twenty-one of them rather than one with prefixes: three in the lease store, eighteen in the fleet store. The lease store is the sharpest illustration: crewlet_leases has an age, and that age is the lease TTL — a renew rewrites the key and restarts the clock, so a node that stops renewing stops holding and nothing has to notice it died. crewlet_epochs sits beside it with no age at all, because a fence that restarts is not a fence. And crewlet_duties holds the fleet singletons and every other lease apart from the seats, because a duty’s TTL follows its own tick, up to three hours, and a tracker walk’s claim its own heartbeat, and a bucket whose age is the 45-second seat TTL refused every one longer than that. Three buckets, three retentions, for the same subsystem.


One node is the design’s degenerate case, not a lesser path: it runs the API, every seat and every duty, with an embedded stream and local coordination. Nothing about the code path changes when a second node appears — what changes is that ownership starts being contested, and there is already a lease for that.

compare-and-set theactivation pointer

each node polls, applies,reports its own status

node-aingress · seats · workers

node-bingress · seats · workers

sat-euseats · labels: zone=eu

Streamsone durable consumer per seat inbox

Coordination KVleases · activation pointercounters · ledgers · secrets

Load balancer / ingress

PUT /configan operator, on any node

One NATS estate — the company

A fleet — every node is the same binary

The presence lease is the membership service. Every node claims node:{ID} on the same heartbeat as its seats, carrying its roles, its labels and its status. Reading node:* back is the whole of fleet discovery — no gossip, no coordinator, no registry to configure — which is why adding a node is starting a process and removing one is stopping it. Dropping that row is the first thing a drain does.

Placement is deliberately dumb. Every node greedily claims up to a fair share — ceil(seats / live nodes), live nodes being the presence leases of nodes that run seats and have not withdrawn from placement because they cannot serve them, and computed per placement group rather than once fleet-wide, because one global ratio strands the seats that are pinned somewhere. Each group’s share bounds that group alone — a node may not spend room it has in one group on another’s seats, or unpinned seats crowd out the pinned one only it can run — and a seat whose teardown failed is charged against the least constrained groups first. Every node computes the same numbers from the same table and stops there; two nodes racing for the last seat is settled by the lease, not by the arithmetic. It converges in both directions, because claiming alone only converges for a fleet that shrinks. A preferred hint orders the attempt and never gates it, so a rolling deploy tends to land seats back where their MCP children are already warm — and a seat whose preferred node is gone is still taken by somebody.

Acquire, equip, then attach — and release in reverse. A seat is not serving until its instance is spawned, its budget loaded, its MCP children are up and its sandbox runs recovered. Only then is the mailbox attached. The release order is the mirror, so a seat is never left consuming work it no longer owns. That ordering is the one behaviour none of the composed packages can express alone, and it is why internal/node exists.

Every ownership question is three-valued. Held, definitively not held, and unknown — and treating unknown as loss tears a healthy company down over a two-second store blip. Unknown is bounded rather than trusted, on two different clocks: a seat stops admitting work once its last successful renew is older than the heartbeat interval, and is only given up once that renew is older than the lease TTL. Stop taking new work early, hand the seat over late. See Coordination.

/health and /ready answer different questions. /health stays 200 through a drain — an orchestrator that killed a node for reporting unhealthy mid-drain would destroy the turns the drain exists to finish — and it is where the posture is visible. /ready steers traffic, and fails on shed and stuck only: wait and isolated stay ready on purpose, because failing readiness on ordinary rollout lag makes the fastest node the cause of a fleet-wide outage, and stepping out of rotation when no peer has the epoch is not shedding, it is stopping.

History is read from every node, not copied to every node. Each node’s event store holds only what it published, so a fleet has no one store of its turns. A history read — the event log, a turn, a trace, the list of turns — is scattered to every live node at query time over the broker’s ephemeral request/reply and merged by the node serving it, and every such answer carries a coverage naming any node that did not answer inside the two-second fleet read budget (ADR-0021). Replicating the detail would put every prompt and response on every node’s disk to answer a question asked a few times a minute; the price is stated rather than hidden — a node that leaves takes its detail with it, while the aggregates survive it in the replicated usage domain. See Reading the fleet’s history.

A draining node keeps answering both. Its listener stays up until the drain has completed, because the probes are what an orchestrator reads while the turns finish. The door it closes instead is the one to new work: every webhook and every write is refused with 503 from the drain’s first moment. See Graceful shutdown.

One config, agreed by pointer. An activation is a compare-and-set on an append-only pointer whose own revision is the epoch, so two operators activating at once get two revisions rather than overwriting each other. Each node polls it, applies, and writes its own status; a node that is behind picks a posture — serve, wait, shed, isolated or stuck — from what it can see of the pointer and of its peers. A shedding node refuses at trigger admission: the delivery goes straight back to the broker and this node stops consuming, on a seat’s inbox and on the ingress topic alike, so the work moves to a peer that has the epoch rather than being run against a company this node is no longer sure of. Lag alone never sheds — every rollout produces lag, and shedding on it would make the first node to apply the cause of a fleet-wide outage. See Control plane.


Every box in the diagrams above is somebody else’s subject in full. This table is the index; each package also states its own rationale in its package doc, so go doc ./internal/coord is the authority on coordination rather than any page here.

BoxPackageRead
The org chart, seats, handlesinternal/orgOrganization model · Humans in the org chart
The two config tiers, the applyinternal/configConfiguration
Webhook routes, verification, parsersinternal/api/webhooks, internal/whsecJira · Confluence · GitHub · GitLab · Slack · Mattermost
Routing a delivery to a seatinternal/notifyEvent system
Subjects, streams, delivery semanticsinternal/queue, internal/eventsEvent system
Inbox batching and coalescinginternal/queue (the drain and the partition), internal/agent/inbox (the guard order), internal/notify (the merge, and where BOTH keys — the inbox partition and the durable conversation identity — are defined), internal/engine (which of the three runs when)Event system
Seat leases, placement, acquire and releaseinternal/seat, internal/nodeSeat ownership
Leases, buckets, the three-valued answerinternal/coordCoordination
The activation pointer and node posturesinternal/configplaneControl plane
The executor/reviewer loop, workersinternal/agent/turn, runner, toolloop, subagentTurn engine · Agent runtime
The tool registry, MCP children, A2Ainternal/tools, internal/mcp, internal/a2aTools & MCP · Tool capabilities
Knowledge-base-sourced prompt fragments, and keeping every node’s copy currentinternal/agent/skills, internal/agent/skillsyncTool skills
Live knowledge searchinternal/knowledgeKnowledge system
Diary, episodes, skill induction, profilesinternal/learningAgent learning
What a seat already said in one threadinternal/agent/ledgerConversation sessions
Models, the fallback chain, the key poolinternal/providersProvider layer · Subscription LLM backends
Detached coding runsinternal/sandbox, internal/hostboxCode sandbox
Cron-scoped recurring workinternal/scheduleScheduling
The company’s sealed credentialsinternal/fleetsecretsSecret store
The local database and its migrationsinternal/storeDatabase · Backups & restore
REST, the dashboard, the socketinternal/api, static/dashboardAPI endpoints · Dashboard design
Event rows, live projection, tracesinternal/observe, internal/tracing, internal/tokensDeployment
The fleet’s turn-level history, read from every nodeinternal/eventfanEvent system
Which tracker a company runs, and why one of them keeps no state—The tracker
DACI, and the structured ask that is all the engine records of a decisioninternal/trackerDecision framework · Asking for a decision
More than one nodeinternal/seat/placementScaling out · Running a fleet · Satellite nodes

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.