Configuration
Crewlet splits configuration into two tiers so a founder can evolve their company at runtime — add a role, swap an LLM provider, plug in a new MCP server, rotate an integration credential, update a policy — without redeploying or restarting the engine.
Two-Tier Split
Section titled “Two-Tier Split”| Tier | Storage | Owner | Update model | Contents |
|---|---|---|---|---|
| A | crewlet.yaml on disk | Ops / SRE | Restart-only | The store files and the object store’s backend (store.objects), the stream and coordination slots, this node’s identity and roles, API host/port, auth and the optional published listener (api.public), the secret keyring, the retention estate (store.snapshot_dir, stream.tracker_retention and backup ownership, retention.backup_owner), logging (level, shape and an optional rotating log file) |
| B | The store (company_config, versioned) | Founder | Live, API-editable, validated, versioned | Everything else: name, mission, vision, policies, the company’s one clock (timezone), providers (LLM + embeddings), turn engine, learning, MCP servers, delegate workers, scheduling, skill_variables, the tracker and knowledge backends, notification coalescing and rate limit, integrations (Slack / Mattermost / Jira / Confluence / GitHub / GitLab / Datadog / Atlassian, plus the Forge app id that verifies relayed Cloud events), org roles & units, and token budgets |
Tier A controls how the engine boots. Tier B is what the company is.
The company’s clock is Tier B for that reason: where a company keeps its hours is a fact about the company, not about the hosts running it, and every node must cut the same day from it. So it is one top-level timezone in the company document rather than a host’s own clock — Local is refused — and rather than a zone per subsystem: the tracker’s “today”, a person’s day and the hour a schedule fires on are one calendar. See The company’s clock.
Tier A example (crewlet.yaml)
Section titled “Tier A example (crewlet.yaml)”logging: level: info # debug, info (default), warn, error format: console # console (default), text, json stderr: true # default. false hands the stream to the file below, # and needs one — a node logging nowhere is refused file: # optional — a durable copy, IN ADDITION to stderr path: "/var/log/crewlet/${CREWLET_NODE_ID}.log" format: json # empty follows logging.format level: debug # empty follows logging.level max_size_mb: 100 # rotates here; there is no "never" (default 100) max_backups: 5 # `.1` (newest) … `.5`; 0 keeps none (default 5)
node: id: "node-0" # optional; see below
stream: type: embedded # a JetStream server inside this process; `nats` points # the same slot at an external server or cluster store_dir: "./crewlet-data/stream"
store: path: "./crewlet-data/company.db" # ONE file, owned exclusively
coordination: type: local # one node; a fleet needs `embedded-kv`
api: host: "0.0.0.0" port: 8000 auth: tokens: - id: founder token: "${CREWLET_API_TOKEN_FOUNDER}" - id: ops token: "${CREWLET_API_TOKEN_OPS}"node.id
Section titled “node.id”Names this process within the company. It labels every log line and the
/health payload — the difference between “a config apply failed” and “the
config apply failed on node-2” the moment more than one process is
running, and the only way a caller behind a load balancer can tell which
process answered.
Resolution order: node.id (a ${VAR} reference works here, as in any
Tier A text field, and the id it resolves to is judged by the rules below —
see Environment Variable References
for how one reads in a number or a boolean field) → the CREWLET_NODE_ID environment
variable → node-0. You do
not need to set it to run a single engine. It starts with a letter or a digit
and holds only letters, digits, ., _ and -, at most 64 characters,
because it ends up in broker consumer names and subjects — and an operator
gesture naming a node, such as crewlet retention evict, holds the id it is
given to the same rule.
It must be stable across restarts, which is why it comes from the deployment rather than being generated per boot: anything the process registers under its identity would otherwise be orphaned on every restart. In Kubernetes use the pod name — a StatefulSet ordinal is ideal; under systemd, the host name.
node.roles
Section titled “node.roles”What this process is willing to do. Four roles, and the default is all four — one process running a whole company, which is every single-node deployment:
node: id: "${CREWLET_NODE_ID}" roles: [seats] # a stateless satellite: agents only, no data labels: zone: eu| Role | What it does | What a fleet loses without it |
|---|---|---|
data | Keeps the company’s durable state on this node’s disk: a copy of the replicated estate and the event log. The company’s files are not under it: they are in the object store, which on the default nats backend is a stream the broker’s members keep. It says nothing about the broker (see below). ingress and workers require it | Every seat’s tracker and knowledge tools, which a node without data answers through one that has it |
ingress | Serves the HTTP API: webhooks, the dashboard, the REST endpoints. A node without it serves only its probes on api.port | No integration can reach the company, and there is nothing to look at |
seats | Claims seat leases and runs agents, and serves their agent-mode tool bridge (/mcp/{token}) when CREWLET_MCP_BRIDGE_URL is set | Every trigger queues up unread |
workers | The company-wide singleton duties: the scheduler tick, the maintenance sweep (retention and removed-seat mailbox retirement), the state logs’ trim, the embedding duty, the sandbox waiter, the integration reconcile loop, the object store’s collector, and the learning background passes (episode lifecycle, skill curation, clustering and promotion) | Nothing fires on a schedule, no sandbox run is collected, no table is swept, no state log is ever trimmed, no knowledge is embedded for semantic search, no integration is reconciled, and no unnamed object is collected |
Subtracting a role subtracts it from this node, never from the
company, so the fleet as a whole still needs every role somewhere. That
is a shape no single node’s config is wrong for, and every symptom of
getting it wrong is an absence — so the engine checks it against live
node presence and logs fleet_role_unmanned when nobody is doing a job.
A node that does not run seats is also left out of the denominator its
peers divide the seats by; counting it would strand the difference.
data is the one role that is a promise about the disk rather than about
work. A node without it keeps nothing that has to outlive it, and Tier A
holds it to that whatever its broker: store.scratch: true is required (its
store is deleted at every boot), and store.replicated_path and
store.snapshot_dir are refused. store.objects is not: it names the one store the whole fleet keeps its files in, and a node without data uploads and downloads through it directly. See
Running a Fleet.
The broker is not a role. How a node’s broker takes part in the fleet’s
is its broker kind, derived from the stream block and from nothing else:
| Broker kind | The stream block that makes it | What the broker is |
|---|---|---|
member | embedded, with no stream.leaf.urls (a solo node included) | runs JetStream in this process, holds stream replicas, votes in the metadata group |
leaf | embedded, with stream.leaf.urls | JetStream off: reaches the members’ across a leaf link, holds nothing, votes in nothing |
client | type: nats | a plain client of an external cluster somebody else runs |
Every node advertises its kind on its presence lease (crewlet fleet broker list and the dashboard’s Settings › Nodes screen show it), and a node advertising a
kind this build does not know — a newer build’s — shows as unknown. The rules that are about the broker
are asked of the kind, whatever the roles: a member of a fleet — one naming a
cluster, peers or a leaf listener — must set stream.store_dir, because it
holds the fleet’s streams for every node that reaches it; a leaf carries no
cluster block, no store directory, no leaf listener and no declined
stream.sync — it runs no JetStream, so it has no file store to keep.
Three pairings of roles and broker are refused, and each refusal names
both ways out: a data node on a leaf, a broker member without data, and
ingress or workers without data. Every data node holds the whole estate
as a member of the broker, and a node without data reaches the estate
through one — so a node without data on an embedded stream joins the fleet
as a leaf through stream.leaf.urls. crewlet validate also warns about two
valid shapes: a member whose peer list names more than four other members
(no stream keeps more than five copies, so a sixth member is a voter holding
nothing — a fleet grows by adding leaves), and a fleet node that left
node.roles undeclared and so runs every role.
node.labels
Section titled “node.labels”Free-form facts about where this process runs, matched by a seat’s
role.placement selector. Values are
strings and are compared exactly — so both the key and the value are
trimmed of the whitespace around them before anything reads them, and two
keys that are one key once trimmed are refused rather than collapsed. A
key is at most 63 bytes — Kubernetes’ limit on a label name — and holds no
whitespace or unprintable character anywhere, because it is matched
exactly wherever it is named: a key with a space inside it is refused
rather than accepted as a label nothing could name. Labels are advertised
to peers on this node’s presence lease, so a label change takes effect
one heartbeat after the restart that made it — not at the next config
activation.
Nothing here means anything to the engine on its own: the org decides what to select on.
Tier B example (company.yaml)
Section titled “Tier B example (company.yaml)”Everything that defines the company — see examples/nimbus.company.yaml for a complete working document, and the configuration reference for every field including the ones that example does not use.
Bootstrap Sequence
Section titled “Bootstrap Sequence”The engine boots in this order:
- Read
crewlet.yaml(Tier A only — the store path, the stream, coordination, api host/port/auth, secrets, logging) logging.Configure(level, format, stderr)— once, incmd/crewlet, which is the only thing that sets the console destination; a later command changes how loud it is withSetVerbosityand keeps the sink already installed. The level and format come from the file’slogging:block with any-log-level/-log-format/-debugflag layered on top, and only where the flag was actually given. Lines emitted before this — the file’s own${VAR}warnings, a refused field — come out under the flags alone, which is the best a process can do about a file it has not opened yetlogging.SetFile(…)iflogging.file.path(or-log-file) names one — a second destination by default: stderr keeps every line, and the file gets its own handler, so it can carryjsonatdebugwhile the terminal keeps its columns atwarn.logging.stderr: falsehands the stream over to the file instead, and is refused with no file to hand it to. It is opened here, before the store and the stream, so the failures those can produce are in it. A path that cannot be opened stops the boot, naming the path: every other bad logging value resolves to a default, but a durable record an operator asked for and silently did not get has nothing pointing at why. The same ordering has a corollary — a boot that fails on the Tier A document itself never reaches this step, so stderr is the only record of it,logging.stderrnotwithstanding- Hold the Tier A this process is about to run to Tier A’s rules once more —
the file with every
-roles,-api-hostand-api-portflag applied — and then to the rules that need both tiers. The engine does this itself, before it opens anything, because it is the one thing that runs Tier A: the file was validated as it loaded, but a flag, or a bootstrap a tool built in code, reaches the engine having passed no rule, and a node that opened its store first would have migrated a database for a config it was always going to refuse - Open the store file and start or dial the stream
- Run migrations — every file, in one pass. There is no lock and no phase ordering to serialize: this process owns its file, so nothing can be racing it, and no DDL depends on a value only the config knows. Embedding columns are declared as plain blobs and the vector width is validated in Go against the active revision at write time, so a schema step never has to read the config first (see
crewlet migrate). - Start the API inside this process, bound to
api.host:api.port, wire up auth middleware, register/config/*routes — and, whenapi.public.portis set, the routes outside parties call (/webhooks/*,/otlp/{token},/mcp/{token}) on that second listener instead (see Deployment → Exposing webhooks without the admin API). A node without theingressrole serves only its/healthand/readyprobes onapi.port, and a seats node’s agent-mode tool bridge beside them, or onapi.public.portwhen that is set - Start the control plane — the reconcile loop that polls the activation pointer, plus a broadcast
crewlet.config.revision_activatednudge that wakes it early SELECT payload FROM company_config WHERE is_active <> 0- Row present: apply the payload, which spawns the full company
- No row: engine stays in the unconfigured state — the API keeps serving so an operator can push the first revision via
PUT /configorcrewlet config import
A boot that fails leaves nothing running. Any step above can fail — an
unreachable broker, a keyring the node cannot open, a provider whose model is a
${VAR} nothing sets — and when one does the node unwinds everything the steps
before it started, in the order a shutdown uses: the state log’s apply loops and
its snapshot donor, the shared MCP child processes, every duty loop, this node’s
publish admission, and the store file and stream it opened itself. So a
supervised crewlet run retries onto a clean host rather than contending with
its own previous attempt — which matters most for the store, since one process
owns that file exclusively and a second open behind a leaked handle fails with
a message about locking that names neither the original failure nor the file.
The engine_boot_abandoned line is what says the unwind finished.
Two equivalent bootstrap entry points
Section titled “Two equivalent bootstrap entry points”Option 1 — bootstrap before run (CLI):
crewlet config import company.yaml # one-shot bootstrap of Tier Bcrewlet run # boots from ./crewlet.yaml + the storeOption 2 — run first, configure the running node:
crewlet run # boots in UNCONFIGURED statecrewlet config import company.yaml -summary "initial bootstrap"# Detects the running engine, goes through its API, and every node# reconciles onto the new activation epoch — no restart needed.The same thing by hand, which is what the CLI is doing:
curl -X PUT https://crewlet.example.com/config \ -H "Authorization: Bearer $CREWLET_API_TOKEN_FOUNDER" \ -H "Content-Type: application/yaml" \ -H "X-Summary: initial bootstrap" \ --data-binary @company.yamlThis is also how you change a company that is already running: the same command, against a node that already has one.
crewlet run defaults -config to ./crewlet.yaml and -company to ./company.yaml in the working directory; naming a path is only needed when a file lives elsewhere. A missing ./company.yaml is not an error — the store is authoritative, so a node with no Tier B file boots on whatever the store holds, or unconfigured when it holds nothing.
-company only ever bootstraps an empty store. Once a company exists the file is ignored, with a company_seed_ignored warning naming it and the active revision, so a restart cannot revert a change made live. Two flags cover the other intents:
| I want to… | Use |
|---|---|
| fill an empty store at first boot | crewlet run -company company.yaml (the default) |
| change a running fleet, no restart | crewlet config import company.yaml |
| make a file the company again on restart | crewlet run -import-company company.yaml |
-company and -import-company together are refused: they ask for opposite things, and picking a winner silently would be picking it about the flag that overwrites a running company.
Unconfigured State
Section titled “Unconfigured State”Until the first active row exists, the engine holds an empty Organization (no name, no roles, no units), an empty provider map, no running MCP processes, no integration clients, and no notification transports.
What stays running:
- The Tier A resources — the stream, the store file, the API socket — all up.
- The API’s
/config/*routes and the node’s reconcile loop — which is exactly what wakes an unconfigured node when the first revision lands. - A structured log line carrying
state=unconfigured, so the unconfigured posture is obvious in logs and on the dashboard.
What returns degraded responses:
| Route | Behaviour while unconfigured |
|---|---|
GET /health | 200 {"status": "unconfigured", "node": "node-0", "configured": false, ...} — 200 because the status code is liveness; read configured for readiness |
GET /ready | 503 {"ready": false, "configured": false, "reason": "unconfigured", ...} — an unconfigured node cannot verify a webhook signature, so it stays out of rotation. reason names which of the four took it out |
GET /config | 404 {"error": "no_active_revision"} with a hint |
GET /config/revisions | 200 [] |
PUT /config | Accepted, and creates the first active revision, as long as the FLEET has no activation either: a node that has not caught up with its fleet answers 412 already_configured (with If-None-Match: *) or 409 revision_advanced, naming the revision the fleet is on. Send no precondition, or If-None-Match: * to insist nothing is configured yet; an If-Match names a revision to match, so it answers 412 no_active_revision |
POST /config/revisions/{id}/revert | 404 — no revisions exist yet |
Per-entity routes (PUT /config/roles/{handle}, etc.) and PATCH /config | 409 Conflict — they edit a document, and there is none; initialise via PUT /config first |
GET /agents, GET /tokens/breakdown | 200 with empty lists / zero counters |
POST /webhooks/... | Signature check still runs (a forgery is rejected as a forgery); body logged at WARNING; returns 503 {"status": "unavailable", "reason": "unconfigured"} with Retry-After so the sender retries. A 200 here would tell the sender the delivery was accepted while discarding it — silent, unrecoverable loss the moment one process of several has simply not caught up yet |
Transition out of unconfigured: the first activation moves the pointer → the reconcile tick picks it up → the apply runs → the spawn cascade executes, including everything boot starts only for a company it already has: the native tracker and knowledge base with their tools and their chart, the code sandbox’s coordinator and completion poll, the reflect dispatcher and the inbound edge (see the native, sandbox_runtime, learning, integrations and maintenance stages below) → the engine is fully alive, with no restart. The dashboard carries the unconfigured state in always-on chrome (a caution banner saying inbound webhooks are being refused, an engine pill that says so, and the first row of the overview’s attention queue), and it clears automatically on the next health tick once /health reports configured: true. See the attention queue.
A Company With No Model Provider
Section titled “A Company With No Model Provider”An empty providers.llm is a valid company: an org chart written before its
credentials exist, and what a company created in the dashboard’s
org builder is until somebody
adds a provider. crewlet validate accepts it (reporting 0 LLM providers),
PUT /config and its dry run accept it, and every node applies it like any
other revision: the fleet view reports ok, the agent seats are placed, and
their mailboxes attach and keep what arrives.
What no seat can do is take a turn, and each node handles that the same way:
- It logs
company_has_no_modelswhen the revision becomes current, namingproviders.llm. This is the line to look for on a company nobody has messaged yet. - It holds every delivery to a seat on that seat’s inbox. The inbox is paused
and the delivery requeued, never consumed, and the node logs
seat_inbox_pausedfor the seat, namingproviders.llmagain. - The apply that adds a provider releases every inbox the node paused
(
seat_inbox_resumed), and the held work runs on the new provider. Nothing sent to a seat in the meantime is lost.
Schedules do not fire while the company has no provider: a fire is work on a
seat’s inbox, so firing through the wait would stack every missed standup
behind the hold and run the backlog all at once. The scheduler stays disarmed
and each node logs schedules_waiting_for_a_model; the apply that adds a
provider arms it, and its first tick catches up at most the most recent missed
fire (Scheduling).
The learning work that calls a model is not built for such a company: the persist decider, the skill synthesizer and refiner and the counterparty profiler that run after each turn, and the background compaction, clustering and promotion passes. The apply that adds a provider builds all of them. The skill curator calls no model and runs as usual.
A delivery is held rather than failed on purpose. A turn that cannot build its runner proves nothing reached outside the engine, so the dispatcher would hand it back to the broker, which would redeliver it until its delivery budget ran out and then drop it: every message sent before the provider arrived would be retried pointlessly and then lost.
Live Propagation
Section titled “Live Propagation”When a new revision is activated (via PUT /config, PATCH /config, a per-entity write, a revert, or crewlet config import), the revision is stored and the fleet’s activation pointer is then moved to it; the pointer’s own KV sequence is the epoch, so the append and the flip cannot come apart. Every node polls that pointer and converges onto it; a broadcast crewlet.config.revision_activated event wakes the poll early but carries no work.
The two steps are not one transaction, and they span two stores: the node’s own database and the coordination KV. That ordering is deliberate: a crash between them leaves a revision nothing points at, which is inert and replaced by the next activation, where the other order would point the fleet at bytes no node had stored. The revision is stored as history and becomes the writing node’s own active revision only after the pointer has moved, so a write the fleet refused is never what that node serves or republishes; see Control Plane.
There is no leader, so any node’s API may write. What keeps two operators from silently overwriting each other is that the flip is a compare-and-set against the revision the write was derived from: the loser gets a 409 naming what won, rather than a 201 for a change the fleet never took. See Concurrent writes.
The nudge carries no work, and that is the load-bearing part. A group subscription is a work queue — the contract says exactly one member of a group receives each message — so a revision delivered that way would be applied by whichever process won the message while every other node kept serving the previous company indefinitely. Polling a pointer inverts that: every node reads the same value and converges on it, and a lost nudge costs a poll interval rather than a revision. The full mechanism — the activation pointer and its epoch, what a lagging node does about its own traffic, and the operator surface — is Control Plane; what follows is the apply itself.
The engine half
Section titled “The engine half”Converging applies the payload. Engine.Apply is a straight line with no
comparison and no early return — there is no apply lock, no payload
short-circuit and no rollback of captured state. It rebuilds the whole epoch,
in a fixed order, and names each stage it got through.
Before the first stage it checks the rules that need both configuration tiers — the one today is that a native tracker or knowledge base may not keep its log on an in-memory stream — exactly as a boot does. A revision that breaks one is refused with nothing touched and an empty stage list; before this ran at every apply, a node that booted unconfigured accepted such a company and started the log whose first restart leaves the node unable to serve.
secrets— re-read the secret store and install a fresh resolver snapshot. First, because re-activating an unchanged revision is the documented rotation gesture: the payload has not moved, so the only thing that can have is what its${VAR}references resolve to.company— validate and build the new epoch, resolving${VAR}where each provider is constructed. A refusal here changes nothing: this node keeps serving the previous epoch. A company with noproviders.llmis not refused: it builds with no model registry (see A Company With No Model Provider). The revision’s sandbox catalogue (providers.sandbox) is built straight after, and one that cannot be built refuses the apply there — reportingsecrets, companyand nothing else, since nothing below has touched the node yet. On every node, whether or not it runs a sandbox yet: the alternative publishes a company whose sandbox-enabled seats plan around a box nobody can mint.native— bring up the native tracker and knowledge base on a node running neither, for its first company on them: the state log (joining the fleet by snapshot first when the node is behind), the read and write sides, the lexical index, and — on a node that publishes — the trim, the embedding duty and the change feeds. Beforetools, because the native tools are registered only where the runtime exists, and beforeintegrations, whose inbound edge includes the native parsers. A start that fails refuses the apply. Conditional: reported only on the apply that started it — once per process, and never at all on a node that met its company at boot. The runtime follows the fleet’s logs rather than a revision, so a later refusal of the same apply does not take it down, and which halves it runs are the ones that first company declared: changingtracker.backendorknowledge.backendafter that still takes a restart.sandbox_runtime— bring up the code sandbox’s coordinator and completion poll on a node running neither, for the first revision whoseproviders.sandboxreaches a cell, and prepare the seats this node already holds: each was taken before there was anything to prepare, so each now attaches its completion topic and recovers the runs recorded on it, as a seat taken afterwards does on its way in. A seat whose preparation fails is handed back (a voluntaryunpreparedrelease), so its next acquisition, here or on a peer, runs the whole of it. Beforetoolsfor the reasonnativeis:run_sandboxand an agent-mode executor are offered only where the runtime exists. A start that fails refuses the apply. Conditional: reported only on the apply that started it — once per process, and never at all on a node that met a sandbox company at boot. Like the native runtime it is a fact about the process rather than a revision, so a later refusal of the same apply does not take it down.tools— equip the new epoch with this node’s builtins. An epoch is published, never mutated, so each one gets its own registry; a node that equipped only its first would serve a company whose agents silently lost every builtin at the first config change.learning: rebuild the reflection workers against the new org. Deliberately cannot fail the apply: reflecting against a stale org is a far smaller wrong than not reflecting. The one exception is a node’s first company: a node that booted with none has no reflect dispatcher to swap workers into, so this stage builds it, and a build that fails refuses the apply for the reason it fails a boot (a company served without it learns nothing while looking healthy). What feeds the dispatcher is not this stage’s: each seat’s reflection subject is attached by the node that acquires the seat, beside its mailbox.sandbox— swap the sandbox manager only. The coordinator and waiter hold this process’s busy set and poll loop; rebuilding them would forget which seats are mid-run and start a second loop over the same rows. The swap carries the backend of any cell the revision dropped as retired, for the runs still on it, and a revision with noproviders.sandboxkeeps the last manager for the runs in flight (see Code Sandbox). It cannot fail — the catalogue was built atcompany— and it runs after every stage that can refuse a node already serving a company, so a refused apply never leaves this node launching through a catalogue its current epoch does not have. Conditional: only where this node runs a sandbox runtime, from its boot or from asandbox_runtimestart.parties— rebuild the party index before the epoch is published, so a seat the revision adds is addressable the instant the epoch carrying it is current.integrations: rebuild the inbound surfaces against the new epoch (Confluence, Datadog, Jira, GitLab, GitHub, and the two chat transports, Slack on every apply and Mattermost when a value it is built from moved), so work items route by the new chart rather than the boot-time one. A third-party app the revision retires (its block removed, orenabled: falsefor GitHub and GitLab) has its parser unregistered, so its deliveries route to no seat; GitHub’s and GitLab’s webhook routes then answer503rather than verifying and ingesting a delivery the routing half would drop. Confluence additionally loses its searcher, or every seat would go on searching a wiki the company has removed, with the credential it revoked. Confluence and Jira re-derive a lead map from the org (space and project key to unit lead), which is what an unrouted page or issue falls through to. GitLab and GitHub have no lead map; theirs re-resolves the engine credential and the participants lookup that fans a thread out to the seats on it. On a node that booted with no company there is no inbound edge to rebuild: boot starts one only for a company it already has, because the inbound consumer group is fleet-wide and a node with no parsers would take deliveries its peers can route. So the node’s first apply starts the edge here, through the same function boot runs, and every later apply reconciles it. A start that fails (the broker refuses the subscription) refuses the apply like a refused build, before the epoch is published, and takes down whatever it had brought up, so the retry starts from nothing. The native tracker’s org chart is reconciled at this stage too. Its projects are stamped with the activation being applied (the pointer’s own instant, to the millisecond and identical on every node), never with the time of the apply. So a re-apply or a restart on the same activation writes nothing, and when several nodes apply one activation at once they contend per project and exactly one writes it (see Work tracker).epoch— publish the new epoch. This is the swap; everything before it built, everything after it reads the now-current company.maintenance— rebuild the retention sweep, on a node’s first company (or the apply that started the native runtime). Its job list is read once, and on a node that booted unconfigured once was before there was a runtime to contribute the operation ledgers, the tracker’s repairs and the inbox sweep, or a company to state the conversation horizon — so it swept none of those until a restart. After the swap, because it reads those horizons off the current company. The old sweep stops, its in-flight tick waited out, before the new one starts. Conditional: only on that first apply, and only on a node that publishes.seat_tools: rebuild the registry each seat this node holds runs against. After the swap, because that registry is a clone of the current epoch’s surface: a seat’s per-role children are filed into a copy of the builtins plus the shared servers, so a new epoch leaves the copy stale. The children themselves are deliberately untouched (they belong to the seat’s lease, not to the epoch; see below), so what is rebuilt is the catalogue a turn is built against, never a process. Reported on every apply, including one where this node holds no seat with per-role children and there is nothing to rebuild.mailboxes: ensure a mailbox exists for every seat. After the swap, because it reads the seat list off the current company, and until something creates a new role’s mailbox every event published to it is dropped rather than retained. When the new epoch has a model provider, this stage also releases every seat inbox the node paused while the company had none, after the seat tools are rebuilt, because the first thing a released inbox does is run a turn. In the same place and for the same reason it re-judges every seat parked on its token budget against the ceilings this revision states, and drops the pause of any seat the revision removed, so a seat later added under the same handle does not arrive paused by somebody who paused a different role. Conditional: only where the engine has a node, becausecrewlet validateapplies to nothing.learning_passes: hand the background learning loops (episode compaction, the skill curator, clustered synthesis and promotion) the passes this revision turns on, built from its models, credentials and knobs. After the swap, because the loops walk the current company’s seats: handed over earlier they would run the new revision’s passes over the previous company’s roster, and a refusal later in the same apply would leave them there for a revision this node never served. The loops themselves are armed once per process and keep their clocks across an apply (see Agent Learning), so this is also where a node that booted with no company, or a company that gained its first provider, starts running them. Reported on every apply, including one on a node with no store or no worker role, which has no loops to hand anything to.scheduler: re-arm the cron loop. After the swap too, and for a sharper version of the same reason: the tick reads schedules off the current company, so arming early would open a window in which the loop fires the outgoing company’s crons.
Then a config_revision_applied event is published on
crewlet.config.revision_applied with status, the applied_subsystems list
and any error.
A failure is reported by how far it got, not undone. Every refusal above
happens before the epoch swap, so it leaves the previous epoch current and
serving — that, rather than a rollback, is what makes a failed apply safe. The
returned list is the stages that did complete, in the order they completed,
so “secrets, company” names both what was rebuilt and where the refusal landed.
It travels on config_revision_applied into the audit event log, where it
outlives the fleet view’s one-minute bucket. The fleet view carries each node’s
epoch, revision, status and failure text but not the stage list, so that
detail lives on the event rather than on the operator surfaces reading the
bucket. The active row stays active either way; the control plane records
the outcome so peers can see it (see Control Plane).
Read that list by name, never by number. Five of the fifteen stages are conditional, so an ordinary apply on a node already serving a company that runs no sandbox reports eleven names and the swap is the seventh of them. The numbering above is the order the code runs, not an index into what a node reports.
A stopping node applies nothing. Stopping waits for an apply already
running, which returns quickly on the cancelled context that asked for the
stop, and refuses every later one with error. An apply that ran on past the
teardown would start again what it had just ended: the scheduler, the
background learning passes, and on a node’s first company the inbound edge.
“No rollback” is not “no mutation”. What the build-first ordering buys is
that a revision which cannot be built changes nothing: NewCompany
validates, resolves the org and constructs the providers without reaching the
network, so stage 2 is the cheapest place to refuse and the one that costs
nothing at all — the sandbox catalogue included, which is built there too.
Past it the guarantee narrows. On a node that already serves a company,
tools is the last stage that can refuse (an embeddings width it cannot
serve, below): learning and sandbox cannot fail there, parties and
integrations return no error, and epoch is the swap. By the time a
tools refusal lands, two things are already mutated: the resolver snapshot
(secrets), and what tools did before it refused — any shared MCP child
whose spec moved, and the skill variables. So that refusal leaves this node’s
shared tool processes on the new company while it still serves the
previous epoch, and reports error. A sandbox runtime the same apply brought
up (sandbox_runtime) stays up too, and harmlessly: the previous epoch offers
no seat run_sandbox, so it launches nothing, while the runs it already
polls and the seats it prepared are the fleet’s rather than the revision’s.
The reflection workers, the party index and the trackers are not among them:
they are rebuilt after the last failure point, which is why they are ordered
there. A node’s first company is the one exception, on both sides of
tools: native refuses it when the runtime cannot start, learning when
the reflect dispatcher cannot be built, and integrations when the inbound
edge cannot start. None of those refusals has a previous epoch to protect, so
what it leaves behind (the shared MCP children, a built dispatcher, the
party index, a native runtime catching up on the fleet’s logs) serves nothing
until the retry the refusal earns, which finds it already there. Widening that window is what would make degraded
reachable, which is why everything an apply cannot un-apply stays behind the
swap.
One knob is refused rather than applied live, because applying it would corrupt data rather than merely disrupt it:
providers.embeddings.dimensions— rows already written carry vectors of the old width, and a similarity query across two widths compares nothing. The apply fails naming the declared width and the width the store already holds. Moving a store from one width to another is a restart, because its vector columns are sized when it opens; after it, the knowledge corpus re-embeds itself, since its vectors are kept per model and width, and the node holding each seat re-fills that seat’s diary and its raw episodes, which store the whole text their vectors are of. Similarity recall compares only vectors of the query’s model and width, so until a row is filled it is not reached by similarity: personal memory reaches a note through its recent half,## Similar prior workleaves a turn out, andquery_episodesreaches it withconversationor by recency — and says how many turns its search could not reach. A change ofproviders.embeddings.modelat the same width is live, and for the same reason it needs no refusal: every stored vector names the model it came from, recall compares only rows of the query’s model, and the diary and the raw episodes are re-filled in the new space the same way, never ranked across two spaces meanwhile. Adding or removing the wholeembeddingsblock is live in both directions: a company that drops it degrades on the next turn to keyword-only search, personal memory chosen from each seat’s recent notes alone, and no recall of similar prior work, since a seat’s recent turns are not similar work.
providers.sandbox is live in every direction. Adding it brings the code
sandbox up on the next apply (sandbox_runtime), on a node that booted with no
company or with a company that had none; changing it swaps the manager
(sandbox) with the runs already in flight kept whole, a dropped cell’s
running jobs included; and removing it stops new launches while the runs in
flight finish. A block that cannot be built is refused on every node, at
company. See Code Sandbox.
Token caps need no apply stage at all, because there is no cap set to
maintain. Usage is shared and caps are not: the fleet’s counter stores only
what each scope has spent, and the limit travels in as an argument on every
charge, read straight off the epoch the turn is pinned to. So a revision that
raises a ceiling takes effect on the next turn on every node at once, with
nothing seeded and nothing to drop when a seat goes away — a role removed,
flipped to human, or dropped to 0 (= unlimited) simply stops having a limit
passed for it. The alternative, caps replicated per node, is what makes an org
ceiling of 500 000 quietly become N × 500 000.
Shared MCP children are reconciled per server against the new epoch’s specs,
comparing every field: a change to url or headers restarts only that
server, and every other child keeps running and keeps contributing the tools
it is already serving. The comparison is over resolved values, so a rotated
credential reads as a changed spec — see
Rotation.
A server the revision removed is stopped rather than merely dropped from
the catalogue, and that is a separate step because the reconcile above is
driven by the specs the current config names — so it never visits a name
that is gone. Without it the child kept running until the engine stopped,
holding the company’s credentials, and a rename ran two of them. Retiring
happens before the reconcile, so a rename frees the name and the tools it
published before the replacement files its own. Taking a leaking integration
offline by deleting its mcp_servers entry therefore takes effect on the next
turn, which is what the rest of this section already promised.
Per-role children are not on this path. They belong to a seat’s lease
rather than to the epoch: the apply-time reconcile skips every non-shared
server, and a role’s mcp_env — which carries the per-agent Slack/GitHub
credentials — is never part of a spec an apply compares. Such a child is
spawned when its seat is claimed and torn down when it is released, so an
mcp_env change reaches it when the seat next changes hands, not on the apply
that carried it.
What an apply does rebuild for a held seat is that seat’s registry — the
seat_tools stage above. The two are separate on purpose: the registry is a
clone of the epoch’s surface and goes stale the moment a new epoch is
published, while the child is a running process holding that seat’s
credentials and must not be restarted for a config edit that did not name it.
So an apply re-files the live children’s catalogue into a fresh clone of the
new epoch’s builtins and shared servers, and the child never learns it
happened.
The API half
Section titled “The API half”There is no second projection to keep in step. The API answers every read
through closures over the engine’s current epoch rather than from a cached
copy of the payload — GET /agents, /org and the dashboard’s queries all
resolve against whatever Apply last published, and the inbound webhook
secrets are re-read per delivery the same way. So a rotated signing secret is
picked up by the epoch swap itself; there is nothing that could drift stale and
nothing to refresh.
There is one wiring, because there is one process: the API is served inside the
engine’s, over the engine’s own backends, and every node runs exactly one
reconciler whatever its node.roles. So the two halves cannot disagree about
which epoch they are on.
configured is derived the same way as everything else: it is true when the
engine’s current epoch holds a company, and it flips the moment an apply brings
a node its first revision. A failed apply leaves the node serving the previous
epoch, which is still a configured node. What a failed apply changes is the
posture, and /ready reads that. Collapsing the two would take a
correctly-serving node out of a load balancer’s rotation for being behind.
Versioned Revisions
Section titled “Versioned Revisions”company_config is an append-only table — every change is a new revision. The schema:
CREATE TABLE company_config ( revision_id TEXT NOT NULL PRIMARY KEY, parent_revision_id TEXT REFERENCES company_config(revision_id), created_at INTEGER NOT NULL, -- unix microseconds, UTC created_by TEXT NOT NULL, -- a label: token id ("founder"), a -- login, a node id, "reconcile loop" created_by_kind TEXT NOT NULL DEFAULT '', -- what created_by names: -- "operator" | "node" source TEXT NOT NULL, -- "api" | "file" | "rekey" summary TEXT NOT NULL, -- short human-readable change note payload TEXT NOT NULL, -- the whole document as JSON, or the -- sealed envelope when a keyring is set is_active INTEGER NOT NULL DEFAULT 0, activated_at INTEGER);
-- At most one active revision, enforced by the database rather than by the-- application remembering to.CREATE UNIQUE INDEX company_config_one_active_idx ON company_config (is_active) WHERE is_active <> 0;The types here are the four SQLite has (TEXT, INTEGER, REAL, BLOB)
rather than UUID / TIMESTAMPTZ / JSONB, and a timestamp is unix microseconds
rather than a date type — Turso is SQLite-compatible in both its query language
and its file format, so that is simply what a column can be. It was also, until
recently, the intersection of two drivers’ dialects; the second driver is
retired and the types are unchanged, because they were never the narrow part.
A revert creates a new revision whose payload equals a prior one — the audit chain stays intact via parent_revision_id.
What a stored revision is held to
Section titled “What a stored revision is held to”A stored revision is not a document somebody just submitted. It passed the validation of the build that wrote it, which is not necessarily the build reading it: during a rolling upgrade a newer peer can store a document this build’s rules would refuse. So reading a revision and running one are held to different standards.
- Reading holds a revision to no rule.
GET /config, a revision read, the diff, the reference index, the entity reads,crewlet config show,exportanddiff, and the prior a write restores its masks from or merges onto all decode the stored document as it is. A revision this build would refuse is exactly the one an operator needs to see and replace, so none of these may refuse it. - Applying holds a revision to the runnable rules. A node’s reconcile tick validates a revision before anything on the node changes, so a refused revision leaves the previous epoch serving untouched. Booting from the store validates the active revision and names it when it cannot run, with
crewlet config importas the way out because the node’s API is not up yet. A company file named with-companyor-import-companyis held to the runnable rules while the node only runs it, because most boots write nothing from it: it is already the active revision, or a bootstrap the store’s own company outranks.POST /config/reloadand a revert validate what they re-activate. - A written document is held to every rule.
PUT,PATCH, a per-entity write, a/setupsubmission that changes the document,crewlet config import,crewlet validateand a company filecrewlet runimports as a new revision (-companyinto an empty store,-import-companyover a different company) validate the entire document the write produces, after its masks are restored. A write over a revision this build refuses therefore succeeds exactly when it corrects it.
The difference between the last two is the admission rules: authoring hygiene that nothing about running a company depends on, and the part of the rule set a later build may add to or relax. A revision that breaks one still runs, so an apply does not refuse it: during a rolling upgrade a newer peer may activate a revision it admitted under rules this build does not share, and refusing it would split the fleet’s epoch. Today they are unique seat names and unique unit names, and unique sandbox setup step names within one setup list (providers.sandbox.setup, or one seat’s sandbox.setup): a step’s env and files are credentials restored by the step’s name, so two steps of one name would leave every write carrying that list refused on masks nobody edited. So is a unit: reference on a seat declared inside a unit it does not name: the reference places only a root seat, so there it moves nothing and reads as a placement (see the organization model). A seat’s own GitHub App (integrations.github) on a human seat is one too: an app is the identity an agent acts as on GitHub, a person acts as their own contact.github_login, and nothing creates or reconciles an app for a person, so the block would read as a setting and do nothing. So is a schedule’s cron that has five fields and still breaks the cron grammar (61 * * * *): a company carrying one runs with that schedule never firing and every other schedule untouched. A written document is refused for breaking one. A revision being applied that breaks one is applied, booted on, reloaded and reverted to like any other, and each node logs org_admission_warning once per violation when it applies the epoch. Every other rule is a runnable rule, and nothing applies a revision that breaks one.
Writes and the whole /config surface require Authorization: Bearer <token>. Reads serve without one by default. Tokens are listed in Tier A
under api.auth.tokens and resolved from env vars at API startup. The matched
token’s id is recorded as created_by on each revision the request produces,
with created_by_kind operator, so revision history carries meaningful
attribution (alice, ci-pipeline, ops) rather than generic strings — and
says which revisions the engine wrote itself (node), which no token id could.
See the API reference.
Reading is what a dashboard does, and the page that would prompt for a token is
itself served unauthenticated — the page that asks for a credential cannot
require one — so a guarded-by-default read surface puts a modal in front of
every first load. Be clear-eyed about what open reads expose, though: /events,
/agents/{id}/memory and /ws/stream carry full LLM transcripts — prompts,
tool arguments, diary entries — to anyone who can reach the port. One line
closes them:
api: auth: allow_anonymous_read: false # every route needs a token tokens: - {id: founder, token: "${CREWLET_API_TOKEN_FOUNDER}"}With reads closed the dashboard authenticates its own socket and prompts for a token when the engine refuses it — including a banner that says refused rather than disconnected, since a rejected credential is not an outage that resolves itself.
The API states which posture it took at startup: api_listening carries
anonymous_read and the token count, and an open read posture on an api.host
that is not loopback adds an api_anonymous_read_on_a_reachable_bind warning.
A laptop and an internet-facing bind are not the same decision, and a warning
that fires identically for both is one nobody reads by the third deployment.
The guard is mounted whether or not api.auth is configured. It applies one
rule (auth.Guard.Requires), and what Tier A supplies is the posture, not the
existence of a check. An API built with no Tier A at all therefore has no token
that can match, which means reads serve and every write plus the whole /config,
/secrets and /setup surfaces answer 401. That is the only safe reading of “an app was built
without being told who may write to it”, and it removes the possibility of a
process that serves /config writes with nothing in front of them.
Served without a token in either posture, because they authenticate by other means or must be reachable to obtain one:
| Path | Why |
|---|---|
/health, /ready | Probes. An orchestrator has no token, and a liveness check that 401s is a liveness check that fails. A single trailing slash is tolerated (/health/), because the guard runs before routing — so the router’s redirect to the canonical path only happens if the request gets past the guard first, and a slash must never be the difference between healthy and evicted |
/webhooks/* | Each verifies its provider’s HMAC before doing anything — a stronger check than a shared bearer token. Includes the Slack OAuth landing page, which a browser reaches mid-install |
A route whose secret is unset has nothing to verify with, so it fails closed: 503 + Retry-After, never an accepted delivery. The sender retries and the delivery flows once the secret is configured — a deployment that has not set one is stalled, not damaged, and nothing unsigned is ever recorded, published, or shown on the dashboard | |
/otlp/*, /mcp/* | The signed per-run token in the path is the credential. Both are reached from inside a sandbox, where the API’s own token must never go |
/, /dashboard, /favicon.ico, /static/* | The page that prompts for a token cannot itself require one. It ships no data: every byte it renders comes from an authenticated fetch |
/ws/stream follows the same rule as every other read. When reads are closed it
needs a credential like anything else — and browsers can’t set headers on a
WebSocket, so it accepts ?token=… as well as the Authorization header.
Prefer the header where a client can send one, since query strings tend to land
in proxy access logs.
| Setting | Effect |
|---|---|
api.auth.tokens | The accepted bearer tokens. Needed for writes and /config, whatever the read posture is |
api.auth.allow_anonymous_read: true (default) | GET/HEAD outside /config, /secrets and /setup serve without a token; writes and those three surfaces still require one, and so do the individual reads that describe the deployment rather than the company’s work (/fleet, /integrations) or somebody else’s personal record (/work/my-work, /work/people/{handle}, /work/inbox, /conversations, and ?viewer= on /work/views, each naming a seat other than the caller’s own) |
api.auth.allow_anonymous_read: false | Every route needs a token, /ws/stream included. The lockdown posture for a deployment that terminates traffic somewhere reachable |
api.auth.disabled: true | Local development only. Everything serves unauthenticated including writes, attribution becomes "anonymous", loud WARNING at startup |
api.auth.company_writers | The token ids that alone may change the company document; empty (default) is every token. See Managed configuration |
Two combinations are worth calling out:
- No tokens at all is a legitimate posture, not a misconfiguration: reads serve and writes are refused outright, because no token can ever match an empty list. A deployment that never writes config through the API therefore has no credential to manage — strictly safer than minting one it will not use.
allow_anonymous_read: falsewith no tokens is refused at boot. It guards every route behind a credential that does not exist, which is not a strict posture but an outage whose only symptom is a uniform401that reads exactly like a wrong token.
CORS defaults to same-origin. The dashboard is served by this process so it
needs no entry; list any other browser origin explicitly in
api.auth.allowed_origins. The previous * default let any site a logged-in
operator happened to visit read every endpoint, so * is now refused at
boot rather than honoured — name each site.
An entry is compared against the browser’s Origin header exactly, which is
always scheme://host[:port]: an entry with no scheme, a trailing slash or a
path matches nothing and is refused for that reason, because an allow-list that
looks configured and never matches fails in a browser console this engine never
sees. A permitted origin’s response carries Access-Control-Allow-Origin and
Vary: Origin; an unlisted one carries neither and the browser blocks it — it
is not refused with a status, because a same-origin POST carries an Origin
too and the default allow-list is empty.
Preflights are answered before the auth guard, and deliberately: a browser
sends the OPTIONS itself with no Authorization header, so answering it
behind the guard would be a 401 on every cross-origin write and the real
request would never be sent. Only Authorization and Content-Type are
permitted as request headers, and a preflight is cacheable for ten minutes —
short enough that removing an origin takes effect within one.
The auth middleware compares tokens in constant time (crypto/subtle).
Failed attempts log api_auth_failed at WARNING (never the candidate token
value); successes log api_auth_ok at DEBUG with operator_id and route.
See the API endpoints reference for the per-route auth + status semantics.
Managed configuration
Section titled “Managed configuration”Some deployments do not edit the company document by hand at all: a GitOps pipeline applies it from a repository, or external tooling such as a Kubernetes operator renders it from custom resources and writes it on every reconcile. Such a system replaces the active revision with its own each time it runs, so an edit a person makes through the dashboard or the API lasts only until then — and nothing tells them. Naming the system’s token as the document’s only writer turns that silent loss into a refusal at the moment of the edit:
api: auth: tokens: - {id: gitops, token: "${CREWLET_API_TOKEN_GITOPS}"} - {id: founder, token: "${CREWLET_API_TOKEN_FOUNDER}"} company_writers: [gitops] # only this token may change the company documentWith company_writers set, every write that changes the document refuses
any other credential with 403 config_managed, naming the writers and what to
do instead. Reads are unchanged: every token still reads /config, its
history and its diffs.
| Refused for a credential not listed | Still open to every credential |
|---|---|
PUT and PATCH /config, every per-entity PUT /config/{kind}/{id}, POST /config/revisions/{id}/revert, and the ?dry_run=true check of each — a check answers what the write would, and the write would be refused | POST /config/reload, which re-publishes the active document’s own bytes |
A /setup connect that writes a pointer, a disconnect, and an agent’s GitHub App — each refused before anything is sealed, queued or created at the vendor | A /setup submission that only rotates a value the document already names, and the provisioning pass and its check |
An offline crewlet config import, an offline crewlet config activate of any revision but the active one, and crewlet run -import-company; a -company bootstrap seed is ignored with a company_seed_ignored warning | /secrets, crewlet secrets, crewlet config seal and crewlet config rekey, and crewlet config activate of the active revision |
Why the secret store and a rotation stay open. A leaked credential has to
be replaced now, by whoever holds the new one, and the managing system cannot
know it is needed. Writing the new value and re-publishing the unchanged
document — PUT /secrets/{name} then POST /config/reload, or one rotation
through Settings › Integrations — changes nothing the managing system renders,
so it is the break-glass a person keeps. If the managing system also syncs the
secret store, it will overwrite a rotated name it owns at its next sync; fix
the value at its source as well.
The engine’s own writes are not judged by the list. A reconcile pass that
records what it discovered at a vendor writes as the node, not as a
credential; a managing system that re-renders the document should carry those
values (they are in the active revision it reads back) or set them itself.
That includes a provisioning pass any operator starts on demand from /setup
or Settings › Integrations: where integrations.jira or
integrations.confluence has no cloud_id, the Atlassian pass records the
cloud_id and site_url of the site it found, and the revision is credited to
the node (reconcile loop), not to the operator who pressed the button.
It prevents drift; it is not a privilege boundary. company_writers
stops a person’s edit from being silently overwritten, and nothing more. Every
token still writes the secret store and can reload, so a token that is not a
writer can still change what the managed document does by rewriting a value
it references — a model’s API key, a webhook secret, or any URL or endpoint
kept as a ${VAR} — and publishing it with POST /config/reload. A
deployment that needs to keep a credential away from the company’s behaviour
must not issue that credential a token at all.
The dashboard shows the document as managed. The viewer
answer says whether the caller may change the document and, to an operator,
who manages it; every control that writes the document is drawn disabled with
that sentence, and the org builder opens read-only.
crewlet validate warns about a list that is valid and almost certainly wrong:
an id that names no token in api.auth.tokens (it lets nobody write), a list
naming no configured token at all (nothing may change the document through the
API — a frozen document), and api.auth.disabled: true beside it (every caller
is the unauthenticated one, which no list can name). An empty, mixed-case,
repeated or reserved (anonymous) entry is refused outright. The list is per
node, like the rest of api.auth: give every node the same one.
To take the document back, remove company_writers from every node’s Tier A
and restart. The decision and its boundaries are
ADR-0030.
Secrets
Section titled “Secrets”Crewlet has two secret-handling behaviours for Tier B. Which one is in effect depends solely on whether a Tier A encryption keyring is configured.
Default: ${VAR} references (no keyring)
Section titled “Default: ${VAR} references (no keyring)”With no secrets: block in crewlet.yaml, the DB stores ${ENV_VAR} reference strings verbatim and resolution happens at provider / transport / integration construction time (internal/engine). The company_config table never holds a real secret; the environment is the source of truth. Safe to back up / export, but every deployment must re-provision the referenced env vars, and rotating a key means editing the env + restarting.
A configured keyring also unlocks a second, independent place a ${VAR} can resolve from: the encrypted secret store, consulted ahead of the environment. That is what lets a provisioner hand a minted credential straight to the engine instead of writing a file someone has to source. It is opt-in and inert until a secret is actually stored.
Encrypted at rest (Tier A keyring configured)
Section titled “Encrypted at rest (Tier A keyring configured)”Add a keyring to crewlet.yaml and Crewlet encrypts the entire company_config payload as one opaque blob (AES-256-GCM) before it reaches the DB:
# crewlet.yaml (Tier A) — the keyring is the sole root of trustsecrets: active_key_id: "2026-01" keys: - id: "2026-01" material: "${CREWLET_SECRET_KEY_2026_01}" # base64(32 bytes); crewlet secrets keygenThe whole document is stored as {"__encrypted__": "enc:v1:<key_id>:<base64>"} — nothing about the config’s structure (org chart, policies, model choices, or secrets) is visible in the database. A stolen DB reveals nothing.
- Encrypt on write. Every write path (
PUT /config, per-entityPUT,crewlet config import,crewlet run -company/-import-company) encrypts the whole document before the payload reaches the DB. - Decrypt at the read boundary. The engine and the API it serves, migrations, and the CLI each decrypt the blob (
secrets.Open, thenconfig.DecodeCompany) into the plaintext structure before use, so the Tier A key is required for every config read.${VAR}references inside the config are kept verbatim in the blob and still resolve from the environment at construction time. - Fail closed. If an activated revision is stored encrypted but no keyring is configured (or the key is missing), the engine refuses to boot rather than run with an opaque blob it can’t read.
- One key, not N env vars. After encrypting, the engine needs only the Tier A key in its environment — not a per-secret env var for every LLM key, MCP token, and webhook secret.
Because the key gates every read, keep it as available as the database itself: the API, dashboard, migrations, and CLI all fail closed without it.
Threat model. Encryption at rest defends against data-at-rest exposure: a leaked backup, a copied store file, a stolen volume snapshot, a coordination-store dump on a laptop — the attacker gets one opaque ciphertext blob per revision, no structure and no credentials. It also keeps config egress clean (GET /config, dashboard views, and revision diffs are decrypt-then-redact — no plaintext secrets) and absorbs accidental plaintext (a raw key pasted into company.yaml is encrypted on write, so it never lands in the DB in the clear). It does not defend against a compromised engine host that holds both the DB and the Tier A key (that host can decrypt — it must, to run; keep the key out of the store’s backup domain — that separation is the point), a malicious operator with a valid key and API access, or in-memory extraction from the live process. The property: plaintext config exists only transiently in the encrypt/decrypt path and in the live engine’s memory — never in durable storage.
Migrating from ${VAR} to encrypted
Section titled “Migrating from ${VAR} to encrypted”Encryption is opt-in and backward-compatible — a plaintext config boots with or without a keyring. To migrate an existing deployment:
crewlet secrets keygen --key-id 2026-01 # prints a key + the Tier A snippet# add the secrets: block to crewlet.yaml, export CREWLET_SECRET_KEY_2026_01crewlet config seal # encrypts the active revision as one documentcrewlet config seal writes a new revision holding the encrypted document; afterwards the per-secret env vars are no longer needed at runtime (only the Tier A key). It’s idempotent — a second run on an already-sealed revision is a no-op.
Rotation
Section titled “Rotation”Rotating one secret (e.g. a leaked LLM key) needs no host access: PUT /config (or a per-entity write) with the new value — the whole document is re-encrypted under the active key as a new revision.
Rotating the master key uses the keyring’s multi-key support so there’s no downtime:
crewlet secrets keygen --key-id 2026-07 # mint the new key# add it to secrets.keys AND set active_key_id: 2026-07,# keeping the old key in secrets.keys so its ciphertext still decryptscrewlet config rekey # re-encrypt the document under 2026-07# verify a clean boot, then drop the old key from crewlet.yamlThe document’s envelope carries the id of the key that sealed it, so rekey decrypts under whichever key sealed it and re-encrypts under the active key. crewlet config rekey --dry-run reports whether it would re-encrypt without writing.
If you also use the secret store, run crewlet secrets rekey alongside crewlet config rekey before dropping the old key — its rows are sealed under the same keyring and would otherwise become unreadable.
Reads and export
Section titled “Reads and export”Every configuration read redacts credentials: GET /config (JSON and ?format=yaml), GET /config/revisions/{id}, the revision diff, the entity reads (GET /config/roles/{handle}, /config/units/{name}, /config/llm-providers/{key}, /config/mcp-servers/{name}), the dashboard’s config and config_entities queries, crewlet config show and crewlet config diff. The read path opens the whole document, then replaces every credential value with the literal marker "__redacted__", so the caller sees the config’s shape and never a credential. The anonymous /org view needs no masking because it carries no credential field at all: it is an explicit public projection of the charter and the organization tree (see API endpoints). A revision sealed under a key the node does not hold is refused rather than served.
What counts as a credential is structural. A field is a credential because its Go type carries a secret:"true" tag, never because of how its value or its key reads, and a tag on a list or a map covers every element. The tagged company fields today: LLM api_keys, the embeddings and sandbox api_key, a cli-agent provider’s cli.auth.token, cli.auth.credential_bundle and cli.env, the integration tokens, admin tokens, API and app keys, webhook and signing secrets, a seat’s integrations.slack bot token and signing secret, integrations.mattermost.bot_token and GitHub App private key and webhook secret, every mcp_env value on a seat or a unit, mcp_servers[].env and mcp_servers[].headers, role.sandbox.env, and each sandbox setup step’s env and files (under providers.sandbox.setup and role.sandbox.setup). Everything else (URLs, hosts, flags, model names, the org chart) is served exactly as stored, including every toggle an operator set explicitly: a schedule kept with enabled: false reads as disabled, so sending the read back never re-enables it. A test fails the build when a field whose name reads like a credential, or a map[string]string named env, headers or files, is added without the tag — in either tier, since Tier A’s credentials carry it too (secrets.keys[].material, api.auth.tokens[].token, stream.token and store.objects.s3.secret_access_key, listed under credential positions) — so a new credential field is masked by declaring it rather than by remembering to. The same tag marks each of these positions "x-crewlet-secret": true in the generated JSON Schema, held by a second test to exactly the set a read masks, so tooling that writes a company without the engine (an editor, a CI linter, a Kubernetes operator rendering the document from its own Secrets) knows where a ${VAR} belongs instead of a value.
Only a whole ${VAR} reference is shown. A credential field whose value is exactly one reference ("${TRACKER_TOKEN}") is left visible: it names a credential and carries none, since the value it points at lives in the environment or the secret store and never in this document. Every other non-empty value is masked, including one that embeds a reference beside literal text. "Bearer sk-live-${SUFFIX}" and "sk-live-SECRET-${ROTATION}" are legitimate (the resolver expands embedded references), and their literal half is a credential. A value using brace syntax the resolver ignores (${line#host=}, ${1}, an unclosed ${) is not a reference at all and is masked too. “Is a reference” is the engine’s own grammar (internal/envref). The variable names an embedded reference carries are still listed, with their paths, by GET /config/references, which reads the unredacted document and answers names only.
A redacted GET, an edited field and a full-document PUT round-trip safely: the write path swaps each marker back to the currently stored value before validating, so a round trip never clobbers or exposes a credential. To change one, supply the new value (or a ${VAR}) at that field.
Members are matched by identity, not by position. Matching by position meant a pure reorder handed each seat its neighbour’s credentials, silently, since the lengths still agreed and no marker was left standing to refuse. So every member that can name itself is matched by that name:
- A seat by its handle, and a unit by its name, anywhere in the document. One index covers the whole stored revision, so a seat moved from the root into a unit, from one unit to another, or back to the root keeps its credentials, and so does a unit moved under another unit. Reordering the roster, or adding a seat, restores every other member’s credentials.
- An MCP server, and a sandbox setup step, by name within its own list. Both names are unique within their list: two MCP servers of one name are refused outright, and two setup steps of one name are refused on every write (an admission rule).
- A list of bare credentials (
api_keys) by position, because it has no identity to match on. Change its length and the masks in it are refused rather than guessed.
An identity that does not name exactly one member matches nothing. A renamed seat (a new handle) or a renamed unit carries no prior value of its own. An identity that is empty, or that the stored revision holds twice (two units called Platform in a revision written before names had to be unique), is left out of the match entirely: picking either member, or falling back to position, would hand one member’s credentials to another. In every one of these cases the mask stays standing and the write is refused with a validation error naming the field, so the caller writes the real value, or a ${VAR}, there.
The untyped maps (mcp_servers[].env and .headers, cli.env, a sandbox step’s env and files) are masked whole, every value in them, whatever its key is called. A host or a region set beside a token in one of those maps is masked with it; write it as a separate, untagged setting where one exists, or as a ${VAR}.
crewlet config export runs on the host, where the keyring already is, and prints the revision as YAML with its credentials in the clear: it opens a sealed revision and renders the document, so the output is importable as it stands. That is what a restore or a migration needs, and it is the one command that prints credentials. crewlet config export -redact masks exactly what the reads above mask, for a dump that is safe to share.
One company per engine
Section titled “One company per engine”An engine runs exactly one company. It opens one store file, that file holds one company_config table, and that table has at most one is_active=TRUE row (zero in the unconfigured boot state; otherwise one). There is no tenant column and no row-level scoping: the revision you activate is simply the company the engine runs. The same rule governs the secret store: one company per coordination estate, so a variable name alone is the key.
To run a second company, run a second engine with its own database.
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.