Deployment
Crewlet requires no infrastructure services. The engine is one binary: its event stream is a NATS JetStream server it embeds, and its store is a local file it creates and owns exclusively. A single host runs a whole company with nothing else installed.
One slot changes when a deployment outgrows one node — the stream, which becomes either a cluster of the members the nodes already embed or a NATS cluster somebody else runs — and this page is mostly about that path. The coordination KV is not a second address: it rides the stream’s own connection, deliberately, so that a node cannot end up holding live leases over a link that still works while the one carrying its inbox has dropped — alive to its peers, deaf to its work. The store never becomes shared either: it stays one file per node, which is why everything genuinely shared between nodes lives in the KV instead. See Running a Fleet and Scaling Out.
The single host
Section titled “The single host”# crewlet.yaml (Tier A)stream: type: embedded # a JetStream server inside this process store_dir: "/var/lib/crewlet/stream" # empty = in-memory, nothing survives a restart
store: path: "/var/lib/crewlet/company.db" # ONE file, this process only
coordination: type: local # one node holding its own seat leasescrewlet run -config crewlet.yaml -company company.yamlThat is the deployment. Point a reverse proxy at the API port for inbound webhooks and the dashboard, and there is nothing else to operate. When the webhooks have to come from the internet and the dashboard must not, give them a listener of their own instead: see Exposing webhooks without the admin API.
Give that proxy a read timeout above fifty seconds on the API port.
Nearly every request is answered in well under a second, but a node eviction
or readmission (Retention) may take up to fifty
seconds — twenty to judge it and thirty to write every log — and crewlet retention evict and the dashboard’s evict dialog both wait a minute for its
answer. nginx’s proxy_read_timeout defaults to sixty, which is enough; a
proxy set below fifty cuts those answers off with a 504. Nothing is lost when
it does — the node finishes the gesture whatever happens to the connection,
and both clients read an answer the engine did not write as one that never
arrived, keeping the operation id to finish it with — but the operator then
has to finish it to see what it did.
The room the stream’s volume needs
Section titled “The room the stream’s volume needs”The engine’s own tracker, knowledge base, vector index and usage history each
keep a log on the stream, and each log’s byte ceiling is reserved on the volume
holding stream.store_dir when its stream is created: the embedded broker
grants a ceiling in full, up front, or refuses to create the stream at all.
Its limit is three quarters of that volume’s free space.
The node sizes the four ceilings together to fit half of that limit, and never below 1 GiB each, so:
- A first boot needs at least 5⅓ GiB free on that volume. Three quarters of 5⅓ GiB is the four 1 GiB floors. Below it the node refuses to boot with an error naming the log it could not reserve, the bytes it needed, the bytes the broker had left, and the Tier A field that sets the ceiling.
- More room buys longer logs, up to a point. Unset, the mutation log and the vector changelog each ask for a quarter of the free space (4..64 GiB), the knowledge base’s log for a quarter of the mutation log’s, and the usage log for a fixed 1 GiB — its size is a count of node-days rather than a rate, so more disk buys it nothing. They are scaled down together whenever they ask for more than half the free space, which on a first boot is every volume with less than about 290 GiB free (at their 64 GiB clamps the four ask for 145 GiB); from there up each log gets what it asked for.
- The ceilings are fixed when the streams are created. Moving the node to
a bigger volume, or setting
stream.tracker_log_max_bytes,stream.tracker_vectors_max_bytes,stream.pages_log_max_bytesorstream.usage_log_max_byteslater, changes nothing about streams that already exist;crewlet retention set-capacityis what changes a running log’s ceiling. A log created later — one a new version adds — is sized from what the existing logs’ ceilings leave of that half, so logs that already hold all of it leave it the 1 GiB floor.
Replication has the whole arithmetic and the refusal’s text.
The compose stack starts nothing
Section titled “The compose stack starts nothing”docker-compose.yml in a repo checkout is for the things around the
engine — the local integration loops, and nothing else. There is no broker
service in it, because there is nothing to run: the engine embeds its own.
Every service is behind a profile, so a bare docker compose up brings up
nothing at all:
cp .env.example .env # first time onlydocker compose --profile gitlab up -d # GitLab (code host)docker compose --profile mattermost up -d --wait # Mattermost (chat)| Service | Port | Details |
|---|---|---|
| GitLab | 8929 / 2424 | gitlab/gitlab-ee, configured external_url http://gitlab.local:8929 — the published port and the URL GitLab builds its own links from are deliberately the same number. 2424 is git-over-SSH and optional; HTTPS plus a token is the path the engine takes |
| Mattermost | 8065 | ${MATTERMOST_LISTEN_PORT:-8065}. Mattermost accepts a websocket upgrade only from a browser whose Origin matches MM_SERVICESETTINGS_SITEURL exactly, host and scheme, so move the port and the site URL together or the page loads with the event stream silently dead |
The stream beyond one host
Section titled “The stream beyond one host”The stream slot has a default and one alternative, and they differ in who
runs the broker rather than in what the engine speaks: embedded starts a
NATS JetStream server inside this process, nats dials one somebody else
runs. The client code is identical either way — one implementation, with the
connection as the only branch — so nothing above the queue can tell which is
in use, and no company config mentions it.
A single host needs neither shape below. A fleet needs exactly one of them, because two nodes each embedding a solo server are two companies that cannot see each other: the solo server binds no socket at all, which is a security property as much as a convenience.
Every node embeds a member of one cluster
Section titled “Every node embeds a member of one cluster”The fleet still ships as one binary and there is still no broker to deploy. Three nodes, each naming itself, its route port and the other two:
# crewlet.yaml (Tier A) on the first node. The other two differ only in# node.id and in which peers they name.node: id: crewlet-1
stream: type: embedded store_dir: "/var/lib/crewlet/stream" cluster: name: crewlet # identical on every member port: 6222 # this member's route port peers: # the others' route URLs - "nats://crewlet-2.internal:6222" - "nats://crewlet-3.internal:6222" host: 10.0.0.11 # bind the route port to the private # interface, not to all of them replicas: 3 # a publish is committed by a quorum # before Publish returns
coordination: type: embedded-kv # the leases ride the stream's own # connection; nothing else to setBind the route port to the network the peers are on. cluster.host is the
interface the route listener binds, and leaving it unset binds every
interface. A route port is how a member JOINS — and a member that joins reads
and writes every stream and every coordination bucket — so on a host with a
public interface an unset host publishes unauthenticated access to the
company’s whole event history. Set it to the private address, or keep the port
off the public interface with a firewall; the engine cannot tell which of a
host’s addresses is the private one, so it does not guess.
A cluster block without a name is refused, because it would do nothing.
The embedded server takes its route port, its bind interface, its advertise
address and its peer list only from a named cluster, so anything under
cluster: with no cluster.name starts a solo node that binds no route
listener and forms no cluster — while every other reading of the same file
(the provisioning budgets, the topology validation) calls it clustered. Tier A
names the missing field instead.
A route port something else already holds is refused at startup, by name.
This is the one clustering failure with no other symptom: NATS does not fail
when its route listener cannot bind — it logs the error and carries on serving
clients, so the node comes up, answers /health, and simply never forms a
route to a peer. Left to report itself, that surfaced two minutes later as a
readiness timeout blaming peer reachability, which is a network path that is
fine. The engine now probes the port before it starts the server and refuses:
stream.cluster.port 6222 is already in use on 10.0.0.11, so this member'sroute listener cannot bind and it could never form a route to a peer — freethat port or give this node a different onecluster.advertise is for when what a member binds is not what its peers can
dial. Members learn about each other from the members they already have: when
node 1 accepts a route from node 2 it tells node 3 where to find node 2, and
node 3 dials that address itself. With nothing configured that address is
derived from the connection’s own remote address, which is correct on a flat
network and wrong wherever the address a member is seen from is not one anybody
else can use — a container with a mapped port, a NAT, a member behind a load
balancer. There, set advertise to the address peers should dial (host and
port, or a bare host to keep this member’s own route port):
stream: cluster: host: 0.0.0.0 # inside the container, bind everything advertise: "crewlet-1.internal:6222" # outside it, this is the addressnode.id is this member’s identity in the cluster, and it has to survive a
restart. The engine passes it as the NATS server name, which must be unique
— a route from a server whose name the cluster already knows is rejected —
and stable, because JetStream places replicas by server name. A node that
comes back under a fresh name is a new peer: its old replicas are orphaned on
a member that no longer exists, and the stream sits short of quorum waiting
for a server that will never return. That is why a clustered member with no
name is refused at startup rather than given a generated one — a generated
name is unique, which is only half the requirement.
It is the resolved id, so the environment is enough. The name follows the
same precedence as everything else that identifies this node — node.id in
the file, then CREWLET_NODE_ID, then the default — so an orchestrator
injecting a pod name needs no node: block at all. Writing
node.id: "${CREWLET_NODE_ID}" (or "${HOSTNAME}") works too, and means the
same thing: Tier A expands ${VAR} references before it decodes. Into a text
field like this one a reference expands anywhere in the value; into a number or
a boolean — a port, stream.replicas — it must be the whole value, and what it
resolves to is read as if it were written there (see Environment Variable References).
One node or three, never two. Two embedded-KV members have no quorum
without each other, so the fleet stops serving the moment either restarts —
and a rolling upgrade restarts them one at a time, which makes the outage
certain rather than unlucky. Tier A refuses a two-member config by name,
counting this node and the other members its stream.cluster.peers
names: an entry recognisably this node’s own route, or a repeat, is left out
of the count with a warning, and an alias of this host it cannot recognise
is counted as a member — see Running a Fleet.
A fresh cluster takes seconds to form, and the engine waits it out rather than hanging. Accepting connections is not the same as being able to serve JetStream: a member answers its client port as soon as it is listening, while the metadata group takes seconds to elect a leader — measured at around eight on a quiet three-member cluster — and until it has one, creating a replicated stream blocks instead of failing.
There are therefore two waits at boot, in order, and they fail for different reasons:
| Wait | Budget | What is happening |
|---|---|---|
| Accepting connections | 30s solo, 2 min clustered | The member recovers its file store and, in a cluster, stands up its route listener while its peers are booting too |
| JetStream current | 60s | The metadata group elects a leader and this member catches up with it |
Then placement retries for as long as the cluster answers “no suitable peers”, inside the per-create provisioning budget — 30 seconds on a solo node and 2 minutes on a member with peers, because the two creates are not the same call underneath. See A clustered node is given longer to create them below.
The clustered accept budget is four times the solo one because a member starting alongside its peers is competing with them for the same disk and the same scheduler, and the asymmetry is stark: failing this wait fails the whole boot, so a budget that is too short turns a busy host into a node that refuses to start and then works on the retry — which during a rolling restart is how one slow member takes out the restart. Too long only means a genuinely broken server is reported later, and the wait is cancellable, so Ctrl-C returns immediately. Every other error is returned at once: a bad subject or a conflicting retention does not clear by waiting, and retrying would turn a config mistake into a half-minute hang with the same message at the end.
Set store_dir, or the fleet forgets. Empty selects an in-memory member:
a restart loses that member’s replicas, and the same server holds the KV
buckets carrying the fleet’s shared records — the token counter, the
completion ledger, open agent-to-agent asks, claimed scheduled fires, detached
(and billed) sandbox runs. The rule is asked of the broker, not of the roles:
every member — every data node on the embedded broker — sets store_dir
whatever else it does, because an in-memory member creates every stream it
provisions in memory. A leaf and a client of an external cluster hold no
streams of their own, so both are refused a store_dir; see
Satellite nodes. For a member, an in-memory stream is
tolerable only for a company whose tracker and knowledge base are both a
vendor’s; on either native backend the engine refuses it outright, as the next
paragraph describes.
On the native backends it is the company’s own record, and it is
refused. With tracker.backend: native or knowledge.backend: native (the
defaults), every work item and every page lives in a log on that stream. An
unset store_dir would mean the first restart recreates those logs empty, and
a node whose rows are ahead of a log that restarted from nothing stops serving
for good. So the engine refuses to boot that pairing, and crewlet validate
refuses it when given both documents, naming stream.store_dir. Either
backend is enough: a company on Jira whose knowledge base is the engine’s own,
which is the default without Confluence, is refused the same way. Only a
company on vendors for both can run an in-memory member.
Divide store_max_bytes when several engines share a filesystem. Every
stream ceiling on the embedded broker is a reservation: the broker refuses to
create a stream whose ceiling it cannot back, and the number it compares
against is this limit. Unset, the broker sizes itself from the free space on
store_dir when its JetStream comes up — three quarters of it, measured once —
which is right when it is the volume’s only tenant and wrong the moment it is
not, because free space bounds the sum of the engines on a disk rather than
each of them. Two engines each taking what they can see over-commit the volume;
three over-commit it by half again. The failure is not a disk-full message: on
a single node it is insufficient storage resources available, and on a fleet
— where the limit is applied by the metadata leader placing the stream rather
than by the member creating it — it is no suitable peers for placement, insufficient storage. Either way it names whichever stream that node happened
to provision last, which reads as a problem with that subsystem. So on
a host running N engines against one filesystem — a test runner, a
multi-tenant box, several companies on one machine — give each of them its own
share:
stream: store_dir: "/var/lib/crewlet/acme/stream" store_max_bytes: 68719476736 # 64 GiB of the volume, this engine's shareIt is measured once, at boot, on both paths: a volume that later grows or shrinks does not move the limit, and a node that should see a resized disk is restarted. Half of whatever is in force is what the state logs’ derived ceilings may reserve between them; the other half is for the streams that reserve nothing and simply grow against it — the seats’ mailboxes, the event log, the dead-letter stream, the memory changelog and every coordination bucket. A refusal names the ceiling that did not fit and the Tier A field that sets it on either topology; it adds the limit in force and what is already reserved only where this node can read them, which is the standalone one. On a fleet the room that refused is another member’s disk, and no member can read another’s.
The clustered embedded broker has no authentication and no TLS. Run it on a trusted network.
A clustered member listens on two ports, on every interface: the route port
named in the config, and a client port the server picks for itself. Neither
carries a credential or a certificate. credentials, token and
stream.tls are dial-side options — they configure this process
connecting out to a URL — and an embedded server has no server-side
counterpart in Tier A at all. Anything that can reach those ports can
publish onto a seat’s inbox, read every event the company produces, and
take its leases.
A private subnet, a security group or a WireGuard mesh between the nodes is
what makes this shape safe, and no configuration substitutes for one. A
deployment that needs the brokers themselves mutually authenticated runs
stream.type: nats against a NATS cluster it secures itself.
An external NATS server
Section titled “An external NATS server”The other multi-node shape, and the one to reach for when the broker has to be secured, operated, or shared on a schedule of its own:
# crewlet.yaml (Tier A)stream: type: nats url: "nats://nats-1.internal:4222,nats://nats-2.internal:4222,nats://nats-3.internal:4222" credentials: "/etc/crewlet/engine.creds" # an NKey/JWT creds file # token: "${CREWLET_NATS_TOKEN}" # or a bearer token tls: ca: /etc/crewlet/ca.pem # the private CA to trust cert: /etc/crewlet/client.pem # this engine's own certificate key: /etc/crewlet/client.key
coordination: type: embedded-kvThe URL goes to the NATS client verbatim, so a comma-separated list of a
cluster’s members is one value as far as this config is concerned and the
client fails over between them. store_dir is refused here by name: it is
where an embedded server persists, and an external cluster keeps its own
storage.
Authentication is credentials or token. credentials is a path to a
NATS credentials file, the NKey/JWT pair a NATS account setup issues per
user; token is a bearer token, and takes a ${VAR} reference so the
secret stays out of the file and out of the config revision history. Set
whichever the broker asks for; both are dial options, and the engine stores
neither.
stream.tls is the transport underneath that authentication, and it is a
separate question from who you are: a broker configured tls { verify: true }
— the hardened default every NATS guide recommends — refuses a connection
presenting no client certificate, whatever credentials would have followed.
ca is the bundle the server’s certificate is verified against, and empty
means the host’s root pool, which is right for a public CA and wrong for the
self-signed certificate most internal estates use. cert and key are this
engine’s own certificate: both or neither, because validation refuses half a
keypair rather than letting it dial and be rejected by the broker with an
error naming neither file. There is deliberately no way to skip
verification — that switch is set once during a bring-up and never unset,
and the connection it leaves behind carries every event this company
publishes to whoever answers on that address.
Those files are opened before the dial, and the error names the field. A
missing ca reports tls.ca /etc/crewlet/ca.pem: no such file or directory,
not a connection failure. Left to the NATS client, an unreadable certificate
surfaces as a dial error that reads exactly like the broker is
unreachable — and sends an operator to debug a network path that is fine,
for a file that is simply not there.
A broker blip is not a node restart. The engine dials with unlimited
reconnects and a one-second wait between attempts, and never gives up on the
URL: the coordination layer already distinguishes “unreachable” from “not
mine”, and a node that keeps its seats through a two-second outage is the
entire point of that distinction. The coordination KV rides this same
connection on purpose — one connection, one fate. An outage that outlasts the
lease TTL (45 s unless coordination.lease_ttl_seconds says otherwise) does
hand this node’s seats to a peer, and that is the intended behaviour rather
than something a reconnect policy should paper over.
The account needs more than publish and subscribe. A node creates what it
uses, on every start and idempotently: the seven engine streams
(CREWLET_AGENT, CREWLET_EVENTS, CREWLET_NOTIFICATIONS,
CREWLET_CONFIG, CREWLET_MEMORY, CREWLET_DLQ, CREWLET_CUSTODY), the four state-log
domain streams (CREWLET_TRACKER_LOG, CREWLET_TRACKER_VECTORS,
CREWLET_PAGES_LOG, CREWLET_USAGE_LOG), a stream per extra subject namespace a company
publishes under, one durable consumer per seat mailbox (an ordinary API
call, measured at 1.7 ms), the twenty-one crewlet_* KV buckets:
three in the lease store, holding the seat and presence leases, the duty
leases and the fencing epochs, and eighteen in the fleet store holding the
shared records — the object store’s backend record and its collector’s report
among them — and, on the default nats object store, the crewlet_files
object store bucket (its stream is OBJ_crewlet_files). A credential
scoped to publishing and consuming fails at boot, on the first stream it
tries to create.
A coordination read costs one ordered pass certified against the stream’s
key index, and an account needs the consumer, stream-info, message-get and
purge APIs. A node lists coordination records constantly — several
fifteen-second duty loops on every tick, and the state-log write fence, which
lists the published trim floors on every write at an expectation of zero, a
subject’s first write among them. Each listing is one pass over a temporary
consumer, which on a replicated bucket is two metadata-raft proposals, and
beside it one stream-info request narrowed to the listing’s keys, which names
every key that has a message. A pass alone cannot tell that it is complete — a
key rewritten under it can land behind the marker that ends it, and every renew
is such a rewrite — so a key the index names that the pass missed, or delivered
only as a delete marker, is read by itself from the stream leader
($JS.API.STREAM.MSG.GET), never through the bucket’s direct get, which any
replica — one that has fallen behind included — may answer. A marker needs that
read because on a replicated bucket the pass may be served by a replica that
applied a key’s delete and not yet its re-creation. On a quiet bucket the
certification therefore costs the index read and one leader read per recent
removal the pass meets, and under writes one more per key the pass lost. What
keeps “recent” recent is the maintenance duty, which sweeps the removal markers
older than three hours out of every shared bucket with no age, each with a
leader-side $JS.API.STREAM.PURGE bounded at the marker’s own revision so a
key written again since is untouched (see
Coordination). While a
bucket’s stream is electing a leader, a listing of it answers unavailable
rather than a short list, since every member answers the index from its own
store until one leads — the answer every caller already treats as a store
that did not answer, never as “no records”. A listing’s one bound is
the broker’s page size for that index, 100,000 keys under one filter; see
Coordination for what
happens past it.
The engine deliberately does not use the batched direct get that would
avoid the consumer: it is served by any replica, and this estate has reads
whose answer is acted on with nothing to arbitrate them. So a credential
scoped only to publishing and consuming is not enough; the account needs the
consumer API, $JS.API.STREAM.INFO, $JS.API.STREAM.MSG.GET and
$JS.API.STREAM.PURGE alongside the rest of $JS.API. If the broker’s own
debug logging is on, that consumer churn is what produces a steady stream of
JetStream connection closed: Client Closed lines — see stream.debug, which
is off by default for exactly this reason.
A clustered node is given longer to create them
Section titled “A clustered node is given longer to create them”A clustered node is given longer than a solo one. Every one of those creates is a local file-store setup on a solo node and a raft round trip on a member of a cluster, against a metadata group whose peers are themselves still booting — so the budget branches: 30 seconds per create solo, 2 minutes clustered, with the whole coordination bring-up bounded at 2 minutes and 5 minutes respectively. The flat 30 seconds these replaced was measured failing: a fleet booting together would lose one create, and because each object discovers a slow cluster independently the failure landed on a different stream or bucket every time. A node that exhausts the budget fails to start rather than running against a group it cannot reach, and the error names the object it was creating.
A request the broker never answers is asked again, not waited on. A clustered metadata request does not always come back late — sometimes the reply is never sent at all. nats-server drops a routed request outright in more than one ordinary situation during a bring-up: a member that is not the metadata leader and holds no assignment for the object returns without replying, and the routed API queue is discarded wholesale when it reaches its limit. A node waiting on one of those is waiting for something nobody will send, so a longer budget buys nothing — measured buying nothing, at two minutes a time, on a different object every attempt.
So the engine now distinguishes three answers to “does this object exist”, not two: it is there, it is not there, and nobody said. The first two are answers and are acted on; the third is silence, and silence is re-asked — at a new leader, or past a queue that has drained. An existence probe that goes unanswered for its whole term also simply falls through to the create, which settles the question either way: absent and it is made, present and it comes back as a peer having won the race. A node no longer fails to start because it could not hear.
A dropped stream or consumer lookup is asked again within two seconds; a create after sixteen. The broker decides a read — does this stream, this consumer exist — when it processes it: it answers, or it returns without a word, and nothing answers that request later. Every fleet booting together meets the second case. An object another node has just asked for is in flight, assigned by the metadata leader but not yet applied by the member chosen to lead it, and until that member has applied it every other member drops a lookup of it. So a lookup of one of the engine’s streams or consumers, of a coordination bucket, or of the object store’s bucket, waits one second for its answer and is asked again a second later — every two seconds, for up to thirty in all. A create is the opposite case: the broker keeps its reply until the object it made has a leader, which can take an election, so a create waits fifteen seconds for its answer and is re-sent a second after that. A lookup used to wait the create’s fifteen too, and a fresh three-member fleet was measured idling sixteen seconds of its boot on one dropped stream lookup.
A create that is taking a while says so while it is happening. Provisioning
was otherwise silent — a node opens every coordination bucket and several
streams in a row and logged nothing between them, so one that hung emitted
nothing at all until its budget expired and the log could not say which object
it was on. Any create still running after 10 seconds now writes one WARN
naming it (coord_kv_bucket_slow, natsobj_bucket_slow,
jetstream_stream_slow, jetstream_consumer_slow), and so does the lookup
that precedes it (jetstream_stream_lookup_slow,
jetstream_consumer_lookup_slow) — that lookup is the first call to reach the
metadata group, so a member stalled against a group that has not settled
waits there, where nothing used to report it at all. The read of each state
log’s ceiling that sizes the logs before any of them is created is the same
kind of lookup and says so the same way (statelog_ceiling_read_slow); the
four are asked at once, so a broker that answers none of them costs one
lookup ceiling rather than four in a row. The object store’s
bucket is looked up and created as one step, and writes one line over the two
(natsobj_bucket_slow). A probe that went unanswered is named as such
(jetstream_stream_lookup_unanswered, jetstream_consumer_lookup_unanswered,
coord_kv_bucket_lookup_unanswered, natsobj_bucket_lookup_unanswered)
rather than failing the boot. One line per object, deliberately: whether more
lines follow is what tells a slow bring-up from a wedged one.
Every broker line names the member that emitted it. More than one
embedded broker can run in one process — a fleet test does exactly that — and
without the name every queue.nats.server line from either of them was
indistinguishable, which is the one question a reader has about a fleet that
did not form. Lines carry server= from stream.cluster.name’s member
identity; a solo broker has no name to carry and the attribute is empty.
And it needs room for the state logs. The three logs reserve their byte
ceilings against the account’s JetStream storage limit when their streams are
created, and the node sizes them to half of what that limit has left. An
untiered limit counts every replica, so a replicas: 3 fleet needs three
times the bytes; a tiered one needs its R3 tier. An account that states no
limit leaves the node nothing to size against but its own disk, and a server’s
own cap then refuses what does not fit, by name. See
Replication.
Replication is asked for, not assumed. stream.replicas is the replica
count the engine requests for each of those streams and buckets, and it
applies to an external cluster exactly as it does to an embedded one — set it
to 3 against a cluster of three or more, or the engine asks for one copy and
gets what it asked for: streams and a lease bucket that survive a process
restart but not the loss of the single server holding them.
Tier A cannot check this number for you here. It refuses replicas above 1
on an embedded stream that names no peers, because that file contradicts
itself — but stream.url names an address rather than a member list, so how
many servers answer behind it is yours to know. Asking for more replicas than
the cluster has members fails at boot, on the first stream the engine tries to
create.
On an external cluster stream.replicas also picks which storage limit the
engine is held to. There is no store_max_bytes on that topology — the limit
is the account’s, and the engine reads it back — and a NATS account states
that limit in one of two shapes, never both. An ordinary account has a single
limit, which the server charges replicas × ceiling against, so the engine
divides it by stream.replicas before it sizes anything. A tiered account
states one limit per replica class (R1, R3, R5, …), already counting
replication, so the engine takes the tier for stream.replicas whole.
A tiered account that has no tier for the class you asked for — R1 and
R5 declared while stream.replicas: 3 — is neither, and neither is a tier
that is declared but carries no storage limit: the account’s report lists
every class it holds objects in, whether or not a limit was ever set for one,
so R3 being present is not R3 being declared. The engine reports a budget
of zero under its own source — account_no_tier or
account_tier_no_limit — rather than as an account that is merely full: the
two carry the same number and the opposite instruction, since one clears by
somebody freeing room and these only by a change of configuration. A tier
declared unlimited is neither again: it states its limit as a negative, the
broker creates against it, and the engine sizes from its own free disk as it
does for any broker that states no limit.
The broker refuses every stream, consumer and bucket create either way, but
not with the same words, so the engine’s own refusal carries the ones you
will actually see. A class the account’s limit table has no entry for is never
resolved at all — no JetStream default or applicable tiered limit present,
before a single byte is compared — and that covers both the missing tier and a
tier the report carries only because the account holds objects in that class. A
class it really does declare with no disk, a memory-only tier, resolves
normally and is refused by the byte comparison instead: insufficient storage resources available on every create that reserves bytes.
What that does to a node depends on whether its objects already exist, and the two cases look nothing alike.
A node that has not provisioned them does not start. crewlet run creates
its streams while it is opening its broker client, and its coordination buckets
straight after — all of it before anything sizes a ceiling — so the first
create is refused and the process exits. The first of those two refusals names
the class the account carries no limit for and the two levers that move it: the
engine classifies it rather than reading it as the object having failed to
appear, so what you get is R3 and stream.replicas and not (and it is not there: stream not found) appended to the broker’s bare text.
A node whose objects are all already there boots, on numbers the broker never
agreed to. Nothing is created, so nothing is refused. The
statelog_broker_states_no_limit warning says which shape it is and which
lever to move, and the statelog_ceilings line below it carries
broker_source=account_no_tier (or account_tier_no_limit) with
broker_limit=0 — so the logs are sized against nothing but what they already
hold between them, floored at a gibibyte each. It keeps working until it has to
make something new: a state-log stream a new version adds, the mailbox for a
seat you just added, a coordination bucket. That create is refused as above.
Either way: set stream.replicas to a class the account carries a limit for,
or have the cluster’s operator declare one for the class you asked for.
Running the Engine + API
Section titled “Running the Engine + API”Single Process (embedded API — the single-host default)
Section titled “Single Process (embedded API — the single-host default)”Any api.port > 0 in the Tier A YAML makes crewlet run start an embedded API server inside the engine process — one process runs the engine, the dashboard, and every webhook route:
api: port: 80 # 0 (the default) disables the embedded APIcrewlet run -config crewlet.yaml # engine + embedded API on :80(-api-port 8000 on the command line does the same.) This is the shape every single-host walkthrough in these docs uses, and that embedded server is the webhook target the integrations register (e.g. http://host.docker.internal:8000/webhooks/gitlab). Port 80 buys exactly one thing: a webhook URL with no :port suffix, which matters when the address is pasted into a vendor’s UI by hand or has to survive a proxy that rewrites ports. It costs a privileged bind — as a non-root process on Linux that needs sudo sysctl net.ipv4.ip_unprivileged_port_start=80 (persist in /etc/sysctl.d/) or CAP_NET_BIND_SERVICE.
The two bundled examples land on either side of that trade, which is the clearest way to read it. examples/nimbus.config.yaml pays for port 80: its company registers GitLab webhooks, so the address gets pasted into a vendor’s UI. examples/nimbus-claude-cli.config.yaml takes api.port: 8000, because a chat-only company has no inbound webhook at all and gets nothing for the privileged bind — its port only has to match the CREWLET_MCP_BRIDGE_URL its own seats dial back on. Make sure nothing else already owns the port you pick.
Do not also start a second node on the same host with such a file — both read the same api.port, and the second binder hits EADDRINUSE and kills whichever server came second.
The listener comes up before the seats do. crewlet run binds the HTTP surface first and only then starts claiming seats, so /dashboard, /health and every webhook route answer within a second of boot even on a company whose agents take much longer to come up. That ordering matters because claiming a seat starts that seat’s per-role MCP servers (one subprocess per server per seat, each a spawn, a handshake and a tools/list), and a company with seven seats and three vendors is twenty-one children. Serving after them made the whole inbound edge dark for as long as the slowest vendor took, which reads exactly like a hung process.
While seats are still being claimed the node reports what is true rather than pretending: /health lists the seats it holds so far, and nothing is lost in the meantime because every seat’s mailbox is created before any claiming (see Event System). A seat’s own children still start before its mailbox is attached, so a turn never begins without its tools — that ordering is unchanged; what changed is that they start concurrently rather than one after another, so a seat attaches in the time its slowest server takes rather than the sum of all of them.
Exposing webhooks without the admin API
Section titled “Exposing webhooks without the admin API”Every route an outside party calls shares api.port with everything else by
default: the vendor webhooks, the admin API (/config, /secrets, /setup,
/operator, the fleet and retention gestures), the read surface and the
dashboard. A deployment that must accept GitLab or Slack deliveries from the
internet but keep the rest private would then have to publish that one port
and filter it by path in a proxy — an allow-list kept outside the engine,
where one forgotten route or one prefix matched too widely puts /config on
the internet and nothing here can see it.
api.public splits the two onto separate sockets instead:
# crewlet.yaml (Tier A)api: host: "10.0.0.5" # private: the dashboard, the REST API, the probes port: 8000 public: port: 8443 # published: webhooks and the sandbox endpoints only # (host defaults to every interface)| Route | Served on |
|---|---|
/webhooks/* — every vendor delivery, the Slack OAuth landing, the GitHub App return | api.public only |
/otlp/{token}/v1/{signal} — sandbox telemetry | api.public only |
/mcp/{token} — the agent-mode tool bridge | api.public only |
everything else: /health, /ready, /dashboard and its assets, every REST read, /ws/stream, /config, /secrets, /setup, /operator/*, /backup, /fleet/*, /work/* | api.port only |
api.port answers a public route with 404 no_route and a hint naming
api.public — a vendor pointed at the wrong port learns it from the first
delivery rather than from a delivery quietly accepted on the socket nobody
published. The public listener answers every other route exactly as it answers
a path nothing serves: the same 404, the same generic hint, nothing that says
an admin port exists behind it. It requires no operator token and decides
before the guard runs, so /config there is that same 404 whatever
credential is sent. The split
is the engine’s own, from one list (the /webhooks/, /otlp/ and /mcp/
prefixes, the routes that authenticate by a provider signature or a signed
per-run token rather than an operator credential), so there is no allow-list to
keep in step with the routes a release adds.
Why the sandbox endpoints are public. A remote sandbox (E2B’s cloud) runs
on somebody else’s network and reaches only what you publish; it calls
/otlp/{token} and /mcp/{token} with the signed per-run token in the path and
never holds an operator credential. Leaving them on api.port would force that
port public for every company running remote sandboxes. Point
CREWLET_SANDBOX_OTEL_RECEIVER_URL and CREWLET_MCP_BRIDGE_URL at the public
address.
Why the probes are not. /health describes the node — its version, its
seats, its posture — for whoever runs it, and an orchestrator or a load balancer
can probe a port other than the one it routes traffic to. Point the public load
balancer’s health check, and the orchestrator’s liveness and readiness probes,
at api.port’s /health and /ready; on Kubernetes that is the pod’s own
address, which the kubelet reaches whether or not the port is in the public
Service. Readiness is per process, so a draining node leaves rotation on both
sockets at once.
What to point where:
integrations.public_base_url(Tier B) is where vendors and sandboxes reach this deployment, so it is now the public listener’s address as your proxy or load balancer publishes it. Every webhook the engine registers and every vendor app it provisions uses it.integrations.dashboard_base_url(Tier B) is where people reach the dashboard —api.port’s address as your private network or VPN reaches it. Every link an agent composes for a person (${crewlet_base_url}in a tool skill, a page change’surl) is built on it, and the public listener could serve none of them. On a deployment with one listener the two fields usually hold the same address; they are two fields because here they cannot.- The dashboard and the CLI stay on
api.port:crewletcommands that read a node’s own Tier A file dialapi.host:api.port, which is unchanged. - A node without the
ingressrole serves its probes onapi.portand one route beside them, its seats’ tool bridge, and withapi.publicset it binds the public port for the bridge alone — so one Tier A file still serves every role, and every node’sCREWLET_MCP_BRIDGE_URLnames its own public port. Itsapi.port: 0(or-api-port 0) is still the hard off switch: it binds nothing, the public port included, whatever the shared file says aboutapi.public.
api.host may stay 0.0.0.0 where the network keeps api.port private — a
container whose public Service or load balancer forwards only api.public.port
— or name a private interface where the host itself is reachable.
crewlet validate refuses a public listener beside api.port: 0 on a node with
the ingress role (no probe would answer) and a public port that is also
api.port, stream.cluster.port or
stream.leaf.port on the same address or with either binding every interface.
Draining. Both sockets go through the one drain gate: from the first moment
of a drain the webhook routes answer 503 with a Retry-After on the public
listener, while the sandbox endpoints keep serving the runs already in flight,
and api.port keeps /health at 200. Once the drain completes the public
listener closes first and api.port after it, each waiting up to five seconds
for its requests, so the probes answer until the very end — an open bridge
session can hold the public listener for its whole five seconds, and a
liveness probe must not fail while it does.
Separate processes (a split deployment)
Section titled “Separate processes (a split deployment)”Run ingress as its own node when you want the webhook receiver to stay up across engine restarts, or the two on separate hosts.
Two processes need a stream they can both reach, which a solo embedded
one is not — it binds no socket, so each would have its own. Either shape
from the section above works: a clustered
embedded stream (stream.cluster), or an external NATS server
(stream.type: nats plus stream.url). They also need shared coordination:
Tier A refuses coordination.type: local alongside either of them, by name,
rather than letting two nodes each claim every seat.
Both nodes can read the same Tier A file: against an external NATS server
nothing in it is per-node except node.id, and CREWLET_NODE_ID injects
that without templating anything. A clustered embedded stream is the one
exception — each member also names its own route port and its own peers. So
give each node its roles at the command line, and its own -api-port: on one
host two listeners cannot share a port, and only the ingress node’s carries
the API — the other’s carries its probes and nothing else:
# Terminal 1: the agents and the fleet duties — /health and /ready onlycrewlet run -config crewlet.yaml -roles data,seats,workers -api-port 8001
# Terminal 2: the webhook receiver and the dashboardcrewlet run -config crewlet.yaml -roles data,ingress -api-host 0.0.0.0 -api-port 8000-api-port 0 on the first binds nothing at all, which is fine on a host where
nothing probes it.
Give each node a distinct node.id (or CREWLET_NODE_ID) — two nodes sharing an id miscount the fleet. See Running a Fleet.
If any seat runs in agent mode, the seats node needs its port, and its CREWLET_MCP_BRIDGE_URL set to that port as a sandbox reaches it. Without the ingress role that listener serves the /mcp/{token} tool bridge beside the probes and nothing else (api_probes_listening … tool_bridge=true); with -api-port 0 the node refuses every agent-mode launch, naming api.port.
With a public listener on one host, the two processes would both bind api.public.port — the ingress node for its webhooks, an agent-mode seats node for its bridge — and the second fails at boot. Give each process its own public port through a whole ${VAR}, which a Tier A number accepts:
api: port: 8000 public: port: ${CREWLET_PUBLIC_PORT}CREWLET_PUBLIC_PORT=8443 crewlet run -config crewlet.yaml -roles data,ingress -api-host 0.0.0.0 -api-port 8000CREWLET_PUBLIC_PORT=8444 crewlet run -config crewlet.yaml -roles data,seats,workersThe seats node binds only its public port, for its bridge; one started with -api-port 0 binds nothing, so it needs no public port of its own.
crewlet migrate is idempotent and safe to re-run. Each node also
auto-migrates its own store file on boot, and two nodes starting together
cannot race, because they are not migrating the same file — every node owns
its own. Running the explicit step first turns a schema change into an
observable step rather than a side effect of startup.
Both take the Tier A bootstrap file (crewlet.yaml) — the founder-owned company YAML is seeded separately (crewlet config import, or crewlet run -company).
dataholds the company’s durable state — a copy of the replicated estate and the node’s own event log;ingressandworkersneed it, and a node without it is statelessseatsruns the agents — claims seat leases, boots the instances, processes their turnsingressserves the REST API — receives webhooks (Slack, GitLab, Jira, GitHub, Confluence) and publishes them to the event queueworkersruns the company-wide duties — the scheduler tick, the retention sweeps, the sandbox waiter
They are one command, and they build the same application: every node learns the company from the active config revision and the live picture from the broadcast event stream. Point CREWLET_SANDBOX_OTEL_RECEIVER_URL at whichever node is externally reachable: an ingress one, which serves the /otlp/{token}/v1/{signal} receiver. Its tokens are per-run and signed, so the node that mints and the node that verifies need no shared memory, and signing uses the Tier A keyring, so a split deployment needs one configured (crewlet secrets keygen); without it each process signs with an ephemeral key, logs sandbox_otel_signing_key_ephemeral, and every token one process mints is forged as far as the other is concerned. CREWLET_MCP_BRIDGE_URL, if any seat runs in agent mode, is the opposite: a bridge session lives in the process that opened it, so each seats node sets it to its own address and serves /mcp/{token} itself, on its own -api-port, even without the ingress role.
Point liveness probes at /health (stays 200 through a drain) and load-balancer readiness at /ready (503 while draining or before the first config revision applies, with the cause in its reason field). Every node answers both, whatever its roles — see Probes for what /ready means on a node that takes no traffic. A draining node keeps its listener until the drain completes, so both probes answer throughout, and it refuses any request that would start new work with 503 and a Retry-After; see During a drain. A node with nothing in flight drains in milliseconds, which is also the whole of an ingress node’s drain, so give such a pod a preStop sleep of a few readiness periods if you need the load balancer to have acted on that 503 before the listener goes. The engine will not sleep on its own: a delay long enough to matter would eat the terminationGracePeriodSeconds the drain itself has to finish inside, and only the deployment knows how much of that grace its longest turn needs.
Both communicate through the stream, and through the coordination KV riding
the same connection — never with each other. Both accept -debug for verbose
logging.
Probes
Section titled “Probes”Every node with an api.port answers /health and /ready on it, whatever
its roles, so an orchestrator can probe each one the same way. What the two
answer is fixed; what /ready means follows from whether the node takes
traffic:
/health (liveness) | /ready | |
|---|---|---|
A node with ingress | 200 while the process is alive, through a drain | Should traffic come here: 503 while draining, unconfigured, or on a shed or stuck posture |
| A node without it | The same, with the node’s half of the health envelope | Is it doing its work: also 503 while its broker link is down, it does not hold its presence lease, or — running seats — it has not been admitted to claim them |
A node without ingress serves nothing else on that port — no API, no
dashboard, no webhook — and a seats node’s agent-mode tool bridge only when
CREWLET_MCP_BRIDGE_URL is set. The reasons, their precedence and both bodies
are in Probes on a node without ingress.
On Kubernetes that is one probe block for every node, with api.host: 0.0.0.0
so the kubelet can reach the pod’s address, and api.port: 8080 — the port the
image’s EXPOSE documents:
# the engine container of any node — ingress, seats, satelliteports: - {name: http, containerPort: 8080}livenessProbe: httpGet: {path: /health, port: http} periodSeconds: 10 timeoutSeconds: 5 # /health reads the coordination plane under a failureThreshold: 6 # budget well inside this; a minute of silence # is a wedged process, not a slow brokerreadinessProbe: httpGet: {path: /ready, port: http} periodSeconds: 5 timeoutSeconds: 5Liveness tolerates a minute because the engine ends a wedged process itself:
the watchdog
exits once its watched duty stalls for a seat lease TTL (45 s), so a liveness
failure is the backstop for a process that cannot even do that. Readiness on a
node without ingress is what a rolling update waits on before it replaces
the next pod, so a satellite that cannot reach its fleet holds the rollout
rather than letting it take the next one down too. Give every pod a
terminationGracePeriodSeconds longer than its longest turn, for the drain.
A StatefulSet of clustered broker members
needs podManagementPolicy: Parallel. The default, OrderedReady, starts a
member only once the one before it is ready, and no member can be ready alone:
the company revision it applies and the presence lease it holds are both in a
coordination store that writes through a quorum of the members — so the first
waits for peers the StatefulSet will not start until it is ready, and the set
never comes up.
Replica count
Section titled “Replica count”Run one crewlet run, and scale up before you scale out. A single
engine handles many concurrent turns — agent handlers are
goroutines, so the practical ceiling is LLM provider rate limits and host
memory, not process count. One node is the design’s degenerate case, not
a lesser path: it holds every lease, and everything a fleet does works
exactly the same way.
Multi-node is supported and certified by a chaos suite that kills nodes
mid-turn under load. Reach for it when a node’s failure is not acceptable
downtime, when you need to terminate traffic separately from running
agents, or when some seats have to run somewhere specific — not as a
throughput lever, because max_concurrent is per process and N nodes is
N × that ceiling whether you wanted it or not.
Running a Fleet is the guide: node roles, seat placement, draining, and rolling upgrades. The two things that bite hardest:
A fleet needs shared coordination.
Seat leases live in the coordination slot. coordination.type: local is a
per-process store, so every node would believe it owns the whole company,
which is why a Tier A file that pairs it with a clustered or external stream
is refused at load on coordination.type. A fleet needs
coordination.type: embedded-kv; see Running a Fleet. The slot
governs the leases only — the fleet’s shared records are on the KV
regardless, because they have to survive a restart as much as a peer, and
the KV rides the stream’s own connection whichever value the slot holds.
And, when the nodes are the broker, a quorum to keep it on.
One node or three, never two: two embedded members have no quorum without
each other, so the fleet stops serving the moment either restarts, and Tier
A refuses that config by name, counting stream.cluster.peers.
stream.replicas: 3 is the other half — one replica count covers the
engine’s streams and the coordination buckets, because both live on the
same broker, and at 1 the loss of a member takes a seat’s mailbox or the
fleet’s leases with it. Against an external NATS cluster the quorum is that
cluster’s to provide rather than the engine’s to count; see
An external NATS server.
Raising it on a fleet that already ran needs the existing objects resized.
Nothing the engine provisions is ever rewritten by a booting node — a
shared stream’s configuration has one writer and it is not whichever node
started last — so a rolling restart after raising stream.replicas finds
every stream and bucket already there at the old count and adopts it. A node
that is short refuses to start and says so, naming both counts: a stream
because an acknowledged publish would be proving fewer copies than
stream.replicas promises, and a coordination bucket because the leases,
the fencing epochs and the company’s secrets would be on fewer disks than
the config claims. Resize the objects deliberately (nats stream update --replicas=3, which covers the buckets too — a bucket is a stream), or
stand the fleet up fresh.
What an acknowledged publish has reached
Section titled “What an acknowledged publish has reached”stream.sync decides, and it defaults to always at every replica count.
Every write is fsynced before the broker acknowledges it, so a publish that
returned is on the disk of the member that took it — which is what the
EventQueue contract’s “durable” means, and what the company’s own records
depend on. The cost is one fsync per write: 1–3 ms on NVMe, and 15–40 ms
at the 99th percentile on a network-attached volume.
It is deliberately not inferred from replicas. The tempting inference —
a replicated member has a quorum instead of a disk, so it can skip the fsync —
is true of one failure class and there are five:
| What fails | Does a quorum survive it? |
|---|---|
| One host loses power | Yes — the other two hold the write |
| The process is killed, or panics | Yes — the page cache is the kernel’s, and the kernel lives |
| An orderly shutdown | Yes — the store is flushed on the way out |
| A rack or an availability zone loses power | No — a majority can go together |
| Correlated power loss across every member | No — three copies of one unflushed page cache is one copy |
A three-node fleet in one rack, which is what a first production deployment usually looks like, is exposed to the bottom two rows by construction.
Declining the fsync is a legitimate trade and it is made explicitly. Set
sync to a duration — 30s — and that duration is the window: the most an
acknowledged write may be behind the disk. Tier A refuses the value in the
four places where it would be recorded and then not honoured:
- against
stream.type: nats, because the field configures the embedded server’s file store and an external cluster stores its own data (setsync_intervalon that cluster instead); - on a leaf (
stream.leaf.urls), because a leaf’s broker runs no JetStream and has no file store at all — the members it joins decide what an acknowledged write has reached, in their ownstream.sync; - below
replicas: 3, because the disk being traded away is the only copy there is, so the window buys nothing; - on a cluster whose peers are all on this host, because the majority the window trades for shares one power supply and one page cache.
Give each node a distinct id — node.id in the Tier A file, or the
CREWLET_NODE_ID environment variable, which is how a container orchestrator
injects a pod name without templating the config. Two nodes sharing an id
miscount the fleet and each compute too small a share.
Each node migrates its own store file at boot; there is no shared schema
to bring up first, and no migration lock, because no two processes share a
file. crewlet migrate applies them
ahead of time when you would rather not do it on the startup path.
What a fleet gets right, each of which was a real defect before:
- Duplicate Slack posts, duplicate Jira comments, two contradictory plans for one webhook. A seat’s inbox is attached only by the node holding its lease, admission is gated on a renew fresh enough to prove exclusivity, and the turn loop re-checks the seat fence at the top of every round and again before each of that round’s tool calls — so a node that loses the seat mid-turn stops before its next call rather than running out the turn beside the seat’s new owner. A turn that finished but whose delivery was never acked is not re-run, because the completion ledger records what shipped.
- Live coding sandboxes torn down mid-run. Recovery is a per-seat step inside the acquire hook, fenced on the claiming node’s epoch, instead of a fleet-wide scan that treated every in-flight run as abandoned.
- Config activation. Delivered by the control plane — a shared activation pointer whose own revision is the epoch, polled by every node — rather than the competing-consumer subscription that used to let exactly one replica apply a revision while the rest ran the previous company.
- Token budgets. Shared counters in the coordination slot, one slot per calendar window, so an org cap of 500 k a day is 500 k a day across the fleet — and they cover every completion the engine makes on a seat’s behalf, the turn loop (the round-cap extension judge included), the coding sandbox, the turn-start context assembly and the auxiliary learning passes alike.
- Duplicate auto-drafted skill pages and N× LLM spend on synthesis. Skill clustering, skill curation and episode compaction are singleton duties (they share one
worker:lease, so a fleet runs each of them on exactly one node), along with the scheduler tick, the sandbox waiter, the seat-subscription walk and the retention sweeps. Each lease is claimed per tick: a node that stops gracefully gives its duties back as it exits, and one that dies mid-duty hands them back by lapsing, which for the longer duties takes up to their TTL (45 minutes for the retention sweep, three hours for the learning passes). - Unbounded table growth.
scheduled_runsandconversation_sessionsboth answer a short-horizon question and are written on every event that asks it. The migrations always said they were swept on a TTL; the sweep exists, behind themaintenanceduty. Most fleet-shared records — the delivery dedupe, the rate valve, the completion ledger, the credential cooldowns and each node’s apply status — are not swept here at all: each lives in a coordination bucket whose own age is its retention, so the broker expires them. Agent-to-agent channels are the exception and are swept by the duty, because a bucket age cannot tell an open ask from an answered one — and so are the removal markers every bucket with no age keeps, which the duty sweeps after three hours so a listing never re-reads every record the company has ever removed. The apply status is the one that hides: it is keyed by node rather than by event, so it does not look short-horizon — but a node that is scaled in, redeployed or crashed would leave its last report behind, which under generated pod names is one per pod that ever ran, and the bucket’s one-minute age is what makes that node vanish instead.
The one thing that is still per-process: max_concurrent. Tier A’s
node.max_concurrent (default 32) is the gate every agent turn takes a slot
from, and it is per node — so an org’s ceiling becomes N × the configured
value. Size it per node, not per company.
For the model underneath all of this — what a node is, what the fleet shares, and where the constants come from — see Scaling Out.
The store
Section titled “The store”One local file per node, opened by Turso — the only driver, so there is no setting that chooses one. The file is in the SQLite format, so any SQLite-compatible client can read it.
Turso keeps a native library cache, and the engine prepares it before the
first query. The driver is pure Go in the sense that matters — no cgo, no C
toolchain — but its engine ships as a ~20 MB native library embedded in the
driver, extracted on first use into $TURSO_GO_CACHE_DIR (default
~/.cache/turso-go) and loaded from there. That cache is shared by every
process on the host and is written without a rename, so two engines starting at
once could leave a half-written file behind that fails verification for good.
Crewlet therefore extracts under a lock in <cache>/turso-go/, and clears and
re-extracts a cache entry that will not verify. Two consequences worth knowing:
- Point
TURSO_GO_CACHE_DIRat a writable, persistent path in an ephemeral container. A read-only or per-restart cache costs a 20 MB extraction on every start; a cache root that cannot be created at all fails the store open with an error naming the directory. - A cache that cannot be repaired names the way out, and there is no
second driver to fall back to: delete that directory by hand, or
point
TURSO_GO_CACHE_DIRat a writable directory of its own. - The linux binaries need glibc, and there is no musl build. The database
engine is a native library loaded with
dlopen, which makes the binary dynamically linked againstlibc.so.6even though it is pure Go and built withCGO_ENABLED=0. On Alpine and other musl systems it fails atexecve, reported asno such file or directoryabout a file that plainly exists. Use a glibc base image — the published one isdebian:trixie-slimfor exactly this reason — or run the engine on a glibc host. macOS is unaffected.
The engine owns the file exclusively. A second process pointed at the same path is not a degraded configuration, it is corruption waiting for a schedule to collide — so nothing that genuinely needs to be shared between nodes lives here. Seat leases, the activation pointer and per-node apply status, the completion ledger, webhook dedupe, the rate valve and credential cooldowns are all in the coordination slot instead.
The load-bearing tables:
agent_diary— each agent’s private observation log, each note with the vector of the model it came from (a per-agent scan the database ranks; there is no vector index). Written by the reflect path, which embeds content on write; the node holding the seat fills any note left without a vector of the current model. The## Personal memoryprefetch reads it via hybrid candidate selection (vector top-50 ∪ recency top-50, deduped by row id) handed to an aux-LLM relevance filter. Shared knowledge is not stored here — natively it is rows in the replicated estate beside the vectors derived from them, and a Confluence knowledge base has no local copy at all; see knowledge system.episodes— one row per completed turn, raw and LLM-compacted shapes in the same table; a raw turn carries the vector of the model it came from and a compacted row none (a per-seat scan the database ranks; there is no vector index). Drained by the episode-lifecycle duty.synthesized_skills+synthesized_skill_versions— auto-drafted skills the agent can load, plus their refinement history.counterparty_profiles— per-(observer, subject, platform)profiles built from observed interactions.agent_onboarding_markers— onboarding bookkeeping, one row per agent.crewlet_events— the observability event store. A phase completion’s token counts are promoted out of its payload into columns, so the spend rollup reads nine narrow values a row instead of hauling every prompt and response across the driver — which is what lets it fold the whole window rather than a capped prefix of it.crewlet_event_parties— which agents each event involves, one row per pair. It is an index of the table above rather than state of its own: the dashboard’s per-seat activity filter matches on it, and it exists because the engine’s planner does no OR-optimization, so the same predicate spread across five columns would scan the log instead of seeking. Swept on the same horizon as the events it points at.conversation_sessions— the conversation ledger: what this seat already said in one thread, rendered back into that conversation’s next turn.company_config— the revision payloads. Which one is current is the fleet’s business, and lives in coordination; see the control plane.secret_values— the bootstrap half of the secret store. The company’s credentials live on the coordination KV; rows written here while the engine was stopped are migrated there at its next start.
Migrations are forward-only: each file in internal/store/schema/ is applied once and recorded by filename, and there are no downgrade scripts. Downgrading the binary below the schema it already migrated is not supported; restore a backup instead. There is no migration lock and no advisory-lock protocol, because one process owns the file — the whole idiom disappears.
Everything else is either:
- YAML config — the org structure and every seat’s definition
- In-memory — agent runtime state, the execution tracker
- An external tool — task state (Jira, GitLab issues)
- The event stream — routing, with a durable per-subscription backlog
Observability
Section titled “Observability”The event store
Section titled “The event store”Crewlet persists every engine event (LLM invocations, task lifecycle, agent
states) to the crewlet_events table in the same file as everything else.
There is nothing to set up: the table is created by the engine’s own
migrations on first start, on whichever path store.path names. No extension
to enable, no managed service to configure, no separate retention system.
The engine registers the event-store writer as a publish listener on the
event queue. Every event is written at publish time, inline on the node that
published it — no queue round-trip and no consumer group, which is precisely
why two nodes can never write the same row and a group rebalance can never
lose one. Events land with dedicated columns for the common filterable
dimensions (event_type, source, category, agent_id, agent_role,
task_id, channel_id, sender, trace_id) plus a JSON column for
everything else.
That inline write is also why a fleet’s event store is per node: each holds
what it published. A history read is therefore asked of every live node
at query time — the node serving the dashboard reads its own store and
scatters the same question to its peers, merges what comes back, and names
any node that did not answer in the answer’s coverage (see
Reading the fleet’s history).
A node that has LEFT the fleet takes its turn-level detail with it; the
aggregates — spend, turn counts, page reads — survive it in the replicated
usage domain. A deployment that wants to keep every node’s detail past its
departure exports to an external sink over OTLP rather than pointing the
nodes at one database, which the exclusive file ownership rules out by
construction.
Custody: the rows of a node without data
Section titled “Custody: the rows of a node without data”A node without the data role writes none of its own. Its store is deleted
at every boot, so rows written there would be gone at its next restart with
nothing able to read them in between (the API runs only where the data is).
So it hands every event it would have written to the data nodes, and exactly
one of them keeps each one:
- It publishes them durably. Events are buffered for at most a second —
a turn never waits on its own audit trail — and published in batches of up
to 256 events or 1 MiB onto the custody topic,
crewlet.custody.records, on theCREWLET_CUSTODYstream. Once the broker acknowledges a batch it survives the node that published it; a batch it could not publish is sent again under the same id, which the broker collapses. The buffer holds 4 096 events while the broker is unreachable, and an event that finds it full is dropped and counted (event_custody_dropped) rather than holding a turn. - The data nodes take them from one group,
event-custody, and each writes the batch it takes into its own event log, where itsGET /eventsand the fleet’s history reads find it. - One data node keeps each batch. The group delivers a batch whose acknowledgement was lost again, to whichever data node asks next — so after writing a batch a data node claims it in the coordination store, create-only. The claim’s winner keeps the batch; any other node that wrote it deletes its copy. A batch is acknowledged only once its keeper is decided, and a data node that crashed between writing and claiming settles its copy the same way at its next boot.
That last step is what keeps every figure derived from a node’s own log
honest: the usage domain sums each node’s day, so a batch two data nodes
kept would bill a stateless node’s spend twice.
Until the node that lost the claim has settled its copy — at once when its claim is answered, and otherwise at its next pass, a minute later or as soon as it can reach the coordination store again — both logs hold the batch’s rows. The fleet’s history reads hold such a row once, wherever they list rows and wherever they count them — the event axis’s bars and category counts, a trace’s or a turn’s total, a page of turns’ tokens, and the integrations’ drop and merge counts: each data node counts the rows it keeps and names the ones it has not settled, and the node you asked counts each named row once (Reading the fleet’s history).
Upgrade the data nodes first. A data node writes each event in a batch by
its own build’s category map, and leaves out a type its build does not place —
so a node without data on a newer build can hand a data node still on this
one events it cannot keep, and they reach no event log. With every data node
upgraded before the nodes without data, nothing is left out.
What gets stored, and under which category
Section titled “What gets stored, and under which category”category is the one column with a closed vocabulary, and it is what the
dashboard’s filter and GET /events?category= group by. It is a property of
the event type, fixed in internal/events, and this table is generated
from that map — a guard test fails if the two drift.
| Category | Event types |
|---|---|
a2a | a2a_channel_closed, a2a_channel_opened, a2a_message_sent |
decision | contribution_received, contribution_requested, decision_requested, decision_resolved |
learning | compaction_completed, compaction_requested, counterparty_profile_updated, episode_written, knowledge_read, persist_decider_completed, prefetch_summary, reflection_completed, skill_archived, skill_promoted, skill_refined, skill_revived, skill_staled, skill_synthesized, skill_used, turn_completed |
lifecycle | backup_requested, config_revision_activated, config_revision_applied, operator_acted, org_started, org_stopped, seat_paused, seat_resumed |
notification | external_notification, notification_skipped, notifications_coalesced, turn_trigger_skipped |
system | agent_phase_completed, agent_phase_started, agent_turn_completed, agent_turn_started, agent_turn_steered, agent_turn_stopped, auxiliary_spend, budget_exhausted, llm_unavailable, phase.tool_skill_blocked, prompt.size, provider_fallback, skill_telemetry_write_failed, subagent_batched, turn.guard_breach |
task | sandbox_clarification_requested, sandbox_run_answered, sandbox_run_completed, sandbox_run_failed, sandbox_run_started, scheduled_task_fired, task_assigned |
webhook | inbound_delivery — one row per delivery presented to one seat, from a webhook route or a Mattermost socket. The row is filed under the delivery’s own label (webhook:<event>, forge:<event>, socket:posted) rather than under inbound_delivery, with the provider’s exact bytes as its payload |
The map is also the admission list. A type that is not in it is not written and does not reach the activity feed — so the exclusions below are deliberate and each one says why, and a new type that nobody placed fails a test rather than vanishing quietly.
| Excluded type | Why |
|---|---|
agent_turn_progress | Fires as each LLM round opens, answers and runs its tools, as a live-only signal; the matching agent_phase_completed is its durable record, so persisting this would fill the log with intermediate states of rows it also holds finished. It still drives the live projection. |
agent_spawned | Placement moves a seat between nodes on every rebalance, so a durable row per claim would fill the log with a fact about scheduling rather than about the company. It still drives the live projection, which is what asks “is this seat running, and where”. |
agent_terminated | The counterpart, excluded for the same reason. It is what takes a released instance’s call off a live screen rather than leaving it showing whatever it last did; whether the seat still runs anywhere is the seat leases’ to say. |
raw_webhook | The delivery is already a row (inbound_delivery, in the webhook category above). This event is the wake an inbound edge publishes for the transports to route, so categorising it too would store every delivery twice — once as what arrived and once as what was forwarded. |
a2a_request | The ask is already a row: a2a_channel_opened and a2a_message_sent record the same exchange under the ids the audit trail is keyed on. This event is the wake it puts on the target seat’s inbox — same reason as raw_webhook. |
a2a_message | The answer is already a row (a2a_message_sent). This event is the wake it puts on the requester’s inbox. |
sandbox_answer_given | The wake an answer by turn puts on the seat’s inbox, and never a turn. What the answer became is already a row (sandbox_run_answered), and that a person gave it is their operator_acted row — same reason as a2a_request. |
tool_skill_page_changed | A nudge between nodes that one tool-skill page moved, so every node’s registry re-reads it rather than only the node that won the webhook. The delivery that caused it is already a row (the webhook category above), and what the change did is a log line on each node, so a durable row would record one wiki edit once more per member of the fleet. |
reflection_due | The wake that puts a finished turn in front of post-turn reflection on the seat’s holder. The turn is already a row (turn_completed), and what reflecting on it did is its own (reflection_completed) — same reason as a2a_request. |
custody_batch | A carrier, not an event: a node without the data role keeps no event log, so it publishes its events in batches and one data node writes each event inside as the row it is (custody). A row for the batch would describe the transport and repeat every event in it. |
budget_meters | A snapshot of the shared token counters, published by every node on a 15-second tick and at once when a budget window first refuses a charge, so a durable row per report is about two million a year per node to answer a question the live projection and GET /budgets answer for free. What the audit log holds instead is the spend the counter is charged with, recorded per phase in the agent_phase_completed rows and per auxiliary purpose in the auxiliary_spend rows every spend query folds, so “what did we spend last month” is answerable and “what was the counter reading at 14:03:15” is not a question anybody asks. It still drives the live projection. |
Stored is not always fed. One class of type is written like every other row
and kept out of the dashboard’s activity feed — a ring of the whole company’s
last few hundred events, since every node’s feed is fed fleet-wide. It is in
GET /events, a turn’s history and a trace like any row, and filterable by its
category.
| Stored, not in the activity feed | Why |
|---|---|
auxiliary_spend | Accounting, not activity: what the auxiliary model cost for one key, coalesced per flush (Budgets and spend § Auxiliary spend). A turn writes several beside its own phases, and a compaction a burst, so in the feed they would push the turns, failures and deliveries a reader is watching out of the ring. |
Querying events
Section titled “Querying events”The dashboard’s Event log is the
intended reader — filters, traces and event detail, over the same
/ws/stream query channel the REST routes use, so both surfaces answer from
one implementation.
For ad-hoc SQL, point any SQLite-compatible client at store.path while the
engine is stopped, or use the read-only endpoints under
/events while it runs. Do not open the
file with a second writer against a running engine.
Tracing
Section titled “Tracing”Crewlet uses OpenTelemetry for distributed tracing. Every event carries W3C Trace Context fields (trace_id, span_id, parent_span_id) that propagate automatically through the system.
How Traces Flow
Section titled “How Traces Flow”Five span names, and that is the whole set. webhook.receive,
agent.turn (plus agent.turn.resume), agent.turn.<phase>, llm.round and
tool.call, alongside schedule.fire for cron-started work.
Span attributes are deliberately thin: the seat, the phase, the model, the round, the tool and its outcome, and token counts. Everything about what a turn did — prompts and responses verbatim, tool arguments and results, the decision — is already in the event store, and a span carries what no event does, which is duration.
Only the LLM round is spanned, not the fallback chain, each member’s backend and the credential pool beneath it — on a three-member chain over a four-key pool that would nest a dozen spans per round and tell you nothing you could act on. Which member and which credential answered is on the phase event.
A suspended run is two spans, not one. A code sandbox run detaches: the phase returns, the process may exit, the seat may move node, and the resume can be days later. A live span cannot survive that, so the suspending span ends and the resume opens a new one under a reconstructed parent — the wait shows up as the gap it actually is.
Dashboard Trace View
Section titled “Dashboard Trace View”The dashboard groups a trace’s events into a tree — reach one from any row that carries a trace_id, or paste the id into the search box:
- Root event (e.g., webhook) shown as the trace header
- Child events nested underneath with connecting lines
- Click
inspect →on LLM turn events to view the full prompt/response - Notification skip reasons shown inline (e.g., “not following this thread”)
OTLP Export
Section titled “OTLP Export”To export traces to Jaeger, Grafana Tempo, or any OTLP-compatible backend:
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318That is the collector’s base URL — the engine appends /v1/traces itself.
Do not include the signal path here; use OTEL_EXPORTER_OTLP_TRACES_ENDPOINT
when you need to give the full URL. Getting this wrong also reaches the sandbox
forwarder, which appends a signal path of its own, and the collector sees
/v1/traces/v1/traces.
When either is set the engine installs a batching exporter at startup and flushes it during shutdown, after the drain, so the spans a shutdown itself produces are exported rather than dropped. Without it, spans are still created and their ids still reach every event, the event store and the dashboard’s trace view — there is simply nothing shipping them anywhere.
The full set of variables, including the protocol, the service name and the sampling ratio, is in Environment Variables. The engine’s exporter and the sandbox OTLP forwarder read the same endpoint and headers on purpose: a coding agent’s spans land in the same backend as the turn that started them, nested underneath it.
Correlating logs with traces
Section titled “Correlating logs with traces”Every log line emitted inside a span carries trace_id and span_id, in all
three log formats. So a span you are looking at in Jaeger and the
lines the engine wrote while it was open are joined by the same identifier:
crewlet run -log-format json | jq 'select(.trace_id == "4bf92f35…")'Lines emitted outside any span carry neither field rather than carrying an
empty one, so a shipper indexing trace_id never sees a placeholder.
Querying Traces in the Event Store
Section titled “Querying Traces in the Event Store”The crewlet_events table stores trace_id, span_id, and parent_span_id as first-class columns, so trace queries are a direct column filter:
-- All events in a specific traceSELECT event_time, event_type, source, summary, payloadFROM crewlet_eventsWHERE trace_id = '<trace-id>'ORDER BY event_time ASC;The dashboard API also provides GET /events/trace/{trace_id} which returns all events in a trace ordered by timestamp.
Logging
Section titled “Logging”How loud a node is, in what shape, and where it writes, is Tier A:
# crewlet.yaml (Tier A)logging: level: info # debug, info (default), warn, error format: console # console (default), text, json stderr: true # default. false needs a file below — see "The log file" file: # optional: a durable copy, IN ADDITION to stderr path: "/var/log/crewlet/crewlet.log" format: json # empty follows logging.format level: debug # empty follows logging.level max_size_mb: 100 # rotate at this size (default 100) max_backups: 5 # rotated files kept beside the live one (default 5)That block is the only way the file says it.
The same settings, on the command line, for one run:
crewlet run -debug # shorthand for -log-level debugcrewlet run -log-level debug # what -debug is shorthand forcrewlet run -log-level info -log-format json # for a log shippercrewlet run -log-file /var/log/crewlet/crewlet.log # a durable copycrewlet run -log-file "" # and no file, for one run, # whatever the Tier A saysA flag overrides the file only when it is actually given. A flag carries
its default whether or not anyone typed it, so crewlet run distinguishes
“the operator asked for info” from “nobody said anything” — otherwise
logging.level: warn in a file would be dead on arrival behind the flag’s own
default. -debug only ever raises: to quieten a node whose file says
logging.level: debug, pass -log-level info. -log-file obeys the same
rule in both directions: it replaces logging.file.path for one run, and an
explicit -log-file "" is how a node with a file configured is asked to write
none. It moves only the path — the shape and the rotation caps describe the
disk this deployment runs on rather than this invocation, so they stay the
file’s.
The first lines of a run come out in the flag’s shape, not the file’s — and
before any log file exists. The ${VAR} warnings a Tier A document produces
are emitted while it is being read, so a node configured format: json writes
those few lines as console, on stderr only, before switching. The log file is
named by the document that is still being read, so it cannot be open yet.
That is the best a process can do about a file it has not opened, and it is
the right way round: -debug is turned on most often to watch the config load
itself fail, so the flags have to take effect first. If a boot fails on the
document itself, stderr is the only place it is recorded — start there before
the log file.
A value the build does not recognise is treated differently in the two
places, on purpose. In a flag it resolves to the default — a bad log level
must never be why a company will not boot. In the file it is refused, with
the field path, by crewlet validate and at boot: a flag is typed by someone
watching the process start, and a file is written once and deployed for
months, so a misspelled level there would run quietly at info for as long as
nobody looked. Either way the fallback is never silent: an unrecognised
-log-level / -log-format, or $CREWLET_LOG_LEVEL / $CREWLET_LOG_FORMAT,
logs a log_level_unrecognised / log_format_unrecognised warning naming what
was written, what the build used instead, and what it accepts.
A log file is the one logging value that does not fall back at all. A
level is an enum with a sane default; a path is not. A node that could not
open the file it was told to write refuses to start, naming the path and the
error, rather than running on stderr alone — an operator who configured a
durable record and silently did not get one has nothing anywhere pointing at
why. The same
applies to $CREWLET_LOG_FILE on the other commands. Once the node is up, a
file that becomes unwritable — a full disk, a volume pulled away — is the
opposite case and is handled the opposite way: the failure is announced on
stderr once, the console sink keeps every line, and the engine keeps running.
It is announced again when the file starts taking writes, so the gap has two
ends.
The three formats
Section titled “The three formats”| Format | For | Shape |
|---|---|---|
console (default) | A person watching a terminal | Fixed columns — time, level, component, event — with attributes dimmed, and ANSI colour when the stream is a live terminal |
text | Grepping without a parser | slog’s key=value: time=… level=INFO msg=seat_claimed component=seat.host seat=eng.alice |
json | A log shipper | One JSON object per line |
console adapts to its sink. Colour appears only on a live terminal, so a
redirected stream carries no escape codes it cannot render — and because a
redirected stream is read later, its lines carry the full date where a
terminal’s carry the wall-clock time alone. CREWLET_LOG_COLOR=always|never
overrides the detection (for a CI viewer that renders ANSI without being a
terminal), and NO_COLOR suppresses it the way it does for every other tool.
A log file is never a terminal, so a console-format file is never
coloured and always carries the full date, whatever CREWLET_LOG_COLOR says —
the variable describes the screen someone is looking at, and nobody is looking
at a file.
The log file
Section titled “The log file”logging.file.path adds a durable copy of the log. It is a second
destination, not a redirect: stderr keeps every line it had. That is
deliberate — stderr is the only sink that exists before the document naming
the file has been read, it is what a container platform captures, and it is
where a boot failure and the watchdog’s exit
notice are written. A node that fell silent
there the moment a path was configured would look exactly like one that had
stopped. A deployment that genuinely wants the file alone redirects stderr in
its unit file or its container spec.
Because the two are separate destinations rather than one stream tee’d in two, each carries its own shape — which is the point:
logging: format: console # columns and colour, for whoever is watching file: path: "/var/log/crewlet/crewlet.log" format: json # one object per line, for the shipperLeave file.format out and the file follows logging.format, so a node that
says nothing writes one log in two places.
The level splits the same way, and both directions are real. file.level
is how loud the file is; unset, it follows logging.level, so -log-level
and -debug move both destinations at once.
logging: level: warn # what a person watching sees file: path: "/var/log/crewlet/crewlet.log" level: debug # what the incident is reconstructed fromA debug file behind a warn console keeps the detail an incident needs
without burying whoever is watching; a warn file behind a debug console
keeps the durable record small while somebody works. The process admits the
louder of the two and each destination filters, so log.Enabled(…, debug)
answers “will this be recorded anywhere” — meaning a debug file costs the
work at every debug call site whatever the console says. That is the price of
asking for a debug file, and it is paid whichever destination reads it.
Turning stderr off
Section titled “Turning stderr off”logging.stderr: false silences the ordinary log stream on stderr once a file
has taken it over:
logging: stderr: false file: path: "/var/log/crewlet/crewlet.log"Use it where the platform already captures stderr and you keep a file — journald plus a log file, or a container with a log driver plus a mounted volume — because there every line is otherwise stored twice. Without it the default stands: a file never silences stderr.
The tradeoff it buys you: with stderr off the file is the node’s only destination, so a file that becomes unwritable loses log lines rather than diverting them. The engine still says so on stderr — the notice names the file, the error, and that the lines are lost rather than continuing elsewhere — but the lines themselves are gone until the file takes writes again. Leave stderr on if a gap in the record is worse for you than storing it twice.
It is not 2>/dev/null, and the difference is the point. Three kinds of
line reach stderr without passing through the configured handler, and this
field keeps all three while a shell redirect throws them away:
| What | Why it bypasses the handler |
|---|---|
| Everything before the Tier A document is read | The log file is named by that document, so it cannot be open yet |
| The seat watchdog’s exit notice | It writes to stderr directly and calls os.Exit(75); a wedged process has not earned a configured handler |
| “log file X: no space left on device” | A sink cannot report its own failure through itself |
A terminal that is about to go quiet says so. Before the switch takes
effect the engine writes one plain line to stderr naming the file it is
handing over to — not through the logger, and so not subject to
logging.level. That matters: the structured log_file_opened record is
ordinary telemetry at info, so on a logging.level: warn node it is
filtered, and without the plain line crewlet run would print nothing at
all with no way to discover the log was in a file.
A node with neither destination is refused, by name: logging.stderr: false with no logging.file.path fails validation and fails the boot, and
so does a -log-file "" that takes the file away from a document that had
switched stderr off. Silence is never what configuring logging meant.
The other commands are unaffected — they read no logging: block, so their
stderr always stays on.
Rotation is built in, and it cannot be turned off. A log file with no
ceiling fills the disk the store is on, and it does it on exactly the
deployments nobody is watching — so there is no “never rotate” spelling, only
a size you will not reach. The live file rotates at max_size_mb (default
100) and the rotated ones are kept as crewlet.log.1 (newest) through
crewlet.log.N, max_backups of them (default 5). Together the defaults
bound the estate at roughly 600 MB. max_backups: 0 is a setting rather than
an absence: it keeps no history at all, which is what a small disk with a
shipper already tailing the live file wants.
The size is checked before the record that would cross it, never after, so a record is never split across two files — half a JSON object at the end of one file and half at the start of the next is a parse error in whatever is shipping it. A single record larger than the whole cap is written whole into an empty file rather than rotating forever around something that can never fit.
Restarting appends; it does not rotate. A restart loop is precisely when the previous incarnation’s last lines are the evidence, and rotating on every boot would push the first failure off the end of the stack by morning.
Missing directories are created, 0700, and the file is 0600. A log line is
redacted but it is not a public document, so a shipper running as another user
needs a chmod you make deliberately.
Already running logrotate(8)? Point it at the same path with
copytruncate — which keeps the descriptor this process holds — and give
max_size_mb a value this node will never reach. Both caps are bounded above
as well as below (max_size_mb at 1 073 741 824, a pebibyte; max_backups at
1000) and a value past either is refused by name: the size is held in bytes,
so a larger one wraps and would rotate on every line — the exact inverse of
what a huge number asks for — and the backup count is a rename per rotation,
so a huge one stalls the rotation instead of keeping more history. A rename-based logrotate
rule moves the file out from under the engine’s open descriptor, and there is
no reopen signal to send it: the engine owns its signals for the graceful
drain (see crewlet run), and a third tier
of signal handling is not worth a mechanism this file already has.
One file per node. Two processes on one host — a split ingress / seats
deployment, or a node beside a crewlet migrate — pointed at one path will
interleave their lines and rotate each other’s file, and nothing detects it.
Put the node id in the path, which resolves like any other Tier A ${VAR}:
logging: file: path: "/var/log/crewlet/${CREWLET_NODE_ID}.log"Every line is structured whichever format is installed, and carries a
component attribute naming the subsystem that emitted it (agent.turn,
mcp.client, seat.host) — the field console promotes into its own column
— so a debug run stays filterable rather than becoming a wall.
The operator commands are quiet by default: they open a store, which logs a
line per migration, and that is noise on a one-shot command whose output is
meant to be piped or diffed. They take no logging flags — only crewlet run
does — so CREWLET_LOG_LEVEL, CREWLET_LOG_FORMAT and CREWLET_LOG_FILE are
their levers: the third appends a crewlet migrate or a crewlet validate to
the same durable record the node writes, which is the reason a CI step wants
any of them. It writes only what the command logs; whatever the command
prints for its caller still goes to stdout, so a piped or diffed output is
untouched. crewlet run ignores all three — its level, shape and file come
from Tier A and its own flags. See
Environment Variables.
Nothing silences a warning.
The embedded broker’s own logs
Section titled “The embedded broker’s own logs”The NATS server the engine embeds logs through the engine’s logger, under the
component queue.nats.server, so what the broker said is always
distinguishable from what the engine said about it. Anything it reports as
wrong — a JetStream write error, a slow consumer, stream recovery after an
unclean shutdown, cluster election trouble — keeps its own severity and
reaches the log whatever else is configured. Its boot narration (“Starting
nats-server”, the JetStream storage line, “Server is ready”) is debug: a
dozen lines describing infrastructure you deliberately did not deploy.
Its own debug output is a separate switch, stream.debug, and it is off
by default:
stream: debug: true # only when the BROKER is what you are diagnosinglogging.level: debug and -debug say how loud the engine is. They are
what you want to watch a turn — the prompt, the tool calls, the review — and
they deliberately do not turn this on, because nats-server’s debug output is
per internal client rather than per event, and the engine’s own coordination
reads manufacture those continuously. Every coordination key listing is an
ordered consumer created and then deleted, and deleting one writes two lines
like:
DEBUG queue.nats.server nats_server detail="JETSTREAM - JetStream connection closed: Client Closed"Two of a node’s fifteen-second duty loops list keys on every tick, so that is
a constant background stream on a node doing nothing at all. Client Closed
is the graceful close reason and nothing is leaking; it is simply the
broker narrating its own housekeeping.
Both switches have to agree for these to appear: stream.debug decides
whether nats-server produces them, and a destination at debug decides
whether anything records them. crewlet validate warns when the first is set
and the second is not. stream.debug is refused for stream.type: nats —
an external cluster logs wherever its own operator configured it to, so a flag
here would reach nothing.
Per-Agent Token Tracking
Section titled “Per-Agent Token Tracking”Every LLM completion records prompt, completion and total tokens plus a call count, per agent and per model, with cache reads and writes broken out so cached prefixes are visible rather than folded into the total.
Read them from the Spend & budgets screen in the dashboard, or over the
socket’s query channel — tokens for the rollup and budgets for the caps
beside the durable counters the engine enforces against. Both are the same
functions the REST routes call, so the two surfaces cannot disagree.
crewlet budgets show # the durable counters, read from a running nodeIt talks to a node rather than to a file: the counters are the fleet’s, and on
the default topology they live inside the running engine. -url and -token
name another node; without them they are taken from the api block of the
config on the command line.
Token Budgets
Section titled “Token Budgets”Set budgets at two levels, each a mapping of ceilings per calendar window —
day, week and month, each optional — cut on the company’s
clock:
- Org-wide —
token_budgetin the top-level YAML config - Per-agent —
token_budgeton each Role definition
token_budget: {day: 3000000, month: 40000000}An absent window is uncapped, and a ceiling of 0 is refused rather than read
as unlimited. See Configuration § Token budgets
for the rules and for the ceilings crewlet validate warns can never bind.
Every model round is charged against both the moment its reply arrives (its
size is known no sooner) and before any tool it asked for runs, in every
window at once — the day, the week and the month it falls in on the company’s
clock — and it is admitted only while every capped window of both had room for
it. A round that does not fit is refused: its tool calls do not run, the turn
stops, and the engine publishes a budget_exhausted event naming the scope that
refused, the window (period, window, resets_at) and its figures, beside
the turn’s own agent_turn_completed. The refused round is counted all the
same, on the seat’s counter and the company’s, because the vendor has
already billed it: the refusing window reads past its ceiling by the round that
crossed it, and every later round is refused against that figure until the
window turns over or the ceiling is raised. Each scope’s check is atomic, and
nothing a charge counted is taken back: a seat write that fails after the
company’s landed stops the round as an outage and leaves it on the company,
whose record of a billed round is true, so only that seat’s own counter is
short of it — logged as coord_kv_budget_spend_uncounted — until the window
turns over. In a fleet the counters live in the coordination
slot, so an org cap of 500 k is 500 k across every node rather than per
process.
A window’s allowance comes back when the window turns over — at local midnight, on Monday, on the 1st — rolled inside the first charge after the boundary, so nothing has to run for it and no node has to be up at midnight. There is no reset: room before a window turns over is made by raising its ceiling, which takes effect on the next turn. See Coordination § Token budgets are windows.
A seat out of room waits; its mail is not lost. Before a delivery is
handed to a seat, the node asks whether one of that seat’s capped windows —
its own or the company’s — is refusing. If one is, the seat is parked:
its inbox is held, the delivery goes back to the broker for one of its
deliveries, and it is delivered again when the window turns over (the one
that ends last, where several refuse) or at once when an applied revision
changes the ceilings. The node logs seat_budget_parked with the window and
when it resets, and seat_budget_park_released when the mail flows again. A
turn refused part-way through is parked the same way, unless it had already
written outside the engine, in which case it is recorded and not run again.
See Agent Runtime § The budget park.
Every seat is counted, capped or not. A company that sets no ceiling
still has its spend on the counters, in every window, so a ceiling added
part-way through a day judges what the day has already spent rather than
starting from zero, and GET /budgets shows an uncapped company’s spend
rather than nothing.
A refusal is also recorded beside the counter, as when the gate last turned a
call away in that window (refused_at on GET /budgets
and on the live token meter),
and the next charge the scope admits clears it, as does the window turning
over. The gate turns work away in two ways, and both are recorded: it refuses
a round’s charge, or — once a window is already full, because a context
assembly, a rewrite, a coding run, a person’s answers or an earlier round took
it there — the work is turned away before its first call, since that call
would be billed and then refused. That second way is every refusal of a full
window the engine can see coming: a turn stops its next call, the
budget park defers a seat’s
delivery, a person’s answer_knowledge question is refused, and the
reflection stage declines a pass and a conversation entry’s rewrites. Each is
recorded on the scope a charge would be refused by (the company’s before the
seat’s), a turn’s once per window for each part of the turn, so a window that
is refusing everything sent its way never reads as one that has refused
nothing. Because the counter carries every round it refused, GET /budgets,
the live meter and the park all read a window that refused a round over its
ceiling by the round that crossed it, never just short of it; refused_at is
when the gate last said no. A window filled with nothing asked of it since — a
background pass, a coding run whose turn has not resumed, and no delivery,
question or pass for that scope after it — carries no refused_at until
something is, though its state is already refusing.
No model call can be checked by its own size first, because its size is known only once it has happened. A turn’s round is judged when it is charged, as above, because a verdict still has something to stop: the tools it asked for and every model call of the turn after it, each refused before it is sent rather than billed and refused in turn. Four other spends have nothing left to stop by the time their size is known, so each is post-charged — added to the counters in the windows it is recorded in, without a verdict — behind a gate that reads the room left before it starts:
- A coding run. Its box spends while the turn is suspended, so its tokens
are known only when the run is collected, and they reach both the seat’s
counter and the company’s in the windows the run is collected in. A run that
takes a counter past its cap is logged as
sandbox_spend_over_budget. - An auxiliary pass — the reflection, profiling, compaction and summary calls the learning subsystem makes on a seat’s behalf. A pass does not start for a seat or company with no room left, and each completion it makes is recorded in full on the seat’s counter and the company’s, past the ceiling included. The same gate stands in front of the rewrite of a conversation entry’s long tool payloads, made after a turn under the same reflection stage: with no room left the entry is still written, each such payload named by its size and digest.
- A turn’s own auxiliary calls — the memory filter, the knowledge query and the episode summary the turn-start prefetch asks the seat’s auxiliary model for before the first phase opens, and every rewrite the turn’s conversation block, ledgers, judge and tools need. Their gate is the turn’s own: a delivery to a seat whose window has no room left is parked before it runs, so no prefetch starts for it; each completion the turn does make is recorded on the seat’s counter and the company’s through the turn’s own meter, which then holds any window the completion filled; and the turn makes no further call, auxiliary or not, while that window is full — except its task card’s rewrite, the turn’s record, made after its last round.
- A person’s knowledge answer — the dashboard’s ⌘K answer
(
answer_knowledge). A person has no seat budget, so it is gated on the company’s windows alone — refusedbudget_exhaustedbefore any model call when one of them has no room — and recorded on the company’s counter alone.
In each case the next round the seat or the company attempts is refused against the recorded figure — and for a turn’s own auxiliary calls, and for a coding run its resumed turn collects, that round is refused before it is sent: the turn’s meter holds what the post-charge answered, or reads it once before the resumed turn’s first call.
Structured Logging
Section titled “Structured Logging”Every significant operation emits structured log entries:
{ "time": "2026-03-12T10:30:00.412Z", "level": "INFO", "msg": "onboarding_phase_complete", "component": "agent.onboarding", "agent": "sarah-chen", "turn_id": "9f3c1e70-…", "marked": true, "rounds": 6, "chain": "…", "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736", "span_id": "00f067aa0ba902b7"}trace_id and span_id are present on any line emitted inside a span — which
is every line a turn produces — and absent entirely on lines that are not, so a
shipper indexing them never sees an empty placeholder. They are the same ids the
tracing section exports, which is what lets you pivot from a slow
span in Jaeger to the lines the engine wrote while it was open.
The JSON key for the message is msg, slog’s own, and its value is the
short, machine-parsable event name every line carries in place of a sentence.
Reacting to events
Section titled “Reacting to events”Every state change in the engine is an event on the stream, so anything that
wants to react — dashboards, alerting, an external audit sink — subscribes
rather than polls. /ws/stream is the read surface for a client; nothing runs
inside the process to hook them, because the engine loads no plugins.
Security Boundaries
Section titled “Security Boundaries”- Scope isolation — agents can only access knowledge within their permitted scopes
- Tool availability — all registered tools available; per-role MCP tools carry role-specific credentials
- Communication permissions — agents can only post to channels they’re members of
- Manager handoffs — agents identify their manager from their identity prompt and reach them through the colleague-surface tools (Slack/Jira/Confluence/A2A); engine-detected failures surface to the operator dashboard as the seat’s
last_error, and an unreachable provider as the seat statestopped/provider - LLM sandboxing — tool execution results are validated before returning to the agent
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.