Coordination
Crewlet keeps state in two places, and which one a fact belongs in is decided by a single question: does the company have to agree on it, or does one node?
- The store is the node’s own database. One file, one process, exclusively owned. Nothing in it has to be safe against a peer.
- The coordination store is the fleet’s shared slot. Everything that must be true for the company rather than for a process lives here.
Getting a fact into the wrong one is not a crash. It is a subsystem that looks correct on one node and is quietly wrong on four — which is how every entry in the table below arrived.
The three-valued answer
Section titled “The three-valued answer”Every question this layer answers has three outcomes, not two:
| Answer | Meaning |
|---|---|
| yes | The store said so. |
| no | The store said so. |
| unknown | The store could not be reached. |
Collapsing the third into no is the single most incident-hardened lesson in the engine. “I do not hold this seat” and “I could not ask whether I hold this seat” are different facts, and treating the second as the first tears a healthy company down over a two-second blip in a database.
So every call returns (value, error), never a bare bool — and each contract states which direction is safe for it, because the safe direction is not the same for all of them:
| Contract | On an unreachable store | Why that direction |
|---|---|---|
| Rate valve | Fails closed — refuse | A valve that opens when it cannot be read is not a valve. The cost of refusing is a delayed notification. |
| Delivery dedupe | Fails open — do not suppress | A claim that cannot be read has not suppressed anything. Suppressing on a read failure drops the delivery entirely, and nothing redelivers it. |
| Completion ledger | Fails open — do the work | The pre-ledger answer. Failing closed parks real work during an outage; the redundant turn is bounded and visible. |
| Lease renew | Holds, briefly | Ambiguity is not loss. The watchdog is what bounds it — see Seat Ownership. |
| Budget charge | Fails closed — stop the round | Money leaves the building for every token, and a counter that cannot be reached must not un-cap a company. An error is not a refusal, though: the caller fails the turn rather than telling an agent it is out of budget. |
A listing obeys the same rule, and it is the place it is easiest to lose.
Reading a whole bucket — the fleet’s node statuses, the open channels, every
node’s log position — either answers every key that was live throughout the
read or answers unknown, never a shorter list, on one broker or a cluster of
them alike; its one bound is a size, the 100,000 keys under one filter spelled
out below. The alternative is not
hypothetical: a listing that reports what it managed to read, with no error,
hands every caller “there are no more records” when the truth is “the store
stopped answering”. For the trim’s published floor that reads as a fleet
needing nothing, which deletes records a node is still replaying.
A listing is one ordered pass, certified against the stream’s own key
index. The pass carries each key together with its value — never a name
list followed by a fetch per name, which is what it was. That shape costs a
round trip per key on top of an ephemeral consumer created and destroyed per
call — paid continuously, since several of a node’s fifteen-second duty loops
read a bucket on every tick, and the state-log write fence lists the trim’s
published floors on every write at an expectation of zero, a subject’s first
write among them. That listing is the floors’ own key class in the
positions bucket, not the per-node positions beside it: the fence compares
this node’s checkpoint against the higher of the published floor and the log’s
first surviving sequence, which it reads from the stream itself.
A pass on its own cannot tell that it is complete. The marker that ends it
is the client’s guess from a pending count, and every bucket here keeps one
revision per key, so rewriting a key the pass has not reached yet removes the
revision it was about to deliver and appends the replacement behind the
marker. A renew is exactly that rewrite, and every live lease takes one on
every heartbeat: measured against the embedded broker, with five presence
leases and three of them renewed in tight loops, 848 of 3,634 membership reads
missed at least one lease that was live the whole time — a live node that
looks gone to placement, or a positions row the trim’s minimum never saw. So
every listing also reads the stream’s per-subject index, narrowed by the same
filter and answered by the stream leader from the index its store keeps under
its own lock, and reads by itself any key the index names that the pass never
delivered. That read is the stream leader’s too — $JS.API.STREAM.MSG.GET
for the last message on the key’s subject, which only the leader answers —
never the bucket’s own Get: on these buckets that is a direct get, served by
any replica in the direct-get group, and a replica that has fallen behind stays
in that group and answers “not found” for a key the leader’s index just named,
which silently undid the certification it was there to perform. An index read
while the stream has no leader is refused rather than trusted: during an
election every member answers from its own store, and one that is behind names
fewer keys than the quorum holds, so a listing during a bucket’s leader
election is unavailable — the third answer — rather than short.
One copy that is behind still answers as the leader, and the bound on it is stated rather than removed. A member cut off from its peers while it leads a stream goes on answering that stream’s leader reads from its own copy until it notices it has lost them — about ten seconds, the broker’s lost-quorum interval — while the majority elects a leader of its own and moves on. So through a node on the minority side of a partition, a bucket its member led reads up to that far behind, and a record the majority wrote reads as absent. Nothing written through that node lands, since a write needs its quorum, and a lease is fenced by its epoch, so what the window can mislead is a decision that writes nothing. Closing it would take a read that proves a quorum — a barrier append per read, as the state log’s linearizable read is — which would turn every lease heartbeat’s read into a write.
A tombstone the pass delivered is not taken as the key’s answer either. The pass is served by whichever replica the broker placed its consumer on, so on a clustered bucket a key that was deleted and re-created before the listing began, read by a pass on a replica that has applied the delete but not yet the re-creation, comes back as a delete marker although it was live throughout. So a marker the pass delivered, for a key the index names, is read again from the stream leader exactly as a key the pass lost is, and the leader’s answer — the live value, a marker, or nothing — is what the listing holds. A marker for a key the index does not name needs no read: the leader held nothing on that key at an instant inside the listing, so the key was not live throughout, and absent is a correct answer for it.
A pass can also stop short and never end, and the certification is what makes ending it safe. The client sends the marker that ends a pass on a delivery, so a pass whose pending messages are all removed before it reaches them — the last counters of a window ageing out, a marker another node’s sweep purged — receives nothing and would wait for ever: measured as one listing of an idle budget bucket that held its caller for nine minutes. So once the index has answered, a pass that goes five seconds without a delivery is ended there (the broker’s own idle interval for this kind of consumer), and every key the index named that it had not delivered is read from the leader like any other key it lost. Every delivery restarts the five seconds, so only a pass the broker has stopped feeding is ever ended this way. The marker sweep, which nothing certifies and which owes nobody a complete answer, ends a pass that goes quiet the same way and sweeps the rest on its next tick.
The one watch the store hands its callers — the seat pauses every node follows — begins with exactly this listing. Its first answer is the one a node replaces its copy with, so it is certified like any other: a pause the pass had not delivered when it ended is read from the leader rather than left out, which would lift that seat’s hold, and a pass that goes quiet is ended after the same five seconds rather than leaving the node with no answer — and deferring every delivery it receives — until somebody next pauses or resumes a seat. Its changes then come from the same consumer, so none can fall between the answer and the first change, and a revision of a pause no newer than the one the answer holds is not passed on as a change. It costs what a listing costs, once for each watch a node starts: at boot, and again whenever the watch has to be re-opened.
A key live from before the listing began until after it ended is therefore in it, however often it was rewritten in between and whichever replica served the pass. A listing that fails at any step hands its caller nothing rather than the half it read, and each key is handed over once, at the newest revision the listing read. What a listing may do is lag — a key claimed or deleted while it ran can be seen late, which is a placement that converges one tick later.
The certification costs one index read — two API requests, run beside the pass rather than after it so their round trips overlap — and one leader read for each key the pass lost or delivered as a marker. The lost keys are zero without concurrent writes and bounded by the keys written while the listing ran. The markers are bounded by the marker sweep: a bucket with no age would otherwise keep the marker of every record it ever removed, and every listing whose pass met one would read it again, so without the sweep a listing’s cost grew with every record the company had ever removed. It is never a read per key, unless the broker stopped delivering a pass for a whole five seconds and the rest of it had to be read key by key.
Its one bound is the broker’s page size for that index: 100,000 keys under one filter. Past it the client reads the index in pages by offset into a list the server re-sorts for every page, so a key removed before a page boundary between two page reads shifts a later key across that boundary and leaves it uncertified — which loses the key only if the pass lost it too.
The batched direct get that would remove even the consumer is deliberately not used here: it is served by any replica, so a follower behind an acknowledged write can hide a row, and the trim floor is a minimum across rows — a row it cannot see raises the floor and deletes records a node still needs. Measured against a single-node broker it also returned empty answers for populated key classes, because a KV bucket keeps one message per subject and that churn leaves the per-subject last-block index it is resolved through stale. An empty answer is the one this estate cannot survive, since “no rows” is legitimate everywhere it is asked.
That is why a resource name is segmented. A lease is named
seat:{handle}, node:{id} or worker:{duty}, and the part before the colon is the class; the key it becomes carries that class as a
subject token of its own, so seat is a wildcard and the seats are addressable
without the nodes. The reads that pay for it run on a ticker: the fleet’s
membership read asks for the presence leases — one class instead of every
lease in the fleet — and the sweep’s
placement hints come from the epochs bucket — the one with no expiry at all,
holding a record for every resource the deployment has ever leased, which used
to be read whole every five seconds to find one node’s seats.
Every read of one key is the leader’s too, not only a certification’s. A claim reads its own write back to learn the store’s deadline, a renew and a release read the lease before they write, and the fleet’s records are read by key — and each of those, answered by a replica that had not applied the write yet, told its caller the write was never made. On a three-member cluster with a hundred nodes claiming at once, 743 of 10,000 claims on fresh seats answered “not held” for a lease that was then theirs until its TTL — a singleton duty dark on every node, since the one that won believed a peer had it — and a renew through a member that was behind answered “no longer yours”, which a node acts on by shedding the seat. So every single-key read goes to the stream leader, and one the leader cannot answer — none elected, none reachable — is unavailable, never “absent”. It costs the hop to the leader: measured in-process, 40–220 µs at the median where a replica’s read took 25–135 µs, and about half the read throughput through one node’s connection — a few percent of it spent by ten thousand seats renewing on the fifteen-second heartbeat.
There is deliberately no all-classes listing. A class is one segment of a
name, so the empty one addresses nothing, and a read of it would answer with
an empty result rather than an error — which reads to a caller exactly like a
class with no members. Asking for a class that cannot address a key is
refused instead. For the same reason a resource may not have an empty segment:
seat: builds a key nothing can decode, so the lease would be written and
then returned by no listing at all, which every node reads as a free seat.
What the fleet shares
Section titled “What the fleet shares”| Slot | Answers | Documented in |
|---|---|---|
leases | Which node runs which seat, and which nodes are alive at all | Seat Ownership |
duties | Which node holds which singleton duty, and every other lease that is not a seat or a node’s presence — the work tracker’s claims on a running move, merge or bulk edit among them. Its own bucket because those and a seat want opposite TTLs: see The lease TTL bounds seats and presence only | Seat Ownership § Singleton duties |
epochs | The monotonic fencing counter each seat’s, duty’s and node’s tokens are minted from. Its own bucket because it is the one thing here that must never expire (see the retention table below) | Seat Ownership |
config | Which company revision is current. The key’s own revision is the fencing epoch | Control Plane |
status | What each node managed to apply, and when it last said so | Control Plane |
ledger | Has this trigger already been worked — read before a turn, written after one | The completion ledger |
claims | Has this inbound delivery been seen — the dedupe that used to be a per-process map, so a third-party app’s retry to a different ingress node woke the same seat twice | Event System |
rate | The notification valve. Four nodes ran four of them, so a seat capped at five a second emitted twenty | Event System |
cooldowns | Which provider credential is cooling after a 429. Per-process monotonic values are not even comparable across nodes | Deployment |
budgets | Org and per-seat token spend per calendar window — one record per scope with a slot for the day, the ISO week and the month on the company’s clock, each carrying the window’s label, its spend and when the gate last turned a call away in it (cleared by the next charge the scope admits, and by the window turning over). Caps stay config-derived in memory; only usage is shared, because a counter per node makes an org cap of 500 000 into N × 500 000. The refusal is kept beside the counter because it is the gate’s own decision and every node reports it: a refusal one node remembered would flicker on a dashboard as different nodes reported. See Token budgets are windows | Deployment § Token budgets |
channels | Who is asking whom, and whether the ask is still open. The record authorizing an answer is read by the node that owns the answering seat — never the one that opened it | Event System § Agent-to-agent |
fires | Has this scheduled dispatch already been claimed. The scheduler is a singleton duty, so it moves — and a successor reading its own database found an empty ledger and gave every company two standups | Scheduling § At-most-once |
rebases | Which instant a unit of work’s tracker and knowledge-base writes are minted at, for the work that began longer ago than the operation ledger remembers — a trigger dispatched a month late, a turn resumed a month after it parked. The attempt that finds its work’s start past that horizon mints at its own instant and records it here, and every later attempt at the same work — a crash re-run, a retried resume, the next half of the turn — inherits it. Company-wide because the next attempt runs wherever the seat’s delivery is taken next: a successor reading its own database found nothing and wrote every write of the attempt before it a second time | Turn Engine § A turn’s two identities |
sandbox runs | Every detached coding run: its box, its suspended conversation, its owner and fencing epoch, and whether its tokens are already charged. A run outlives its turn, its process and sometimes its node, and is recovered by whichever node owns the seat next | Code Sandbox |
secrets | The company’s credentials, one sealed envelope per ${VAR} name. Coordination holds bytes it has no key for; the Tier A keyring opens them at the edge. It was the last kind of company-wide state living in a node’s own database, so crewlet secrets set reached one node and a rotation half-landed | Secret Store |
integrations | Where each external surface’s reconcile pass got to: its phase, its findings, the address it was set up against, and whether a disconnect has been asked for. It is company-wide because the loop is a fleet singleton and moves — a status in a node’s own database would be a screen that changed answer depending on which node served the page | Integration Reconcile |
mailboxes | Which seat mailboxes may exist, and since when a seat has been missing from the active revision. Every node records a handle before it creates the seat’s durable subscription, because a removed seat’s handle is gone from the org every node derives names from and the retirement’s absence stamp and mark have nowhere else to live. A mailbox that escaped the record is found by listing the broker’s subscriptions. Every change is a compare-and-set, since a returning seat’s registration and the sweep that retires a mailbox write the same record | Seat Ownership § Singleton duties |
follows | Which chat threads each seat is following, one record per (backend, seat, channel, thread). It is company-wide because an inbound chat message is claimed and parsed by ONE node — notify-inbound is a competing consumer group — and the next reply in the same thread by whichever node wins that time: a follow only one node could see made a non-mention reply reach its seat by chance, less often the more nodes ran. Every follow is written here and nowhere else; no node’s own database holds one | Slack |
custody | Which data node keeps each batch of a stateless node’s events. A node without the data role keeps no event log, so it publishes its events in batches and a data node writes each batch it takes — and since the group can deliver one batch to two data nodes, the writer then claims the batch here, create-only, after writing it: the winner keeps the batch and any other writer deletes its copy. Company-wide because the two writers are different nodes, and claimed after the write because a claimant that died before writing would hold a batch nobody wrote | Deployment § Custody |
objects | Which store the company’s files are in — nats, or s3:<endpoint>/<bucket>/<prefix> — recorded create-only by the first node to boot and compared by every node after it, which refuses to boot on a mismatch rather than split the files between two stores; and what the object-collector duty last found, because the duty moves and every node’s /fleet reports the latest pass whoever ran it | Object Store § One store per fleet, § Collection and audit |
positions | What the state log may delete, in four key classes: where every node stands per domain, the live pins a backup or a joining node holds, what each owner’s newest backup covers, and the floor the trim itself published with the term that is holding it. Four classes in one bucket because all four answer one question and all four need the same retention, which is none. Four more share it for that retention alone — a log’s capacity operation, each node’s admission to publish, each node’s acknowledgement that it restarted for the operation, and the seat pauses (seat_pause): which seats a person has paused, by whom, why, and whether they also stopped the running turn — because each must outlive any clock: an expiring operation admits publishers, an expiring admission hides one, an expiring acknowledgement un-seals a barrier that has already run, and an expiring pause is a resume nobody chose. Every listing filters by class | Retention, Changing a log’s ceiling, Agent Runtime § Pausing a seat |
The page and work-item embeddings are not here. They were a slot of their own once, on the argument that a derived thing wants a lifecycle of its own. They are now rows in the replicated estate, derived once by a fleet-singleton duty and applied on every node through the state log like every other replicated fact — which is what makes the company pay the provider bill once rather than once per node. See Knowledge System.
A fleet is not configured — it is discovered from these, which is why adding a node is starting a process and removing one is stopping it.
Credential cooldowns, in practice
Section titled “Credential cooldowns, in practice”The cooldowns slot is the one an operator sees behave differently the moment a second node joins, so it is worth stating what it actually does.
A rate limit belongs to the key, at the vendor — not to the process that discovered it. Without sharing, four nodes each pay their own 429 to learn what the first one already knew, and with a two-key bag that is eight wasted calls and eight slowed turns for one quota window. The two halves of the fix are deliberately on different clocks:
| Half | When | Why there |
|---|---|---|
| Publish | Synchronously, on the bench that caused it | That is the only moment the fact exists. Deferring it to a tick leaves a window in which every peer rediscovers it. The write is detached from the caller’s context and bounded at two seconds — the call it belongs to has already failed and is about to be retried on another key. |
| Pull | Every 15 seconds, per node | A cooldown runs for a minute at the very least (60 s is the configurable floor), so reading one a few seconds late costs nothing — while a coordination read in front of every model call would put the store’s latency under every turn and its availability under the whole company. |
Three properties follow from that, and each is load-bearing:
- A record is an instant, not a duration. A peer that received “cool for an hour” would restart the hour whenever it happened to read the record, so a key benched once would stay benched as long as anyone kept pulling.
- A pull extends, never shortens. A peer’s record is evidence a key is refused; the absence of one is not evidence a key works. So a node whose own 429 no peer heard about is never talked out of it — and an unreadable store is a no-op rather than a mass un-benching.
- The record is scoped by the provider entry, and carries a hint rather than the key. One credential listed under a
fastentry and asmartentry is two rate-limit buckets at the vendor; an unscoped record would turn one model’s burst into a company-wide outage. And the ledger is a shared store, so what goes in it is 12 hex characters of SHA-256 — enough to tell a handful of keys apart in a log, not reversible.
A node that has just started pulls immediately rather than waiting out its first interval: a fresh process has an empty bench, and the fleet may have a key cooling for the next hour. When one arrives that way the node says so — credential_cooled_by_peer, naming the provider, the key’s hint and the time left — which is the only answer to the question this creates: why is a key benched on a node that never saw a failure?
A single node shares nothing, because there is no peer to tell. Cooldowns stay in its own process, exactly as they did before any of this existed.
Retention is a bucket’s age
Section titled “Retention is a bucket’s age”Every slot above except epochs, config, channels, sandbox runs, secrets, integrations, mailboxes, objects and positions forgets on a horizon, and the horizon is a property of the bucket, not of the write. leases is in that group and is the load-bearing case: a lease does not expire because something deletes it, it expires because the bucket’s age is the lease TTL, which is exactly what makes a dead node’s seat reclaimable with nobody around to release it. duties is in it too, with one difference that matters: its age only reaps a record, and a duty — or a tracker claim — ends at the deadline its own record carries, judged by every reader against the broker’s clock (see The lease TTL bounds seats and presence only). claims has the same difference: a delivery claim lapses at the deadline its record carries, judged against the clock of the node claiming next, so a five-minute webhook claim and a thirty-minute socket claim share one bucket. That deadline is data in the record rather than a per-key TTL, so the create-only rule below does not reach it.
That is a constraint rather than a preference. On the default embedded backend a per-key TTL is create-only: an update clears it, leaving the key immortal. A rate window that is incremented four times would therefore never expire — the one key in the system guaranteed to be written more than once. So each retention is fixed when its bucket is created, which is why they are separate buckets rather than prefixes in one:
Fixed when the bucket is created means fixed by whoever created it. Every node opens every bucket at boot, and a node that finds one already there adopts it rather than rewriting its configuration — including the leases bucket, whose age is the lease TTL. So on a fleet the retentions in force are the ones the first node to boot asked for, and a peer configured differently logs coord_kv_lease_ttl_differs naming the value actually in force and then runs at it — its acquires, its heartbeat, the budget it spends giving seats back on a drain and the lag at which its event-loop watchdog ends the process all derive from the live TTL rather than from its own file, because the bucket’s age is what expires a lease and a node claiming longer than the bucket allows would have every acquire refused and hold no seats at all. The alternative — every node asserting its own Tier A on every boot — is N writes against a metadata group that is still electing, resolved by boot order, so the node that came up last would silently redefine how long every other node’s leases lived. Changing a retention is therefore an operator gesture, not a restart: align the config across the fleet and delete the bucket while the fleet is down, so the next boot re-creates it.
Replication is the one adopted difference a node refuses to run with. Everything else a running bucket can disagree with is reported — the lease TTL above is warned about and then honoured — because every other difference changes how the store behaves, and behaviour is visible. A replica count changes nothing until a node is lost, and then it changes everything: raise stream.replicas from 1 to 3 on a fleet that already ran and a rolling restart finds every bucket still there at one replica, adopts it, and goes on holding every lease, every fencing epoch and the company’s secrets on a single disk while each node reports itself correctly configured for three. So a node that is short refuses to start, naming both counts. Equal or higher passes, so a single-replica development node against a replicated fleet’s buckets still runs. Resizing is the same operator gesture as any other bucket change — a bucket is a stream, so nats stream update --replicas=3 covers it.
| Bucket | Age | Sized from |
|---|---|---|
leases | the lease TTL (45 s by default) | The expiry is the mechanism: a renew rewrites the key and restarts the clock, so a node that stops renewing stops holding, and its seats become claimable without anything having to notice it died |
duties | 3 hours, the longest duty’s TTL, and never lowered | Each record carries the TTL its duty or claim asked for and is judged by that deadline against the broker’s clock, so the scheduler’s 30-second duty moves within 30 seconds and an abandoned move’s claim frees a minute after its last heartbeat. The age only has to outlive the longest duty, and it is only ever raised, because a lowered age would have the broker reap a long duty a peer still holds |
epochs | none | The fencing counter, and a fence that restarts is not a fence: a deleted key would hand the next owner a token a zombie is still writing under. It is a separate bucket from leases and duties for exactly this: they want opposite retentions |
rate | a few multiples of the window | A closed window must age out, and must never outlive its successor |
claims | 30 minutes, the longest claim’s TTL | Like duties, each record carries the TTL its claim asked for and lapses at that deadline: a webhook delivery is claimed for 5 minutes — a third-party app’s redelivery and an operator’s replay, not its full retry schedule — and a Mattermost post for 30, because a seat that reconnects replays the last 15 minutes and every node holds every seat’s socket (see Running on a fleet). The age only has to outlive the longest claim; a claim asking for more than the age in force is refused rather than clamped |
ledger | 7 days | Must outlast the queue’s redelivery horizon and the scheduler’s catchup ceiling — expiring a completion a tick could still evaluate lets that fire run twice |
cooldowns | 24 hours | The longest cooldown anything sets. A cooldown stores its own end instant, so the bucket only has to outlive the longest one |
status | 4 reconcile intervals (~60 s) | A node that stops reporting must vanish from the fleet view rather than linger as a healthy row nobody is writing |
config | none | The pointer is the fencing sequence, and a fence that restarts is not a fence |
budgets | 32 days | The longest calendar window a counter has a slot for is a month — 31 days, and at most an hour of clock change — and every charge rewrites the record, so a record older than that counts nothing a current window can be refused against. The age is not the reset: a window’s allowance comes back when the window turns over, rolled inside the charge that crosses the boundary (see Token budgets are windows) |
fires | 7 days | Must outlast the scheduler’s catchup ceiling, for a sharper reason than the ledger’s: a completion that expired early makes a turn re-run, while a claim that expired early makes the catchup pass dispatch a fire the fleet already ran |
rebases | 30 days | The operation ledger’s own retention, a day past the horizon a recorded instant is inherited within: an attempt inherits one only while it lies within that horizon of the attempt’s own clock, and the age counts from the write, which is no earlier than the instant recorded — so the bucket forgets a record exactly when no attempt could still inherit it. Forgotten sooner, the next attempt would mint anew and write again what the attempt before it wrote |
sandbox runs | none | The sharpest version of the channel case: a run parked on a person’s answer waits days, and its record is the only thing that knows a billed box exists. Its own pause reaper and its terminal delete are what end it: a run’s record is deleted the moment the run settles, done or failed, once its box is reclaimed, and a removed seat’s runs are ended when its mailbox is retired. No settled run is left for the completion poll and every seat recovery to read again |
channels | none | A bucket’s age cannot tell an open channel from a closed one, so a TTL would reap the authorization record of an ask still waiting for its answer. Closing an idle channel and deleting a closed one are decisions instead, taken by the maintenance duty |
secrets | none | A credential is not short-horizon state, and an expiring secret is an outage on a timer — one that arrives at the moment a vendor rejects a token every node believes it still has. A secret leaves when an operator unsets it |
integrations | none | A status is standing state, not a recent event: it says what the last pass found, and it is true until the next one. One that expired would make a converged surface read as never-reconciled and send the loop to re-provision what is already there. It is bounded by the number of surfaces a company has rather than by a horizon, and a row leaves when its block leaves the company document |
follows | 90 days | The one aged slot whose horizon is a last-activity stamp rather than a window: every re-assert — a mention, a collective address, the seat posting into the thread — rewrites the record, so the bucket’s age tracks the conversation. Ninety days is where a chat thread stops being live on every backend that ships one, and the asymmetry makes it safe: a dropped follow costs at most one missed non-mention reply, which the next mention re-establishes, while keeping every follow for ever grows a record read on the hot path of every inbound message |
custody | 32 days | A day past the event log’s own retention: a data node that wrote a batch and crashed before learning whether it kept it asks at its next boot, so the record has to outlive every row it can be asked about — and nothing longer, since a batch older than the log’s retention has no row left to settle |
mailboxes | none | A record’s age cannot tell a seat that is still in the company from one that left, so an age would forget a mailbox that still exists and leave it retaining mail for a seat nobody runs. It is bounded by the handles a company has ever used, and a record leaves when the maintenance duty retires its mailbox |
objects | none | The backend record is standing state for the life of the deployment, and one that expired would let the next node to boot record a different store — the company’s files split between two places with nothing failing. The collector’s report beside it is replaced by every pass, so it needs no age either |
positions | none | The sharpest case in the table. A node’s position is what the trim reads to decide what every other node may delete, so a key that expired would read as a node that has applied nothing — which either pins the trim for ever or, read the other way round, lets it delete records that node still needs. A node stops being counted by an operator’s audited eviction, never by a clock — and its row stays even then, because a readmission is judged by the position it holds |
Putting two of those in one bucket gives one of them the other’s retention, and every such mistake is silent — a cooldown that expired in a second, a fleet view showing a node that died last week.
This is also why the retention sweep in the maintenance duty has no jobs for the aged buckets: the broker expires those records, so there is nothing left for a sweep to delete, and a job that swept an empty table every tick would only report that it had. Every ageless bucket is the exception, for the reason its row gives — nothing expires them, so removal is a decision somebody takes, and each names a different somebody. channels and mailboxes are the maintenance duty’s own decision, the one closing an idle ask and the other retiring a removed seat’s inbox; integrations is the reconcile loop’s, which forgets a surface whose block has left the company document on the tick that notices; secrets wait for an operator’s unset; the objects backend record is removed only by an operator moving the company to another store, and the collector’s report beside it is replaced by each pass; a node’s own positions row is never removed at all — an audited eviction stops the trim counting it, and the row is kept because a readmission is judged by it; and sandbox runs end at their own pause reaper or a terminal delete.
Removal markers are swept
Section titled “Removal markers are swept”Removing a record does not remove it from the bucket’s stream: a delete or a purge appends a marker that says the key was removed, and on these buckets — one revision per key — the marker replaces the value. An aged bucket takes its markers with everything else. An ageless one keeps each for the life of the deployment, and every listing whose pass meets one reads it again from the stream leader, so without a sweep a listing’s cost would grow with every record the company had ever removed.
So the maintenance duty sweeps them, as the job coordination_markers: on one node for the whole fleet, every 15-minute maintenance tick, from every shared-state bucket with no age — config, channels, sandbox runs, secrets, integrations, mailboxes, objects and positions. That set is derived from the retention each bucket is opened with rather than listed, so a new ageless bucket is covered without anybody remembering to add it, and one that never removes a record costs a pass that finds nothing. epochs is the one ageless bucket outside it, and needs nothing: a fencing counter is never removed, so it holds no marker. The sweep removes the markers written more than three hours (coord.MarkerRetention) before the tick, and a bucket it could not sweep is retried on the next tick without stopping the others.
Three hours is not what keeps an answer right. Nothing in the engine takes a marker as an answer: a listing asks the leader about each one its pass delivered, and a create over a removed key asks the leader what the key holds before it writes — a live value is the key existing, a marker is the revision to write over, and nothing at all is a key never written. The contract suite sweeps every marker, however new, and holds each record family to the answers it gave before. The horizon decides how long a removal stays visible in the bucket against what that costs: a listing reads again each marker its pass meets, so it pays for the removals of the last three hours and a quarter — the horizon plus one tick — where without the sweep it paid for every removal the deployment had ever made; and three hours is the longest claim any operation on these records can hold (coord.MaxDutyTTL), so nothing that read a record before it was removed can still be acting on that read when its marker goes.
The sweep removes only what it saw. It purges each key through the marker’s own revision, and the broker applies that bound, so a record created again after the sweep’s pass saw the marker has a later revision and is untouched. The client library’s own sweep (PurgeDeletes) purges a key outright, and would take that new record with the marker.
The lease TTL bounds seats and presence only
Section titled “The lease TTL bounds seats and presence only”A seat lease and a duty lease want opposite TTLs. A seat is renewed on a heartbeat, so its TTL is a few heartbeats and a dead node’s seats move within a minute; a node’s presence is renewed on the same heartbeat. A duty is claimed once per tick of the work it guards, and a tick runs from ten seconds (the scheduler) to an hour (the learning passes), so a duty’s TTL has to outlive several of its own ticks: the scheduler’s is 30 seconds, the integration reconcile’s four and a half minutes, the retention sweep’s 45 minutes and the learning passes’ three hours. The work tracker’s claims are the same kind of lease: a cross-project move or a merge holds a claim for as long as its walk runs, renewed every 15 seconds and lapsing a minute after its last renewal, and a bulk edit holds the fleet’s one bulk admission for twice the time its rows are projected to take to apply — about two minutes for the largest.
A bucket’s age is the longest TTL it can keep, so one bucket cannot serve both. While duties shared leases, every duty longer than the seat lease TTL was refused, and on a fleet running coordination.type: embedded-kv the retention sweep, the mailbox retirement, the integration reconcile, the skill curator and every integration setup pass never ran, with one warning per attempt (maintenance_duty_claim_failed, integration_duty_unknown) as the only sign. Moving the duties out fixed the duties alone: the tracker’s claims stayed in leases, so at the shipped 45-second lease TTL every cross-project move and every merge failed before it wrote anything, with an error naming a TTL, and a large bulk edit was let through with no admission at all. A single node running local coordination was unaffected by either.
So leases holds seats and presence and nothing else, duties holds every other lease — every worker: lease and every class a caller claims under, the tracker’s move:, merge: and bulk: included — and the rules are these:
- Any lease other than a seat’s or a presence lease may ask for any TTL up to three hours (
coord.MaxDutyTTL), on every backend, whatever the seat lease TTL is. One that asks for more is refused with an error naming the ceiling, on the in-memory backend as well, so a duty or a claim too long for a fleet fails in a single-node test rather than only in production. - The ceiling is the longest duty’s TTL, and an engine test holds the two equal, so neither can move without the other. Every other lease it bounds is far shorter.
- Changing
coordination.lease_ttl_secondschanges seats and presence only. A duty’s TTL comes from its own cadence and a tracker claim’s from its own work, so shortening the lease TTL for faster failover refuses neither, and lengthening it permits neither anything longer. - Neither holds a newer build back. The mixed-version gate counts presence and seat leases only, so it never reads this bucket. A duty is claimed ungated and can outlive a crashed holder by up to three hours. Counted, it would keep a fleet from placing seats for that long after the last older node died mid-upgrade.
Token budgets are windows
Section titled “Token budgets are windows”A token budget is a set of ceilings per calendar window — the day, the ISO week from Monday and the calendar month, cut on the company’s clock — and the budgets slot counts spend per window (ADR-0019). Each scope — the company, and each agent seat — has one record with a slot per period: the label of the window the slot counts (2026-09-23, 2026-W39, 2026-09), what has been spent in it, and when the gate last turned a call away in it.
- A charge is admitted only while every capped window of both scopes has room, and it is still one compare-and-swap per scope — the company first, then the seat. Every window of a charge is counted, capped or not, so a month ceiling added to a budget that capped only the day finds the month’s spend already there.
- A refused charge is counted too. A round is charged once its reply has arrived, because that is when its size is known, so by the time it is judged the vendor has billed it: what a refusal stops is the round’s tool calls and every round after it, not the spend. The round goes on the seat’s counter and the company’s, whichever of the two refused, and the refusing window reads past its ceiling by the round that crossed it. A counter that dropped the refused round read short of what had been paid for, and its room was a lie: at 98 of 100, a 3-token round refused and then a 1-token round admitted on top of the 3. Counted, every later charge is refused until the window turns over or its ceiling is raised.
- A refusal made without a charge is recorded too. Once an answer has shown a window with no room left for a single token, every later charge in it is certain to be refused, so the engine turns work away before its first call rather than pay for a call to be told so: a turn’s meter stops the turn’s next call (a round, the round-cap judge, a rewrite), the budget park defers a seat’s delivery, a person’s
answer_knowledgequestion is refused, and the reflection stage declines a pass and a conversation entry’s rewrites. The call each one stops is the one whose charge would have stamped the window, so each records the refusal itself (coord.Budgets.Refuse): one compare-and-swap on the scope a charge would be refused by, the company’s before the seat’s, stamping every capped window that has no room left for a single token — judged against the counter as it stands, never against what the caller remembers, so a window that turned over or that the caller’s ceilings leave room in is not stamped. It counts nothing and writes nothing where it would stamp nothing. A park, a refused question and a declined pass are one refusal each; a turn’s meter, asked before every call, records once per window for each part of the turn — a turn resumed from a coding run records again in a window still full, since its first held call is a refusal of its own. - Nothing a charge counted is taken back. A seat write that fails after the company’s landed leaves the round on the company, whose record of it is true, and the charge answers an error naming the seat’s missing share (
coord_kv_budget_spend_uncountedin the node’s log), so the round stops as an outage. A caller that merely hangs up between the two writes leaves no such partial: once the company’s write has landed, the seat’s is finished on a context that outlives the caller’s, because the round is spent whatever the caller does next. Taking the company’s half back used to leave a billed round on neither counter — nothing that charges asks again — and the next round was admitted against room it had used. What the partial costs is the seat: its own counter is short of that one round until the window turns over, so its own ceiling can trip a round late. The company, which every seat is judged against, stays exact. A post-charge — spend recorded after the fact with no verdict: a collected coding run, an auxiliary call — leaves the same partial the same way, and the one caller that offers such spend again (a collected run whose resume failed) records the seat’s share alone the second time, so the company is not counted twice. - The roll is the reset. A slot still labelled with an earlier window is moved onto the current one — spend and refusal cleared — inside the same write that counts the charge. Nothing is scheduled and nothing runs at midnight: the allowance comes back on whichever node charges the scope next, and two nodes racing to roll one slot write the same label. There is no reset command and no reset route; room before a window turns over is made by raising its ceiling.
- A slot never rolls back. A node whose clock trails a peer’s across a boundary finds the slot already on the next window and counts there, rather than handing the new window its allowance back. A read behind such a slot answers the same way: it states the later window, its spend and its refusal, never the earlier window unspent — every read of the counter (a turn’s headroom, the learning gate,
GET /budgets, the live meter) sees exactly what the next charge would be judged against. That lasts a few seconds behind a peer’s clock, and up to a day after the company’stimezonemoves west. - The calendar is the caller’s. Every charge and every read carries the windows it is about; the store holds labels and never reads a clock. A turn charges each round in the windows current when the round is charged, on the clock of the epoch the turn is pinned to; a detached coding run is counted in the windows it is collected in.
- A contended counter waits before it retries. The company’s record is written by every round of every seat, so a charge that loses the compare-and-swap waits a jittered 1 ms, doubling to 32 ms, before trying again — about a third of a second across its sixteen attempts before the round fails closed. Retried at once, 32 concurrent charges on one record ran out of attempts in four runs of ten.
What a node says about itself
Section titled “What a node says about itself”Every node’s presence lease is renewed on its heartbeat, and each renewal carries the node’s status beside its roles and labels: turns in flight, whether it is draining, its config posture, when it started, how far its replicated state has come up — and how its MCP servers started:
mcp— one row per configured MCP server: whether it is shared, how many of its instances started and how many did not, how many tools one serves, and one failure’s reason (clipped to 240 bytes, with the seat it belonged to). One row per server, not per child, because a per-role template has a child for every seat the node holds and the status is re-sent on every beat. A status with no rows is a node that started none. A child that dies after starting is not observed here; its next call fails and says so.
Freshness is the heartbeat interval, the same as every other column of the fleet view. A node whose status hook overruns its share of the beat publishes no status for that beat, and a reader treats it as “did not say”, never as zero. A key in the status this build does not know is ignored rather than refused, so a successor sharing the fleet can publish what it adds and this build still reads the rest of that node’s status.
What stays node-local
Section titled “What stays node-local”The node’s own database holds everything a single node is the only reader of. The test is not “is it durable” — all of it is — but “would a peer reading this change any answer?”
- The event log. It records what this node saw; a peer’s copy would claim this node had seen it too.
- The diary, episodes, counterparty profiles, synthesized skills — as a CACHE, not as the only copy. A seat’s memory is read by the node running that seat, and that node changes: placement moves seats, so the honest test is not “would a peer read this” but “would the seat’s next owner need it”, and the answer is yes. Every memory row is therefore also published to a compacted changelog on the stream, and a node hydrates a seat’s rows into its own store before the seat takes work — see seat memory. The local copy stays local because recall computes vector distance in the database, which is the one thing a KV cannot do.
- Conversation history. Read only by the seat’s owner — but carried on the same changelog as the rest of its memory, because ownership moves and a seat that forgets what it already said repeats itself in the thread.
- The company payload. Bulk that every node holds its own copy of. Only which revision is current is shared — see Control Plane.
- The secret store’s bootstrap half, and only that. The company’s credentials are a shared slot (
secrets, above); what stays in a node’s own file is the rowscrewlet secrets setwrites against a stopped node, which that node migrates onto the fleet at its next start and deletes locally. The keyring that opens either is Tier A on disk, never a shared record — see Secret Store § Propagation. scheduled_runs— this node’s dispatch history, for the dashboard and the retention sweep. Not the claim; that is thefiresslot above.
And two things stay per-process deliberately:
max_concurrent. Tier A’snode.max_concurrent(default 32) is the gate every agent turn passes through, and it is per node — so an org’s ceiling is N × the configured value. Size it per node, not per company. This is the one knob a fleet genuinely changes the meaning of.- A seat’s MCP subprocesses. They are children of the node that claimed the seat, and they die with the release. Only their status is shared, on the heartbeat (above).
Backends
Section titled “Backends”| Topology | Coordination store | When |
|---|---|---|
| Embedded (default) | The engine’s own in-process NATS JetStream KV | One node, or a fleet whose embedded servers cluster with each other |
| External NATS | The same KV, on a cluster this node dials | A fleet that wants the stream and everything riding it to outlive any single engine |
| Memory | An in-process twin | Tests |
Coordination is never an estate of its own, and that is a construction rule rather than a convenience: the coordination store rides the stream’s own NATS connection. The engine starts or dials exactly one broker, and the leases and the shared records above are opened on the connection the queue is already using. A second dial would work and would be worse — two connections to one broker fail independently, so a node could go on renewing leases over the one that still works while the one carrying its inbox has dropped. Alive to its peers, deaf to its work, and holding every seat it owns while doing none of it.
The twin is not a lesser implementation: it is held to the same certified suite as the real backends (internal/coord/coordtest), because a twin that agrees only with itself proves nothing.
The leases follow coordination.type. The shared state does not.
Section titled “The leases follow coordination.type. The shared state does not.”coordination.type: local is about leases, and only about leases. A single node has no peer to fence against and re-claims every seat at boot, so an in-process lease table is the honest implementation there — which is exactly what that setting selects.
Everything else above is a record, and a record has to outlive the process, on one node as much as on four. So the shared slots always live in the KV, whatever the coordination slot says, and what persistence they get is the same choice as the event log’s: stream.store_dir.
That distinction was not always drawn, and each consequence was silent. The token counter — which then had no retention at all, because a cap was a ceiling for the life of a deployment — went back to zero on every restart of a default single-node engine. A turn completion no longer suppressed the redelivery it exists to suppress. A detached sandbox run, which is a billed box, was forgotten by the engine that launched it. Leaving stream.store_dir empty still selects an in-memory server and has all of those effects, but it says so on the tin: it is the same switch that makes the event log itself disposable.
On the embedded backend the coordination store lives inside the running engine. It exists while the engine runs. That is the correct trade for a single node — nothing else to install — but it has two visible consequences.
An offline crewlet config import — one run while the engine is stopped — cannot move the activation pointer, because there is nothing running to move it in. It marks the revision active in this node’s own database and says so; a node that starts holding an active revision the fleet has no pointer for publishes it at boot, so a restart converges without any operator action. Against a running node the same command takes the other route entirely: it detects the held store and goes through that node’s PUT /config, which moves the pointer immediately.
And the operator commands that read or act on this state talk to a running node rather than to a file: crewlet budgets show, crewlet backup and crewlet secrets are clients of that node’s API. Opening the store from outside would either find nothing (the engine is down, and an embedded broker exists only while it runs) or corrupt it (the engine is up, and a second broker on the same store directory is accepted rather than refused).
See also
Section titled “See also”- Seat Ownership — leases, the fencing epoch, singleton duties and the watchdog
- Control Plane — the activation pointer and per-node apply status
- Scaling Out — the five kinds of coupling a fleet had to resolve, and which one a lock actually fixes
- Deployment — running more than one node
- Event System — the queue this sits beside, and what it is not
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.