Skip to content
You are reading documentation for unreleased main. This page is not in 0.1 yet.

Seat Ownership

A seat is a role in the org chart, addressed by its handle. Seat ownership is how a fleet of Crewlet nodes decides which node runs which seat — and, more importantly, how it guarantees that no two of them run the same one.

The rule the whole design serves is one sentence: a seat is not a thing you can half-own. A node either holds a seat’s lease, runs its agent, consumes its inbox and answers its sandbox completions, or it does none of those things.


Every seat has a durable inbox topic, crewlet.agent.{handle}.inbox, consumed under a durable subscription named agent-{handle}. A subscription is a competing-consumer group: each message goes to exactly one attached member.

That is exactly right when the members are one node’s consumer. It is catastrophic when two nodes both attach: the broker splits the seat’s traffic between them, so one agent’s conversation runs as two interleaved turn streams on two processes, each unaware of the other. Turn exclusion is in-process state; neither node can see the collision, and nothing raises.

So attachment has to be exclusive, and exclusivity has to be provable across processes. That is what a lease is for.


Ownership is a lease in the fleet’s coordination store (internal/coord): a record with a TTL and a monotonic epoch. On a fleet (coordination.type: embedded-kv) the record lives in the crewlet_leases KV bucket, whose age limit is the lease TTL, and the epoch counter in the untimed crewlet_epochs bucket; a single node runs the in-memory twin of the same contract.

seat:{handle} owner=node-a:9f3c1e70 epoch=7 acquired_at=… expires_at=… preferred=node-a

Three properties carry everything above it:

  • The owner is a process incarnation, not a machine. {node_id}:{random}, minted fresh at boot. A live lease is renewable by its own owner string, so two processes sharing an identity would both hold the seat at the same epoch — and the default node id is the shared constant node-0. The stable node id goes in preferred, where restart-stability is what you actually want.
  • The epoch is a fencing token, monotonic for the resource’s lifetime. It is kept apart from the lease record, in a bucket with no age limit, because a KV deletes a key when it expires: a counter stored on the lease would restart at 1 and hand the next owner a token its predecessor is still using.
  • A lapsed lease cannot be renewed, only re-acquired — and re-acquiring bumps the epoch even for the same owner, because during the gap that owner’s in-flight work was unprotected and must be fenced against its own past self.
  • A claim that fails part way leaves nothing held. On a fleet a new tenure is three writes — the lease record in a claiming state, the counter, then the token into the record — and a claim that fails after the first (the counter’s bucket between leaders, a write whose answer was lost, a caller that gave up) gives its record straight back, so its owner or a peer can claim at once. It used to leave the record claiming until its TTL, which every node read as held and its own owner read as a sibling claim about to commit, so the seat or duty sat dark for a lease’s length. The give-back takes only the record that call wrote, never a sibling claim’s or a tenure that did commit, and a counter the failed claim advanced is a harmless gap. Only a claim that dies, or whose store refuses the give-back as well, leaves its record to lapse on the TTL, as a dead owner’s lease does.

Beside them the record remembers when its tenure began. acquired_at is the coordination store’s own timestamp on the write that won the lease, stamped when the epoch is minted and carried unchanged by every renewal — so it moves exactly when the epoch does, and “node-2 · since 08:02” in the Since column of Settings › Nodes means the seat has not moved since 08:02. A heartbeat restamping it would report every seat as “since a few seconds ago”. It is on the store’s clock for the reason expires_at is, and nothing decides ownership by it: it is a fact for a person, and the fence is still the epoch. Every lease carries one, and a renewal never restamps it — the moment it would stamp is the renewal’s, not the claim’s.

Placement is deliberately dumb, and lives in internal/seat:

  • Every node holds a node:{id} presence lease, renewed on the same heartbeat as its seats. Counting the live ones is how a node learns the fleet size. It cannot be inferred from seat ownership: a fleet where nobody has claimed anything yet would read as zero nodes, and every node would then take every seat.
  • A node claims up to ceil(seats / live nodes), its fair share, and never more than seat.ClaimLimitPerSweep (4) per pass, because each takeover costs an MCP spawn. It tries the placement groups with the fewest matching nodes first, and within a group the seats whose preferred hint names it, for stickiness.
  • With role.placement in play the share is computed per placement group — the seats that share one placement — as ceil(group size / live nodes placing seats that match it), and each share bounds its own group only. A node holds at most its share of each group it matches; room in one group is never room in another. A satellite labelled for one pinned seat that could spend its total on any seat it matches would fill up on unpinned seats and leave the pinned seat, which nobody else may run, unserved — with nothing reported, because the fleet’s totals still covered it. Trying the groups with the fewest matching nodes first means a claim-limited pass spends itself on the seats with the fewest other homes.
  • A node holding more than its share of a group hands that group’s excess back, at most seat.ReleaseLimitPerSweep (2) per pass — even when its total is within capacity, because the surplus is room an eligible peer is waiting for. Claiming alone converges only for a fleet that shrinks: a node that booted alone holds every seat, and a peer joining later computes a share it can never reach. Without the give-back, scaling out does nothing until something dies. It gives back the least constrained groups first, so when the release limit cuts a pass short a pinned seat is the last to go.
  • A pass runs every seat.SweepInterval (5 s), which is the cadence for what changes in the fleet, and at once when a config apply changes the seats themselves: a role added or removed, a seat’s placement moved, or a node’s first company. Those would otherwise wait for the next tick, a new seat run by nobody until then. The pass runs at the end of the apply, once everything a seat’s acquisition reads is current, and at most one such pass per interval comes early, so a node never runs more than two passes’ claims in one interval however many revisions arrive.
  • A draining node claims nothing — and that holds for a pass already under way when the drain began. Each claim checks again, under the seat’s own lock, whether the drain has begun, and a lease taken while it did is handed straight back before anything is attached to it, so a pass cannot take back the seat the drain has just given up and leave it on a node that is leaving.
  • A seat whose teardown could not be proven (below) keeps its lease and counts against capacity wherever it sits, even in no group because the node may no longer run its role. It is charged first, and what capacity is left is handed to the node’s groups most constrained first, each up to its share — one number per group, which claiming and giving back both read. There is no total kept beside it: a node holding an unpinned seat and a stuck one would read “full” against a total while its pinned group had room, claim nothing, and shed nothing because no group was over. Charged first, the stuck seat squeezes the unpinned group instead, the unpinned seat goes to a peer, and the pinned seat is claimed.
  • A node that cannot serve its seats’ work at all gives every seat back (unserviceable, below) and marks its presence row withdrawn (seats_withdrawn) until that clears. Its lease stays live, so without the mark its peers would go on counting it in every share and leave its part of each group free for a node that will never claim it.

The share is a ceiling, so a group’s shares over the nodes that match it sum to at least its size, and a node at its share of a group has no room in it to re-claim what it just released. Rebalancing converges rather than oscillating, and every group some live node placing seats matches — running seats and not withdrawn — is covered by those nodes’ shares. A group no such node matches is the only one the shares cannot cover, and it is what seats_unplaceable reports, read off the same arithmetic the claims are bounded by. Two states leave a share unclaimed for a while without that warning, each with its own signal: a node that is not ready to claim yet (behind, below) still counts, because it keeps serving what it holds and takes its share once it is level; and a node whose teardown keeps failing holds its share down by the stuck seats, which it alarms on until the teardown succeeds or the process restarts.

node-bleases (coordination)node-anode-bleases (coordination)node-aalone — share is 3share is now 2spawn, budget, MCP, sandbox recovery,THEN attach the inboxacquire node:node-aacquire seat:ceo, seat:eng, seat:opsacquire node:node-bListLive(ClassNode) → 2release seat:ceo (voluntary)expire seat:ceo in placeacquire seat:ceo → epoch 2

The acquire hook (node.Node.OnAcquire, preparing the seat through Engine.prepareSeat) establishes the seat in a known state and attaches the inbox consumer last: the per-role MCP children, the seat’s memory hydrated from the changelog, the seat’s reflection subscription, the sandbox control subscription and the interrupted sandbox-run recovery, then the inbox. A seat that starts receiving work before its MCP children are up runs its first turn with an empty tool surface. The sandbox half has one more way in: on a node whose code sandbox arrives with an apply, after it already holds seats, that apply gives every held seat its control subscription and run recovery then, under the same per-seat lock an acquisition and a release take, so neither is raced; a seat that cannot be prepared is handed back voluntarily (unprepared), and its next acquisition runs the whole hook. The release hook is the mirror: the seat’s children die with its lease, because the credentials in one are that seat’s identity and a child left running would let this node keep acting as an agent a peer now serves. See Tools & MCP.

Establishing is an edge, and it is reported. The seat admits no turn while its acquire hook runs, and something the hook starts can be refused for exactly that reason: the sandbox recovery retries the resume of an answer the seat’s last holder recorded and never resumed with, and that retry is told the seat is not yet established. Nothing used to say when that lifted — an establishment is not a renew’s admission edge — so the retry slept until the clock said a heartbeat had passed, 15 s at the default TTL. The host now reports the moment the hook returns and the seat starts admitting turns (seat.Hooks.OnEstablished), the node passes it on together with the renew that proves ownership again after a store blip (node.Config.SeatAdmitted), and the engine re-checks the answers waiting on ownership at once.

A seat a person paused is attached already held. A pause holds the seat’s inbox under a named hold, and a release drops every hold with the attachment — so the node that acquires a paused seat (placement moved it, or its holder restarted) takes the pause hold before it attaches the inbox (node.Config.AttachHolds). A hold taken after the attach would be taken after the first delivery can arrive, and the mail the pause was holding would be the first thing the new holder ran. See Agent Runtime § Pausing a seat.

Releasing has two modes, because losing a lease and choosing to let go are opposites:

ModeWhenWhat happens
Voluntarydrain, capacity rebalance, role decommissioned, placement moved, this node unable to route its seats’ calls to the estate (unserviceable), a held seat that could not be prepared for a code sandbox an apply brought up (unprepared)quiesce → let the in-flight handler finish under a bounded wait → detach → release the lease
Fencedrenew returned false, the TTL grace expired, an acquire hook failed, config posture went shed/stuckdetach first, abandon in-flight work, republish nothing

Fenced release never republishes. A peer may already be running the seat, and a republished event is a new message: a second copy of work the successor is already doing, carrying none of the identity the completion ledger’s idempotency and the batch layer’s aging both key on — so nothing downstream can collapse the two. Handing the delivery back unacked keeps that identity, and the successor gets exactly what this node never finished.

A teardown that cannot be proven does not release the lease. A lease held too long costs latency; one released too early costs correctness. So a seat whose release hook fails (on_release returns an error or panics) goes undead: out of the held set, so this node starts nothing new on it, and still renewed, so no peer can take a seat this process may still be consuming.

Undead is a state, not a grave. The teardown is retried on every heartbeat, and the lease is released the instant one succeeds — the usual causes are transient (a consumer mid-delivery, an MCP child that has not finished dying), and the retry is what returns the seat to the fleet. A retry that keeps failing keeps the seat, and re-raises its alarm every twenty heartbeats with the elapsed time, because the failure itself is not news but still failing is.

Only a restart of that process can free a seat whose teardown never succeeds — its leases lapse at the TTL and peers pick them up. That is an operator’s call rather than an automatic one: it also moves every healthy seat on the node, which is the wrong trade for one stuck MCP child and the right one for a process that has stopped being able to close anything.

A handler has two ordinary outcomes: queue.Ack or queue.Nak (which asks for the message back and goes on consuming). Seat handoff needs a third, so the queue protocol has one:

return queue.Defer(fmt.Sprintf("seat %q is not owned here", handle))

The delivery goes back to the broker at once — a NAK is how a JetStream client returns one, and the seat’s next owner sees it in about a millisecond, where letting the ack window lapse instead would park that seat’s mail for thirty minutes on every lease movement — and the attachment then stops consuming. The second half is what a bare NAK does not have. A node that handed the message back and kept fetching would be handed it straight back, refuse it again, and spend one of the message’s twenty-five deliveries on every lap, until an event nothing is wrong with dead-letters on a node that was never entitled to it. One handoff costs one delivery, which is what that budget is sized for.

Three paths use it, and they are the three ways this node can be the wrong one to run a delivery it was just handed: the seat is not owned here, the in-turn fence tripped mid-dispatch, or the config posture went shed/stuck. Each also records the deferral, because a deferral quiesces the consumer and the resume is edge-triggered on the next successful renew — without it the seat is owned, attached and deaf.

“Do I hold this seat?” is a question about a local snapshot refreshed on a 15-second heartbeat against a 45-second TTL, so the honest answer can be a full TTL stale — precisely the window an ownership check exists to close. A membership check cannot meet its own exit criterion.

What is provable is that a successful renew at time t bought exclusivity through t + ttl. So seat.Host.MayStart returns the epoch only when the last successful renew is inside one heartbeat interval, and reports false otherwise (seat_admission_stale). Every turn that starts is then certified owned for at least ttl - heartbeat.

That also gives the right answer during a database blip. The lease row is untouched by an unreachable store, so the seat is kept — shedding on a two-second outage would tear a healthy company down — but new turns stop at the first failed renew. The consumer is quiesced, and un-quiesced when a renew succeeds again. Both edges matter: without the second one the node comes back healthy, still owning the seat, still attached to it, and never reads from it again.

A copy that is behind, and a copy that is wrong

Section titled “A copy that is behind, and a copy that is wrong”

A data node answers its seats out of its own copy of the company’s records — the tracker and the knowledge base, applied from the shared log into this node’s database (see Replication) — through the same router every node’s seats use. That copy can be in two quite different bad states, and the engine treats them as opposites.

What it meansWhat the node does
BehindRecords are on the log that this node has not applied yet. It is catching up, and it will.Keeps every seat it holds, and claims no new ones until it is level. The sweep logs seat_claims_withheld at debug, and the fleet view counts how many of this node’s replication loops are current
WrongThe copy cannot become current by applying more recordsTakes its copy out of service and keeps every seat. Its seats’ calls go to the other data nodes, exactly as a node holding no data is served, and other nodes asking it are told its copy is out_of_service. The node logs estate_copy_not_served when the copy goes out of service, and serves again once a reading finds the copy sound

Six states are wrong, and each is a fact about the rows rather than about how far along they are:

  • the applier has stopped at a record it cannot apply;
  • the node has been evicted from the fleet, so its peers drop everything it writes;
  • its rows are below the log — records it never applied have been deleted, and the hole will never fill. A floor that has been unreadable for four heartbeats counts here too: an unread floor is not a floor that is satisfied. A node that is only below the published trim floor, while the log still holds every record it lacks, is not here: it is replaying them, which is behind;
  • its checkpoint names a stream that is not this one, which is a log deleted and rebuilt underneath it (crewlet retention reanchor is the repair) — or a history that log no longer continues from, because a peer re-anchored it past this node’s generation (the node repairs that itself, by adopting a snapshot from the new generation);
  • its applied position has stopped moving for a minute while records wait — a stall, which is the one of these that looks like lag and is not: the node owes progress and is not making it;
  • it has held a record it cannot decode for longer than the deferral grace, which means it is running a build that cannot read what its peers write.

Lag is never one of them, at any size. A node a million records behind keeps its seats and works through the backlog; a node one record behind — which is every node for the moment between a write landing on the log and its own applier reaching it — is not in a different state from one that is level. The distinction is what the apply_lag alarm says: being behind is worth looking at, and it is not what moves work.

Why not give the seats back, as a node whose copy was wrong once did: the seats were never the problem, the copy was, and every call they make is routed (how a request reaches the estate) — so a peer that took them over would be served by the very same data nodes, at the cost of every seat’s processes and memory moving. On a single node there is no other holder, and the company stops working until the state clears: every call is refused naming the estate nobody serves, which is the right trade for rows that are wrong — the alternative is agents acting on them.

What a node does give its seats back for is not being able to route at all: its seats’ calls need another data node — it holds no data, or its own copy is out of service — while its view of who serves the estate has been unreadable for longer than the 60-second bound every cached coordination fact is held to (unserviceable, seats_shed_unserviceable). A data node whose copy is sound answers from it without the view, and keeps its seats whatever the view says. The release is voluntary, like a rebalance: the in-flight turn finishes, and the seat leaves when it goes idle. While it lasts the node is withdrawn from placement and says so on its presence row, so its peers divide the seats without it and take them up, and a seat only it may run is reported unplaceable rather than left waiting.

Fencing: what it protects, and what it cannot

Section titled “Fencing: what it protects, and what it cannot”

The epoch is threaded into the sandbox run state: every mutation on a live run record is refused when the record’s owner_epoch outranks the writer’s. Beside it, the seat fence (seat.Host.Fence) is checked in the turn loop at the top of every round and again before each of that round’s tool calls. A zombie’s late write to a run it no longer owns bounces; a zombie’s turn stops before its next tool call.

The fence exists because admission and the work it admits are minutes apart. MayStart proves the seat was held when the turn started; the turn then calls models, fires tools, posts to chat and writes to the tracker long after that, and the queue contract is explicit that nothing else closes the gap — a detach “does NOT wait for a running handler”. The commonest trigger is not even a lost lease: the placement sweep hands a seat back voluntarily when the fleet grows, a peer claims it within a sweep interval, and without the fence the previous owner goes on being that seat until its turn happens to end.

It closes on the epoch, not on membership, because a seat can be lost and re-claimed here within one heartbeat and the re-claim is a different grant. Three things deliberately do not close it:

Why not
An unreachable coordination storeSays nothing about ownership, so the heartbeat keeps the seat. Closing here would tear a healthy company’s turns down over a two-second blip during which no peer could have claimed anything — the three-valued rule paying for itself
A seat whose teardown could not be provenThe lease is kept and renewed precisely so no peer can take it, so the grant has not moved and the in-flight turn is racing nobody
A detached run — a coding CLI in agent mode, or a sandbox jobIt outlives its turn by design; its placement is on its own row and the process that collects it is often not the one that launched it. What is fenced is the resume, under the grant the collecting node holds

The same fence carries a person’s stop. A pause that asked for the running turn to stop closes it at the same round boundary with turn.ErrStoppedByPerson — ownership is checked first, since a turn on a seat that moved is the successor’s to stop or not. That turn is not deferred: the trigger is recorded in the completion ledger and acked, because the person stopped it so that it would not go on. See Agent Runtime § Pausing a seat.

A turn the fence stops because the seat moved is deferred, not failed: the delivery is healthy and the seat’s new owner is entitled to it, so it goes back the way the three screening paths above send one — immediately, and with this attachment quiesced. That is what a deferral buys here; it is not cheaper. One handoff costs one delivery whether the turn never started or was a phase in — on every backend, the in-memory twin included — which is why the budget is sized for handoffs as well as failures (see Event System). The exceptions are a turn that panicked and a turn whose own record proves it already wrote outside the engine — each is acked and recorded instead, because a successor would reach the same defect or repeat the write.

It is not on every seat-scoped write, and the honest inventory is narrower than “the learning tables are unfenced”. What a duplicate write actually does, per table:

TableA second writer today
episodesCollapsed against the reader that matters. One row per unit of work in the node’s own store, which is the only one its recall reads — see Keying a write on the work below
counterparty_profiles.interaction_countCollapsed. The increment is skipped when the last counted work key repeats
agent_onboarding_markersUpsert plus learning.Onboarding.Claim, a cross-process single-flight claim: already exclusive
agent_diaryByte-identical content collapses on write. Two turns that word the same fact differently still land twice

The last one is deliberate. Nothing can key a differently worded diary entry to its twin — that needs the duplicate turn not to happen, which is the completion ledger’s job, not a write guard’s.

The instinct is to fence these the way a sandbox run is fenced, on owner_epoch. That works for a mutation of an existing row and fails here twice over: an insert has no prior row to hang the condition on, and a fence loses data in the case where nothing went wrong — a node that completes a turn, acks the delivery and only then lapses would have its episode refused. The turn happened; the memory of it is gone.

So the write is keyed on the work rather than fenced on the writer. Every turn dispatched from a ledgerable trigger carries a work_key derived from its constituent event ids — the same identity the completion ledger uses, and the one thing that is stable across a re-run (a turn_id names ONE RUN, so anything keyed on it records the duplicate instead of collapsing it — and a re-run is ordinary rather than exceptional: two nodes mint two runs for one trigger, and a turn that breaks before reaching outside the engine is redelivered and runs again. See a turn’s two identities).

epoch fencework key
Zombie and owner both complete, same nodeone rowone row
Owner completes, acks, then lapsesrow lostone row
Ledger fails open, turn legitimately re-runstwo rowsone row

Exclusion is a unique index on (agent_handle, work_key) plus INSERT … ON CONFLICT DO NOTHING — one statement, no read-then-write, so two writers racing inside one process cannot both see “not there” and both insert. It is per agent rather than global: two seats legitimately act on one trigger (a broadcast, a task assigned to a unit) and each one’s episode is its own memory. The column is nullable and an empty key maps to SQL NULL, which the index treats as distinct from every other NULL — so an unkeyed turn is never deduped against another, which is the whole reason it is nullable rather than defaulting to ''.

Placement moves seats — a node joining or leaving, a drain, a rolling upgrade — and a seat’s memory is written to the node’s own database. So the memory has to move too, or a seat that changes node reads a store that has never heard of it and runs its next turn having forgotten everything it learned elsewhere.

Every memory row is published to a compacted changelog on the stream: one subject per row, on a stream that retains exactly one message per subject, so it holds the current value of every row rather than a log of every write. A node acquiring a seat replays that seat’s subjects into its own store before the mailbox attaches, and a hydration that fails refuses the seat — a peer that can hydrate should take it instead, and a seat serving with amnesia produces work its own history contradicts.

What travels: the diary, episodes, counterparty profiles, synthesized skills and their versions, onboarding markers, and the conversation ledger. What does not: this node’s own audit log and dispatch history, which are records of what this node did rather than of what the seat knows.

A row that arrives for one this node already has is resolved by what kind of row it is. An append-only row — a diary entry, an episode, a skill version, a ledger entry — is immutable once written, so the carried copy is discarded and the local row stands, with one exception: a diary note’s or an episode’s vector, which the holder fills after the insert for a row written without one or under a model the company has since left. Setting the vector stamps the row with a fresh change sequence, so the changelog carries it like a new row, and a node already holding the row takes the carried vector over a stored one of no model, another model or another width — never the reverse, so a stale copy replayed later cannot erase a filled vector. A small mutable row — a counterparty profile, a skill’s state and use count, an onboarding marker — is one whose update is the content, so the carried copy wins. That is safe rather than merely convenient: a seat is held by one node at a time, so the changelog’s latest value for a subject is by construction the value its current owner wrote.

What is new is found by that change sequence: one counter per node, which every inserted diary note, episode and skill version — and every vector set on a note or an episode — is stamped from, and which only ever increases. Each cycle carries a seat’s rows past the highest value the last cycle published. It is not the row’s rowid, because these tables are keyed on text and a rowid is handed out again once the newest row is deleted: after the diary’s expiry or its trim took the newest note, the next note took the same rowid, below what had already been published, and was never carried while the node held the seat.

A row can also collide with one this node already has under a different name — two episodes for one work key, written under two ids on two nodes. That is skipped rather than returned as an error, and the distinction matters more than one row: hydration runs inside seat acquisition, so an error refuses the seat, and a single duplicated episode would make a seat unplaceable across the whole fleet.

A read is answered by the holder. Every node that ever held a seat keeps a copy of its memory, and only the node holding it now keeps that copy current. So a screen reading a seat’s memory or its conversation ledger (agent_memory, conversations) is never answered from whichever node served the request: that node reads the seat’s lease and answers from its own store if it is the holder and has the seat attached, asks the holding incarnation on an ephemeral scatter if a peer is, and answers empty — naming no holder — if no node holds the seat, because a copy of unknown age is not the seat’s memory. A holder that is silent, or still hydrating the seat, is an unavailable answer rather than an empty one — a holder that is draining or whose heartbeat lapsed is asked like any other, and its silence is that same answer after the two-second budget. Every node answers for the seats it holds, including a node that serves no API. The list of every agent’s memory (memory_overview) follows the same rule in one round rather than one per seat: the asker lists every seat lease once, reads its own seats from its store, and puts one request on the broker naming each holder’s seats — so a fifty-agent company is one scatter, not fifty — and a holder that did not answer is named in the answer’s coverage while its seats say so rather than reading as empty. See internal/learning/memread.

Deletes are deliberately not replicated. The learning lifecycle drops rows constantly, and carrying a tombstone for each would double the protocol to keep a table converged that already converges itself: a hydrated node may resurrect rows its predecessor had swept, and the next lifecycle pass drops them again by the same rules. A crash loses at most one cycle of memory (30 seconds); a graceful handoff flushes on release and loses none.

What the seat learns is learned where it is held. Post-turn reflection is the largest writer of a seat’s memory, so it runs on the seat’s holder rather than on whichever node a fleet-wide group handed it to — there, its rows would land in a store the seat’s holder never reads and no flush ever carries. The node that ran a turn puts a reflection_due wake on the seat’s own reflection subject, crewlet.agent.{handle}.reflect; the acquire hook attaches that subject after hydration, so a pass reads what the seat already knows, and the release hook lets it go before the last flush, so nothing new writes the seat’s memory there after it. A wake published while the seat is between holders waits on the subject for the next one. A pass still running when the release lets the subject go is not waited for — the release is fenced and an auxiliary model call is not — so what it writes after the flush is the same bounded loss as a crash.

The index is the node’s, like the table it is on, and that is the honest scope of the guarantee. episodes is a seat’s memory and lives in the node’s own database — read by the node running that seat, never by a peer. So a duplicate written on two different nodes is two rows in two databases, each under its own id — and memory replication carries both onto the changelog rather than collapsing them there. What collapses them is this index, on the node that next replays the seat: it imports the first and skips the second. The guarantee is therefore eventual and reader-scoped rather than fleet-wide at write time, which is the honest shape of it. What the index collapses is the case that actually recurs against one reader: one node writing twice — a redelivery it worked again, or a legitimate re-run after the completion ledger failed open.

A turn with no ledgerable trigger — a scheduled fire, a worker, a sandbox resume — carries an empty key, skips the guard entirely, and writes exactly as it always did. It has no cross-node duplicate to collapse.

Fencing protects database state. It cannot protect outbound effects. run_sandbox makes this concrete: it acquires a real, billed sandbox before the pending row is written, so no epoch-fenced insert can undo a box that is already pushing commits. The property the design offers is bounded duplication, not none — and what bounds it is the in-turn fence.

Fencing and admission bound the window in which two nodes could be working one seat. They do not close a narrower one: a turn finishes, its outbound effects ship, and the node dies before the delivery is acked. At-least-once then hands that trigger to the seat’s next owner, which re-runs the whole turn — the same Slack reply, the same Jira comment, from an agent with no idea it already spoke.

The completion ledger records what finished, and is read before the next turn starts. It lives in the fleet’s coordination store, not in a node’s own database — a redelivery that lands on a peer has to find it, and a ledger each node kept privately would find nothing and run the turn again, which is the exact failure it exists to prevent.

It is deliberately not a claim, and the absences are the design:

  • No in_progress state. The seat lease is already the mutual exclusion — one consuming node, serial within it — so a claim’s only honest disposition for a stale in-progress row is “supersede and re-run”, which is exactly what you do with no row at all. An earlier design had one; five of five reviewers rejected it, because every other defect they found existed only to service that state.
  • No expiry, no supersede rule. A record means the work is done, and done does not lapse. Records age out on the bucket’s own seven-day retention — garbage collection, not semantics, and its floor is the scheduler’s catchup ceiling rather than a round number: forgetting a completion a tick could still evaluate lets that fire run twice.
  • Keyed on constituent event ids. A multi-event partition is merged into one digest before the turn runs, and that digest is minted fresh on every coalesce, so a key taken from it would differ on every redelivery and match nothing. Recording constituents also means a redelivery that overlaps a previous partition only partially — A+B ran, then A+B+C arrives — skips A and B and runs C.
  • And on the chat messages a later turn answered without being woken for them. A turn that fails before anything leaves the engine is NAK’d, and a NAK’d message comes back behind its conversation’s newer mail — so the next message in the thread runs first, its turn reads the thread and sees the failed one still waiting, and it answers the thread as it stands. That turn reports the waiting messages its thread block showed it (somebody else’s, after the seat’s own last reply, before the message that woke it — never one a bound dropped), and they are recorded beside its own triggers under a second key: the message’s identity on its chat backend. A trigger is dropped if either key is recorded, so the failed message’s retry is skipped — turn_trigger_skipped says a later turn answered it in its thread — rather than answered a second time, out of order. Nothing about this is the node’s: the redelivery can land wherever the seat has moved, which is why it is the ledger’s fact and not a node’s memory.
  • And on a parked coding run’s answer. A chat reply recorded as the answer to a parked run’s question is recorded here too, so a copy of it — a redelivery whose acknowledgement was lost — is never run as a turn the run has already been resumed with. So is an answer given by turn (sandbox_answer_given), once its route has settled it and again when its run spends it, and that route reads the ledger before it reaches the run: a copy that comes round after the run used the answer, let it go or ended is acknowledged rather than announced a second time.

Both directions fail open, and that is the whole failure policy. An unreadable ledger cannot tell you whether work was done, and the only safe answer to that is the one the engine gave before the table existed: run it. Failing closed would park real work during a database blip — and the seat’s own admission gate already refuses new turns within one heartbeat of a store it cannot reach. The write happens after the side effects shipped, so failing to record them cannot un-ship them.

Two notes on coverage:

  • Only trigger types that run a turn are consulted (inbox.Ledgered: task_assigned, external_notification, and the two A2A wakes a2a_request and a2a_message) — besides an answer by turn, which runs no turn and whose own route consults it. Everything else that reaches a seat’s inbox, an observability event or a wake already answered, is logged and dropped, so recording it would be bookkeeping about nothing. The set is closed: a type outside it contributes nothing to the work key and is never recorded, so its redeliveries re-run, which is why a new trigger type is added there rather than merely published.
  • A suspended sandbox turn IS recorded, at the suspend. Past that point the pending run’s own at-most-once flip is the authority for the rest of the work, and the trigger itself is finished with.

A2A was exempt while its content rode a process-local queue that the old A2A inbox handler drained destructively: a re-run found an empty channel whatever the ledger said, so neither branch of the choice could be honoured. The content rides the durable wake event now, and the exemption is gone. The hop that carries the answer back is the one that needs it: the responder is guarded twice over (it replies and closes, and a closed channel refuses a second answer), but the reply reaching the asker lands on a channel that is already closed by design, so the ledger is the only thing between a redelivery and a second turn spent acting on the same answer.

A short-circuited trigger publishes turn_trigger_skipped. Without it, “the agent never answered” and “the agent already answered, on a node that has since died” are the same observation.

The ledger above covers a turn that finished. The harder case is one that did not: a phase breaks in round two, after round one has already commented on an issue, woken a colleague or started a coding run.

That used to hand the delivery straight back — a NAK, on the reasoning that “nothing was recorded, so a redelivery runs it cleanly”. Nothing was recorded, and that is exactly the problem: the completion ledger is skipped on the error path, and the two writes a work_key guards are an agent’s episode row and its counterparty profile. Every MCP write, every chat post, every a2a_ask (a fresh channel per call) and every run_sandbox is keyed on nothing at all. So one deterministic mid-turn failure replayed round one’s external effects across the broker’s whole delivery budget — 25 attempts a second apart, each one a real comment on a real issue.

A broken turn is now two cases, and the engine tells them apart from its own record rather than from err != nil:

what the turn’s record proveswhat happens
nothing reached outside the engineredelivered, exactly as before — a provider that never answered, a runner that could not be built, a refused budget, a seat handed to another node mid-call. None of them wrote anything, and every one is worth trying again.
a call reached outside the enginerecorded and acked. The rest of the turn is lost; its writes are not un-doable, and only one of those two is recoverable by trying again.
the turn panickedrecorded and acked, whatever the record proves. A panic is a defect in the engine, so a redelivery runs the same code on the same input and panics again, having repeated whatever came before it. The panicking round’s own record is lost with it, so “nothing reached outside” could not be established anyway. Before panics were recovered, one unwound into the queue backend’s handler guard, which NAKs, and the trigger came back for the whole delivery budget.

The proof is deliberately narrow, because a true answer spends a trigger. A call counts only when it is MCP-backed and not positively annotated read-only — the same rule the delivery gate has always used — or when the tool’s own annotations say it is open-world, which is how a2a_ask and run_sandbox count despite being first-party. A tool nobody classified proves nothing. In particular the executor’s own submit_work, the discovery pair and the sub-agent spawner carry no annotations at all, so they cannot trip it; a predicate written as “not proven read-only” would fire on every turn that closed a single round.

Giving up on a trigger is never silent. The turn has already published its own completion marked failed, and a TurnTriggerSkipped beside it says the trigger behind it will not come back, and why. A panic also publishes turn.guard_breach(kind="unhandled_exception"), which is what puts the failure on the seat as its last_error, and the log line that recovered it (turn_phase_panicked, dispatch_panicked or sandbox_resume_panicked) carries the stack.

The same decision guards the other path a turn can arrive by. A resumed turn re-enters the executor’s suspended conversation, so a redelivery repeats every call the resumed round made — and a turn coming back from a coding box is the one most likely to have pushed a branch already. A resume that broke after acting, or that panicked, therefore keeps its claim, which is what stops a retry winning the flip, and its run is settled like one that finished: the box is reclaimed and the seat is free for its next turn. Every other resume failure still un-claims and comes back, because the suspended conversation is the expensive thing there and a resume that proved nothing has lost nothing by trying again.

Two bounds worth stating plainly. The record is per turn, in one process: two nodes that both run one partition — possible if a turn outlives the 30-minute ack window — each hold their own, so neither sees the other’s. And the completion row is only written for ledgered trigger types; a type outside that set is never recorded, so its redeliveries re-run as they always have.

A seat is unowned during a lease gap, a claim ramp, a rebalance, or a full fleet restart. Its mail must survive all of them.

It does, because the durable subscription is what retains messages, and the subscription exists whether or not anything is attached to it:

  • Every seat’s subscription is created at boot, by every node, at the earliest message — and for every seat in the company rather than this node’s share, because a mailbox is a fact about the company and the node that ends up serving a seat may not be this one. Creating one is a plain client call that attaches nothing (1.7 ms, idempotent), so it can neither take a share of a peer’s live traffic the way creating one by subscribing would, nor cost anything when a peer got there first. A config apply that adds a role runs the same walk again.
  • Detach is non-destructive: the subscription, its cursor and everything it retains survive, so nothing is lost and whoever attaches next gets all of it. That is a promise about the mail, not about the price: a message the detaching consumer had already taken and not settled goes back like any other hand-back and spends one of its deliveries, and it may come back behind messages that were never delivered rather than ahead of them. One handoff costs one delivery, which is what the 25-delivery budget is sized for.
  • Deleting one is explicit (delete_subscription) and reserved for a seat that has left the company, whose mailbox must not accumulate undeliverable events for ever. The maintenance duty does it, after a grace period: see The removed seat.

Creation is a boot step because the alternative is a silent drop, not a slow one.

The agent and notification streams retain by interest: a message is kept while a durable consumer that has not acked it exists, and a message published to a subject no subscription covers is discarded at the publish. That is the queue contract’s stated behaviour rather than a broker surprise, and it is the same rule that makes an unowned seat’s mail safe — the subscription, not a consumer, is what holds it.

So the window that loses mail is the one before a seat’s subscription exists, which is why every node creates every seat’s mailbox at boot and why the cost of doing so had to be a millisecond. While the seat is in the company nothing reaps it: a durable consumer carries no inactivity threshold, so a seat can stay unowned for as long as a rebalance, a failed teardown or an operator takes.

A seat that leaves the company, because its role was deleted, renamed to a new handle or changed to a human seat, leaves its mailbox behind. Nothing consumes it again, and an interest-retained subscription keeps every event still addressed to the handle. Left alone, that mail is retained for the life of the deployment, and a seat later added under the same handle attaches to the old backlog and works it under a role definition that never wrote it.

So the mailbox is retired: once the seat has been absent from the active revision for 24 hours, the maintenance duty deletes its inbox, its sandbox control subscription and its reflection subscription, and the mail they hold with them.

What a removed seat hadWhat happens to it
Its mailbox (the inbox, the sandbox control subscription and the reflection subscription)Kept, with its mail, for 24 hours after a sweep first sees the seat missing, then deleted
Its coding runs (running, parked on a question, re-seeding, or mid-resume)Kept for the same 24 hours, then ended as part of the retirement and before the mailbox is deleted: each box is reclaimed, each loss is announced as a sandbox_run_failed event with reason seat_removed, and each run’s record is deleted
Its memory (diary, episodes, counterparty profiles, onboarding markers)Kept. Memory is keyed by the handle or by the agent id derived from the company name and the handle, so a seat added again under the same handle reattaches to it
Its seat leaseReleased by the node that held it, on that node’s next placement pass after it applies the revision, which the apply asks for at once (seat_released_role_gone)

Why a grace period rather than deleting on the apply. A delete is the one change here that cannot be undone, and seats are removed by mistake: an edit that is reverted, a builder operation that is undone, an import of an older file. Twenty-four hours is long enough for a seat restored within a working day to come back to the mail it was sent while it was gone, and short enough that mail for a seat nobody runs is not kept for more than a day. The clock starts when a sweep first observes the absence, so a retirement is never early: at worst it is one maintenance tick (15 minutes) late, and later still while no node that runs worker duties has applied the current revision.

How the fleet knows a mailbox exists. A mailbox’s name is derived from its handle, so nothing needs to remember it while the seat is in the company. A removed handle is gone from the org, though, so every node records each seat in the coordination store’s mailboxes bucket before it creates the subscription. That record carries what the retirement runs on: when the seat was first seen missing, whether a retirement is in flight, and the version every write is conditional on. A registration that fails is logged as seat_mailbox_unregistered and the mailbox is created anyway, because a seat in the company losing mail is worse than a mailbox the next sweep registers.

The broker is the backstop. A mailbox can exist with no record: a node’s registration failed, and the node created the mailbox anyway. So every sweep also asks the broker which seat mailboxes it holds (the queue contract’s subscription listing), and registers each one belonging to a seat that is neither in the active revision nor in the registry, logged as seat_mailbox_discovered. Its absence is stamped on the tick that finds it, so it is retired 24 hours later like any other removed seat’s. Only a subscription the mailbox grammar produces counts (the seat’s inbox under its inbox group, its control subject under its control group, or its reflection subject under its reflection group); a consumer on a seat’s inbox under any other group is somebody else’s and is left alone. A listing that fails only postpones the discovery; the registry’s own records are judged regardless.

The sweep, one tick at a time. For each record:

  • The seat is in the active revision: an absence recorded earlier is cleared (seat_mailbox_returned), and a seat that returns and is removed again starts a new 24 hours.
  • The seat is missing and no absence is recorded: the sweep stamps one (seat_mailbox_absent, naming when the mailbox will be retired).
  • The absence is older than the grace period: the mailbox is retired (seat_mailbox_retired).

A seat in the active revision with no record at all, a mailbox created by a node whose registration failed, is registered by the sweep so it can be retired if the seat is ever removed. A seat outside the active revision with no record, whose mailbox only the broker’s listing still finds, is registered and stamped absent in the same tick.

Unknown is never absent. The roster is the agent seats of the revision the fleet’s activation pointer names, read on a node that has applied that revision. A pointer that cannot be read, a node still converging on the current epoch, a registry or a seat lease that cannot be read: each makes the tick stamp nothing and retire nothing, and the error is logged as the reason. A node that is behind must not be able to judge a seat absent that the fleet has just added back.

A seat still held is kept, and a seat being retired cannot be claimed. The seat host releases a seat whose role is gone, so a live lease on a removed seat is a node still serving an older revision, still attached to the mailbox. The retirement waits until that lease is released, and says so as seat_mailbox_retirement_held; the same line covers a claim the protocol gate refused, while a node on a lower lease protocol still holds a presence or seat lease in the fleet, and its refused field says which of the two it was (held or protocol), so the line sends an operator to the node still serving the seat or to the lower-protocol node that has not left yet, never to both. To retire, the maintenance duty claims the seat’s lease itself, under an owner of its own, and holds it until the subscriptions and the record are gone. A node that installs a revision adding the seat back claims seats from that revision without consulting the mailbox record, and attaching the seat’s consumer creates the very subscriptions the retirement is deleting; the lease is the one thing that node’s claim loses to. While it is held the fleet view shows the seat’s lease owned by the node running the duty, normally for the few milliseconds a retirement takes on a healthy broker. A retirement never acts for longer than half the lease TTL, so its last delete lands while the claim still excludes every node, and the lease is released the moment it finishes.

A retirement ends the seat’s coding runs first. A detached run’s completion is routed to the seat’s control subscription and a parked question’s answer arrives on its inbox, so once those are deleted nothing can ever reach the run again, and its record would sit in the fleet’s run store for good. So after the mark and before any subscription is deleted, the duty ends every run the seat still has, under the seat lease it holds: a node adding the seat back cannot claim it and recover a run while the run is being ended. The runs are kept through the grace for the reason the mail is, so a seat restored within a day comes back to its parked questions and its running jobs. A run that cannot be ended (a box that cannot be reached within the retirement’s budget, a run record that cannot be read) unmarks the retirement, which is retried on the next tick with the runs already ended gone. A node that runs no sandbox coordinator, because its company configured no providers.sandbox when it started, cannot reach a box, so it refuses to retire a seat whose runs are still recorded and leaves that seat to a node that can.

Every write is a compare-and-set, and a retirement is marked before it deletes. Two sweeps can overlap while the duty moves between nodes; they read the same record, and exactly one wins the mark that starts the retirement. A node that adds the seat back while the retirement is deleting its subscriptions finds the mark and waits for the retirement to finish before it creates the new inbox, so the delete cannot land on it; the new mailbox starts empty. If the retirement does not finish within twice its 30-second budget, the node takes the record over (seat_mailbox_retirement_taken_over), and the retiring sweep, finding its record changed, restores the inbox in case its delete landed after the node’s create. A sweep that dies after marking leaves a mark a later sweep resumes once it is 15 minutes old (seat_mailbox_retirement_resumed), and a retirement that cannot delete a subscription is unmarked and retried on the next tick.

Every failure above assumes a node either works or dies. The one that is neither is a process whose duties have stopped turning while the process stays alive — a deadlock, a duty blocked on something that never returns — and it is the worst case, because the two halves of ownership come apart:

  • Its seat leases lapse, because nothing is renewing them. Peers take the seats over, correctly.
  • Its stalled handler is still holding a delivery. Pull consumers prefetch nothing, so what a stopped loop holds is bounded by the batch already in its hands rather than by the seat’s backlog — but those messages are ack-pending for a corpse, and the clock on them is the broker’s: JetStream returns a fetched-unacked message when the 30-minute ack window elapses, and closing the connection does not shorten it. The successor serves everything published after it claims the seat; it is that one batch that waits.

Nothing can be scheduled out of that state either: the duty that stalled is the one that would have to run the recovery, so anything queued behind it waits on the very blockage it is reacting to. What a watcher can do unilaterally is end the process — and it is worth being exact about what that buys, because it is not the held mail back:

  • It removes the actor. A wedged node that later resumes acts on a seat it no longer owns, with its MCP children still spawned and its credentials still live — and a turn’s outbound effects, a chat post or a work-item comment, are not something an epoch fence can reach.
  • It lets the node come back. A process that neither works nor dies goes on answering /health, so nothing restarts it: the seats it was handed are covered by peers, but its capacity stays lost until a person notices.

So every node runs a watchdog: each duty stamps a beat as it turns, a separate goroutine compares, and it calls os.Exit(75) when the lag passes the lease TTL. Five things about it are deliberate:

  • The threshold is not a config knob of its own. It is the same number the lease TTL is. Past it the node is provably not the owner, and letting the two drift is how a process gets to be simultaneously “not the owner” and “still holding the mail”. So it follows the lease TTL — the one this deployment is running, whether that came from coordination.lease_ttl_seconds or was adopted from the bucket a peer created, not the shipped 45 seconds. A node leasing its seats for twenty seconds and watching to forty-five would hold mail it did not own for the difference; one leasing them for two minutes would be killed while it still provably did.
  • The exit is the crudest possible one — os.Exit, which runs no deferred function, rather than a panic, a signal, or a graceful shutdown. The duty that stalled is the one that would have to run the shutdown, so anything waiting on it hangs; trying is how a watchdog ends up wedged too. The notice goes straight to stderr rather than through the logger for the same reason: a configured handler may batch, format, or ship lines somewhere — a log file among them — which is more machinery than a wedged process has earned. So this notice is not in logging.file, only on stderr. That is exactly why logging.stderr: false silences the ordinary log stream and not the stream itself — a node that handed its whole log to a file would otherwise lose the one line explaining its own death. A deployment that discards stderr at the shell or the unit instead (2>/dev/null, StandardError=null) does lose it, and its only remaining evidence is the orchestrator’s exit code 75. Exit code 75 is distinct from any ordinary failure, so an orchestrator’s restart log says what happened.
  • It is disarmed for the whole of a normal shutdown. Teardown is the one part of the process that legitimately blocks the loop — reaping MCP subprocesses, joining threads, tearing sandboxes down — and exiting through the middle of it would abandon the seat release that makes a drain graceful. A shutdown that hangs is a SIGKILL away; a shutdown that exits without releasing costs every peer a full TTL of dark seats.
  • The beat cadence is scaled to the threshold, not set independently. A beat slower than the threshold makes a perfectly healthy node shoot itself, so both the stamp interval and the poll interval are ceilings derived from it.
  • A duty that is gone is not a duty that is wedged. From the watcher the two are indistinguishable — the beat simply stops refreshing — and they are opposite situations. A wedged duty is still alive inside a live process: still sitting on a fetched batch, and still able to act on a seat it no longer owns. A duty that has finished took its work with it, so there is nothing left holding anything. The watchdog therefore stands down when nothing it watches is live any more, rather than exiting. Without that check, every engine abandoned rather than stopped arms a suicide timer that fires one TTL later on a perfectly healthy process — which is not hypothetical: it killed this repository’s own test suite at 63%, exit 75, with zero test failures.

Single node or fleet, it is armed the same way. With one node no peer is waiting on anything the wedged process still holds, but a wedged engine is a dead engine either way, and leaving is what lets a supervisor notice.

A detached coding run outlives the node that started it, so its completion has to reach whichever node owns the seat now. Each seat has a control topic, crewlet.agent.{handle}.control, attached and detached alongside the inbox — so routing emerges from who subscribes, exactly as it does for the inbox, rather than from any “which node” computation.

It cannot ride the inbox itself: while a run holds the seat, every inbox delivery is parked (requeued and acked), and a completion riding the inbox would be parked behind the very busy state it exists to clear.

The run record carries owner and owner_epoch, so a run is recovered by the node that owns the seat, under that node’s epoch, as a step inside the acquire hook. The record lives in the coordination store, which is what makes that possible at all: on the node’s own database the successor’s recovery pass listed nothing, and the run’s box was neither resumed nor reaped.

Routing is org-derived, never instance-derived

Section titled “Routing is org-derived, never instance-derived”

Agents exist only on the seat’s owner, so any code that resolves a recipient through the local agent pool is broken in a fleet. A fleet-wide consumer group hands a delivery to an arbitrary node; that node looks the recipient up in its own pool, finds nothing, and drops it — (N−1)/N of the time.

Routing needs only handle → (inbox topic, agent id), and both are derivable from the org every node has in full:

topic := topics.AgentInbox(handle)
seat := org.AgentSeatByHandle(handle)
agentID, ok := org.AgentIDFor(seat) // uuid5 over (org name, handle)

Both are derived from the ORGANIZATION, which every node holds in full. The running agent is an execution detail; it must never be a routing one — a lookup among the seats this process happens to be running answers “is this agent here?”, and a miss means “not on this node”, never “does not exist”.

Some work belongs to the company rather than to a seat. Running it on every node is not merely wasteful, it races — N reapers deciding independently to expire the same paused sandbox, N clustering passes writing N sets of near-identical auto-drafted skills.

Each sits behind a worker:{duty} lease, claimed per tick rather than held, so there is no handoff protocol: a node that stops gracefully gives every duty back once its duty loops have finished their last tick (duty_released), and a peer takes it on its next tick; a node that dies releases it by lapsing, and a peer takes it on the first tick after the lease runs out. A duty’s TTL is sized from its own work rather than from the seat heartbeat (most outlive three of their own ticks, so one slow claim never moves them, and the integration reconcile’s outlives one whole pass), which makes it anything from 30 seconds to three hours, and duty leases live in a bucket of their own for that reason: the seat lease TTL does not bound them. They are:

DutyWhat it doesWhy once
sandbox-waiterPolls live sandbox boxes, keeps them alive, reaps expired pausesEach poll is a reconnect, so N nodes means N reconnects per box per tick — and N racing reapers
schedulerEvaluates every schedule and fires what is dueThe fleet’s fire claim already makes a dispatch at-most-once, so peers are not wrong — they lose the race on every fire, having walked the whole org to get there
learningAll four learning background passes: clustering skills out of episodes, the active → stale → archived lifecycle, episode compaction, and cross-agent promotionClustering reads every agent’s episodes and writes skills, so N nodes produce N sets of near-identical pages and N× the LLM spend; the curator publishes a lifecycle event per transition and races its own optimistic-concurrency guard. One lease for all four because they run on one loop
integration-reconcileRuns every connected third-party app’s reconcile pass on its own cadenceN nodes would each sweep every surface on every tick. What stops two writers at one third-party app is a different lease, worker:setup-provision-<kind>, which every writer at a surface holds around one pass and gives back when it ends, so it is mutual exclusion rather than a duty; this one keeps the sweep itself on one node
maintenanceRetention sweeps for every short-horizon table in the node’s own database (events, scheduled_runs, conversation_sessions, the expired and over-cap agent_diary rows, counterparty_profiles), plus both halves of the A2A channel sweep: the idle-close of an ask no turn ever answered, and the delete of one closed long enough. And the retirement of a removed seat’s mailbox and its coding runs once the seat has been absent for 24 hours. And the coordination store’s removal markers, from every shared-state bucket with no age, once they are three hours old — each purged through its own revision, so a key written again since is untouched. The channel and mailbox records and the markers are the shared things swept here, and the coordination store says why: its other slots expire on a bucket’s age, which cannot tell an open channel from a closed one, or a present seat from a removed oneIdempotent range deletes, so peers are harmless, just N times the write amplification and vacuum churn. The mailbox retirement is not idempotent in the same way, so it takes its own compare-and-set on every record and claims the seat’s own lease while it deletes, rather than relying on the duty’s lease
retentionThe state-log trim: every 15 minutes it evaluates each domain’s floor across every data node’s position, the backups and the holds, purges what they all permit, and publishes what it concluded and which term is holding it (see Retention)Two nodes purging and publishing at once would publish two conclusions about one log, and how long a term has been holding is a property of a series of ticks that must be one node’s
embeddingsComputes the embedding of every page and task whose vector is missing or stale and publishes it to the vectors log, which every node appliesEach node embedding the corpus would be one provider bill per node for one company — the cost the vectors log exists to pay once
object-collectorRuns the object store’s two passes against the one store the fleet’s files are in: hourly, deletes every object stored and keyed more than a day ago that no row names (and stops deleting when this node’s estate is incomplete), and abandons uploads that never finished; daily, audits that every object a row names is in the store and whole, raising objects_missing when one is not. Records each pass in the coordination storeTwo collectors are safe — a key is named only by the write that uploaded it, and deleting an object that is gone is not an error — but each would list the whole store and read every row naming an object, for one answer

Without a placement host — the single-node case — the answer is always yes: there is no fleet to be a singleton within. A duty claim that fails (an unreachable lease store) skips the tick rather than proceeding: unknown ownership is not ownership, and assuming otherwise is how every node decides it is the singleton at once.

A claim is not the only way a tick declines. The integration-reconcile loop asks a second question first — is this node’s config posture current? — and on shed or stuck it declines without claiming, so the lease lapses and a peer on the current revision takes it. That question is about this node’s own configuration rather than about ownership, and it is asked only there: the loop is the one duty whose work reaches outside the deployment, so a stale document means accounts and webhooks converged to a revision the fleet has replaced.

Not everything periodic is a duty. The tool-skill walks (at boot, on an apply that moves the skills source, and every 10 minutes) look like one and are not: they populate a process-local registry, so every node has to run them or its agents have no tool skills at all, and the news that one page changed is broadcast to every node rather than claimed by one (see Keeping every node current). The test is whether the work produces shared state (a duty) or warms a local cache (not one). The reflect dispatcher is a third shape: every node has one, and what feeds it is the reflection subject of each seat that node holds, so the seat lease already decides which node reflects on a turn and a duty lease would add nothing. The seat-mailbox walk is a fourth, and the sharpest: it is company-wide state, and it still runs on every node, because creating a durable consumer is idempotent and costs a millisecond, while putting it behind a lease would leave a seat added by a config apply with no mailbox until the duty’s holder got round to it, and mail published to a subject with no subscription is dropped rather than queued.

A rolling upgrade puts a vN and a vN+1 node on the same lease table and the same topics at the same time. That is fine as long as both agree on what holding a lease means, and catastrophic when they do not.

The rule is asymmetric: a node refuses to claim anything while a live presence or seat lease is held at a lower protocol version. The check only ever looks down, so lower-protocol nodes keep working; higher ones wait, visibly, until the last lower lease lapses. A rolling deploy converges because that is what a rolling deploy does.

Those two kinds are the whole question. A presence lease is what says a node of the older build is alive; a seat lease is what it is still running, and a draining node gives its presence up first and keeps serving its seats until each is handed over, so the seats have to count as well. A duty lease does not count, and neither does any other claim (the tracker’s walk claims). A duty is claimed ungated, so it never waits for the gate, and it outlives its holder by its own TTL, up to three hours for the learning passes. If it counted, an older node that crashed during the upgrade would have its presence and seats lapse within one lease TTL (45 seconds as shipped), but its duties would keep every newer node from claiming a seat for hours.

Three consequences worth stating plainly:

  • Lease schema evolution is additive-only. A field the vN build does not read is invisible to it; one it requires is a crash.
  • A downgrade across a protocol bump needs a full drain. The gate only ever refuses a higher-protocol claim beside a lower lease, so a lower-protocol build takes over a higher node’s expired leases unchecked. Stop the whole fleet before rolling back.
  • The wait is an outage window, and it is the point. Higher-protocol nodes claim nothing until the last lower presence or seat lease lapses or is released — at the shipped 45-second TTL plus however long the lower nodes take to drain. A duty an older node still holds does not lengthen it, even if that node crashed. Plan the rollout for it rather than being surprised by it: the alternative is two builds disagreeing about what a lease obliges them to do, which is silent and unbounded rather than visible and finite.

What the gate costs does not grow with the fleet. “Is any live presence or seat lease held at a lower protocol” is a question about every such lease in the fleet, and the lower-protocol builds it is for say their protocol nowhere but inside each lease record — so it used to be answered by reading every lease, twice per claim. A node converging from cold claims its seats one after another, so its claims cost the square of the seats: on a three-member cluster with about 9,300 seats held a claim took 0.3–0.5 s, and a hundred nodes claiming ten thousand seats took the brokers out of quorum. A node now keeps a view of the seat lease bucket while it is claiming — a watch that takes in every write to it — and judges the gate from it, after reading the bucket’s last write from its leader and waiting for the view to reach it, so the answer is the one a full read would have given. A claim costs the same with ten leases held as with ten thousand (measured on one member: about 1 ms, where it was 6 ms and 230 ms). The view stops a minute after the node last judged a gate, so a node that is not claiming does not pay for the fleet’s heartbeats reaching it — and a claim on a seat a peer holds judges no gate at all, because a refusal names the rule that made it (the peer’s hold or the protocol gate). So a node with room for one more seat and nothing free to take — the steady state of most of a fleet whose seats do not divide evenly — reads no protocol floor and keeps no view: it used to read the floor after every sweep that took nothing, which kept its view running and took in every lease write the fleet made (measured with 2,000 seats a peer renewed at the ten-thousand-seat heartbeat rate: 7,426 messages to the idle node per five-second sweep, now 4,010). The floor is read only after the gate has refused one of a pass’s claims, when the claim has just judged the same view, and a gate refusal ends the pass, since every other claim in it would be refused the same way.

A node with room and nothing free tries no seat at all. It used to learn that nothing was free by trying every seat it may run — a read of each seat’s lease from its stream’s leader, plus a walk of the epochs bucket to order them — every sweep, which was the rest of that cost: measured with 2,001 seats and a peer holding 1,001 of them, 4,021 messages to the idle node per sweep. Every node now advertises how many seat leases it holds (seats_held, the seats it runs and the ones whose teardown it could not prove, counting only seats some live node may run) on its presence row, which each sweep already lists to size its share; when those counts together cover every seat some live node may run, the pass claims nothing and says so (FleetFull on the sweep’s record) — 10 messages per sweep on the same fleet. Nothing is claimed on the counts, so a wrong reading costs time or reads and never a seat held twice, and they err toward trying: a count a row does not carry readably reads as zero and adds nothing to the sum, a node whose presence lapsed is not listed, so its seats read as free the moment it goes, and a roster the pass could not read concludes nothing. A lease on a seat that is not placeable — a stuck seat whose role is gone, or whose placement no live node now matches — is left out of the count, because the sum is compared with the placeable seats: counted, it stood in for a free one, and every node read the fleet as full for as long as the teardown kept failing. A node that takes seats advertises its new count at the end of that pass, and one that gives a seat back advertises it in the release itself, whichever path releases it — so a seat given back is taken at the next sweep rather than after the giver’s next heartbeat.

The current protocol is 4. It is raised only when holding a lease comes to mean something a build at the previous protocol could not honour — never for a new field. At 4, holding a seat lease means three things:

  • The completion ledger. The holder consults and settles the completion ledger, so a seat taken over never re-runs a turn whose effects already shipped.
  • Placement. The holder satisfies the seat’s role.placement. The lease itself is only a mutex and knows nothing about where a seat belongs, so a holder that ignored placement would claim a seat pinned to a node id or a label it does not carry — and succeed, silently violating the operator’s pin.
  • Windowed token counters. The holder charges the seat’s rounds to the windowed token counters — a slot per day, week and month on the company’s clock. Two builds charging different records would each admit against a figure missing what the other spent, and every cap would bind late by exactly that much.

GET /health lists the seats this node holds (seats). The Settings › Nodes screen and the fleet query (GET /query/fleet) show every live node with the seats it holds, its roles and labels, its lease protocol and when its presence expires. The log carries the rest: seat_sweep reports each pass’s held count against the computed capacity, seat_claimed and seat_attached / seat_detached carry the seat and the epoch (a detach also carries its reason), seats_unplaceable names seats no live node can run, and seat_claims_blocked_by_older_protocol names the protocol floor a lower-protocol peer is imposing.

unproven_seconds on GET /health is the number to alert on: a map of seat handle to how long its teardown has been failing, present only when one is stranded, and absent from seats for as long as it is. The seat_still_unproven log line carries the same alarm with the attempt count, repeating every twenty heartbeats. Alert on the duration, not on the field appearing or on the first seat_release_unproven: a teardown that fails once and succeeds on the next heartbeat retry is a working system (seat_release_recovered), while a seat still stranded minutes later is a seat nothing in the fleet is running.

Everything above runs unchanged on one node: it is the degenerate case, not a second code path. On the default coordination.type: local the leases never leave the process, which means it believes it owns the whole company. That is correct for one node and catastrophic for two, which is why the Tier A file refuses the combination at load: local coordination beside a clustered embedded stream or an external NATS stream fails validation on coordination.type, naming embedded-kv as the fix.

A fleet sets coordination.type: embedded-kv and gives the nodes one stream to share, either a clustered embedded server or an external NATS. The two go together by construction: the coordination KV rides the stream’s own connection, so refusing the mixed pairing keeps the halves from drifting apart. See Running a Fleet.


  • Scaling Out — the model this sits inside: what a node is, what the fleet shares, and where these constants were measured
  • Control Plane — how nodes converge on one company config, and the posture a lagging node takes
  • Event System — topics, groups, inbox batching
  • Code Sandbox — the detached run whose completion this routes
  • Deployment — running more than one node

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.