Skip to content
You are reading documentation for unreleased main. This page is not in 0.1 yet.

Read Consistency

Every node of a Crewlet fleet holds its own copy of the company’s work. The copies are built by applying one ordered log, so they converge — but at any instant they are at different points along it. This page is about the one question that follows: how fresh does this particular answer have to be?

There are four answers, and picking one is the whole of it.

LevelWhat it promisesWhat it costs
linearizableNo answer from before this read arrived.An append to the log and a wait for it to come back.
sessionNo answer from before your own last write.Nothing at all when you are caught up, which is the common case.
staleWhatever this node holds, with its lag reported.No broker call.
consistent_prefixA coherent point in the log’s own order, possibly behind.No broker call.

linearizable is for a decision, which is why it is what an agent reads at. Use it when the answer decides something — an admission check, a gate, a number somebody is about to act on irreversibly. It is the only level that establishes a position at the log’s current end before answering. Every seat tool read is one of these: a create that refuses a project the company does not have, a hand-off that names a colleague, a turn that reports what it found. A seat has no screen on which to notice it was reading a stale copy, so it never reads one.

stale is what the dashboard and the read API use. A board that is two hundred milliseconds behind is a board, and the answer carries its own lag so a reader can tell. A tile that took a barrier append to redraw would put the fleet’s whole read rate on the log to remove a staleness the next redraw removes anyway.

session is the write path’s level, and a read surface offers it only with a position. It waits for your own high-water mark, which you have to supply. A surface that accepted the level bare would wait for the zero position, serve whatever that node happened to hold, and label the answer session: a wrong label rather than a weaker answer. So the read grammar accepts read_level=session only beside min_position — the position the write you are reading back answered with — and refuses it without one, naming the key and the two other honest asks, linearizable or stale with max_lag_seq. A seat’s tools carry no position and choose no level, so there the ask is unreachable. Inside the engine the level is real and used: a write waits for its own last write to be applied before it opens the snapshot it decides from.

That wait is per log, not per object. A node that has just published a bulk update waits for it to apply locally before its next write on that log — any subject — and may be refused behind, naming its own record. That is what read-your-writes means on the write path: the alternative is a write that cannot see the write before it. It costs nothing when the node is caught up, which is every ordinary write.

consistent_prefix is the default for nobody, and it is either asked for or resolved to. It promises a named prefix of the log — everything up to a stated position, with nothing from after it — and makes no statement about age. That makes it weaker than session in a way the answer cannot show, so nothing defaults to it: a surface that quietly downgraded to it would give every reader that did not know to ask for more an answer they could not tell apart. The dashboard may ask for it, and a replication answer resolves to it when the lag is unmeasurable — which is not a downgrade but the honest name for what is left when there is no age to claim, and the answer says so.

The level is a property of the surface asking, never of the caller who happened to omit the key. There are four:

SurfaceDefaultMay the caller choose?
A seat’s own tools, inside a turnlinearizable, floored at its node’s own writes and at the change that woke the turn (below)No
The operator MCP, about tracker contentlinearizableNo
The dashboard and the REST read pathstaleYes — linearizable, stale, consistent_prefix, or session beside a min_position
Any answer about replication — the retention report, Settings › Backups & retention’s lag, whether a purge landedstale, weakening to consistent_prefixNo — it is derived, not chosen

Only the screen chooses, and the reason is that only the screen can see what it got: the level and the lag are rendered beside the rows, so a person who asks for a weaker answer is shown the one they were given.

A replication answer is the one row where the SUBJECT decides the level, not the caller, and it holds on every surface including the operator MCP. linearizable means every mutation committed anywhere in the company before this read was issued is in the answer, and it is established by appending a barrier and waiting through its position. But the question these answers ask is how far behind that same log this node is — the barrier is the instrument and its health is the subject. A node cannot produce a linearizable answer to a question about its own replication, so that is not a stronger answer costing more; it is a level that cannot be served, refusing in precisely the incident somebody opened the page for.

It weakens one step further when this node could not measure its own lag at all — the broker unreachable, or coordination — because stale is a claim about age and there is then no age to claim. crewlet retention status and Settings › Backups & retention say so in a sentence rather than printing the same figures under the stronger name.

An agent cannot choose because the level is not a model’s to pick — a tool argument for it would be a model trading correctness for latency it cannot perceive. The operator’s reads cannot choose because nothing is wrong and the person is deciding something about their own company: a knob that only ever weakens the answer is one somebody turns once, forgets, and then reads a stale board from for a year.

Every answer reports the level it was actually read at, so a caller that asked for one and got another can tell.

stale on its own accepts an answer of any age. A caller that will not says so with max_lag_seconds, max_lag_seq, or both — and the read refuses too_stale past whichever is reached first. They are two readings of one distance rather than two distances: the record count is what the broker actually answers, and the duration is derived from it through this node’s own drain rate, so a caller who can say “at most 250 records behind” is naming the measured quantity instead of an estimate made from it.

Both are refused at every other level, because they are a staleness bound and nothing else is — asking for linearizable&max_lag_seq=250 is a caller who believes they asked for something they did not.

A zero bound is not a bound: it accepts anything, which is what makes declaring one the caller’s own decision rather than a default somebody inherits.

Every write answers with the position its record landed at — <stream>@<generation>:<sequence> — on the tool answer, the operator MCP’s answer and the retention routes alike, and every read takes it back as min_position. The answer is then served from no earlier than that position, at whatever level was asked:

  • At linearizable it costs nothing: the barrier the read appends is already past any position a write on the same log answered with.
  • At session it is the high-water mark the level waits for, which is what makes the level offerable to a caller outside the engine at all.
  • At stale and consistent_prefix it turns “whatever this node holds” into “whatever this node holds, from here on”. The lag is still checked against any bound and still reported on the answer; what changes is that a node which has not reached the position within the read budget refuses behind — with the same derived retry hint — rather than serving rows from before the write.

A position on another domain’s log is refused wrong_stream, not waited for.

Two spellings, one triple: inside an answer a position is the object {stream, generation, seq} (seen_through, incomplete.from, a listing’s position), and as a parameter — min_position, a cursor, since — it is the token <stream>@<generation>:<sequence>, because a query string carries no object. A tool answer’s position is already the token, since that is what its reader pastes back.

This is the wire’s half of read-your-writes. A seat needs none of it, because its reads are linearizable; a person’s assistant filing an item over the operator MCP and redrawing a board over the REST route holds nothing else it could ask that board to include.

It is a barrier append, not a field check.

The reader publishes one record onto the domain’s barrier subject and waits for its own applier to reach it. The acknowledgement is what carries the proof: at replicas: 3 a PubAck means a majority of the raft group has the record durably, so any position at or below it is one the whole group agrees on. At replicas: 1 there is no group and the acknowledgement proves only that the one member has it — which is exactly as strong as that deployment is.

The barrier is a real record with a real cost. It is not free for most reads and cheap for a few: every linearizable read waits for exactly one append to come back. Measured on an idle loopback cluster that is about 1.5 ms from a follower at three replicas, and about 0.4 ms solo with stream.sync: always — plus whatever the applier takes to reach it. Treat those as a floor rather than a budget.

The reason it is an append and not a STREAM.INFO field is that a field read can be served by a member that has been partitioned away from its own group: an isolated former leader answers with a last sequence it believes and the majority has moved past. An append cannot be served that way, because there is nothing to commit against.

A read that cannot be served at the level asked for is refused with a code rather than downgraded. Each code names a different thing to do.

CodeWhat happenedWhat to do
behindThis node has not reached the position the read needs — including a node below the published trim floor whose missing records the log still holds, which it is replaying.Wait — the answer carries a retry hint derived from this node’s measured drain. It clears on its own.
too_staleThis node’s lag is past what the read said it would accept.Same, or accept more staleness.
stalledThis node’s applied prefix has stopped moving — its applier halted, or has been retrying a failure it cannot get past for longer than the retry budget.Its rows are frozen, so a short answer would be wrong rather than old. Check the applier — crewlet retention status names the domain, its position and the error it is retrying. A retried failure clears on its own the moment an attempt succeeds. A consistent_prefix read, which makes no statement about age, is still answered by a stalled node at or above the published trim floor — a frozen prefix is still a coherent one — and refused by one below it.
no_quorumThe barrier did not commit: the broker answered and a majority did not agree.Retry after the hint (4 s, the broker’s own minimum election timeout). If it persists, a member is down or partitioned.
broker_unreachableThe broker did not answer at all.Retry. Not the same as no_quorum, and the difference is where to look.
log_fullThe log is at the byte ceiling its ordinary appends are held to — on the tracker and pages logs a sixteenth below the broker’s, the rest being kept for gate records — so no barrier can be written.Raise the ceiling with crewlet retention set-capacity, or unblock the trim — crewlet retention status names the term. stale keeps answering, so a full log costs linearizable reads — every seat tool read among them — rather than every read.
deferredThis node holds a record it cannot decode covering what this read is about.Ask another node, or upgrade this one. No amount of waiting changes it.
deferred_scope_unknownThe deferred record’s own scope could not be read, so nothing can be said about what it covers.It blocks the whole domain, which is why it is a different code. Upgrade the node that is behind on the record version.
below_floorRecords this node never applied are gone from the log.Its rows are missing state no replay can supply. The node adopts a peer’s snapshot on its own; ask another node meanwhile. See Retention.
floor_unknownThe published trim floor could not be read.An unreadable floor blocks: guessing here keeps a node serving over a hole it cannot see. Check coordination.
evictedThis node has been removed from the fleet.Nothing it holds is authoritative. Readmit it, or route elsewhere.
maintenanceThis node runs in a capacity window (-mode maintenance or -mode seal), which publishes nothing — and a linearizable read proves the log’s end by appending a barrier to it.Ask again once the fleet is back in normal mode. Every other level appends nothing and is still answered. Every node of a fleet in a window is in one, so no other node answers it either.
wrong_streamThe position this read was asked to reach is on another stream — including a min_position naming another domain’s log, which is refused at every level rather than quietly dropped. Or this node’s own log is not the one its rows are keyed to: its checkpoint is past the log’s end — what a broker restored from an older copy looks like, since it keeps the stream’s creation instant — or the log holds, at that checkpoint, another record than the one it consumed there, which is the same restore written past this node’s rows, or the stream was deleted and rebuilt — found at boot against the checkpoint, or under a running node by the position heartbeat and by any write at an expectation of zero, from the broker’s own creation instant — or a peer re-anchored the log past this node’s generation, so it continues from that peer’s rows rather than this node’s. Every one of these findings refuses the node’s writes with the same word.A caller bug, a cursor from before a reanchor, a recreated stream or a restored broker; see Retention. A peer’s reanchor clears on its own, once this node has adopted a snapshot from the new generation.

Four of them are worth coming back to this node for — behind, no_quorum, broker_unreachable and stalled. The rest are not, and the distinction is in the code rather than in a retry loop’s guesswork: a caller that retried deferred would loop forever.

Completeness is a different fact from freshness

Section titled “Completeness is a different fact from freshness”

Every set answer carries complete beside its level, and the two fail independently.

An answer can be perfectly fresh and incomplete. A record this node cannot decode is retained rather than applied, and if it touched something the read is about, a row that would have entered the set has no local trace. The level says nothing about that — so complete: false is its own field, with an incomplete block saying how many records could intersect, the lowest position among them, and the scope they declared.

It deliberately does not name the objects. Enumerating them would disclose neither the direction of the difference — a row that would have entered the set or one that would have left it — nor reliably the right identifiers, and it would truncate. direction is always "unknown", and it is a field rather than an omission so a reader meets the fact instead of inferring it.

Three things are outside this vocabulary entirely, and asking for a stronger level does not touch them:

  1. An apply bug. N nodes applying one ordered log deterministically produce N identical copies including of a bug. Every level agrees; they agree on the wrong thing.
  2. A record this build cannot decode. It is retained, not applied. linearizable establishes a position past it and still cannot show you what it would have written.
  3. What another company’s system did. A level is a statement about this log. A Jira issue that changed a second ago is not in it.

Read-your-trigger is a floor, not the mechanism

Section titled “Read-your-trigger is a floor, not the mechanism”

An agent woken by a change to the engine’s own tracker or knowledge base sees that change. That is guaranteed by the wake carrying the record’s own position — trigger_position, where on its log the change was committed (see the event system) — and not by the level the turn’s tools then read at:

  1. Before the turn’s first read, the node hands that position to its read-your-writes floors: one table per node, the same one every write the node makes raises.
  2. Every tracker or knowledge-base read the turn makes carries the floor on its own domain’s log to whichever node answers it — a tracker read the tracker log’s, a page read the knowledge base’s — this node’s own copy included.
  3. That holder waits up to two seconds to have applied it, or answers behind and the next holder is asked.

The floor travels with the request rather than being waited for before the turn opens, because the node running the turn may hold no data at all and have no applier of its own to wait on. A floor on one domain’s log never holds up a read of the other: a tracker read does not depend on the knowledge base’s rows.

What the floor does not reach. The ranked searches — search_knowledge, search_work_items and the turn-start knowledge block — read an index each node builds behind its own rows, so no log position describes them and they carry no floor: a seat woken by a page created a moment ago may not find it by search yet, and finds it by listing or by its id. A vendor’s wake — a Slack message, a Jira issue — carries no position at all, because what changed is in another company’s system.

So a seat’s linearizable reads are not what makes it see its own trigger. What they buy is the other half: that an answer the turn decides on is not one from before the read arrived. The operator’s reads are linearizable too and are not settable: the person asking is deciding something about their own company, and operators are few.

  • Replication — what the two positions in every answer mean, and what the log does and does not promise.
  • Retention — the trim, the terms that block it, and what a full log costs.
  • API endpoints — the level header and the refusal shape on the wire.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.