Skip to content
You are reading documentation for unreleased main. This page is not in 0.1 yet.

Metrics

Generated from metrics.Catalogue(). Do not edit — change the catalogue.

The engine exports OpenTelemetry metrics through the same OTEL_* environment its traces use, because a company has one telemetry backend and two settings would let you split signals across two of them where a correlation resolves on neither:

VariableWhat it does
OTEL_EXPORTER_OTLP_ENDPOINTThe collector. Metrics are posted to <endpoint>/v1/metrics.
OTEL_EXPORTER_OTLP_METRICS_ENDPOINTOverrides the base for metrics alone.
OTEL_EXPORTER_OTLP_PROTOCOLhttp/protobuf (default) or grpc.
OTEL_EXPORTER_OTLP_HEADERSSent with every export — an API key, a tenant.
OTEL_METRICS_EXPORTERnone switches the export off and leaves traces on.
OTEL_METRIC_EXPORT_INTERVALExport period in milliseconds. Default 60000.

The provider is installed whether or not you are collecting. Every instrument exists either way, so no code path in the engine branches on whether metrics are “on”, and the numbers on crewlet retention status are read from the same recorder a collector would be reading. With no endpoint set the measurements are taken and not shipped.

Every attribute below is a closed set. No metric carries a task key, a subject id, a page title or a node id as an attribute — the node is the resource, set once — so the number of time series is bounded by this page rather than by how much work your company has done.

A distribution, exported with the engine’s own bucket boundaries: twenty-one powers of two from 64 µs to 64 s, which resolves a percentile to within a factor of two at every scale here — from a 40 µs index probe to a 16-second bulk apply. The SDK’s default boundaries stop at 10 s, so a slow apply would land in an overflow bucket and read as “at least 10 s” for ever.

MetricUnitAttributesWhat it makes visible
crewlet.statelog.publish.durationmsdomain, outcomeA write path slowing down before it starts refusing. The outcome dimension separates the three answers a write has, so a rise in pending reads as an applier falling behind rather than as a broker getting slower.
crewlet.statelog.publish.rounds1domainContention on one subject, which the compare-and-set round cap bounds and nothing measured. A distribution creeping toward the cap is a hot object; reaching it is a refusal a model reads as a colleague editing the same thing.
crewlet.statelog.write.session_waitmsdomainThe wait a write pays for its own previous write to apply. Zero when the caller is caught up, which is the common case, and the shape of the load that is not.
crewlet.statelog.barrier.durationmsdomainThe broker round trip under every linearizable read, and the first number a drifting fsync or a degrading quorum moves. It was a benchmark’s p50 on an idle loopback cluster and nothing in production.
crewlet.statelog.read.waitmsdomain, levelHow much of the read budget a barrier or session wait actually spends. A p95 approaching the budget is reads about to start refusing, which is the warning the refusal itself is too late to be.
crewlet.statelog.apply.latencymsdomainTHE COMMIT-TO-APPLY GAP: from the broker’s own timestamp on a record to this node committing it. Every read level is a policy about this quantity and nothing measured it.
crewlet.statelog.apply.tx.durationmsdomain, bound_byHow long one apply transaction holds the store’s writer, and which budget ended it. A transaction is what every waiter behind it pays, and rows were only ever a proxy for the duration.
crewlet.statelog.apply.record.durationmsdomain, kindOne record’s apply. A single record past the time budget is still one transaction, so this is the real ceiling on how long a read can be delayed — a sentence in a design document until it was measured.
crewlet.statelog.apply.batch.rows1domainRows per apply transaction, which is what the row budget bounds and what the drain rate divides.
crewlet.backup.durationms—How long a backup took, which is the window the trim hold covers and the I/O the copy spends competing with the applier’s own commits. It is what turns the retention guide’s worked example into a number for THIS hardware.
crewlet.store.pool.waitmsfileHow long a reader waited for a connection. It is what says the reader pool is too small on this node, which nothing could say before.
crewlet.tracker.search.scan.durationmspath, rungOne search’s scan, split by whether it ran for a turn’s prefetch or for somebody’s deliberate search, and by the ranking it actually served (rung: hybrid, keyword, semantic, or none). Only the prefetch had a published percentile, and the interactive path is the one with a target; a keyword scan is a fraction of a fused one’s cost, so reading them as one series hides a slow meaning half.

A value that goes both ways, sampled at each export.

MetricUnitAttributesWhat it makes visible
crewlet.statelog.drain.rows_per_second1domainThe applier’s observed drain, which every retry hint divides by. Seeded from a benchmark and then measured, so a hint on real hardware stops being an extrapolation from somebody else’s.
crewlet.statelog.drain.commits_per_second1domainCommits per second, which is the fsync rate under synchronous = FULL and the number a device budget is spent by.
crewlet.statelog.apply.lag.seq1domainHow many records this node is behind the log’s head.
crewlet.statelog.apply.lag.secondssdomainHow OLD the oldest unapplied record is. Seconds are what a stall grace, a pending outcome and a copy taken out of service all turn on; sequences are not, and a lag of 4 000 says nothing about whether anything is wrong.
crewlet.statelog.applied_through1domainThe prefix this node has actually applied, which is lower than its checkpoint whenever a record was retained.
crewlet.statelog.deferred.count1domainRecords this build could not read and kept. Non-zero is a rolling upgrade in progress; non-zero and not falling is one that stopped.
crewlet.statelog.deferred.oldest_age_secondssdomainHow long the oldest retained record has been retained, which is what decides whether this node’s copy stays in service.
crewlet.statelog.waiters1domainCallers blocked on the applier right now. It is the depth of the queue a slow apply is making.
crewlet.statelog.log.bytesBydomainWhat the log actually holds, against its ceiling below.
crewlet.statelog.log.max_bytesBydomainThe ceiling, read from the running stream rather than from this node’s own configuration — the two differ, and the running one is what refuses the append.
crewlet.statelog.log.headroom_fraction1domainHow much of the ceiling ordinary writes are held to is left — on the tracker and pages logs, the ceiling less the gate reserve kept above it for evictions. At zero the log refuses every write AND every linearizable read, and the remedy is a fleet-wide maintenance cycle, so this is the one number worth alarming on long before it is small.
crewlet.statelog.trim.blocked_secondssdomain, termHow long one retention term has held the trim, named. A trim blocked for weeks is a log walking toward its ceiling with a cause an operator can act on.
crewlet.backup.ages—How long ago the newest COMPLETE backup finished, read from the manifests on disk rather than from a counter this process keeps. A counter records that a process believed it took a backup; the disk records that one exists, and they differ in exactly the cases the alarm is for — a copy deleted, a volume never mounted, a schedule pointing at a path nobody ships from. A node holding no manifest reports NO SERIES rather than a value: zero is the freshest backup imaginable and every other number is an age nobody measured, so neither can stand for an absence. The absence is what the backup_age alarm says in words, and an absent() rule over this gauge is what a collector alerts on.
crewlet.backup.holds1—Live trim holds. A pin that outlives its owner stops the trim until the stale bound expires it, so a count that does not return to zero is a backup that crashed mid-copy.
crewlet.store.wal.bytesByfileA write-ahead log a checkpoint cannot pass grows, and this is the only way to see it before the volume fills.
crewlet.store.bytesByfileThe store’s size on disk, which the snapshot’s free-space precondition and the provisioning rule are both derived from.
crewlet.tracker.search.concurrency1—Searches in flight, which is the row of the supported-corpus table this node is actually on: inside the one-second budget, about 345 000 sources with one running and about 136 000 with eight through the full scan, and about 545 000 and 183 000 through an index probing half its lists.
crewlet.tracker.vector.coverage1—The fraction of sources carrying a current vector. It is how a stalled embedding backlog is reported, since it never drops a seat.
crewlet.alarm.active1kindWhether each named alarm is firing right now, 0 or 1. It is the same table the operator record renders and the CLI exits non-zero on, so a collector and a person see one answer.

Only ever rises. Rates and totals are your collector’s arithmetic, never this engine’s.

MetricUnitAttributesWhat it makes visible
crewlet.statelog.publish.outcomes1domain, outcomeThe three-valued write outcome, counted, and only the three — a refusal is on publish.refusals instead, because it says the write never happened at all. unknown is the one that matters most and had no counter: a broker flapping into ambiguity was visible only to the model that received the answer.
crewlet.statelog.publish.conflicts1domain, subject_kindWrites that spent their whole round budget losing races on one subject, BY KIND. The refusal counter beside it says a conflict happened and not what it was about, and the remedy differs entirely: one contended object is a design question and a contended kind is a hot subject.
crewlet.statelog.publish.rejections1domain, subject_kindHow often a write loses a race, per kind of subject. It is what says whether a counter, a rank order or an ordinary object is the contended one.
crewlet.statelog.publish.refusals1domain, reasonWrites refused before or instead of an append, by reason, and EVERY value it carries is here, each with its remedy. The node: evicted (this node is removed from the fleet — run the write on another), deferred (it holds a record it cannot decode — a newer build serves it), behind (it has not applied a position the write needs — clears on its own), below_floor (it is below the log and must adopt a snapshot — another node writes meanwhile), floor_unknown (the floor or the log’s ends could not be read — clears when coordination answers). The log: log_full (at its ceiling — raise it or unblock the trim), skew (a store and a stream restored out of step — restore both from one backup), wrong_stream (not the log this node’s rows came from: rebuilt, ending below its checkpoint, holding another record there, or re-anchored past it by a peer — re-anchor, or adopt where a peer holds the generation), log_truncated (it lost records a peer’s rows hold — settle which history the fleet keeps). The object: deleted (a permanent deletion marker — nothing recreates it). The operation: op_reused (its id names a record on another object) and superseded (its record landed and a later one undid it) — both answered by a NEW operation under a fresh id, never by a retry. A record that landed and a gate dropped is counted under the gate that dropped it — evicted, deleted, abandoned (written in a generation a reanchor abandoned) or overtaken (written after a restored reanchor, below its generation) — and is never re-decided, because republishing makes another record nothing applies. Such a record holds its operation id for the log’s duplicate window, so the same id sent inside it, by any node, is collapsed onto the record and refused the same way: the write is retried under a fresh id, or under the same one once the window has passed. Beside the refusals: conflict for a write that lost every round, exists for a create whose object already exists, and error for a write that failed before it could answer. A refusal is not one of the three outcomes: it says the write never happened, and each reason has a different remedy, so one counter with an outcome dimension would hide them all.
crewlet.statelog.read.refusals1domain, level, codeEvery refusal code, counted. The codes, each with its own remedy, had no counter between them, so an operator had no rejection rate for any of them.
crewlet.statelog.read.served1domain, levelReads answered per level, which is the denominator every refusal fraction needs and the check on the assumed read rate the log’s own size is derived from.
crewlet.statelog.barrier.appends1domainBarrier records appended. Against reads served it is the single-flight ratio, which says whether coalescing is doing anything at all.
crewlet.statelog.barrier.applied1domain, streamBarrier records applied from each log — every node’s linearizable reads on it, since every node applies every record. It is the rate census_drift holds against the census, per log: a node’s own appends are only its share of the fleet’s reads, and on a fleet of several serving nodes they describe a company that many times quieter than it is. The alarm’s 24-hour window files each barrier at the hour the broker stored it, so a node replaying a backlog does not read days of reads as one; this cumulative series counts them as they are applied, so it jumps by a backlog’s barriers when one is replayed.
crewlet.statelog.linger.yields1domainHow often a waiter cut a batch short. It is the batching the applier gives up to answer a read promptly, and without it that trade is invisible.
crewlet.statelog.apply.records1domain, resultRecords consumed, by what happened to them: applied, retained, gated, skipped, or reprocessed by a build that could read what an earlier one retained — a retained record a gate drops when it is reprocessed is counted gated, as it would have been live, since it wrote no row. A node applying nothing while its position advances is healthy on lag alone.
crewlet.statelog.apply.retries1domainAttempts the apply loop retried in place after a failure that was not a stop — a fetch the broker did not answer, a transaction the disk refused, an applier that errored. A rate that stays up past the retry budget is a node whose rows have stopped moving, and its health says so.
crewlet.statelog.apply.tx.aborts1domainApply transactions whose body the store ran more than once, because an attempt failed transiently after it began. It reads zero by construction: this database detects write conflicts per file, so every write transaction takes the file’s lock at its BEGIN and queues for it, and no commit elsewhere in the file can abort an apply. A non-zero count on the operator’s own hardware means the retry budget is being spent rather than held in reserve.
crewlet.statelog.records_gated1gate, subject_kindRecords an apply gate dropped, by the gate that dropped each — the domain’s own (evicted, deleted) or the framework’s (abandoned, overtaken). Each node counts a record once, where its own applier drops it, when the transaction that drops it commits — never once per attempt at that transaction. A write refused because its record — or another node’s copy it was collapsed onto — was dropped is a refusal, counted under crewlet.statelog.publish.refusals, and not a second drop. A dropped commit is recoverable by nothing, and this is the only place anyone would see that it happened.
crewlet.history.answers1question, coverageWhat each fleet history read covered: complete when every live node answered inside the fleet read budget, partial when one did not. Turn-level detail lives only on the node that published it, so a partial answer is missing that node’s rows — and the history_partial alarm is a fraction of this counter.
crewlet.tracker.search.answers1coverage, semanticWhat each answer actually covered: whether every bucket of the corpus was scanned, and whether the semantic half ran (full), was asked for and lost (skipped), or was never asked for (off — a keyword search, or a company with no embeddings). Both alarms below it are a FRACTION of this counter, and a short answer is indistinguishable from a short corpus without it.
crewlet.tracker.feed.unreadable1sourceChange records this build could not translate into a wake. Both domain consumers are deliberately uncapped, so such a record is never dropped — it redelivers for ever at the head of the consumer with every wake behind it waiting, which has no other symptom at all.
crewlet.tracker.bulk.calls1resultHow often a bulk edit is issued, by result: admitted under the fleet’s bulk lease, refused because another is applying, and fail_open — let through with NO lease, because the coordination store did not answer the admission or the mixed-version gate refused it. The refusal arithmetic rested on an assumed ten a day, a number with no counter behind it; this is that number. A fail_open is a bulk the fleet-wide bound did not cover, and counted as admitted it left no trace: a store that refused every admission looked exactly like a fleet whose bulks never collided.
crewlet.tracker.bulk.apply_secondss—Seconds of applier occupancy bulk edits projected, summed. Over 24 hours it IS the fleet-wide read-degradation budget: every second here is a second in which reads are behind and writes are pending on every node.

Rates. Every counter is a monotonic total and your collector divides. A rate computed in this process would be a rate over a window nobody chose, disagreeing with the one on your dashboard.

A /metrics route. OTLP reaches Prometheus through the collector you already run for traces. A second wire format would be a second thing to authenticate on an API whose guard and whose exemptions are both load-bearing.

Per-object series. See the attribute rule above.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.