Metrics
Generated from metrics.Catalogue(). Do not edit — change the catalogue.
The engine exports OpenTelemetry metrics through the same OTEL_* environment
its traces use, because a company has one telemetry backend and two settings
would let you split signals across two of them where a correlation resolves on
neither:
| Variable | What it does |
|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT | The collector. Metrics are posted to <endpoint>/v1/metrics. |
OTEL_EXPORTER_OTLP_METRICS_ENDPOINT | Overrides the base for metrics alone. |
OTEL_EXPORTER_OTLP_PROTOCOL | http/protobuf (default) or grpc. |
OTEL_EXPORTER_OTLP_HEADERS | Sent with every export — an API key, a tenant. |
OTEL_METRICS_EXPORTER | none switches the export off and leaves traces on. |
OTEL_METRIC_EXPORT_INTERVAL | Export period in milliseconds. Default 60000. |
The provider is installed whether or not you are collecting. Every
instrument exists either way, so no code path in the engine branches on
whether metrics are “on”, and the numbers on crewlet retention status are
read from the same recorder a collector would be reading. With no endpoint set
the measurements are taken and not shipped.
Every attribute below is a closed set. No metric carries a task key, a subject id, a page title or a node id as an attribute — the node is the resource, set once — so the number of time series is bounded by this page rather than by how much work your company has done.
Histograms
Section titled “Histograms”A distribution, exported with the engine’s own bucket boundaries: twenty-one powers of two from 64 µs to 64 s, which resolves a percentile to within a factor of two at every scale here — from a 40 µs index probe to a 16-second bulk apply. The SDK’s default boundaries stop at 10 s, so a slow apply would land in an overflow bucket and read as “at least 10 s” for ever.
| Metric | Unit | Attributes | What it makes visible |
|---|---|---|---|
crewlet.statelog.publish.duration | ms | domain, outcome | A write path slowing down before it starts refusing. The outcome dimension separates the three answers a write has, so a rise in pending reads as an applier falling behind rather than as a broker getting slower. |
crewlet.statelog.publish.rounds | 1 | domain | Contention on one subject, which the compare-and-set round cap bounds and nothing measured. A distribution creeping toward the cap is a hot object; reaching it is a refusal a model reads as a colleague editing the same thing. |
crewlet.statelog.write.session_wait | ms | domain | The wait a write pays for its own previous write to apply. Zero when the caller is caught up, which is the common case, and the shape of the load that is not. |
crewlet.statelog.barrier.duration | ms | domain | The broker round trip under every linearizable read, and the first number a drifting fsync or a degrading quorum moves. It was a benchmark’s p50 on an idle loopback cluster and nothing in production. |
crewlet.statelog.read.wait | ms | domain, level | How much of the read budget a barrier or session wait actually spends. A p95 approaching the budget is reads about to start refusing, which is the warning the refusal itself is too late to be. |
crewlet.statelog.apply.latency | ms | domain | THE COMMIT-TO-APPLY GAP: from the broker’s own timestamp on a record to this node committing it. Every read level is a policy about this quantity and nothing measured it. |
crewlet.statelog.apply.tx.duration | ms | domain, bound_by | How long one apply transaction holds the store’s writer, and which budget ended it. A transaction is what every waiter behind it pays, and rows were only ever a proxy for the duration. |
crewlet.statelog.apply.record.duration | ms | domain, kind | One record’s apply. A single record past the time budget is still one transaction, so this is the real ceiling on how long a read can be delayed — a sentence in a design document until it was measured. |
crewlet.statelog.apply.batch.rows | 1 | domain | Rows per apply transaction, which is what the row budget bounds and what the drain rate divides. |
crewlet.backup.duration | ms | — | How long a backup took, which is the window the trim hold covers and the I/O the copy spends competing with the applier’s own commits. It is what turns the retention guide’s worked example into a number for THIS hardware. |
crewlet.store.pool.wait | ms | file | How long a reader waited for a connection. It is what says the reader pool is too small on this node, which nothing could say before. |
crewlet.tracker.search.scan.duration | ms | path, rung | One search’s scan, split by whether it ran for a turn’s prefetch or for somebody’s deliberate search, and by the ranking it actually served (rung: hybrid, keyword, semantic, or none). Only the prefetch had a published percentile, and the interactive path is the one with a target; a keyword scan is a fraction of a fused one’s cost, so reading them as one series hides a slow meaning half. |
Gauges
Section titled “Gauges”A value that goes both ways, sampled at each export.
| Metric | Unit | Attributes | What it makes visible |
|---|---|---|---|
crewlet.statelog.drain.rows_per_second | 1 | domain | The applier’s observed drain, which every retry hint divides by. Seeded from a benchmark and then measured, so a hint on real hardware stops being an extrapolation from somebody else’s. |
crewlet.statelog.drain.commits_per_second | 1 | domain | Commits per second, which is the fsync rate under synchronous = FULL and the number a device budget is spent by. |
crewlet.statelog.apply.lag.seq | 1 | domain | How many records this node is behind the log’s head. |
crewlet.statelog.apply.lag.seconds | s | domain | How OLD the oldest unapplied record is. Seconds are what a stall grace, a pending outcome and a copy taken out of service all turn on; sequences are not, and a lag of 4 000 says nothing about whether anything is wrong. |
crewlet.statelog.applied_through | 1 | domain | The prefix this node has actually applied, which is lower than its checkpoint whenever a record was retained. |
crewlet.statelog.deferred.count | 1 | domain | Records this build could not read and kept. Non-zero is a rolling upgrade in progress; non-zero and not falling is one that stopped. |
crewlet.statelog.deferred.oldest_age_seconds | s | domain | How long the oldest retained record has been retained, which is what decides whether this node’s copy stays in service. |
crewlet.statelog.waiters | 1 | domain | Callers blocked on the applier right now. It is the depth of the queue a slow apply is making. |
crewlet.statelog.log.bytes | By | domain | What the log actually holds, against its ceiling below. |
crewlet.statelog.log.max_bytes | By | domain | The ceiling, read from the running stream rather than from this node’s own configuration — the two differ, and the running one is what refuses the append. |
crewlet.statelog.log.headroom_fraction | 1 | domain | How much of the ceiling ordinary writes are held to is left — on the tracker and pages logs, the ceiling less the gate reserve kept above it for evictions. At zero the log refuses every write AND every linearizable read, and the remedy is a fleet-wide maintenance cycle, so this is the one number worth alarming on long before it is small. |
crewlet.statelog.trim.blocked_seconds | s | domain, term | How long one retention term has held the trim, named. A trim blocked for weeks is a log walking toward its ceiling with a cause an operator can act on. |
crewlet.backup.age | s | — | How long ago the newest COMPLETE backup finished, read from the manifests on disk rather than from a counter this process keeps. A counter records that a process believed it took a backup; the disk records that one exists, and they differ in exactly the cases the alarm is for — a copy deleted, a volume never mounted, a schedule pointing at a path nobody ships from. A node holding no manifest reports NO SERIES rather than a value: zero is the freshest backup imaginable and every other number is an age nobody measured, so neither can stand for an absence. The absence is what the backup_age alarm says in words, and an absent() rule over this gauge is what a collector alerts on. |
crewlet.backup.holds | 1 | — | Live trim holds. A pin that outlives its owner stops the trim until the stale bound expires it, so a count that does not return to zero is a backup that crashed mid-copy. |
crewlet.store.wal.bytes | By | file | A write-ahead log a checkpoint cannot pass grows, and this is the only way to see it before the volume fills. |
crewlet.store.bytes | By | file | The store’s size on disk, which the snapshot’s free-space precondition and the provisioning rule are both derived from. |
crewlet.tracker.search.concurrency | 1 | — | Searches in flight, which is the row of the supported-corpus table this node is actually on: inside the one-second budget, about 345 000 sources with one running and about 136 000 with eight through the full scan, and about 545 000 and 183 000 through an index probing half its lists. |
crewlet.tracker.vector.coverage | 1 | — | The fraction of sources carrying a current vector. It is how a stalled embedding backlog is reported, since it never drops a seat. |
crewlet.alarm.active | 1 | kind | Whether each named alarm is firing right now, 0 or 1. It is the same table the operator record renders and the CLI exits non-zero on, so a collector and a person see one answer. |
Counters
Section titled “Counters”Only ever rises. Rates and totals are your collector’s arithmetic, never this engine’s.
| Metric | Unit | Attributes | What it makes visible |
|---|---|---|---|
crewlet.statelog.publish.outcomes | 1 | domain, outcome | The three-valued write outcome, counted, and only the three — a refusal is on publish.refusals instead, because it says the write never happened at all. unknown is the one that matters most and had no counter: a broker flapping into ambiguity was visible only to the model that received the answer. |
crewlet.statelog.publish.conflicts | 1 | domain, subject_kind | Writes that spent their whole round budget losing races on one subject, BY KIND. The refusal counter beside it says a conflict happened and not what it was about, and the remedy differs entirely: one contended object is a design question and a contended kind is a hot subject. |
crewlet.statelog.publish.rejections | 1 | domain, subject_kind | How often a write loses a race, per kind of subject. It is what says whether a counter, a rank order or an ordinary object is the contended one. |
crewlet.statelog.publish.refusals | 1 | domain, reason | Writes refused before or instead of an append, by reason, and EVERY value it carries is here, each with its remedy. The node: evicted (this node is removed from the fleet — run the write on another), deferred (it holds a record it cannot decode — a newer build serves it), behind (it has not applied a position the write needs — clears on its own), below_floor (it is below the log and must adopt a snapshot — another node writes meanwhile), floor_unknown (the floor or the log’s ends could not be read — clears when coordination answers). The log: log_full (at its ceiling — raise it or unblock the trim), skew (a store and a stream restored out of step — restore both from one backup), wrong_stream (not the log this node’s rows came from: rebuilt, ending below its checkpoint, holding another record there, or re-anchored past it by a peer — re-anchor, or adopt where a peer holds the generation), log_truncated (it lost records a peer’s rows hold — settle which history the fleet keeps). The object: deleted (a permanent deletion marker — nothing recreates it). The operation: op_reused (its id names a record on another object) and superseded (its record landed and a later one undid it) — both answered by a NEW operation under a fresh id, never by a retry. A record that landed and a gate dropped is counted under the gate that dropped it — evicted, deleted, abandoned (written in a generation a reanchor abandoned) or overtaken (written after a restored reanchor, below its generation) — and is never re-decided, because republishing makes another record nothing applies. Such a record holds its operation id for the log’s duplicate window, so the same id sent inside it, by any node, is collapsed onto the record and refused the same way: the write is retried under a fresh id, or under the same one once the window has passed. Beside the refusals: conflict for a write that lost every round, exists for a create whose object already exists, and error for a write that failed before it could answer. A refusal is not one of the three outcomes: it says the write never happened, and each reason has a different remedy, so one counter with an outcome dimension would hide them all. |
crewlet.statelog.read.refusals | 1 | domain, level, code | Every refusal code, counted. The codes, each with its own remedy, had no counter between them, so an operator had no rejection rate for any of them. |
crewlet.statelog.read.served | 1 | domain, level | Reads answered per level, which is the denominator every refusal fraction needs and the check on the assumed read rate the log’s own size is derived from. |
crewlet.statelog.barrier.appends | 1 | domain | Barrier records appended. Against reads served it is the single-flight ratio, which says whether coalescing is doing anything at all. |
crewlet.statelog.barrier.applied | 1 | domain, stream | Barrier records applied from each log — every node’s linearizable reads on it, since every node applies every record. It is the rate census_drift holds against the census, per log: a node’s own appends are only its share of the fleet’s reads, and on a fleet of several serving nodes they describe a company that many times quieter than it is. The alarm’s 24-hour window files each barrier at the hour the broker stored it, so a node replaying a backlog does not read days of reads as one; this cumulative series counts them as they are applied, so it jumps by a backlog’s barriers when one is replayed. |
crewlet.statelog.linger.yields | 1 | domain | How often a waiter cut a batch short. It is the batching the applier gives up to answer a read promptly, and without it that trade is invisible. |
crewlet.statelog.apply.records | 1 | domain, result | Records consumed, by what happened to them: applied, retained, gated, skipped, or reprocessed by a build that could read what an earlier one retained — a retained record a gate drops when it is reprocessed is counted gated, as it would have been live, since it wrote no row. A node applying nothing while its position advances is healthy on lag alone. |
crewlet.statelog.apply.retries | 1 | domain | Attempts the apply loop retried in place after a failure that was not a stop — a fetch the broker did not answer, a transaction the disk refused, an applier that errored. A rate that stays up past the retry budget is a node whose rows have stopped moving, and its health says so. |
crewlet.statelog.apply.tx.aborts | 1 | domain | Apply transactions whose body the store ran more than once, because an attempt failed transiently after it began. It reads zero by construction: this database detects write conflicts per file, so every write transaction takes the file’s lock at its BEGIN and queues for it, and no commit elsewhere in the file can abort an apply. A non-zero count on the operator’s own hardware means the retry budget is being spent rather than held in reserve. |
crewlet.statelog.records_gated | 1 | gate, subject_kind | Records an apply gate dropped, by the gate that dropped each — the domain’s own (evicted, deleted) or the framework’s (abandoned, overtaken). Each node counts a record once, where its own applier drops it, when the transaction that drops it commits — never once per attempt at that transaction. A write refused because its record — or another node’s copy it was collapsed onto — was dropped is a refusal, counted under crewlet.statelog.publish.refusals, and not a second drop. A dropped commit is recoverable by nothing, and this is the only place anyone would see that it happened. |
crewlet.history.answers | 1 | question, coverage | What each fleet history read covered: complete when every live node answered inside the fleet read budget, partial when one did not. Turn-level detail lives only on the node that published it, so a partial answer is missing that node’s rows — and the history_partial alarm is a fraction of this counter. |
crewlet.tracker.search.answers | 1 | coverage, semantic | What each answer actually covered: whether every bucket of the corpus was scanned, and whether the semantic half ran (full), was asked for and lost (skipped), or was never asked for (off — a keyword search, or a company with no embeddings). Both alarms below it are a FRACTION of this counter, and a short answer is indistinguishable from a short corpus without it. |
crewlet.tracker.feed.unreadable | 1 | source | Change records this build could not translate into a wake. Both domain consumers are deliberately uncapped, so such a record is never dropped — it redelivers for ever at the head of the consumer with every wake behind it waiting, which has no other symptom at all. |
crewlet.tracker.bulk.calls | 1 | result | How often a bulk edit is issued, by result: admitted under the fleet’s bulk lease, refused because another is applying, and fail_open — let through with NO lease, because the coordination store did not answer the admission or the mixed-version gate refused it. The refusal arithmetic rested on an assumed ten a day, a number with no counter behind it; this is that number. A fail_open is a bulk the fleet-wide bound did not cover, and counted as admitted it left no trace: a store that refused every admission looked exactly like a fleet whose bulks never collided. |
crewlet.tracker.bulk.apply_seconds | s | — | Seconds of applier occupancy bulk edits projected, summed. Over 24 hours it IS the fleet-wide read-degradation budget: every second here is a second in which reads are behind and writes are pending on every node. |
What is deliberately not here
Section titled “What is deliberately not here”Rates. Every counter is a monotonic total and your collector divides. A rate computed in this process would be a rate over a window nobody chose, disagreeing with the one on your dashboard.
A /metrics route. OTLP reaches Prometheus through the collector you
already run for traces. A second wire format would be a second thing to
authenticate on an API whose guard and whose exemptions are both load-bearing.
Per-object series. See the attribute rule above.
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.