Skip to content
You are reading documentation for unreleased main. This page is not in 0.1 yet.

Backups & Restore

What durable state a Crewlet deployment holds, where each piece lives, how to back all of it up so a later restore actually works, and which losses are survivable without one.

The short version: crewlet backup takes a verified copy of a running node, without stopping it. It goes through the engine because it has to — the store file is locked to that process and the embedded broker binds no socket, so nothing outside the engine can read either estate. The cold runbook further down remains the belt to that braces.

Terminal window
crewlet backup -dir /var/backups/crewlet/2026-08-30T18-00
Backup written to /var/backups/crewlet/2026-08-30T18-00 on node-0 in 1.412s
WHAT FILE SIZE CONTENTS
store (node) store.db 252.0 KiB 20 migrations
store (replicated) store-replicated.db 1.2 MiB 3 migrations
stream CREWLET_AGENT streams/CREWLET_AGENT.snapshot 1.1 KiB 5 messages
bucket crewlet_token_windows streams/KV_crewlet_token_windows.snapshot 512 B 3 messages
…

The path is on the engine’s host, not yours — this writes files where the node runs and downloads nothing. Any node produces one, whatever its roles: every node holds its own store and, on the embedded topology, its own broker. A node that dialled an external NATS cluster is the exception, and it says so: its copy carries the store estates alone, and the stream half is backed up at the cluster (nats account backup) from the same moment.

A deployment’s durable state lives in six places:

EstateWhereWhat it holds
The node’s own store filestore.path, with its -wal sidecarThe seat’s memory — diary, episodes, counterparty profiles, synthesized skills, onboarding markers, the conversation ledger — which is also replicated onto the stream, so this file is a cache of it rather than its only copy; and, held here only: the audit event log (30 days), scheduled-run history, the company-config revision history, the secret store’s bootstrap rows, and this node’s own record of any snapshot it has adopted
The replicated estatestore.replicated_path, with its -wal sidecarEverything a state log’s applier derives from the fleet’s own records — the work tracker, the knowledge base’s pages, the embeddings that search them and every node’s daily usage — together with the checkpoint that says how far this node has applied. Derivable by replay only while the log still holds the records: past the trim floor, a node with no copy of this file adopts a peer’s snapshot instead
The object storestore.objects: the OBJ_crewlet_files stream on the broker’s members (nats, the default), or an S3-compatible bucket (s3)The bytes of the company’s files, one object per upload. One store the whole fleet shares, not a directory of any node’s: on nats a stream like any other, on s3 somebody else’s bucket. The rows naming the files are in the replicated estate
The stream estatestream.store_dir per embedded member, or the external NATS clusterAgent mailboxes (unacked in-flight work), the shared event and config streams, one ordered log per state-log domain — which is the record of truth the file above is derived from — and every coordination KV bucket: seat, presence and duty leases and fencing epochs, the activation pointer with the current company payload, the completion ledger, delivery dedupe, budget counters, scheduled-fire claims, detached sandbox-run records, the sealed credentials
Tier A, on diskcrewlet.yaml and the environment it readsThe keyring (CREWLET_SECRET_KEY_*) — the sole root of trust for everything sealed — plus API tokens and any NATS credential/TLS files
cli-agent homesPer-seat state directories on the engine hostSubscription CLI logins (portable via crewlet llm export)

Classify before you size the job:

  • Rebuildable, safe to lose: every TTL’d coordination bucket — the rate valve, delivery dedupe and node status regenerate, credential cooldowns re-learn at the cost of some rate-limit errors, and the rebase records re-form at the next attempt, at the cost that a turn whose work began more than twenty-nine days ago, retried after the loss, writes again what its earlier attempt wrote — and the leases and epochs provided the whole fleet cold-starts together (they re-form from nothing).
  • Held twice: the learning tables and the conversation ledger. Each row is also on the memory changelog, which is how a seat’s memory follows it to a new node — so a store file lost with the stream estate intact costs at most the last sync cycle, and the seat re-hydrates the rest on its next acquisition.
  • Kept by the store, as many copies as it keeps: the object store’s objects. On nats they are a stream at stream.replicas copies, so they survive exactly what the logs survive, and a member that fails is caught up by the broker; on s3 they have whatever durability the bucket’s provider promises. An object the store has lost or changed anyway is gone, and the collector’s daily audit raises objects_missing naming the file. Only a backup covers that case.
  • Derived, and rebuildable only within the replay window: everything in the replicated estate. A node that loses that file replays the domain logs from the beginning and arrives at exactly the same rows — but only if the logs still hold them. Past the trim floor the records are gone, and the node fetches a peer’s verified snapshot instead, which it does automatically at boot. That fallback needs a peer: a single-node company that loses this file and whose logs have been trimmed has lost the trimmed history, which is the case retention.backup_max_age exists to keep from arising.
  • Authoritative, with no other copy: the event log’s history, the config revision history, the sealed credential bucket, the budget counters, and each detached sandbox-run record, which is the only thing that knows a billed box exists.

One directory, and the manifest is the claim: a directory holding manifest.json is a complete backup, one without it is the debris of a run that did not finish. Nothing else in the directory says so, which is exactly why the manifest is written last.

2026-08-30T18-00/
├── manifest.json what was captured, from which node
├── store.db the node estate, self-contained
├── store-replicated.db the replicated estate, self-contained
├── objects/ s3 only: every object that copy names,
│ └── files/ one file per object, under its key —
│ └── 0199a3c2-… the bucket's own layout under its prefix
└── streams/
├── CREWLET_AGENT.snapshot a mailbox stream
├── KV_crewlet_secrets.snapshot a coordination bucket
├── OBJ_crewlet_files.snapshot nats only: the company's file objects
└── … one per stream and bucket found

What is worth knowing about it:

  • Each store copy is taken with VACUUM INTO and then verified — reopened, integrity-checked, its schema compared against the database it came from, and a sha256 of the finished file recorded in the manifest — before it is renamed into place. A copy that will not open is a failed backup rather than a surprise on the worst day of the deployment’s life; the digest is what tells a copy that was truncated in transit from one that was bad when it was made. Each is self-contained: no -wal travels with it.
  • A node is two databases, and a backup carries both. The node estate holds the audit log, memory, the config revisions and the secret bootstrap; the replicated estate holds everything a state log’s applier writes. They are separate files because a snapshot for a joining node is a copy of the second one alone, and it must not carry the donor’s audit log or the bootstrap half of its secret store. Restoring one without the other gives a company whose halves are from different moments. The replicated estate is something the node holds — every node with the data role holds it from the moment it starts, whether or not its company runs the native tracker or knowledge base, because the rows its log derived stay on disk either way — and the backup copies what the node holds rather than whatever happens to be open. An estate the node holds that is closed at that instant (an adoption replacing its file, or one whose reopen failed) refuses the backup, naming it, rather than writing a manifest without it; take it again once the node reports the estate serving. A node without data holds none and backs up its own file alone.
  • The objects come from the object store, not from this node’s disk. On nats they are already a stream — OBJ_crewlet_files — and the stream snapshot the backup takes anyway carries them, so no object is copied on its own: the manifest’s objects.stream names the stream. Once the snapshot is taken, the backup asks the stream’s leader about every object the replicated copy names, and objects.objects and objects.bytes count the ones it holds — a key is never reused, so an object the store holds after the snapshot finished is one the snapshot holds too. On s3 the backup reads every object the replicated copy names — read from the copy itself, for the position’s reason below — from the bucket, four at a time so a backup never saturates the store every upload and download also uses, and streams each to disk, checking it against its file’s size and SHA-256 as it is written. Each lands under objects/files/<key> — the bucket’s own layout under its prefix — first under a temporary name, renamed only once it has checked out.
  • A backup takes what its previous one holds (s3). A key is minted for one upload and never names other bytes, so an object this node’s previous backup already holds is the same object: when that backup’s directory is still on this host (the dir this node last announced), each object in it is linked into the new backup — copied where the filesystem refuses a link, another mount among them — and read back against its file’s size and SHA-256 before it counts, so a copy that rotted in the earlier artefact is fetched again rather than carried forward. Only what is new since is read from the bucket. Each one is a complete file of the new backup, never a reference into the old one: shipping or deleting either directory leaves the other whole. The manifest’s objects.reused and objects.reused_from say how many and from where.
  • An object the store has lost is recorded, an object it could not answer for fails the backup. An object the store answered it does not hold, or holds at another size or under another digest than its file records, is gone whatever the backup does: the backup carries everything else, lists it in the manifest’s objects.lost as {"object", "named_by"} — the key and the file, PROJECT/path — logs backup_objects_lost, and is announced like any other. Refusing it would refuse every later backup too, and with them the trim of every log in the fleet, over one lost file. The objects_missing alarm is what sends somebody to replace the file. An object the store did not answer about may be intact there, so it fails the backup (POST /backup answers 503 objects_unreachable), naming it: take it again once the store answers. The manifest’s objects records dir or stream, objects, bytes, reused, reused_from and lost; a copy naming no file has none.
  • Streams are enumerated, not listed. A namespace stream is created on first publish and a coordination bucket’s name depends on a configurable prefix, so what gets captured is what is actually there.
  • Every estate or none. A failure anywhere leaves the directory without a manifest. A backup missing an estate is not a partial backup, it is an unrestorable one: the store alone loses every lease, ledger and credential; the streams alone lose this node’s audit log, its scheduled-run history and the config revision history — everything the store holds that the memory changelog does not carry.

It is a copy of a moment, not an instant — the engine keeps working throughout, and the pieces are separated by however long the copy took. The store is copied first, and that order is now mandatory rather than preferable.

The store carries a position: the tracker’s rows are derived from an ordered log by an applier that commits its checkpoint in the same transaction as the rows, so a store copy is a claim about what has already been applied.

  • Store first leaves the artefact holding a store at position P beside a log that has since moved past it. A restore replays the difference. The gap is bounded and replayable, and it costs a few minutes of work being applied twice — which is free, because the applier’s guard is monotone in the position. The replay starts from the restored rows, not from the broker’s record of what the node read: the node’s reader of each log was copied with the streams, after the store, and may already have acknowledged past P. The node rebuilds such a reader at P when it boots and logs jetstream_domain_consumer_rebuilt (see Replication).
  • Store last would leave a store at position P beside a log whose newest record is below it. Every subsequent record then lands at a sequence the store has already marked applied, and the version-guarded write drops it silently. That is not a gap, it is a permanent hole nothing reports — a restored company quietly missing whatever was written during the copy.

So the order trades a bounded, replayable gap against a permanent, silent one.

The gap is only replayable while the log still has the records, and the fleet’s own trim deletes a record once every counted node has committed past it. A backup is not a counted node, so two things close that window and they are different kinds of thing:

  • A trim hold is taken before the first byte is copied, at the position this node’s appliers stand at, and released when the manifest is written. It is what makes the race not happen. It is heartbeated: a pin that outlived its owner would stop the trim for ever and the log would grow to its ceiling, so the fleet ignores a hold nobody has renewed. A node that cannot write the hold refuses the backup rather than taking one whose gap may be trimmed away while it runs.
  • An assertion — the log’s first surviving sequence must be at or below the copy’s position plus one, checked per domain after the stream snapshots from bounds those snapshots already captured. It is what makes “restorable” a checkable inequality rather than a hope. A backup that cannot assert it writes no manifest, which is how a reader tells debris from a backup.

The manifest records the position read from the copy itself, not from the live database: the checkpoint commits with the rows, so the position inside a file is the only one that describes that file, and the applier ran throughout the copy.

Every backup taken over POST /backup is audited. Once the copy begins, the node writes a backup_requested event to its own event store — who asked (the token’s name, and the person it is bound to), which node, which directory, and whether it finished — a failed run included, since it can leave files behind. A directory refused before the copy began (relative, not empty, or one the host cannot create or write because of the path itself) wrote nothing, answers 400, and is not recorded. A disk that fails or fills while the directory is prepared is the node’s failure rather than the path’s: it answers 500 and is recorded, and so is a refusal that arrives once part of the copy is already in the directory. Filter the event log on source=operator to see it beside every other change a person made through the engine; see the runtime audit.

Settings › Backups & retention takes one too: Take a backup asks the node serving the page (POST /backup) for a directory on that node’s host — nothing is downloaded — offering a fresh directory beside that node’s last copy when it has one. The copy can take minutes on a large store; closing the dialog does not stop it, and the outcome arrives as a notice either way.

The same screen shows what the fleet has backed up (GET /backups): each owner’s newest copy — its directory, size, reach and whether the trim counts it, with the one the trim reads marked newest — and the backup history, every backup a person asked any node for over the event log’s 30 days, failures included and each with the host that holds it. The register keeps only each owner’s newest point, so the history is where a failed night shows.

crewlet backup runs against one node and copies what that node can reach — which is not the same as what that node owns.

The stream estate is the fleet’s, so any node’s copy of it is the whole company’s: the mailboxes, the leases and epochs, the activation pointer with the current company payload, the completion ledger, the budget counters, the sealed credentials — and, since a seat’s memory is replicated onto it, every seat’s diary, episodes, profiles, skills and conversation ledger, no matter which node wrote them. On a clustered embedded stream you are snapshotting a replicated stream, so one member’s snapshot carries what its peers hold too.

A node is two database files, and they answer differently.

The replicated estate is a copy of state every node holds: the tracker’s projects, tasks, comments and history, derived from an ordered log by an applier that runs identically everywhere. Any healthy node’s copy of it is the company’s, in the same sense the stream estate is.

The node estate is that node’s alone, and what only lives there is what only that node did: its audit event log, its scheduled-run history, its share of the config revision history. Those exist nowhere else and no peer’s backup contains them.

So:

  • For the company’s state — one node is enough. Everything a restore needs to bring the company back is on the stream estate and the replicated estate, and a single node’s backup captures both.
  • For the complete audit trail — take one per node, on the same schedule, and keep them together. The event log is per-node history with no second copy, so a fleet-wide audit trail is the union of every node’s.

There is no fleet-wide barrier and deliberately no attempt at one: each node’s copy is its own moment, and the estate that has to be internally consistent — the stream estate — is captured as one set within a single run.

The backup lands on the engine’s host, which is not a backup until it leaves that host: ship the directory to object storage or another machine as a second step. Schedule it the way you schedule anything else against a node (cron, a systemd timer, your orchestrator) — one directory per run, named by timestamp, since a destination that already holds something is refused rather than merged.

The schedule is yours, and so is its cost. Every run is a full copy, so the footprint is arithmetic rather than a judgement:

backup storage = copies retained x replicated estate bytes

A flat every 6 hours, retained 14 days is 56 copies, which at a year-five estate of ≈ 44 GB is ≈ 2.4 TB — plus a full VACUUM INTO of that estate four times a day on a live node, competing with the applier’s own commits for the same disk. That is a real cost nobody quotes, so the shipped guidance is tiered:

TierKeptCovers
every 6 hours8 copiesthe last two days, at the granularity an incident needs
daily14 copiesthe fortnight, at the granularity a discovered problem needs

Twenty copies rather than fifty-six — ≈ 875 GB at the same year-five estate, for a recovery point that is worse by nothing anybody has ever needed. crewlet.backup.duration is what measures the copy’s own cost against your hardware — the window the trim hold covers and the I/O the copy spends competing with the applier’s own commits; start from the table and move it once you have that number.

How stale is too stale is a separate setting. retention.backup_max_age is what the trim reads, and it is deliberately not derived from the schedule: a company that never backs up never trims, loudly and by design, so the engine has to know what “recent enough” means to you rather than inferring it from how often a cron happened to fire. The age itself is read from the newest complete manifest on disk rather than from a counter the engine keeps — a counter records that a process believed it took a backup, and the disk records that one exists. They differ in exactly the cases the alarm is for — a copy taken and then deleted, a volume never mounted, a schedule pointing at a path nobody ships from — and in every one of them the counter says the fleet is protected. It is published as crewlet.backup.age, by the node that holds the artefact, because that is the only process whose disk the manifests are on; a node that has never taken one publishes nothing rather than a zero, since zero is the freshest backup imaginable.

The backup interval IS the recovery point for history below the trim floor. Above that floor the log holds every record on R replicas and every node holds the applied rows, so losing a node loses nothing. Below it the log holds nothing, and each node’s own database file is the only copy of that history — N of them, independent, none replicated. A schedule of six hours is therefore a six-hour RPO for that half of the company’s past, and no replica count changes it. See Retention.

A finished backup announces itself to the fleet. When the manifest is written, the taker publishes what the copy reaches — per stream, with the generation — so the trim’s backup term can see it from whichever node holds the duty. That node is often not the one that took the copy, and can never see a directory on another host. A backup that ran and announced nothing leaves a fleet with a working nightly schedule whose log grows for ever, so a failure to announce is logged rather than silent — and it is bounded and self-correcting: the log keeps a longer window than it needed to, and the next backup announces again.

The announcement, the trim hold and the manifest’s node_id are all keyed on the node’s resolved id: node.id when the file sets it, else CREWLET_NODE_ID, else the default. A node named only through the variable is the shape a container orchestrator runs, and keying on the raw field would give every such node one shared, blank key.

A pin that outlives its owner is visible. A backup takes a trim hold before the first byte is copied and releases it when the copy ends; crewlet.backup.holds is how many the fleet is carrying, and a count that does not return to zero is a backup that crashed mid-copy. Until the stale bound expires that pin the trim does not advance, which has no other symptom at all until the log walks into its ceiling.

Name the owner. retention.backup_owner is free text — a person, a team, a scheduler’s name — and crewlet backup records it in the manifest, which is where it stops being configuration and becomes durable evidence of who was responsible for the artefact somebody is now restoring. crewlet validate warns when it is unset.

Two rules carry over from the cold runbook and are worth repeating because this path makes them easier to forget: the directory holds every credential the company has, so treat it exactly as you treat the secret store; and the keyring must not travel with it, or the sealing is undone.

Still the belt to the online path’s braces, and the one whose restore exercises no recovery code at all. Use it when you want a copy that involves no running engine — before a risky upgrade, or when taking the deployment down anyway.

  1. Drain and stop every node — SIGTERM or Ctrl+C once, and let the drain converge; see graceful shutdown.
  2. Copy, per node: both store files — store.path and, on a node with the data role, store.replicated_path — each together with its -wal sidecar — committed data lives in both, while the -shm and .lock sidecars are transient — and stream.store_dir for every embedded member. Copying them all out of one instant is what keeps the node’s local state, the replicated estate and the fleet’s shared state telling one story. On the default nats object store the company’s files are a stream, so stream.store_dir already carries them.
  3. Copy the bucket, on s3: the objects under the configured prefix’s files/, with the provider’s own tool (aws s3 sync s3://acme-files/crewlet/files/ objects/files/). Nothing writes them while the fleet is stopped, so this copy is of the same instant as the rest; a bucket left out is a backup whose rows name files it cannot restore.
  4. Copy Tier A: crewlet.yaml and any NATS credential/TLS files it names — and record where the keyring material comes from. Keep the keyring out of the data’s backup domain (Secret Store § Backups): a backup that carries both the ciphertext and its key has undone the sealing.
  5. Export subscription logins, if any seats run on a coding CLI: crewlet llm export <key> packs each into one portable bundle.
  6. Start the fleet again.

A stream.store_dir left empty selects an in-memory stream server: nothing survives a restart and there is nothing to back up. Set it before backups are worth discussing at all.

Restore is an operator procedure against a stopped fleet, not a command: every hazard below is about ordering and identity, and a tool that hid them behind one verb would be hiding exactly what has to be got right. What crewlet backup produces is what these steps move.

The store half is two file copies: put store.db at the node’s store.path and store-replicated.db at its store.replicated_path (by default crewlet-replicated.db beside store.path), with no -wal beside either — each copy is self-contained, and a stale sidecar from the old database is the one thing that would corrupt it. Both, from the same backup set: they are one node’s state, and a restore holding one of them has an audit log and a tracker from different moments. The object half depends on the backend. On nats there is nothing to do apart from the streams: the objects are the OBJ_crewlet_files snapshot and restore with every other stream below. On s3 the bucket is not part of the fleet and usually needs nothing at all — it kept its own copies while the fleet was down. If it lost data, or the company is restoring into a new bucket, copy the backup’s objects/ directory into the bucket under the configured prefix: it is laid out as the bucket is, each object at files/<key>, and a key never names two different objects, so it is one sync, merged into whatever is there —

Terminal window
aws s3 sync objects/ s3://acme-files/crewlet/

— with the bucket and prefix from store.objects.s3, and the provider’s own tool or --endpoint-url for a store that is not Amazon’s. An object the restored rows name and the store does not hold is reported by the collector’s next audit (objects_missing); crewlet objects status lists the files. The stream half is restored into a broker with nats stream restore per snapshot for an external cluster; for the embedded topology, restore into a fresh stream.store_dir on a node started for that purpose. Then:

  • Restore whole estates together, then cold-start the whole fleet. The fencing epochs and the activation pointer must never move backwards while any live node remembers newer values — gaps in those counters are harmless, resets are not. A KV estate restored under a running fleet hands out epochs that live leaseholders outrank. Every node down → restore store files and the stream estate from the same backup set → start everything.
  • Keep node identity. A clustered embedded member’s replicas are placed by server name, which is the node’s resolved id (node.id, else CREWLET_NODE_ID): a node restored under a fresh name is a new peer, its old replicas are orphaned, and the stream sits short of quorum waiting for a server that will never return.
  • Expect bounded duplicates, not loss. Mailboxes hold exactly the unacked backlog, and a restored completion ledger and dedupe window are older than the outside world — so some already-handled triggers re-run. That is the same at-least-once posture the engine holds after any crash. The reverse skew is the one to avoid: a ledger newer than the mailboxes it acquits writes off work that never ran, which is why the embedded topology’s one-directory, one-instant copy is the paved path. On an external NATS cluster, nats stream backup / nats account backup against the cluster are the equivalent, taken across the streams and KV buckets as one session.
  • Config converges on its own. The current revision’s payload rides the coordination store beside the activation pointer, so a node restored with a stale store picks up the live revision; crewlet config export from any running node round-trips the document, sealed or not.
  • A store file lost with the stream estate intact is nearly free. The learning tables and the conversation ledger re-hydrate from the memory changelog when the seat is next acquired, so what is actually lost is that node’s audit log, its scheduled-run history and its config revision history. Start the node with an empty store; it migrates fresh and its seats arrive remembering. A lost replicated database is re-derived as above describes, and the node rebuilds its reader of each log at the checkpoint its new rows hold, so the records its old reader acknowledged are delivered again (see Replication).
  • Total loss of the stream estate without a backup is survivable by re-provisioning — secrets resolve store-first-env-second so a brand-new node starts from the environment, and every stream, bucket and mailbox is created idempotently at boot — at the price of the non-rebuildables above: the token counters lose the current day’s, week’s and month’s spend (every capped window re-arms silently — a company or seat that had spent its ceiling gets the whole allowance back until that window turns over), sandbox-run records vanish (a billed box leaks until its own TTL), and the completion ledger forgets (bounded duplicate turns).
  • Do not copy the store file while the engine runs. A live WAL database copied mid-write is a torn copy; the engine’s exclusive lock and the driver’s one-process rule exist precisely because there is no safe second opener. crewlet backup is the supported way to copy a running node, and it works by asking the engine to copy its own database.
  • Do not treat a directory without a manifest.json as a backup. It is an attempt that did not finish, and the manifest’s absence is the only thing that says so.
  • Do not point sqlite3, Litestream, or any other SQLite tooling at the live file. The file format is SQLite’s, but the live coordination is not: the store’s engine does not support mixed-tool multi-process access. Reading a cold copy with sqlite3 is fine; writing one is not a supported path back.
  • Do not restore a KV estate into a running fleet — the epoch rewind above.
  • Do not back the keyring up beside the data it seals.

Filesystem snapshots — the other online option

Section titled “Filesystem snapshots — the other online option”

An atomic volume or filesystem snapshot (LVM, ZFS, btrfs, EBS) that captures the store file, its -wal, and stream.store_dir at one instant is a crash image: restoring it recovers exactly as if the node had lost power at that moment — WAL replay on the store, unacked work redelivered from the mailboxes. A non-atomic copy of a live tree is not this, and gets no such guarantee.

Treat snapshots as defense in depth rather than the copy you must be able to trust: restoring one is a crash recovery, which is the least-proven surface of a pre-1.0 database engine, and the store’s own vendor recommends keeping independent backups. crewlet backup is the copy to trust — it is verified at the moment it is taken — and the cold runbook is the one whose restore exercises no recovery code at all.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.