Scheduling
The Scheduler (internal/schedule) lets an agent — or a whole unit —
own recurring work without an external cron emitting webhooks into
the engine. A QA Engineer can run a smoke-test pipeline every morning, a
team can hold an async standup at 9:30, a Knowledge-Base agent can audit
Confluence weekly — all declared in the YAML org config.
Schedules are a first-class part of the org model: they hang off a
role (a roles: entry’s schedules) or a unit (a units: entry’s
schedules), share one Schedule model
(org.Schedule), and change live with the rest of the organization when a
config revision is applied.
How it works
Section titled “How it works”A single loop ticks on a short interval (default 10s). Each tick:
- Enumerates every
Scheduleacross the current organization (all roles and all units, read fresh each tick, so an applied revision takes effect on the next tick with nothing to rewire). - Works out which schedules are due since the previous tick.
- Resolves the runner(s) for each due fire.
- Claims the fire in the fleet’s
firesslot (at-most-once) and publishes aTaskAssignedto each runner’s inbox topic (crewlet.agent.{handle}.inbox).
Because it reuses the existing TaskAssigned event, the agent runtime
path is unchanged — a scheduled turn runs executor → reviewer like
any other, and the full learning loop (episodes, diary, reflection)
runs on it. Periodic work is more worth learning from, not less.
The Schedule model
Section titled “The Schedule model”schedules: - name: morning-smoke # unique within the role/unit; part of the idempotency key cron: "0 9 * * 1-5" # standard 5-field cron (see below) timezone: Europe/Amsterdam # IANA tz; falls back to the company's `timezone` task: "Run the smoke-test pipeline and triage failures" target: each # unit schedules only — each | lead enabled: true # set false to keep it in config without firing timeout_seconds: 180 # hard wall-clock cap on the turn (default 180 s — # roomy for a multi-tool turn, stops runaway loops) catchup: true # fire one recent missed tick on restart (see Catchup)| Field | Default | Meaning |
|---|---|---|
name | — (required) | Identifier, unique within the role/unit. Renaming lets a same-minute fire re-run once. It is also what a fire is labelled by — the feed’s line and the turn’s label read swe was assigned scheduled work morning-smoke — so every fire of one schedule reads alike, and is recalled alike; the fire’s own run id stays in the event’s detail. |
cron | — (required) | 5-field cron expression, evaluated in timezone. |
task | — (required) | The task prompt handed to the runner agent. |
timezone | the company’s timezone | IANA timezone the cron is evaluated in. |
target | each | Unit schedules only. Who runs it (see Delivery). Ignored for role schedules. Human seats never run schedules: each fans out to direct agent roles only, an enabled lead schedule under a (possibly inherited) human lead is a config error, and human seats cannot define role schedules. |
enabled | true | false keeps the schedule in config but never fires it. |
timeout_seconds | 180 | Hard wall-clock cap on the scheduled turn. |
catchup | true | Whether to fire a recent missed tick on (re)start. |
Config load (and so crewlet validate) checks each schedule’s shape: a
non-empty name and task, a cron with exactly five fields, a
timezone that loads and is not Local or localtime, a non-negative
timeout_seconds, a target of each or lead, and names unique within
the owning role or unit. A written document is also held to the cron
grammar — the same parser the scheduler evaluates it with — so an
expression with five fields and an invalid value (61 * * * *,
0 9 * * MON-FRY) is refused at that schedule’s cron, naming the field
and the value. The grammar is an
admission rule: a
revision being applied that carries one (a newer peer’s, admitted under other
rules) still applies, its node logs org_admission_warning, and that one
schedule is skipped on every tick with schedule_parse_failed naming it while
the rest of the company runs.
Which clock a schedule fires on
Section titled “Which clock a schedule fires on”A schedule that names no timezone fires on the company’s clock — the
top-level timezone
of the company document, UTC when absent — which is the same clock every
“today”, due band and overdue mark in the tracker is cut on. So a
0 9 * * 1-5 standup on a Berlin company fires at 09:00 in Berlin, on the day
the board calls today. There is no scheduler-wide default zone: one would be a
second company clock, and a standup that fired at 09:00 on one clock while the
board cut its days on another disagreed with it about which day that 09:00 was
on.
A schedule’s own timezone is that one piece of work’s wall clock — the
Tokyo team’s 09:30 standup on a Berlin company — and nothing else is cut on
it. Local and localtime are refused for both: each is whatever zone the
host reading it is set to, and the scheduler is a singleton
duty that moves between hosts, so the
standup would move with it.
The company’s clock is read at every tick, like the org is, so an apply
that changes timezone moves the next tick’s fires without re-arming the
loop.
Cron syntax
Section titled “Cron syntax”Standard five fields — minute hour day-of-month month day-of-week —
with *, ranges (1-5), steps (*/15, 0-30/10), lists (9,17), and
three-letter month / day names (JAN, MON). Day-of-week accepts 0-7
where both 0 and 7 are Sunday. Vixie semantics: when both
day-of-month and day-of-week are restricted, a day matches if either
matches.
"30 9 * * 1-5" 09:30, Monday–Friday"0 9,17 * * *" 09:00 and 17:00 every day"*/15 * * * *" every 15 minutes"0 14 * * 5" 14:00 every Friday"0 2 1 * *" 02:00 on the 1st of each monthRole vs unit schedules
Section titled “Role vs unit schedules”units: - name: Backend type: team lead: Backend Lead channel: C_BACKEND schedules: # target defaults to `each` → every direct member posts their own update - name: daily-standup cron: "30 9 * * 1-5" timezone: Europe/Amsterdam task: "Post your standup: shipped yesterday / on today / blockers." # opt-in lead-coordinated variant (gather + summarize) - name: weekly-report cron: "0 16 * * 5" timezone: Europe/Amsterdam target: lead task: "Collect the week's progress from your reports and post a summary to the team channel." roles: - name: Backend Lead - name: Backend Dev schedules: # role-scoped — runs as this role - name: morning-smoke cron: "0 9 * * 1-5" task: "Run the smoke-test pipeline and triage failures"A role schedule always runs as that role (its target is ignored). A
unit schedule resolves its runner(s) from target.
Runners are resolved from the org, never from the agents running in
the ticking process. A fire is addressed to the runner seat’s inbox and
consumed by whichever node owns that seat — which is rarely the node
whose tick won the ledger claim. The seat’s agent id comes from
org.Organization.AgentIDFor, the same UUIDv5 over (company name, handle)
every node derives, so the TaskAssigned a scheduler publishes names
exactly the identity the turn will run under.
Schedules are not inherited. Unlike lead and channel, a
schedule on a department does not cascade to child units — that
would silently multiply a standup across every squad. Declare schedules
explicitly on each unit that should run them.
Delivery: who runs it
Section titled “Delivery: who runs it”target | Runner(s) | Use it for |
|---|---|---|
each (default) | every direct role of the unit | async standups, “everyone files their own status” |
lead | the unit’s effective lead (org.Organization.EffectiveLead) | gather-and-post standups, weekly reports, “review overnight Jira for the team” |
These are the only two unit-schedule targets. There is intentionally no
“specific role” target: a static role pin is exactly what a role
schedule is, so to run a recurring task as one person, put a schedule
on that role. lead is kept because it’s dynamic — it follows whoever
leads the unit (including inherited leads), so the schedule survives a
leadership change untouched.
each fans out one independent task per member — each with its own
dedup identity — so a slow or failing member never blocks the others.
each resolves direct roles only, never descendant units.
A standup is just a scheduled task: the runner does any fan-in with the
colleague tools it already has (a2a_ask, Slack reads, the roster). There
is no dedicated “standup” primitive — the same shape covers retros, sprint
kickoffs, demo-prep, on-call handoffs, and manager 1:1s
(an each schedule where each report opens a private A2A review with its
manager).
Guarantees
Section titled “Guarantees”At-most-once
Section titled “At-most-once”Every fire is claimed in the coordination store’s
fires slot before publishing — a first-writer-wins key per dispatch
identity, so a restart, a slow tick, a re-evaluated minute, or the
scheduler duty moving to another node can never fire the same run
twice.
scope_type · scope_id · schedule_name · fire_label · target_handleThe claim is shared rather than node-local because the scheduler is a singleton duty and therefore moves: a lease lapse, a drain or a rolling upgrade hands the tick to a peer, and a peer reading its own database would find an empty ledger and let its catchup pass re-fire everything the previous holder claimed.
fire_label is the schedule’s local wall-clock stamp
(YYYYmmddTHHMM in its timezone), not the UTC instant — so the dedup is
DST-correct (see DST & cron edge cases). The
identity is five separate components, and the key they are rendered into
escapes each one, so a delimiter inside a unit or schedule name cannot
make two different fires claim one key.
The claim fails closed: a coordination store that cannot be read yields no dispatch and the tick is retried on the next one. That is the opposite polarity to the completion ledger, deliberately — that one asks “has this work been done”, whose safe answer is to re-run, while this one asks “may I start”, whose safe answer is to wait.
scheduled_runs — the node’s own table — stays as this node’s audit
record of what it dispatched, which is what the dashboard reads and what
the retention sweep purges. It is the same split the token counter made
when it moved to the shared budgets slot:
what the fleet has to agree on is “may I start”, and nothing more. Its
outcome is fired, skipped_catchup or skipped_paused; the downstream turn result
(done, failed or timed out) lives in the normal turn telemetry
(agent_turn_completed, and turn.guard_breach when a guard fired) under
the same trace, because every fire starts a trace of its own and the turn
restores it.
The at-most-once guarantee lives in the coordination store, not in the
node’s file: a node whose scheduled_runs table cannot be read still
schedules correctly, and what it loses is its own dispatch history in the
dashboard.
Missed-tick catchup
Section titled “Missed-tick catchup”Catchup is evaluated only on the first tick after (re)start. If the
engine was down across a scheduled fire, the single most-recent missed
fire is run if it falls inside the catchup window — half the schedule’s
period, clamped to [catchup_min_seconds, catchup_max_seconds] (default
120s–7200s). Older misses are never backfilled; a missed fire outside the
window is recorded as skipped_catchup for audit. Set catchup: false on
a schedule to opt out entirely.
A paused seat’s fires are skipped
Section titled “A paused seat’s fires are skipped”A fire that comes due while a person has its runner seat
paused is claimed and recorded
skipped_paused, and not sent. It is not queued behind the pause: the
seat’s inbox is held while it is paused, and a fire waiting there would run
whenever somebody resumed it — a standup days late, one for every day of the
pause. Claimed under the fire’s own identity (its runner included), so a peer’s
tick of the same minute, or this node’s after the resume, does not send it
after all; a resumed seat picks up at its next fire.
Each node reads the pauses from its own watched copy of the coordination record, so the answer follows the duty wherever it moves. A node that has not read the pauses yet (the seconds after a boot) dispatches the fire, the opposite polarity from the duty’s own unknown: a paused seat’s inbox holds a fire rather than running it, while a fire skipped on a read that failed is a standup lost for a seat nobody paused.
Hard wall-clock timeout
Section titled “Hard wall-clock timeout”Each scheduled turn carries timeout_seconds (default 180), on the fire’s own
task_assigned payload. The turn engine checks it between rounds: a turn
past the cap starts no further executor and reviewer round and ends failed,
with a turn.guard_breach event of kind scheduled_timeout and
error_kind: scheduled_timeout on its agent_turn_completed event, so a
runaway loop can’t monopolise the runner. A failing or timed-out run never
blocks the next tick.
Between rounds, never mid-phase. A phase that is running may already have fired side effects — a post, a comment, a status change — and abandoning it there would leave one half-written with nothing recording that it happened. That is the same reason a drain lets a running turn finish. So the cap refuses to start the next round rather than interrupting the current one, and a single round that outruns the cap is bounded by the phase timeouts beneath it. Round one always runs: a cap that refused before any work started would report a turn as timed out having done nothing, which is a misconfiguration rather than a slow turn.
A resumed turn — one re-entering a suspended code sandbox run — carries no cap. The cap bounds the turn a fire started; the resumed half is finishing work the box already did.
Configuration
Section titled “Configuration”System-level knobs live under a top-level scheduling: block:
scheduling: enabled: true # master switch tick_seconds: 10 # scheduler poll interval jitter_seconds: 0 # max per-schedule deterministic spread (see below) catchup_min_seconds: 120 # lower clamp on the catchup window catchup_max_seconds: 7200 # upper clamp on the catchup windowThere is no zone here: a schedule that names none fires on the company’s clock (see Which clock a schedule fires on).
The scheduler auto-enables when enabled is not false and the org
actually declares at least one schedule (see When the loop runs):
orgs with no schedules never spin up the tick loop.
Thundering herd (jitter)
Section titled “Thundering herd (jitter)”When many schedules share a popular minute (everyone writes
0 9 * * *), they all become due at once. Each node’s turn concurrency
gate (node.max_concurrent) already queues the burst, so this is a
smoothing concern, not a correctness one. Set scheduling.jitter_seconds to a non-zero value to
spread firing: each schedule gets a deterministic offset in
[0, jitter_seconds] derived from its scope + name, so the 9am wave is
fanned out across that window. The canonical fire minute still forms the
idempotency key, so dedup is unaffected. Default 0 fires exactly on the
minute.
When the loop runs
Section titled “When the loop runs”The engine arms the tick loop when four things hold, and re-checks all four on every config apply:
scheduling.enabledis notfalse,- this node reaches the fleet’s coordination store and the stream: the at-most-once claim lives in the coordination store, because a scheduler with a process-local claim looks identical to a correct one until there are two nodes, and then every company gets two standups. The node’s own store is not a condition; without one the node only loses its local dispatch history,
- the company actually declares at least one role or unit schedule,
- the company configures at least one model provider under
providers.llm. A company with none is valid, but its seats hold every delivery on their inboxes until a provider exists (A Company With No Model Provider), so a loop firing through that wait would stack every missed standup behind the hold and run the whole backlog when the provider arrived. Each node logsschedules_waiting_for_a_modelinstead.
Because these conditions are re-evaluated live, adding the first schedule
to a running company, or the first provider to a company that has schedules,
arms the loop on that apply, and removing the last one disarms it and releases
its fleet duty; neither needs a restart. A freshly armed loop starts with the
missed-tick catchup described above, so a company whose provider arrives late
catches up at most its most recent missed fire, never the whole wait. The tick
knobs (tick_seconds, the catchup clamps) are read when the loop is armed,
so changing those lands at the next arm, like the retention sweep’s horizons.
Whether this node has the loop armed is visible on the engine as
Engine.SchedulerRunning, and the scheduler_armed / scheduler_disarmed
log lines say when it changed. An armed loop fires only on the node that
holds the scheduler fleet duty.
Observability
Section titled “Observability”- Dashboard. Agents › Schedules (
#/agents/schedules) lists every configured schedule — name, scope (a role’s seat by its name, or a unit), cron with what it means in words, task, whom it wakes, when it next fires and how it last went — and the recent dispatch ledger below it, both sortable. One schedule’s own page (#/agents/schedules/{scope_type}/{scope_id}/{name}) adds its whole definition, the fires after the next worked out in the schedule’s own zone, and every fire this node’s ledger still holds for it (theschedule_runsquery), because the company-wide ledger is only its newest fifty. The list is backed byGET /schedules, which servesschedules(the resolved schedule list with next-run times, projected from the current organization on each request) andrecent_runs(the 50 most recentscheduled_runsrows). The next-fire times tick as you watch: every relative time in the product reads one shared clock rather than being baked at render. ScheduledTaskFiredevent (crewlet.events.scheduled_task_fired) is emitted per dispatch withscope_type,scope_id,schedule_name,target_handle, andscheduled_at— surfaced in the dashboard / event store.- The coordination store’s
firesslot is the at-most-once claim; the node’sscheduled_runstable is its own dispatch history. - Structured logs:
scheduler_armed,scheduler_disarmed,schedules_waiting_for_a_model,schedule_fired,schedule_catchup_skipped,schedule_no_runners,schedule_parse_failed,schedule_claim_failedandschedule_publish_failed(a claimed fire whose publish failed is not retried, because the broker may already hold it).
DST & cron edge cases
Section titled “DST & cron edge cases”- Fall-back (clocks repeat an hour): a wall-clock cron time in the repeated hour maps to two UTC instants, but both share one local fire-label, so the run fires once, not twice. (The claim identity uses the local wall-clock minute, not the UTC instant.)
- Spring-forward (clocks skip an hour): a cron time in the skipped
hour has no instant that day, so it does not run that day — and
nothing was “missed” from the window’s view, so there’s no
skipped_catchuprow. If a daily run is critical, usetimezone: UTC, or keep it out of the skipped hour — which is not the same hour everywhere: it is 01:00–02:00 inEurope/London, 02:00–03:00 across most of the rest of the EU and in the US, and something else again in zones that shift by other amounts or on other dates. Check the hour for the zone you actually set rather than assuming 02:00; a30 1 * * *schedule onEurope/Londonis in the gap and silently skips one day a year. - Day-of-week ranges don’t wrap:
sat-sunis rejected as a descending range (Sunday is0). Write6-7,sat,sun, or0,6for weekends. - Catchup window is a bounded heuristic (half the period, clamped to
[catchup_min, catchup_max]); for irregular cadences (e.g.0 9 * * 1,5) it’s an approximation of the inter-fire interval, not an exact value.
Out of scope
Section titled “Out of scope”- Multi-platform delivery. The runner already owns its outbound surfaces (Slack channel, Jira project), so delivery routing is the agent’s job, not the scheduler’s.
- Script-only monitors. A cheap monitor runs as a low-budget role; the scheduler always dispatches to an agent turn.
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.