Configuration Reference
Crewlet companies are defined across two tiers — see the Configuration concept page for the full design.
- Tier A (
crewlet.yaml, restart-only): this node’s identity and roles, the store file, the stream and coordination slots, API host/port/auth, the secret keyring, logging (level, shape and an optional rotating log file). - Tier B (
company.yamlimported into the store, live-editable): everything else — identity, providers, integrations, MCP servers, roles, units, turn engine, learning, budgets, scheduling.
This page documents the Tier B fields below. For Tier A see Configuration concept page §“Tier A example”.
Machine-readable version. crewlet schema company emits the JSON
Schema for everything on this page, generated from the models
themselves — the same files the repository commits under schema/,
which make schema regenerates whenever a model changes, so a field
named here and a field the schema accepts cannot drift. Point your
editor at it for autocomplete and typo
squiggles, or hand it to an AI assistant — see
Authoring with an AI assistant. Check your file with
crewlet validate <file> (add -json for located, classified problems
and warnings, see the validation loop);
it reads no environment, so it works before any secret is exported.
Credential positions are marked. Every field, list member or map
value that holds a credential carries "x-crewlet-secret": true in the
schema, so an editor or a tool writing this file knows where a ${VAR}
reference belongs instead of a value — see
Credential positions.
Unknown keys are rejected. Every config model forbids extra
fields, at every level including roles and units — a mistyped
backstroy: fails validation naming the exact path rather than being
silently ignored.
Top-Level Fields (Tier B)
Section titled “Top-Level Fields (Tier B)”name: "Acme AI Corp" # required — company namemission: "Build intelligent products" # optional — company missionvision: "Lead the AI industry" # optional — company visiontimezone: Europe/Berlin # optional — the company's ONE clock, an IANA zone # (default UTC). See "The company's clock" below.token_budget: # optional — org-wide token ceilings, one per calendar day: 3000000 # window on the company clock, each optional; an month: 40000000 # absent window is uncapped. See "Token budgets" below.notification_rate_limit: 10 # optional — inbound notifications one seat may be woken by # per second. 0 (the DEFAULT) is unlimited, so the valve is # off unless you set it. Drop-based safety valve against # webhook storms and notification loops — burst handling is # inbox coalescing below, which batches instead of dropping. # Fails OPEN: a valve that cannot reach its counter passes # the notification rather than swallowing it.notification_coalesce_window_seconds: 0 # optional — inbox linger window (seconds) absorbing bursts # before an idle agent's turn starts. 0 (default) adds no # latency: backlog that piled up while the agent was BUSY # still coalesces into one digest turn per conversation. # Max 60 — the window counts against the broker ack-timeout # budget bounding a message's unacked lifetime. # See concepts/event-system.md § Inbox batching.notification_coalesce_max_batch: 20 # optional — events collected into one DRAIN before it is # partitioned by conversation, 1..100 (default 20). It # bounds a digest as a consequence — a digest cannot exceed # a drain — and it bounds the drain, which has to fit the # ack budget alongside a whole turn. A drain spanning # several conversations shares the cap between them, so # raise it for seats that routinely serve several busy # threads at once.
policies: # optional — org-wide policies (full text renders into the executor's prompt) - "All code must be reviewed before merging" - "Communicate decisions in writing"
workers: {...} # optional — reusable delegate templates (see below)
roles: [...] # optional — root-level org-wide agents (see below)units: [...] # optional — org structure tree (see below)
turn_engine: # optional — executor/reviewer turn config max_iterations: 3 # hard cap on self_iterate loops per turn max_tool_rounds: 24 # max tool-call rounds within one executor run # (the reviewer has no knob: it holds one # submission tool, so its budget is structural) onboarding_max_tool_rounds: 10 # dedicated first-turn onboarding pass (0 = disabled) delegation: # bounds on every `delegate` call — see Turn Engine § Workers max_parallel: 3 # workers running at once within one call max_tasks_per_call: 8 # tasks one call may contain; over it the call is REFUSED max_turns: 20 # tool rounds one worker may run (a request above it is clamped) max_turns_ceiling: 40 # highest max_turns a worker template may declare budget_fraction: 0.2 # ONE slice for the whole call, shared by every task in it min_tokens_per_task: 500 # call refused up front if the per-task share falls below this task_timeout_seconds: 300 # wall-clock cap on one worker call_timeout_seconds: 900 # wall-clock cap on the whole call, dependency waves included sandbox_min_budget_tokens: 2000 # refuse a coding run below this remaining budget delegation_depth_limit: 3 # max colleague-handoff chain depth before depth_cap guard breach extension_enabled: true # round-cap extension judge (executor + onboarding) execute_max_tool_rounds_ceiling: 48 # hard ceiling for executor rounds across extensions (2x base 24) onboarding_max_tool_rounds_ceiling: 20 # hard ceiling for onboarding rounds across extensions (2x base 10) extension_round_step: 8 # max rounds the judge may grant per extension call conversation_session: # what this seat already said in ONE thread / issue / PR, # carried into that conversation's next turn # (see concepts/conversation-sessions.md) enabled: true # feature gate — a live kill switch; off restores the # pre-ledger prompt exactly max_entries: 20 # entries KEPT per conversation, trimmed at write # time. A second bound applies at RENDER time: # the injected block drops whole oldest entries, # saying how many, once it would exceed 24k # characters retention_days: 30 # matches the event store's horizon; applied at next start
learning: # optional — agent-learning subsystem # Every `*budget_tokens` / `*_max_tokens` cap here # bounds the ANSWER of a call that does not think. A # thinking model spends its thinking from the same cap, # so a call to one is sent the model's own output cap # instead, and these passes ask for `low` effort to keep # it short (see "Claude models" under Providers) enabled: true # master switch (auto-disables without DB + embeddings) episodic: retrieval_limit: 5 # default hits query_episodes returns when the # model names no limit (1-20) reflect: enabled: true # ReflectEngine + reflect_and_persist tool persist_decider: true # run the post-turn PersistDecider on every turn budget_tokens: 5000 # soft cap on the decider's LLM call (0 disables) summarize_episodes: true # cheap-model briefing of the turn-start # Similar prior work hits (query_episodes is # never summarised) summarize_max_tokens: 400 # soft cap on that briefing, in tokens counterparty: enabled: true # CounterpartyProfiler + auto-inject + lookup inline budget_tokens: 3000 # soft cap on the profiler's LLM call per turn skill_synthesis: enabled: true # SkillSynthesizer (single-turn + clustered) min_tool_calls: 5 # single-turn trigger threshold budget_tokens: 4000 # soft cap on the synthesizer's LLM call max_skills_per_agent: 50 # hard cap; once reached the synthesizer no-ops duplicate_jaccard_threshold: 0.7 # reject near-duplicates of existing skills # Clustered synthesis (opt-in). A fleet singleton: one node runs the tick. scheduler_enabled: false # set true to enable the background tick scheduler_interval_seconds: 3600 # seconds between ticks. Hourly is cheap: # a tick is one indexed scan per seat # and only calls the model when a NEW # cluster qualified cluster_window_hours: 168 # look-back window (7d). Bounds the QUIET # seat, whose last 200 turns can span a # year; episode_fetch_limit bounds the # busy one cluster_min_size: 3 # turns that must converge before a # cluster earns a skill. Two is a # coincidence cluster_jaccard_threshold: 0.6 # similarity that pools two turns. Lower # than duplicate_jaccard_threshold on # purpose: pooling asks "same kind of # work", rejecting asks "same skill" episode_fetch_limit: 200 # max episodes pulled per agent per tick skill_refinement: enabled: true # both halves: the post-turn refiner AND # the refine_skill tool. false withdraws # the tool too — they write the same rows auto_refine_on_success: true # append "Observed in practice" on done auto_refine_on_failure: true # append "Counter-example" on failed. # self_iterate is never refined: the turn # is not settled yet. Both false leaves # the refiner unbuilt budget_tokens: 3000 # soft cap on the refiner's LLM call max_body_chars: 20000 # a refinement whose result exceeds this # is refused, not truncated max_versions_kept: 10 # history retention per skill (older pruned) skill_promotion: enabled: true # daily cross-agent promotion pass. Needs a # knowledge base (integrations.confluence) # — a promoted skill is a draft # page a unit lead reviews min_sibling_count: 3 # distinct SEATS that must converge. One seat # with four similar skills is a catalogue # to curate, not a team practice jaccard_threshold: 0.6 # similarity that pools two seats' skills budget_tokens: 4000 # soft cap on the promotion LLM call personal_memory: max_refreshes_per_turn: 3 # cap on distinct context_hint values per turn # (idempotent repeats of the same hint are free) episode_lifecycle: # Trigger: the background pass counts each seat's raw rows on its own # tick — one indexed query per seat — and compacts whoever is over. max_raw_episodes_per_agent: 500 # threshold that fires CompactionRequested # Action 1: drop non-terminal mid-state rows past this age. non_terminal_max_age_days: 14 # Action 2: drop skill-consolidated rows past this grace. consolidated_grace_days: 30 # Action 3: cluster + LLM-compact remaining raw rows past this age. compaction_min_age_days: 30 compaction_min_cluster_size: 3 # singletons / pairs left raw compaction_jaccard_threshold: 0.6 # tool-sequence similarity to pool a cluster compaction_batch_size: 200 # max raw rows pulled per pass compaction_budget_tokens: 4000 # soft cap on the compactor LLM call exemplar_count: 2 # raw rows kept per cluster for drill-down # Action 4 (optional): evict ancient compacted entries for hard storage caps. compacted_max_age_days: 0 # 0 = disabledPer-role auxiliary model (for reflection + episode summarisation) and extension-judge model:
roles: - name: Engineer llm: claude-sonnet # the seat's own model — what the executor runs on llm_auxiliary: gpt-4o-mini # cheap/fast model for reflection llm_judge: gpt-4o-mini # cheap/fast model for the round-cap extension judgellm_judge is invoked when the executor or the onboarding pass exhausts its tool-round cap;
it decides whether the agent is making progress (extend) or thrashing
(fall through to rescue). Falls back to llm -> "default" if unset.
See the Turn Engine extension judge
section for details.
See the Turn Engine and Agent Learning docs for what each field controls.
The company’s clock
Section titled “The company’s clock”timezone: America/Los_Angeles # an IANA zone name; absent is UTCtimezone is the company’s one clock, and every calendar edge the engine cuts is cut on it:
- where today, this week and this month begin — for a
due=filter, the due bands on a board, the overdue mark on every row, a person’s own day (my_work) and the workload’s overdue counts; - what a relative date (
tomorrow,eow,+7d) or an all-day due date resolves to, whether a seat, an operator’s assistant or the dashboard wrote it; - the wall clock a schedule that names no
timezoneof its own fires on — socron: "0 9 * * 1-5"is 09:00 in this zone.
It is a clock for authored instants and calendar boundaries only. No duration — a lease, a timeout, a retention horizon — is measured against it, because a duration measured on a wall clock changes length twice a year. And there is only one: the tracker and the scheduler take no zone of their own, because a company whose board, whose people’s days and whose standups each kept a clock had one “today” per subsystem. A schedule’s own timezone is that one piece of work’s wall clock — the Tokyo team’s 09:30 standup — and nothing else is cut on it.
The value must be a real IANA name (Europe/Berlin, America/New_York, UTC), and Local and localtime are refused: each is whatever zone the host reading it is set to, so two nodes would cut two different days from it. The anonymous org projection carries the resolved clock — UTC where none is written — so the dashboard cuts its days where the engine does. An apply that changes it moves the next answer, the next scheduler tick and the next tool call; a turn already running keeps the clock of the epoch it started on.
Token budgets
Section titled “Token budgets”token_budget: # on the company, and on any agent seat day: 3000000 # one local day, midnight to midnight week: 10000000 # one ISO week, from Monday month: 40000000 # one calendar monthA token budget is a ceiling per calendar window, and every window is cut on the
company’s clock: a day runs from local midnight to the next, a week
is the ISO week from Monday, a month is the calendar month. Each key is optional and the
windows are independent — a model round is admitted only while every capped window has
room for it, and each window opens again on its own when it turns over. There is no single
number any more: one number was a ceiling for the life of the deployment, which nothing
reset but an operator zeroing a counter, and token_budget: 10000000 is refused with the
mapping to write instead.
The company’s token_budget caps every seat’s spend together. A seat’s own caps that seat
alone, on top of the company’s: every round is charged to both, so a seat is stopped by
whichever ceiling it reaches first. A human seat spends nothing, so a token_budget on one
is refused.
An absent key is the only way to leave a window uncapped. A ceiling of 0 or less is
refused (must be at least 1 token … remove token_budget.day for no daily ceiling): 0
used to mean “unlimited”, and it is also what a ceiling of nothing would be, so it means
neither. Stopping a seat on purpose is not a budget’s job.
crewlet validate warns — it never refuses — about a ceiling that
can never refuse a turn, because another ceiling is always reached first:
- a day ceiling at or above the week’s or the month’s (a day lies inside both);
- a week ceiling that seven days at the daily ceiling already reach, or a month ceiling that 31 days at the daily one, or six weeks at the weekly one, already reach;
- a week ceiling at or above the month’s, which can then bind only in a week that straddles two months;
- a seat’s ceiling at or above the company’s for the same window, or for a longer window that holds it (a seat’s day against the company’s month) — everything a seat spends is the company’s spend too.
Spend is counted in the fleet’s coordination store, one figure per window, so it survives restarts and is one number for the whole company however many nodes run it — and a window nothing caps is counted too, so a ceiling added mid-window judges the spend already in it. A seat whose capped window is refusing is not handed work: its mail waits on its inbox until the window turns over or a revision raises the ceiling (see the budget park). There is no reset: a window’s allowance comes back when the window turns over, and room before then is made by raising its ceiling. See Deployment § Token Budgets for how a refusal is recorded and reported.
Scheduling
Section titled “Scheduling”System-level knobs for the Scheduler — the
cron analogue that fires role/unit schedules:. The scheduler
auto-enables when enabled is true, a database is configured, and the
org declares at least one schedule.
scheduling: # optional — role/unit scheduled work enabled: true # master switch tick_seconds: 10 # scheduler poll interval jitter_seconds: 0 # max per-schedule spread to smooth a shared cron minute catchup_min_seconds: 120 # lower clamp on the missed-tick catchup window catchup_max_seconds: 7200 # upper clamp on the missed-tick catchup windowA schedule that names no timezone of its own fires on the company’s
clock; there is no scheduler-wide default zone.
See the Scheduling concept doc for delivery
modes (each / lead), at-most-once semantics, catchup, and the
per-task wall-clock timeout.
Providers
Section titled “Providers”providers: llm: default: # named provider (referenced by roles via `llm: default`) type: openai # openai | anthropic | openai-compatible | cli-agent model: gpt-4o # required, and supports ${ENV_VAR}. A reference that # resolves to nothing is REFUSED naming the variable — # unlike a missing api_key, which still builds: every # call then comes back a clean 401 that names the # provider, where a missing model has no such tell and # the request is simply malformed api_keys: # one or more keys; multiple enables rate-limit rotation - "${LLM_API_KEY}" # supports ${ENV_VAR} references # - "${LLM_API_KEY_BACKUP}" # add more for rate-limit rotation # an empty list takes the conventional variable # (OPENAI_API_KEY / ANTHROPIC_API_KEY, from the secret # store, then the environment); with neither, the # provider still builds and sends no key, and a vendor # that needs one refuses its calls as unauthorized # (401), each failure naming the provider. # A NAMED key that resolves to nothing stays # nothing: the entry never borrows the # conventional variable in its place. # Settings › Models & keys shows each key by # the variable it names, whether it resolves, # and when a benched one comes back cooldowns: # optional — TTL when a key is marked exhausted rate_limit_seconds: 3600 # 429 / 402 default cooldown (a Retry-After / x-ratelimit-reset auth_seconds: 300 # 401 / 403 default cooldown header on the error overrides it; # repeated auth failures on one key back off exponentially) # a bench is SHARED across the fleet, so a peer's 429 benches the # key here too — see concepts/coordination.md base_url: "${LLM_BASE_URL}" # optional — custom endpoint; supports ${ENV_VAR} references. # REQUIRED for openai-compatible and REFUSED when it is # empty — including when a ${VAR} resolved to nothing. # `openai-compatible` means "not OpenAI", and an empty # base_url would take the OpenAI backend's own default: # the company's whole model traffic sent to # api.openai.com under a key that is not an OpenAI key, # with a 401 naming a vendor nobody configured as the # only symptom. # On `openai` / `anthropic` it is genuinely optional and # points the vendor's own wire format at a gateway or # proxy instead of the vendor host timeout_seconds: 600 # optional — bounds one HTTP attempt (default: 600). A unary call IN TOTAL; # a streamed call (every anthropic call, and an executor's rounds # elsewhere) by its SILENCE — the longest # wait with nothing arriving, first byte included — never its length, # so a round thinking for many minutes is not cut off half-way # (the cli-agent backend drives a subprocess and uses cli.timeout_seconds instead) reasoning: false # optional, openai only — send reasoning_effort to a reasoning # model (default: false). REFUSED on anthropic, where thinking is # not a switch (see "Claude models" below), and on openai-compatible. # gpt-4o is not a reasoning model, so this entry sets no # reasoning_effort: with reasoning off it is refused (see `reasoner`) # There is NO temperature field, on any type: the phases of a turn # send none, so every round runs at the vendor's own default # (1.0 on OpenAI), and only a call that needs one names it — the # extension judge asks for 0, the knowledge answer and the auxiliary # passes for 0.2. openai sends it on a call that is not reasoning; # anthropic only where the model samples (see "Claude models" below) reasoner: # an OpenAI reasoning model type: openai model: gpt-5 api_keys: - "${OPENAI_API_KEY}" reasoning: true reasoning_effort: high # optional — how hard the model thinks: low | medium | high | xhigh | max # openai: sent while `reasoning` is on (unset: nothing is sent, the # endpoint's own default, and no call can lower it, since that # default is not a level the engine can compare). Refused with # reasoning off, where it would do nothing # anthropic: output_config.effort on EVERY call (unset: high), on a # model that takes effort at all; a level the model does not take # is refused (see "Claude models" below) # A CEILING for every call on this entry: a call may ask for less (the # extension judge, the knowledge answer, the turn-start filters and every # learning pass ask for `low`) and never for more claude: type: anthropic model: claude-sonnet-5-5 # the request SHAPE follows the model — see "Claude models" below reasoning_effort: high # optional (default: high) claude-gateway: type: anthropic model: fast # a gateway's own alias for a model... claude_model: claude-haiku-4-5 # optional, anthropic only — ...named here, so its requests are # shaped as that model's. An exact id from the table below; # refused beside a model the table already reads base_url: "${CLAUDE_GATEWAY_URL}" reasoning_budget_tokens: 8000 # optional, anthropic BUDGET-ERA models only (claude-haiku-4-5, # claude-sonnet-4-5, claude-opus-4-5 and older): the thinking # budget, at least 1024 and below the model's output cap. # Unset: that model does not think. Refused on every other model # and every other type # reasoning, reasoning_effort and reasoning_budget_tokens are all # REFUSED on a cli-agent entry: a coding CLI driven headlessly takes # no per-call reasoning flag and carries its own configuration. # Pick the reasoning model with `model` instead budget: # multiple providers supported type: openai model: gpt-4o-mini api_keys: - "${OPENAI_API_KEY}"
subscription: # a coding CLI you already subscribe to, # driven headless — NO API key. # See concepts/subscription-llm-backends.md type: cli-agent model: sonnet # whatever the CLI's --model accepts cli: agent: claude-code # claude-code | codex | gemini-cli | qwen-code # | opencode | cursor-agent | copilot | grok # | muse-code | kimi-code | hermes | pi # | custom state_dir: "" # optional — credential dir + per-seat CLI homes. # Empty: $CREWLET_LLM_CLI_HOME/<key>, else # ~/.crewlet/llm-cli/<key>. Use a persistent # volume in an ephemeral container. # Entries naming the SAME dir share one login # (how per-phase models run off one # `crewlet llm login`); they must then also # share the same `agent`. timeout_seconds: 300 # optional — one CLI invocation, wall clock. # Separate from the entry's HTTP # timeout_seconds: this covers process # launch + the model call + the CLI's retries max_concurrent: 4 # optional — CLI processes at once. Each is a # 200-400 MB runtime, and subscription plans # throttle concurrency hard env: {} # optional — extra child env, ${ENV_VAR}-resolved. # The child gets an ALLOWLISTED environment, # never the engine's, so declare anything else # the CLI needs here auth: mode: subscription # subscription | api-key | inherit-env. # api-key puts the entry's ONE api_keys value # in the CLI's key variable, and a write is # refused for a CLI with none (hermes, pi, # opencode, kimi-code); api_keys is refused in # any other mode, since nothing would read it token: "" # optional — ${VAR} holding a headless # subscription token; empty falls back to the # profile's own var (CLAUDE_CODE_OAUTH_TOKEN). # subscription mode only, and only on a CLI # that mints one (claude-code); a write is # refused elsewhere, where it would reach # nothing credential_bundle: "" # optional — ${VAR} holding a `crewlet llm export` # blob; empty falls back to # CREWLET_LLM_CLI_<KEY>_CREDENTIALS overrides: {} # optional — replace any profile field when a # vendor renames a flag OR moves the field the # answer lives in (`text_paths`); validated # here, so a typo is refused by `crewlet # validate` and by every API write. # A moved answer field is the drift that # passes validation and fails at a seat's # first turn — `crewlet llm doctor` names it
embeddings: # optional — the semantic half of a native # knowledge search (the company's pages and # work items, embedded once for the fleet) # and similarity search for the # agent-learning subsystem (agent_diary # candidate selection AND episode recall). # Omit it and search is keyword only, # the diary's candidates are its recent # notes alone and episode recall renders # nothing (recent turns are not similar # work); nothing else changes. type: openai # openai | openai-compatible model: text-embedding-3-large # required, and it DECIDES THE WIDTH: this # one emits 3072, text-embedding-3-small # 1536. A new company gets the large one, # because the width is what search quality # rests on and it is not a knob worth # guessing api_key: "${OPENAI_API_KEY}" # supports ${ENV_VAR} references; empty falls # back to OPENAI_API_KEY, the same variable # the chat backend reads base_url: "..." # optional — custom endpoint. This is the only # difference between `openai` and # `openai-compatible`, so a local embedding # server needs nothing else # dimensions: 1024 # OPTIONAL OVERRIDE, 64..4096. Unset takes the # named model's OWN width, which is what # you want unless you are deliberately # shortening it: third-generation models # truncate on request, and a shorter vector # is a smaller index and a worse answer. # A model this build does not know is # REFUSED with `dimensions` unset rather # than given a guess — name a known model # or state the width yourself. embed-v4.0 # takes no override: Cohere's endpoint # does not support the parameter # max_input_tokens: 512 # THE MODEL'S LIMITS, in tokens, and unset # max_batch_inputs: 32 # takes the ones its vendor documents. # max_batch_tokens: 16384 # Required where none is documented — a # model this build does not know, and the # request limits of gemini-embedding-001 # and embed-v4.0, whose OpenAI-compatible # endpoints document none. A stated value # may only LOWER a documented one (a # gateway that accepts less); raising one # is refused naming the fieldThe width belongs to the model. text-embedding-3-large emits 3072,
text-embedding-3-small 1536, gemini-embedding-001 3072 and embed-v4.0
1536, and leaving dimensions unset takes whichever the model named here
emits. There is no global default, because a number that was right for one
model is silently wrong for the next — and a model this build does not know is
refused rather than guessed at, naming both ways to fix it.
Every request asks for the width, in the OpenAI dimensions parameter, where
the endpoint takes it — so a shortened width is the width that comes back —
and every answer is checked against it either way:
| Model | dimensions at its endpoint | What the engine does |
|---|---|---|
text-embedding-3-large, text-embedding-3-small | Supported: the vector is shortened to the width asked (OpenAI’s reference) | Sends it, so a width below the model’s own is the width stored |
gemini-embedding-001 | Not mentioned by Google’s compatibility page, whose examples send none | Sends it, as it always has; an endpoint that ignored it would answer 3072, which the width check refuses at any other width |
embed-v4.0 | Listed as unsupported by Cohere’s Compatibility API, which takes no other way to choose a width either, so it answers at the model’s default, 1536 (the model page) | Never sends it, and refuses any other dimensions, naming the field |
| any other model | Unknown | Sends it |
So do its limits. How many tokens one input may hold, how many inputs one request may carry and how many tokens one request may carry in all are facts about the model, and the engine carries what each vendor documents for the endpoint it calls:
| Model | Tokens an input | Inputs a request | Tokens a request |
|---|---|---|---|
text-embedding-3-large, text-embedding-3-small | 8 192 | 2 048 | 300 000 |
gemini-embedding-001 | 2 048 | state it | state it |
embed-v4.0 | 128 000 | state it | state it |
Google’s and Cohere’s OpenAI-compatible endpoints document no limit per
request — Cohere’s native API takes 96 texts a call, but that is a different
endpoint’s — so those two are refused until max_batch_inputs and
max_batch_tokens are stated, and a model this build does not know states all
three, exactly as it states its width. A stated limit may only lower a
documented one, for a gateway or proxy in front of the model that accepts less;
raising one is refused, naming the field, because the provider would refuse
what the engine then sends.
A company that needs a limit states it. A stored company naming
gemini-embedding-001, embed-v4.0 or a model this build does not know,
without the limits it needs, is refused at boot as well as at an apply,
exactly as a model with no known width is — and, the same way, embed-v4.0
at any dimensions but 1536, which its endpoint cannot produce. No node
starts on that revision, so state the limits in the same edit that names the
model.
The engine counts bytes against those token limits rather than shipping a tokenizer per vendor: every tokenizer these models use emits at most one token per byte of the text it is given, plus the few special tokens a server wraps an input in (none for OpenAI, whose count is the input’s own; sixteen are set aside for every other model). So an input of 8 192 bytes is always inside OpenAI’s 8 192-token window, whatever language it is in — at the price that ordinary prose, three to four bytes a token, is held to about a quarter of the window. An input past that bound is refused before anything is sent, and a batch is sent in as many requests as the model’s limits need. What a text too long for one input becomes is each caller’s choice: the knowledge corpus embeds a document’s opening, while a turn’s ask and an episode are split between words into pieces that each fit, embedded in one call, and pooled into one vector — so a long ask is still searched by meaning, all of it, rather than refused.
The width is a contract with the store, not with the model. Vectors of two different widths cannot be compared, so a row written at the wrong one is not a degraded search — it is a row that can be written and never read back, silently and permanently. Two checks follow from that, and both refuse rather than adapt:
- At the apply. A revision whose
dimensionsdiffers from the width this store already holds is rejected, with an error naming both. Changing the width is a restart, which is a decision for an operator who is watching rather than a silent divergence discovered at the first recall weeks later: after it the knowledge corpus re-embeds itself and each seat’s diary and episodes are re-filled by the node holding the seat. - Per model, at recall. Two models of one width are two spaces —
text-embedding-3-smallandembed-v4.0both answer 1 536 floats — so every diary and episode vector is stored with the model it came from and recall compares only rows of the query’s model. Changingmodelat the same width is therefore an ordinary apply: each seat’s diary and episodes are re-filled in the new space by the node holding it — about a million bytes of text a minute across the company — and until a row is, it is reached by recency and by conversation rather than ranked against a space it is not in, andquery_episodessays how many of a seat’s turns its search could not reach. - On every call. A vector that comes back at the wrong width is refused rather than stored — on every call and not just the first, because a gateway or aggregator can move models mid-deployment.
What a failure costs depends on who asked, so nothing in the provider
retries. For a turn starting — diary and episode recall, a search’s query
vector — a timed-out or refused call is no vector: personal memory keeps its
recent half, ## Similar prior work renders nothing and query_episodes says
its search could not run (a recent turn is not similar work), and a hybrid
search serves its keyword half — and a retry would be spent inside a prefetch
somebody is waiting on. For the knowledge corpus, a failure is
a backlog, and the duty asks again on its next tick. What the provider gives
every caller instead is which failure it was, in three classes: refused
(HTTP 400, 413 or 422, or an input past the bound — sent again unchanged it
will be refused again, so a caller can set that one input aside rather than
resend its whole batch for ever), transient (408, 409, 425, 429 and 5xx, a
timeout or a network failure — asking again later may succeed) and
configuration (401, 402, 403, 404 and the remaining 4xx, or a vector of the
wrong width — nothing will succeed until providers.embeddings is fixed).
One call’s ceiling is not another’s. A single embedding is held to 15
seconds where nothing tighter bounds it — an episode embedded after its turn,
for one. What a person or a turn waits on is held to two seconds instead: a
search’s query vector, whether a person or a seat’s search_knowledge asked; a
turn’s ask as the turn starts; the hint query_episodes or refresh_memory
passes; and a note reflect_and_persist keeps, which the seat’s model is
waiting on mid-turn (a note the reflection after the turn keeps is held to the
same two seconds). Past it each degrades rather than waits. A hybrid search
serves its keyword half and a semantic one answers nothing, each saying why.
Personal memory is chosen from the recent half of its candidate pool alone. Similar prior work renders nothing, and
query_episodes says its search could not run — by design, since a recent
turn is not similar work and the block would offer it as precedent. The note
is kept without a vector until the node holding the seat fills it. So a slow
embeddings server shows up as embedding_failed searches, personal memory drawn
from recent notes only and no similar prior work, not as slow searches, slow
turn starts and stalled tool calls. A batch request carries up to
the model’s request total (300 000 tokens on OpenAI) and is held to 60 seconds,
a fifth of the five minutes a corpus tick may go without progress. Nothing has
measured how long a server takes over a full request — OpenAI’s or a self-hosted
one on CPU, the deployment most likely to need longer. If batches time out
against a slow server, lower max_batch_tokens: a smaller request is a shorter
one, and the ceiling a stuck call is held to does not move. Keep it at or above
the model’s per-input window (max_input_tokens, or the model’s own) — one
input must fit one request, so a request total below the window is also the most
one input may hold, and the knowledge corpus would then embed a shorter opening
of every long source, each under a new digest and so embedded again.
Claude models: thinking, effort and sampling
Section titled “Claude models: thinking, effort and sampling”An anthropic request’s SHAPE is a function of its model, because the Claude
generations accept different requests and a field a model does not accept is
an HTTP 400 — which the fallback chain does not
retry, since a refusal of the request itself is not something another attempt
fixes. So the engine reads what each model accepts from one table
(internal/providers/llm/anthropic/claudemodel) and sends nothing the model in
front of it would refuse:
| Models | Thinking | reasoning_effort | reasoning_budget_tokens | Output cap |
|---|---|---|---|---|
claude-fable-5-1, claude-mythos-5-1, claude-fable-5, claude-mythos-5, claude-opus-5-5, claude-opus-5, claude-opus-4-8, claude-opus-4-7, claude-sonnet-5-5, claude-sonnet-5 | adaptive, every call | low · medium · high · xhigh · max | refused | 128K |
claude-opus-4-6, claude-sonnet-4-6 | adaptive, every call | low · medium · high · max | refused | 128K |
claude-opus-4-5 | only with a budget | low · medium · high | 1024 to below the cap | 64K |
claude-sonnet-4-5, claude-haiku-4-5 and older | only with a budget | refused | 1024 to below the cap | 64K (older: 32K–64K) |
| any id the table does not know | adaptive, every call | every level | refused | 64K |
- There is no thinking switch. On a model with adaptive thinking every
call sends
thinking: {type: adaptive, display: summarized}— the summary is what fills the round’s reasoning and the live thinking text on the dashboard, and it costs the same as the hidden default. Turning thinking off is refused outright by Opus 5.5, Fable and Mythos, and on Opus 4.8 and 5 it makes the model write a tool call into its prose instead of making it. How hard the model thinks isreasoning_effort, which defaults tohigh: every phase this engine runs is long-horizon agentic tool use, and Opus 5.5’s own default ismedium, so an unset level would silently think less after a model upgrade. The calls that need little — the round-cap judge, the knowledge answer, the turn-start filters and the learning passes — ask forlowon their own, and no call can ask for more than the entry’s level. A call’s level that falls on one the model skips (xhighon Opus 4.6) is sent as the next level down that it takes. - A budget-era model thinks only with a budget.
reasoning_budget_tokensturns thinking on with that many tokens; unset, it does not think. - No temperature is sent by default. A call that names one (the judge asks for 0, the auxiliary passes for 0.2) gets it only on a model that takes a sampling parameter and on a call that is not thinking — in practice a budget-era model with no budget. Elsewhere it is dropped rather than refused, because those models answer any temperature with a 400.
max_tokensis the model’s own cap. Thinking is spent from the same cap as the answer, so a smaller one truncates a round or empties it; an unused cap costs nothing, and spend is bounded by token budgets instead. A call’s own smaller cap is honoured only on a call that is not thinking.- No forced tool choice and no prefill are ever sent: a phase that must end in a call names it and is asked again when a round ends without it, and a conversation always ends on a user turn or a tool result.
- Reasoning written under an earlier tool set is shed on
claude-fable-5-1,claude-opus-5-5,claude-sonnet-5-5and any id the table does not know. Those models bind each thinking block to the system prompt and the tools of the request that wrote it, and refuse it with a 400 once either has changed. A phase’s tools do change:activate_tooladds one, and a resumed run renders them again. So the rounds before a change are replayed without their thinking, while their text and calls are replayed whole. Every other model is replayed every block, because it runs no such check. See the conversation only grows.
A Bedrock or Vertex spelling of a model (anthropic.claude-opus-5-5,
us.anthropic.claude-sonnet-4-5-20250929-v1:0, claude-opus-4-5@20251101)
and a dated snapshot are read as the model they spell. An id the table does
not know — most likely a model released after this build — gets the current
generation’s shape, and crewlet validate warns about it at claude_model:
if it is a gateway alias for an older model, name that model there, or its
calls may be refused. Every rule above is checked when the config is
validated, and again when the provider is built, where a model written as a
${VAR} is first known.
The table is compiled into the build and the models are not, so
crewlet llm doctor <key> checks an entry against the live model: it reads
the model’s record from the vendor’s Models API, reports every place the
table disagrees with it — a problem when the disagreement puts a field the
model refuses on every call — and sends one real round in a phase’s shape,
which must come back as a tool call. A gateway that does not serve
/v1/models is reported as not served rather than as a failure. See the
CLI reference.
Tier A (crewlet.yaml)
Section titled “Tier A (crewlet.yaml)”Tier A is restart-only and says where this node’s stream, store and API are. Example:
# crewlet.yaml — Tier A bootstrapstream: type: embedded # a JetStream server inside this process: # no listener, no port, no service to # operate. `nats` points the same slot at # an external server — the same client # code either way, so it is a connection # choice rather than a second backend store_dir: "./crewlet-data/stream" # empty = in-memory: right for a test, # and nothing published survives a # restart. A company on the engine's own # tracker or knowledge base (the # defaults) keeps every item and every # page there, so the engine refuses to # boot it on an in-memory stream rather # than lose them at the first restart. # And a MEMBER OF A FLEET's broker — # embedded, no leaf urls, and naming a # cluster, peers or a leaf listener — # needs it WHATEVER its roles and its # company: it holds the fleet's streams # for every node that reaches it. A # leaf is refused one # store_max_bytes: 68719476736 # how much of that directory's volume the # EMBEDDED broker may hold — the ONE number # every stream ceiling on it is compared # against, because a ceiling is a # RESERVATION the broker refuses if it # cannot back it. UNSET, the broker sizes # itself: three quarters of that volume's # free space when its JetStream came up, # plus what it already occupies there, # which is what a single-engine host should # have. MEASURED ONCE, AT BOOT, on either # path — a disk that later grows or shrinks # does not move this limit, and a node that # should see a resized volume is restarted. # SET IT WHEN MORE THAN ONE ENGINE SHARES A # FILESYSTEM and divide it between them: # free space bounds their SUM, so two # engines each sizing themselves from what # they can see over-commit it, and the # failure is `insufficient storage # resources available` — or, on a fleet, # `no suitable peers for placement, # insufficient storage` — naming whichever # stream was provisioned last. Bounds: # 5 GiB..64 TiB — the four state logs' # 1 GiB floors and one for every stream # that reserves nothing — and it must not # be smaller than the ceilings declared # inside it. # REFUSED for `type: nats` — an external # cluster's account limits are its own # operator's, and this node reads them back # rather than declaring them # url: "nats://nats.internal:4222" # required for `nats`, REFUSED for # embedded — an embedded server has no # address, so a url there is read by # nobody. The server it reaches must # accept messages of 8 MiB # (`max_payload: 8MB`, every server in # that cluster): an event, or a node's # answer to another, may be that large, # so a server at nats-server's 1 MiB # default is refused at connect, naming # it. A file's bytes cross in messages # of 128 KiB, which a server's default # already carries # replicas: 3 # 1 solo (the default); 3 across an EMBEDDED # cluster, where it is what makes a publish # quorum-durable before it returns. At most # 5, JetStream's own ceiling, and on an # embedded cluster at most its members — # this node and `cluster.peers` — so >1 # without peers is refused: this node has # nothing to replicate to. It is the # BROKER's copies of its streams and # buckets — the company's files among # them under the default `nats` object # store (`store.objects`) # cluster: # an EMBEDDED server joining its peers, which # name: crewlet # is the fleet topology: every node embeds # one member of one cluster. REQUIRED once # anything else here is set — the server # reads none of these fields from an # unnamed cluster, so a block without it # starts a solo node that forms no cluster # at all, and Tier A refuses that. EMBEDDED # ONLY: an external cluster is formed by # its own operator's config, so this whole # block is REFUSED against `type: nats` # rather than accepted and read by nobody # port: 6222 # `node.id` is # peers: # the member name, so it must survive a # - "nats://node-1:6222" # restart — a name minted at boot orphans # - "nats://node-2:6222" # this member's replicas every time. # `peers` lists the OTHER members, each # scheme://host:port: an entry that is # recognisably this node's own route # (host+port, or advertise) or a repeat # is not counted as a member, with a # warning, and one that is not dialable # is refused # host: 10.0.0.11 # the interface the route port binds. # Unset binds EVERY interface, and a route # port is unauthenticated cluster access — # set it to the private address the peers # reach # advertise: "node-1.internal:6222" # what peers should DIAL for this member, # when that differs from what it binds — a # mapped container port, a NAT, a load # balancer. Members learn each other from # the members they already have, so unset # this address is derived from the # connection's remote address, which # behind a NAT is unreachable or somebody # else's. A bare host keeps this member's # own route port # leaf: # the LEAF LINK, and a node is on exactly # one side of it. WHICH SIDE is its # BROKER KIND, derived from this block and # nothing else: `urls` makes it a leaf, # and any other embedded broker is a # member. `node.roles` is a separate # question: what the node's DISK keeps. # urls: # A LEAF dials the members' leaf # - "nats-leaf://node-1:7422" # listeners here — any one that answers # - "nats-leaf://node-2:7422" # will do — and its embedded broker runs # NO JetStream: no replica, no vote, no # stream store (`store_dir`, `cluster` and # a leaf `port` are refused beside it). # Every stream and bucket its clients use # is a member's, reached across this link. # A node without `data` must be a leaf # and a node with it must not: every data # node holds the whole estate as a member # port: 7422 # A MEMBER's listener, where leaves join. # A member that opens one is in a fleet # even with no peers, so it must persist # (`store_dir`, as above: the leaves that # join it keep nothing, so it keeps # everything they do) and needs # `coordination.type: embedded-kv` # host: 10.0.0.11 # the interface the listener binds. Like # the route port it accepts any connection # that reaches it: the fleet is one trust # domain on a network its operator # controls, so bind it there # advertise: "node-1.internal:7422" # what leaves should dial for this member, # when it differs from what it binds # sync: always # what an acknowledged publish has actually # reached. `always` (the default, at every # replica count) fsyncs every write before # acknowledging it; a duration — `30s` — # declines the fsync and names the window # an acked write may be behind the disk. # It is NOT inferred from `replicas`: "a # replicated member has a quorum instead # of a disk" holds when one host loses # power and fails when a RACK does, and a # three-node fleet in one rack is three # copies of one unflushed page cache. A # window is refused where it would be # recorded and not honoured: against an # external cluster (which stores its own # data), on a leaf (which runs no # JetStream, so the members' own `sync` # decides), below 3 replicas (no quorum # to trade the disk for), and on a # cluster whose peers are all on this # host (one failure domain). Declining # it on a real three-host fleet is a # legitimate trade and costs 1–3 ms per # write on NVMe # debug: true # the EMBEDDED broker's OWN debug logging, # which is a SEPARATE question from # `logging.level` — that one is how loud # the ENGINE is. Off by default because # nats-server's debug output is per # internal-client rather than per event, # and the engine's coordination reads # manufacture those: every KV key listing # is an ordered consumer created and # deleted, and each deletion writes two # `JetStream connection closed: Client # Closed` lines. Two 15-second duty loops # list keys every tick, so this is a # constant background stream on an idle # node — turn it on when the BROKER is # what you are diagnosing. Its warnings and # errors are never gated by it and always # reach the log. What it unlocks are debug # lines, so `logging.level: debug` (or # `-debug`, or a `debug` log file) still has # to be recording them — `crewlet validate` # warns when nothing is. REFUSED for `type: # nats`: an external cluster logs wherever # its own operator configured it to # tracker_log_max_bytes: 17179869184 # the byte ceiling on the mutation log — # the ordered stream the engine's own # tracker writes through. UNSET DERIVES a # quarter of the stream volume's free # space, clamped to 4..64 GiB, because one # fixed number is five years of history on # the modelled rate and one boot on a # small disk. CROSSING IT REFUSES, it does # not shed: there is no age bound on this # stream, so a full log drops no history — # the append is refused, loudly, naming # whatever is blocking the trim. ONE # BUDGET FOR ALL THREE LOGS: the broker # reserves each ceiling in full when it # creates the stream, so the derived # ceilings of this field and the two below # are scaled down together to fit half of # what the broker can grant them, less the # ceiling any of the three that already # exists holds (never below 1 GiB each). # A value you set is never scaled, and a # boot that cannot reserve it fails # naming the field, the # bytes it needed and the bytes the broker # had. Every one of the three is the value # a stream is CREATED with: editing it # later changes nothing until # `crewlet retention set-capacity` does. # WHAT THE BROKER CAN GRANT is # `store_max_bytes` wherever you set one, # and Tier A refuses a limit smaller than # the ceilings declared inside it # tracker_vectors_max_bytes: 68719476736 # the vector changelog's ceiling # (1 GiB..2 TiB). SIZED FOR THE PEAK: the # stream keeps one message per source, so # a week's minting is ~91 MB — but # changing the embedding model rewrites # EVERY source in a few hours, and for the # following week all of them are in the # window: the whole corpus, ~17 MB per # agent seat per YEAR OF HISTORY. UNSET # DERIVES what the mutation log derives # from the volume (a quarter of its free # space, 4..64 GiB), whatever # tracker_log_max_bytes is set to: that # field bounds a trailing window of # records, and this log's peak is the # company's whole history. At 64 GiB it # holds a model change twice over for # ~2 000 seat-years (400 seats in their # fifth year); past that, set ~34 MB per # agent seat per year of history. A # ceiling sized from the steady state # would refuse the one operation it exists # to survive # pages_log_max_bytes: 4294967296 # the knowledge base's log, the ordered # stream every native page write goes # through (1..256 GiB). UNSET DERIVES a # quarter of the mutation log's ceiling, # set or derived (so 1..16 GiB from a # volume alone): a knowledge base is # a few thousand pages against a tracker's # hundreds of thousands of items, so its # log grows at about a quarter of the rate # and a blocked trim fills either one in # the same time. Crossing it refuses the # append, like the mutation log's # usage_log_max_bytes: 1073741824 # the usage log's ceiling (default 1 GiB, # 1..64 GiB): the compacted stream every # node publishes its own company days to — # spend, turns, page reads, schedule fires # — so history is answered fleet-wide and # outlives the node that spent it. SIZED BY # A CENSUS, not a rate: one message per # (node, day, seat or schedule) for 181 # days, about 217 MB for three nodes of # forty seats and twenty schedules. # Crossing it refuses the append and the # log_headroom alarm names `usage` # tracker_retention: # when the log may be trimmed. Every term # here is a statement about the OPERATOR's # estate rather than the company's policy, # which is why it is Tier A. Read only on # a node with the `data` role, which is # where the trim and the snapshots run # min_age: 7d # the age floor NO trim may cross, whatever # the other terms say (24h..90d). It can # only make a trim more conservative, so it # is a lower bound on how long the log # keeps a record and never a ceiling — and # it says nothing about any node's own # store file # backup_max_age: 24h # how stale the newest backup may be before # the trim STOPS ENTIRELY (1h..30d). A # company that never backs up never trims: # the log is the only copy of what no node # has applied yet # backup_floor: engine # whose word the trim takes for what is # backed up. `engine` follows the newest # backup the engine wrote and verified; # `operator` follows an explicit # acknowledgement, for a company that trims # only what has left the host — and then # trims NOTHING until somebody says so # snapshot_interval: 24h # how stale a node's newest snapshot may be # before it takes another (1h..7d). A # snapshot is a full copy of the replicated # estate, so four a day is a day's worth of # I/O to save a joining node a replay it # can do in under a minute. CROSS-FIELD: # `snapshot_interval × 2 < min_age`, or # every snapshot is older than the trim # floor and a node that lost its store has # nothing to resume from # rejoin_window: 30m # the budget for a node to become a # complete replica (5m..24h) — what a join # is measured against and reported on. A # setting rather than a constant, because # the answer is a property of the # operator's disks and network # event_retention_hours: 720 # 0 takes the queue's own default (30 days). # Unbounded is deliberately not expressible: # an event log nothing sweeps grows for the # life of the deployment # credentials: "" # path to a NATS credentials file # token: "${CREWLET_NATS_TOKEN}" # bearer token for an external server # tls: # the TCP layer under that auth: which CA to # ca: /etc/crewlet/ca.pem # trust; empty = the host roots # cert: /etc/crewlet/client.pem # the CLIENT certificate, for a broker # key: /etc/crewlet/client.key # configured `tls { verify: true }`. # Both or neither — half a keypair is # refused at validation. There is no # skip-verify switch, deliberately: it is # set once during a bring-up and never # unset, and a private CA is one file
store: path: "./crewlet-data/company.db" # the NODE estate — the audit log, memory, # config revisions, the secret bootstrap. # Owned exclusively by this process. Not a # shared database, and no DSN: two engines # on one file corrupt it # snapshot_dir: "./crewlet-data/snapshots" # where this node keeps its own snapshots # of the replicated estate — the file a # peer joining the fleet copies instead of # replaying the whole log. Absolute, or # relative to the store's directory. THE # DEFAULT PUTS A FULL COPY ON THE SAME # VOLUME as the live database, which is why # the snapshot loop refuses below 1.1× the # store's size free rather than filling the # disk the applier is committing to. A # separate volume is the production shape # replicated_path: "./crewlet-data/crewlet-replicated.db" # the REPLICATED estate — everything a # state log's applier writes, held whole # by every data node from boot, and by no # node without `data`. Empty puts it # beside `path`, which is what makes "back # up the data directory" true. It is a # separate FILE rather than more tables # because a snapshot for a joining node is # a copy of the estate's file alone; # separate it only to put it on a # different disk, and never onto the same # file as `path` # scratch: false # DELETE this node's store at every boot, # under the store's own lock, and open it # with no replicated estate at all. # REQUIRED on a node without the `data` # role and REFUSED on one with it, so the # deletion is never a surprise either way. # Explicit rather than derived from the # roles because it is the one setting here # that deletes something. The offline # `migrate`, `config`, `secrets` and # `search eval` commands refuse a scratch # store: what they wrote would be gone at # the next boot # max_open_conns: 0 # connection-pool bound for this node's # own database, and the readers of the # replicated estate's: that pool keeps at # least two readers, plus one pinned # connection per state log it carries, # ADDED on top rather than taken out of # them. 0 takes the store's own default # of four. Raise it if the `pool_starved` # alarm fires — see reference/alarms.md — # which means reads are queuing before # they start # busy_timeout_seconds: 0 # how long a WRITE waits for the # database's write lock before giving up # and retrying once; 0 takes the store's # default of 5s, half the dashboard's own # query timeout. It bounds the wait # wherever it happens — in the driver, or # in the engine's own FIFO queue for that # lock. Raise it on a node doing bulk # applies, where `store_tx_retry` in the # log names it; see guides/replication.md # objects: # where the company's FILES are kept — see # concepts/object-store.md. EVERY node # carries the same block, a node without # `data` included: an upload or a download # goes to the store from the node serving # it. The first node to boot records the # store in the coordination store, and a # node configured with another refuses to # boot, naming both # backend: nats # nats (the default, and what an absent # block means): a JetStream object store # bucket, `crewlet_files`, on the fleet's # own broker, kept at `stream.replicas` # copies on its members like every stream # and backed up with them. s3: an # S3-compatible bucket, named below # s3: # REFUSED unless backend is s3 # endpoint: "" # the S3 API's base URL; empty is Amazon # S3's own for the region. R2, MinIO, # Ceph's gateway and GCS's interoperability # endpoint each name theirs here # region: eu-west-1 # REQUIRED — every request is signed for # a region, even where the provider # ignores it (R2: auto, MinIO: us-east-1) # bucket: acme-files # REQUIRED. Checked at boot: a node that # cannot reach it, or whose credentials it # refuses, does not start # prefix: "" # prepended to every key — an object is # `<prefix>files/<key>` — so one bucket # can hold more than one company; e.g. # `acme/` # path_style: false # address the bucket in the path rather # than the host name; MinIO and most # self-hosted gateways need it # access_key_id: ${S3_ACCESS_KEY_ID} # both as ${VAR} references, # secret_access_key: ${S3_SECRET_ACCESS_KEY} # or NEITHER, which takes # the SDK's own chain: the environment, a # shared profile, a web identity or the # instance's role. A private CA is read # from AWS_CA_BUNDLE
coordination: type: local # one node holding its own seat leases; # a fleet needs `embedded-kv`. There is no # address to give it: the coordination KV # rides the stream's OWN connection. A # second dial to the same broker fails # independently, so a node could hold live # leases over a connection that still works # while the one carrying its inbox has # dropped — alive to its peers, deaf to # its work # lease_ttl_seconds: 45 # 0 takes the seat layer's own measured # default of 45s — three heartbeat # intervals, so two consecutive missed # renewals still leave a full interval to # recover in. Shorter speeds failover and # sheds healthy seats on ordinary jitter. # Seats and presence only: a fleet duty # sizes its lease from its own tick and # a tracker move, merge or bulk edit # from its own work. # IT IS THE BUCKET'S TTL, set by whichever # node created the lease bucket first and # ADOPTED by every node after it — a peer # booting with a different value does not # rewrite it, and logs # `coord_kv_lease_ttl_differs` naming the # one in force — AND RUNS AT IT, because # the bucket's age is what expires a # lease and a node claiming longer than # it would have every acquire refused. # Its heartbeat and a stop's allowance are # fractions of the live value too. So # make it agree across the fleet: # changing it requires deleting the # bucket while the fleet is down
api: host: "0.0.0.0" port: 8000 # node.roles decides what it carries. With `ingress`: # the dashboard, the REST API, the webhooks and the # probes. Without it: /health and /ready only (plus a # seats node's agent-mode tool bridge, unless # api.public is set). 0 (the default) binds nothing # at all public: # optional — a second listener for the routes outside port: 8443 # parties call: /webhooks/*, /otlp/{token}, /mcp/{token}. # Once set they are served ONLY here (404 on api.port), # and everything else — the dashboard, the REST API, # /config, /secrets, the probes — ONLY on api.port. # host defaults to every interface. 0 / unset keeps # every route on api.port auth: tokens: - id: founder token: "${CREWLET_API_TOKEN_FOUNDER}"
logging: level: info # debug, info (default), warn, error format: console # console (default: columns and colour for a person), # text (slog key=value), json (for a log shipper) stderr: true # default. `false` hands the stream to the file below and # needs one: a node logging nowhere is refused. It never # silences a boot failure or the watchdog's exit notice, # which is what it has over `2>/dev/null` file: # optional — a durable copy, IN ADDITION to stderr path: "/var/log/crewlet/${CREWLET_NODE_ID}.log" # missing directories are created; empty writes no file format: json # the file's OWN shape — empty follows logging.format level: debug # the file's OWN level — empty follows logging.level. # A debug file behind a warn console, or the reverse max_size_mb: 100 # rotate here. There is no "never": an uncapped log # file fills the disk the store is on (default 100) max_backups: 5 # kept as <path>.1 (newest) … <path>.5. # 0 keeps none — one file, and no history (default 5)
retention: backup_owner: platform-oncall # who owns this deployment's backups: a # person, a team, a scheduler's name. Free # text, read by a human at the moment an # alarm names it. `crewlet validate` warns # when it is unset — a company that never # backs up never trims, so "who is # responsible for this" has a real answer on # every deployment that intends to keep # workingcrewlet validate prints warnings as well as refusals, and they are
separate on purpose: the exit code turns on refusals alone, so a CI step gates
on what cannot run and still prints what its operator should read. A warning is
a configuration that is valid and carries a consequence worth knowing before it
is applied — a declined fsync’s window, a trim that will never advance until
somebody acknowledges a backup, a broker member with nowhere to persist the
streams it holds (never a leaf or a client of an external cluster, which hold
no stream of their own), a stream.cluster.peers entry left out of the member
count because it is this node’s own route or a repeat, a node that declares
the ingress role and binds no port, a unit keyed on a name somebody will
rename.
api.port is one listener per node, on every node, and node.roles
decides what it carries rather than whether it exists. A node without the
ingress role serves its two
probes there
and nothing else, so an orchestrator can probe a satellite exactly as it
probes an ingress node, and one Tier A file serves both shapes. There is
deliberately no separate probe port: on an ingress node the probes are already
on this listener, and a second address for the same two routes would be one
more thing to keep in step. crewlet validate warns about a node whose
declared roles include ingress while its port is 0: the node every
integration, browser and probe looks for would bind nothing. It warns rather
than refuses, because crewlet run -api-port sets the port at run time.
api.host and api.port are what this node binds, which is rarely where it
answers: a fleet behind a load balancer binds 0.0.0.0:8000 and is reached
at https://crewlet.example.com. Those outside addresses are Tier B’s, each
written once: integrations.public_base_url, where vendors and
sandboxes reach the deployment and every provisioner registers it, and
integrations.dashboard_base_url, where people reach the dashboard and every
link the engine composes for one points.
api.public splits the routes outside parties call onto a socket of their
own. With api.public.port set, the vendor webhooks and app landings
(/webhooks/*) and the sandbox endpoints (/otlp/{token}, /mcp/{token}) are
served on that listener and nowhere else, and every other route — the
dashboard, the REST API, /config, /secrets, /setup, /operator, the live
socket and the /health and /ready probes — only on api.port. api.port
answers a public route with a 404 no_route whose hint names api.public; the
public listener answers every other route exactly as it answers a path nothing
serves, with no hint, so the published socket says nothing about an admin port
behind it, and it does so whatever credential is sent, since it requires none.
Publish api.public.port, keep api.port private, point
integrations.public_base_url (and CREWLET_MCP_BRIDGE_URL /
CREWLET_SANDBOX_OTEL_RECEIVER_URL for remote sandboxes) at the public address
and integrations.dashboard_base_url at the address people reach api.port
on — see
Deployment → Exposing webhooks without the admin API.
On a node without the ingress role, whose only route is its seats’ tool
bridge, the bridge moves to api.public too, so one file still serves every
role. crewlet validate refuses a public listener beside api.port: 0 on a node
with the ingress role (that is no HTTP surface at all, and no probe would
answer); on a node without it api.port: 0 stays the hard off switch and binds
nothing, the bridge included. It also refuses any two listeners the
file opens on one port — api.port, api.public.port, stream.cluster.port
and stream.leaf.port — when they bind the same address or either binds every
interface; -roles, -api-host and -api-port are held to the same rules.
Binding a person to a credential crosses the tiers the other way: an
api.auth.tokens[].id is named from the company document’s
roles[].contact.crewlet_operator_id, never from a seat: field on the token,
because Tier A holds the keys to the secret store and may never read Tier B.
A token id is lowercase, and crewlet validate refuses anything else: the
binding is matched case-insensitively while every write made with the token
records its id exactly, so Founder would be a person bound for their writes
and missing from their own reads, and Founder beside founder would be two
credentials one binding admits as the same person.
The event store (LLM observability) is a table in that same file, created by
the engine’s own migrations on first start — there is nothing to configure
beyond store.path. See
Deployment → The event store
for the full layout.
Organization Structure
Section titled “Organization Structure”Roles can live in two places: at the org root (for org-wide agents like a CEO or cross-cutting advisor) or inside units (for team-scoped agents). Both are optional — you can have root-level roles only, units only, or both.
Root-Level Roles
Section titled “Root-Level Roles”roles: # optional — org-wide agents - name: CEO goal: "Set company direction" manages: [VP Engineering, PM Lead] # ... same role fields as unit roles (see table below)Root-level roles participate fully in the manages[] hierarchy and task
routing. Their knowledge is scoped to the org (visible to all agents).
They don’t inherit mcp_env from any unit.
The units key defines your org as a recursive tree of OrgUnit nodes.
Each unit can contain roles directly and/or child units, supporting any
nesting depth — flat teams, departments with sub-teams, divisions, or custom types.
units: - name: Engineering # required — unit name, and what people # read: in a prompt, on a board, in a # channel topic id: engineering # optional — the unit's STABLE IDENTITY. # Lowercase letters, digits, `-` and # `_`, starting with a letter. A name # is prose and gets renamed; an id is # read by nobody and survives, so # everything durable keys on the id # when there is one. Unique across the # chart, and it may not equal another # unit's NAME. Adding one to a unit # that already has work filed against # it rewrites nothing: a filter on a # unit matches its id and its name. # IT DOES NOT STOP A RENAME # RE-ONBOARDING the seats beneath it — # onboarding turns on the NAME, which # is what an agent reads as its team, # so changing the name changes the # context those seats were introduced # with, id or no id type: department # optional — unit type (default: "team") lead: CTO # optional — inherited from parent if omitted purpose: "Build and ship the product" # optional children: # optional — nested child units - name: Backend # required type: team # optional lead: Tech Lead # optional — inherited from parent if omitted goals: # optional - "Ship features on 2-week cadence" channel: backend # optional — the team's chat channel, inherited project: "BACK" # optional — the unit's tracker "home": routing + space: "BACK" # write target. NOT read scope, NOT a credential, # and NOT vendor-specific — the same keys name a # native project/space or a Jira/Confluence one mcp_env: # optional: MCP creds shared by the unit's direct agent roles atlassian: # (real tool credentials only; the chat transport JIRA_API_TOKEN: "${BACK_JIRA_TOKEN}" # identity is per-agent) roles: - name: Tech Lead # required — unique agent identity goal: "..." # optional — individual mission backstory: "..." # optional — personality, background, expertise llm: default # optional — named LLM provider token_budget: {day: 200000} # optional — this seat's own ceilings per window handle: tl # optional — custom handle (default: auto-slugified) manages: [Engineer A, Engineer B] # optional — hierarchy links responsibilities: # optional - "Review all PRs" behavioral_guidelines: # optional - "Be thorough in code reviews" integrations: # optional — per-agent transport identity mattermost: bot_token: "${MATTERMOST_BOT_TOKEN_TL}" username: tl-bot # optional — the bot's Mattermost username channel: backend # optional — a channel this bot is added to mcp_env: # optional — per-agent MCP server credentials atlassian: JIRA_USERNAME: "${TL_JIRA_USER}" JIRA_API_TOKEN: "${TL_JIRA_TOKEN}" mattermost: MATTERMOST_TOKEN: "${MATTERMOST_BOT_TOKEN_TL}" # same token as role.integrations.mattermost github: Authorization: "Bearer ${GITHUB_TOKEN_TL}"Role Fields Summary
Section titled “Role Fields Summary”| Field | Type | Required | Description |
|---|---|---|---|
name | string | yes | Unique seat identity |
kind | agent | human | no | Who holds the seat (default agent). human marks a human seat — addressable, never spawned; rejects every runtime-only field below and requires at least one contact identity |
contact | dict | human seats | External identities: slack_user_id, mattermost_user_id (a username, not an ID), atlassian_account_id (Jira+Confluence), github_login, gitlab_username, crewlet_operator_id. Each accepts a literal ID or exactly one whole-value ${VAR} reference, resolved at use time through the secret store and then the environment, by every consumer alike; values are whitespace-stripped, and a ${VAR} embedded inside a longer string is rejected at validation (see Humans in the Org Chart). No two seats may declare the same identity — an external account belongs to one person, and a duplicate is refused rather than silently sending one of them somebody else’s mail. crewlet_operator_id is the odd one: it names one of Tier A’s api.auth.tokens[].id, binding that credential to this seat so a person writing through the dashboard or the API acts as themselves rather than as a token. It is an attribution and never an address — the engine never sends as itself — so it is left out of rosters and colleague cards, and a seat carrying only this one is reachable through their dashboard queue rather than by an @-mention. It may not be anonymous, the id a disabled auth guard stamps on every caller: the literal is refused, and a ${VAR} resolving to it binds nobody |
availability | string | no | Human seats only — free-text availability rendered into rosters and lookup_colleague results |
goal | string | no | Individual mission statement |
backstory | string | no | Personality, background, expertise |
llm | string or dict | no | Named LLM provider key — the executor’s, and so the seat’s own model. Dict form adds default, review, subagent, auxiliary, judge, sandbox keys for the per-phase model split; there is no execute key, because llm is it |
llm_review / llm_subagent / llm_sandbox | string | no | Per-phase overrides (alternative to the dict-shaped llm) |
llm_auxiliary | string | no | Cheap/fast model used by reflection workers (PersistDecider, episode summariser) |
llm_judge | string | no | Cheap/fast model used by the round-cap extension judge; falls back to llm |
token_budget | dict | no | This seat’s own token ceilings per calendar window — day, week, month, each optional — on top of the company’s. See Token budgets |
learning_enabled | bool | no | Overrides the company’s learning setting for this seat alone. Unset inherits it; false skips reflection for this seat’s turns while still writing its episodes, which are cheap and useful to a person regardless |
handle | string | no | Custom identity slug (default: auto-derived) |
email | string | no | Agent email address |
unit | string | no | Root-level seats only. The name of the unit this seat belongs to; the seat is moved into that unit’s members before anything else reads the org chart, exactly as if it had been written under the unit. It is what PUT /config/roles/{handle} writes, so an API-created seat can land in a team without rewriting the unit’s block; a hand-written file normally nests the seat under its unit instead. A name that matches no unit leaves the seat at the root and is reported by crewlet validate as a dangling reference. On a seat already nested in a unit it moves nothing, and one naming a different unit is refused |
manages | list[string] | no | Names of the seats or units this seat manages; a unit expands to every seat in it and in its descendants |
responsibilities | list[string] | no | Role responsibilities |
behavioral_guidelines | list[string] | no | Behavioral rules |
mcp_env | dict | no | Per-agent MCP server credentials, keyed by server name — env vars for stdio servers, HTTP headers for http servers (e.g. atlassian.JIRA_USERNAME / atlassian.JIRA_API_TOKEN, confluence.CONFLUENCE_USERNAME / confluence.CONFLUENCE_API_TOKEN, slack.SLACK_MCP_XOXB_TOKEN, mattermost.MATTERMOST_TOKEN, github.Authorization: "Bearer …"). The per-agent tool-credential surface only — scope a server via its own filter (JIRA_PROJECTS_FILTER / CONFLUENCE_SPACES_FILTER) if needed. The unit’s project / space identity is project and space (below), not here |
integrations.github | dict | no | This seat’s own GitHub App — one per agent, because an app has exactly one bot identity. A person writes two fields: tier (read_only, the default, review or full_access; a hyphen is accepted for the underscore) and repos (the repositories it works in as owner/name; empty means every repository the installation covers). app_id, app_slug, installation_id, private_key and webhook_secret are written by the engine when the app is created and installed, the last two as ${VAR}s pointing at the secret store — see What lands where |
integrations.slack | dict | no | This seat’s own Slack app: bot_token, signing_secret, optional channel. Both credentials are required together — without the token the seat receives messages it cannot answer, without the secret its route answers 503 while the app’s settings page reports a healthy request URL. crewlet slack provision mints both into the ${VAR}s these fields point at |
integrations.mattermost | dict | no | Per-agent Mattermost transport identity (bot_token, optional username, optional channel). One credential, three readers: the same token is named as mcp_env.mattermost.MATTERMOST_TOKEN for the MCP subprocess, and the inbound websocket for this seat authenticates with it too |
project | string | no | Authored on a unit or root-level role (→ org.Unit.Project / org.Role.Project). The team’s tracker project as identity: an item that names nobody in the org chart routes to the unit lead, it is the project the team files work under, and on the engine’s own tracker it is what an item filed with no unit belongs to — whoever filed it (see the work tracker). Vendor-neutral — it names a native project or a Jira one, whichever tracker.backend the company runs, so switching backends does not rewrite the org chart. Not an MCP credential, and it does not scope knowledge reads. Keys are an upper-case letter plus 1–9 upper-case letters or digits (ENG, PROD), which is the shape every backend accepts |
space | string | no | Authored on a unit or root-level role (→ org.Unit.Space / org.Role.Space). The team’s knowledge container as identity: a page change that names nobody routes to the unit lead, and it is where the team writes. Vendor-neutral and shaped like project, above. It does not scope reads — read scope is the org-wide knowledge.scope only. The reserved containers (knowledge.skills_container, default TS, and knowledge.root_space, default HOME) are refused here |
workers | list[string] | no | Which worker templates this seat may delegate to. Empty means every one — a company that publishes three workers wants its seats using them, and requiring each seat to opt in turns a shared library into per-seat copy-paste. A name no template defines is refused at load |
sandbox | dict | no | The seat’s code sandbox gate. Absent means the seat is never offered the sandbox tool. enabled: true offers it; run_in picks where its code work runs (direct, container, e2b, or self — inside its own agent-mode executor run; empty inherits providers.sandbox.default_run_in); coding_agent (claude-code or opencode) overrides the provider default; pause_ttl_seconds (unset inherits the provider’s, 0 never holds a paused box, negative refused) and max_turns (unset inherits, 0 uncapped, negative refused) tune its runs; env is the environment injected into them and where external-service tokens are declared; mcp.servers names which of the seat’s MCP servers the coding agent gets; setup is per-seat provisioning applied after providers.sandbox.setup. A block with any of run_in, env, mcp or setup but enabled unset is refused, since none of it would be read |
placement | dict | no | Which nodes may run this seat; absent means any node that runs seats (see Placement). Agent seats only: a human seat is never claimed, so one carrying a placement is refused. node pins it to one node id and labels names pairs a node must all carry under node.labels; give both and both must hold. Everything is compared exactly against what a node advertises, so each part takes the shape a node’s own is held to. A label key follows the node-label grammar: not empty, at most 63 bytes, and no whitespace or unprintable character anywhere (zone, topology.example.com/rack). A label value may hold spaces inside it or be empty, but none around it, since a node’s values are trimmed. node is a node id: it starts with a letter or digit and holds only letters, digits, ., _ or -, at most 64 characters — and it is compared as written, so a ${VAR} is not resolved here. A selector outside these shapes could never match a node, so it is refused at validation, naming the field, rather than left for the running fleet to report as seats_unplaceable |
schedules | list | no | Role-scoped recurring tasks — see Schedules |
Worker templates
Section titled “Worker templates”workers: holds the reusable delegates a seat’s executor hands work to with the delegate tool, keyed by the name it types into that call. A template carries the half of a worker that does not change per task — the persona, the tool set, the model, and the shape of the answer — so a seat writes only the task.
workers: researcher: description: reads sources and reports findings with citations system_prompt: | You research things carefully and report only what you can point at. tools: [confluence_search, confluence_get_page] # a REQUEST, not a grant — see below model: fast # a providers.llm key; omit for llm_subagent max_turns: 12 # omit for turn_engine.delegation.max_turns output: # omit for the default {result, notes} type: object properties: findings: {type: string} citations: {type: array, items: {type: string}} required: [findings]| Field | Type | Required | Description |
|---|---|---|---|
description | string | yes | What this worker is for, written for the executor choosing one rather than for the operator. It is the only part of the template that reaches the parent’s prompt |
system_prompt | string | yes | The worker’s persona and standing instructions. The runtime preamble (no nesting, no colleague contact, how to answer) is appended, so a template never restates the boundary and cannot weaken it by forgetting to |
tools | list[string] | no | The tools this worker asks for. Empty means none at all, which is the right shape for a summariser. Naming a tool grants nothing: every name still passes the worker filter — the caller’s own live tools, minus the engine-control denylist, minus shared-surface writes — so workers: is never a privilege-escalation path |
model | string | no | A providers.llm key. An explicit key gets no fallback chain: an operator’s cheap-model choice that quietly ran somewhere else is worse than a refusal. Omit to take the seat’s llm_subagent chain |
max_turns | int | no | Tool rounds this worker may run. Refused at load — not clamped — above turn_engine.delegation.max_turns_ceiling: a template is an edit, and an operator overruled by a number nothing in the file mentions writes it again next time |
output | dict | no | The JSON Schema the worker’s submit_result publishes, so the answer comes back as fields the parent can index rather than prose it re-parses. Must be an object schema with 1–12 named properties, at most 3 levels deep, whose required entries are properties it actually has; keywords the engine does not read are passed through to the provider untouched |
See Turn Engine — Workers for what a delegate call looks like and how dependency waves work.
Schedules
Section titled “Schedules”Roles and units can own recurring work via a schedules: list (a
cron analogue). Each entry fires a task on its cron expression; see the
Scheduling concept doc for the full design.
units: - name: Backend type: team lead: Backend Lead schedules: # unit schedule, target defaults to `each` → every direct member runs it - name: daily-standup cron: "30 9 * * 1-5" # 5-field cron, evaluated in `timezone` timezone: Europe/Amsterdam # IANA tz; defaults to the company's `timezone` task: "Post your standup: shipped yesterday / on today / blockers." - name: weekly-report cron: "0 16 * * 5" target: lead # unit schedules: each | lead task: "Collect the week's progress and post a summary to the team channel." roles: - name: Backend Dev schedules: # role schedule → runs as this role - name: morning-smoke cron: "0 9 * * 1-5" task: "Run the smoke-test pipeline and triage failures"| Field | Type | Required | Description |
|---|---|---|---|
name | string | yes | Unique within the role/unit; part of the idempotency key |
cron | string | yes | Standard 5-field cron (min hour dom month dow), evaluated in timezone |
task | string | yes | Task prompt handed to the runner agent |
timezone | string | no | IANA timezone this one schedule fires in (default: the company’s timezone). Not Local or localtime |
target | string | no | Unit schedules only: each (default — every direct member) or lead (the effective unit lead). Ignored for role schedules; for a per-person task, use a role schedule |
enabled | bool | no | false keeps the schedule in config without firing (default true) |
timeout_seconds | int | no | Hard wall-clock cap on the scheduled turn (default 180) |
catchup | bool | no | Fire one recent missed tick on restart (default true) |
Scheduling requires a database (for the at-most-once scheduled_runs
ledger). Schedules are not inherited by child units.
Integrations
Section titled “Integrations”Inbound / notification integrations live under a single integrations: block — admin credentials, webhook secrets, and outbound transports. These carry only non-tool config; the MCP tool servers are declared in mcp_servers, and each agent’s per-server credentials live in role.mcp_env. It’s the inbound mirror of mcp_servers (outbound tool actions).
integrations: public_base_url: "${CREWLET_PUBLIC_URL}" # where VENDORS and sandboxes reach this deployment dashboard_base_url: "${CREWLET_DASHBOARD_URL}" # where PEOPLE reach its dashboard forge_app_id: "ari:cloud:ecosystem::app/your-forge-app-id" # Jira Cloud's delivery path
jira: url: "${JIRA_URL}" # a Data Center instance, or a Cloud site # cloud_id: "${JIRA_CLOUD_ID}" # an Atlassian Cloud id — give this OR url # site_url: "https://acme.atlassian.net" # with cloud_id: the base for links people open token: "${JIRA_API_TOKEN}" # API token (org read account) email: "${JIRA_EMAIL}" # Cloud only — the account's email, for Basic auth webhook_secret: "${JIRA_WEBHOOK_SECRET}" # Data Center only — HMAC-SHA256 secret
confluence: # the knowledge base url: "${CONFLUENCE_URL}" # a Data Center instance, or a Cloud site # cloud_id: "${CONFLUENCE_CLOUD_ID}" # an Atlassian Cloud id — give this OR url # site_url: "https://acme.atlassian.net" # with cloud_id: the base for links people open token: "${CONFLUENCE_API_TOKEN}" # API token (org read account) email: "${CONFLUENCE_EMAIL}" # Cloud only — the account's email, for Basic auth webhook_secret: "${CONFLUENCE_WEBHOOK_SECRET}" # Data Center only — HMAC-SHA256 secret
slack: # per-seat apps live on each role typing_status: always # working indicator: always (default) | addressed status_phrases: # optional — replaces the built-in wording, per phase execute: ["is nimbusing...", "is thinking very hard..."] review: ["is re-nimbusing...", "is double-checking..."]
mattermost: # self-hosted chat — transport AND inbound fleet enabled: true url: "${MATTERMOST_URL}" # instance base URL (required when enabled) team: nimbus # team slug (required when enabled) typing_status: always # working indicator: always (default) | addressed provisioning: # read only by `crewlet mattermost provision` channels: [town-square, engineering] # channels every agent bot joins
github: enabled: true # url: "https://github.example.com" # Enterprise Server; omit for github.com webhook_secret: "${GITHUB_WEBHOOK_SECRET}" # HMAC-SHA256 secret (required when enabled) token: "${GITHUB_ENGINE_TOKEN}" # read credential for participant fan-out provisioning: # read only by `crewlet github provision` org: nimbus # (organization holding the repositories) repos: [nimbus/api, nimbus/web] # (extra owner/repo entries to hook) org_webhook: auto # auto | true | false
gitlab: enabled: true url: "https://gitlab.com" # instance base URL (required when enabled) signing_secret: "${GITLAB_SIGNING_SECRET}" # 19.1+ whsec_ HMAC — the only verification mode, required token: "${GITLAB_ROUTING_TOKEN}" # optional read PAT → participants-based routing provisioning: # read only by `crewlet gitlab provision` group: nimbus-hq # top-level group agent service accounts join access_level: developer # developer | maintainerpublic_base_url— where vendors reach this deployment from outside, which is rarely whatapi.hostandapi.portbind. A bare origin (scheme://host[:port], no path, query or fragment): every consumer appends its own rooted path to it, so leftovers here land in the middle of every address. It is the base each provisioner’s-public-urldefaults to, the address every registered webhook points at, and the base of every vendor app’s redirect and manifest — never a link a person follows. With a public listener it is that listener’s published address. Write it as a whole${VAR}and staging and production answer at their own addresses off one revision; it is stored verbatim and resolved wherever a registration is built, so a process that cannot read it registers nothing rather than a webhook at the literal text of a variable.dashboard_base_url— where people reach this deployment’s dashboard, a bare origin likepublic_base_urland resolved the same way. Every link the engine composes for a person is built on it: what tool-skill prose reaches as the reserved variable${crewlet_base_url}— which is why a company may not declare askill_variablesentry of that name — and theurla page change’s notification carries. On a deployment with one listener it is usually the same address aspublic_base_url; behind a public listener it isapi.port’s address, since the public one answers the dashboard404. It never falls back topublic_base_url: unset, no link is composed, and a skill renders${crewlet_base_url}literally with askill_variable_unresolvedwarning.forge_app_id— verifies the Forge Invocation Token (FIT) on Cloud webhooks against Atlassian’s JWKS; theaudclaim must match. Required when Jira Cloud delivers through the Forge app.jira— the Atlassian tracker, served end to end. Giveurlorcloud_id, never both — they are two ways to name one instance andcrewlet validaterefuses the ambiguity.tokenis the org read account (an issue’s watchers are the one routing input a webhook never carries);emailswitches authentication to Cloud’s Basic scheme;site_urlis the human base for links when the instance is named by a cloud id.webhook_secretis required for Data Center and unused on Cloud, whose events arrive through the Forge app instead. Each seat’s own credential lives inmcp_env.atlassian(ormcp_env.jira) and is what the engine resolves its account id from — see Jira.confluence— the knowledge base, and the query-time search behind every turn’s “Relevant knowledge” block and thesearch_knowledgetool. Same address rule asjira:urlorcloud_id, never both.tokenis the org read account a seat with no Confluence credential of its own searches under; a seat WITH one searches as itself and Confluence enforces its page permissions natively.webhook_secretis required for Data Center and unused on Cloud. The knowledge backend is single-homed — the engine wires exactly oneknowledge.Searcher, because two would make an agent’s answer to “what do we already know about this” depend on which one was asked. Scope reads withknowledge.scope; publish withcrewlet confluence import. See Confluence.slack— the hosted chat backend. The org-level block carries no credentials at all: Slack gives each agent its OWN app, so the token and signing secret live on each role’sintegrations.slack, and this block holds only the working-indicator settings.typing_statustakesalways/addressedand defaults toalways, and Slack is where that default costs least: its indicator renders TEXT, so a phase change is something the person waiting can actually read, which is also what makesstatus_phrasesworth having here. There is nooff. Inbound events arrive per seat at/webhooks/slack/{handle}, verified against that seat’s own signing secret. Provision withcrewlet slack provision. See Slack Integration.mattermost— the self-hosted chat backend, and the one integration that is both inbound and outbound: enabling it starts the outbound transport and the websocket fleet that holds one connection per agent seat (Mattermost has no usable inbound webhook, so nothing has to reach the engine — no public URL, no tunnel).urlandteamare both required when enabled. Per-agent identity lives on each role’sintegrations.mattermost.bot_token, named again asmcp_env.mattermost.MATTERMOST_TOKENfor the MCP subprocess.typing_statustakesalways/addressedand defaults toalways, which on this backend is the expensive one: the indicator has to be re-asserted every few seconds rather than every 45, so a multi-minute turn costs one to two orders of magnitude more requests than Slack’s for strictly less information. Most Mattermost deployments wantaddressed. There is deliberately nostatus_phrases: Mattermost renders a fixed client-side indicator with no API for the text. Theprovisioning:sub-block is read only bycrewlet mattermost provision, not the engine. A company may run Mattermost and Slack together — they are different workspaces with different people in them, and an org migrating from one to the other runs both for a while. See Mattermost Integration.github— the hosted code host, served end to end.urlis optional: leave it unset for github.com, whose API lives on a different host rather than a path on the web UI, and name an Enterprise Server there — the REST base is derived either way.webhook_secretis required when enabled, and takes any string (GitHub signs with it verbatim, so unlike GitLab’s there is no shape to get wrong). The optionaltokenis a read credential for participant fan-out — a payload carries the author, assignees and requested reviewers but not who has commented or reviewed — and it is whatcrewlet github provisionregisters webhooks with. Each seat’s own credential lives inmcp_env.githuband is what the engine resolves its login from; a human seat is reached bycontact.github_logininstead. Theprovisioning:sub-block is read only bycrewlet github provision:org_webhook: autotakes one organization-level hook where the credential may (covering repositories created later) and falls back to one per repository where it may not. A company may run GitHub and GitLab together — they are two hosts with different repositories on them. See GitHub Integration.check_interval_seconds— how often a converged integration is read back, and therefore how long access somebody revoked by hand at a vendor goes unnoticed. Unset takes 600 (ten minutes); the floor is 60, and a shorter value is refused naming the field rather than clamped. It is the only thing that ever finds a revoked credential, weighed against what a converged pass costs at the vendor — one read per seat and per project, per surface, per interval — so a small company can afford60and a large one should lengthen it. Zero means unset, never “off”. See Integration Reconcile.gitlab— webhook config + boot-time identity resolution.urlandsigning_secretare both required when enabled — inbound webhooks are verified by the GitLab 19.1+ signing-token HMAC only (the plainX-Gitlab-Tokenscheme is unsupported; self-managed < 19.1 is not supported). The optionaltoken(a read-only PAT; the provisioner mints a dedicatedcrewlet-engineaccount for it) enables participants-based routing — comments and state changes reach everyone participating in the issue/MR, not just assignees and mentioned users.webhook_name(defaultcrewlet) names the hooks this deployment owns, so a change of public base re-points them instead of leaving a live orphan at every level; give two deployments watching one instance two names. The GitLab MCP server is ashared: falsemcp_serversentry — by default the officialglab mcp serve(stdio, spawned per-role by the engine, no separate server) — and each agent’s service-account PAT goes inrole.mcp_env.gitlab.GITLAB_TOKEN. Theprovisioning:sub-block is read bycrewlet gitlab provisionand by the engine’s own reconcile loop;mode(groupby default, orinstanceon a self-managed instance) says where a service account is owned and therefore which route creates it, mints its tokens and deletes it — the CLI’s-modeoverrides it for one run. See GitLab Integration.datadog— monitor alerts as inbound events, so a firing monitor wakes a seat the way a comment on a merge request does.webhook_tokenis required when enabled and is compared constant-time againstX-Crewlet-Token: Datadog attaches headers with fixed values only, so nothing varies with the payload to sign and this token is the whole authentication — which is why it must be at least 26 characters, the length the dashboard’s Generate button andcrewletitself mint. A route with nothing to compare against answers 503 rather than accepting a delivery.route_tois required and names an agent seat this company declares, or the literalnoneto dismiss alerts nobody owns on purpose; a handle no seat has, or one naming a human seat, is refused, and no seat may itself be handlednone.provisioningis required too —site(the region the keys were issued in, checked against Datadog’s own list),api_keyandapp_key— because the engine registers the webhook that makes an alert arrive at all, and an enabled block without it reports itself connected and receives nothing.webhook_name(defaultcrewlet) is the name of that definition and therefore the handle a monitor writes,@webhook-crewlet; give two deployments watching one organization two names.handle_tag(defaultcrewlet) is the monitor tag key that names a seat, so a monitor taggedcrewlet:sre-leadreaches that seat. See Datadog Integration.transports— outbound delivery transports (e.g.email). The Mattermost and GitLab transports are auto-derived from the sections above; this list adds any others. Jira is inbound-only by design and has no transport: an agent’s writes to an issue go through its own MCP tools under its own credential, never through the engine. Confluence is inbound-only for the same reason — a page an agent writes is written by its own MCP tools. Slack is the exception among the chat backends: its transport is per seat, derived from each role’s own app credentials rather than from an org-level block, so it is not listed here either.
Knowledge
Section titled “Knowledge”knowledge: backend: confluence # native (default) | confluence | none scope: ["ENG", "HANDBOOK"] # org-wide containers every agent can search (optional) skills_container: TS # tool-skill pages; excluded from routing and search root_space: HOME # the organisation's own pages, e.g. the root Onboarding page vectors: true # fuse semantic recall; unset derives from providers.embeddingsknowledge.backend is which knowledge base this company runs, and there is exactly one: two would make an agent’s answer to “what do we already know about this” depend on which was asked. Leaving it unset derives — confluence when an integrations.confluence block is declared, native otherwise — so a company that has configured nothing gets a wiki, and an Atlassian company that has not read this note keeps the backend it had. Naming native beside an integrations.confluence block is refused: pages would live in two places with nothing keeping them in step. none is a real posture — the ## Relevant knowledge block stays empty and search_knowledge is not registered.
knowledge.scope is the org-wide read scope, materialised onto org.Organization.KnowledgeScope and read by whichever searcher is wired. Empty means unscoped, and what unscoped MEANS differs by backend: natively it is the whole company, because the engine is the boundary — every reader is a seat of one company and there is no second account to launder a read through. On Confluence it is whatever the asking seat’s own account can read, which is why a credential-less seat searching unscoped there gets nothing: an unscoped query on the shared org token is how one seat reads a page its own account never could. Set it only to narrow to a curated floor.
knowledge.skills_container (default TS) holds tool-skill pages and knowledge.root_space (default HOME) holds the organisation’s own pages, starting with the root Onboarding page every seat reads first. Both are reserved: excluded from knowledge search and from routing, and refused as a unit’s own space. skills_container is three-valued — absent takes the default, a name takes that container, and an explicit "" turns tool skills off entirely.
knowledge.vectors fuses semantic recall into the search. Unset derives from whether providers.embeddings is configured — a company already paying for embeddings for its diary gets the better search — and an explicit true with no provider is refused, because there would be nothing to compute an embedding with.
Tracker
Section titled “Tracker”tracker: backend: native # native (default) | jira | noneWhich work tracker this company runs, on exactly the terms knowledge.backend runs on. Unset derives jira when an integrations.jira block is declared and native otherwise; native beside integrations.jira is refused, because work would be filed in two places and a unit’s project key would name two trackers. none is a company with no work tracker at all: nothing is filed or routed as a work item, and a seat is woken by everything else — chat, schedules, an a2a_ask, and the webhooks of the integrations the company has connected, a GitHub or GitLab issue among them.
The two axes are separate on purpose. A company running a native tracker against a Confluence wiki, or Jira against native pages, is an ordinary arrangement rather than a mixture to refuse — they are two products with separate routing and separate lead maps.
The native tracker’s own policy
Section titled “The native tracker’s own policy”tracker: backend: native native: # ONLY on a native company — an inbox # horizon on a company running Jira # describes nothing, and is refused inbox_retention_days: 365 # how long a person's inbox keeps a row # (default 365, 30..3650). THE HISTORY IT # POINTS AT IS UNTOUCHED — this is a # mailbox horizon, not an archive one, and # "what was I told about last year" is # answered by the history either wayEverything else a tracker could be told is either a fact about the operator — how they back up, how long their disk holds a replay window — which lives in Tier A under stream.tracker_retention, a fact about the whole company — the clock its dates mean is the top-level timezone — or a decision the engine makes once for everybody.
A native tracker or knowledge base needs a stream that survives a restart. Their write-ahead logs live on the stream, and an embedded stream with no stream.store_dir keeps its streams in memory, so a restart recreates them empty, and a node whose durable tables are ahead of a stream that restarted from nothing refuses to serve permanently, with no snapshot that helps. crewlet validate refuses that pair when it is given both documents, and so does the engine at boot. Either backend starts the log: a company on Jira whose knowledge base is the engine’s own, the default without Confluence, is refused the same way. Only a company whose tracker and knowledge base are both a vendor’s (or none) starts no log at all and is unaffected, which is why the rule needs both files to see.
MCP Servers
Section titled “MCP Servers”All MCP tool servers are declared here — including the Jira/Confluence (atlassian), Slack, and GitHub servers. A shared: true server runs once for everyone; a shared: false server is a per-role template, and each agent supplies its own credentials via role.mcp_env[name] (env vars for stdio, HTTP headers for http).
mcp_servers: # shared stdio server (one instance for all agents) - name: tavily command: npm args: ["exec", "--yes", "--", "tavily-mcp@latest"] env: TAVILY_API_KEY: "${TAVILY_API_KEY}"
# per-role stdio server — Jira + Confluence share one mcp-atlassian - name: atlassian shared: false # per-role: token from role.mcp_env.atlassian command: uvx args: ["mcp-atlassian"] env: JIRA_URL: "${JIRA_URL}" CONFLUENCE_URL: "${CONFLUENCE_URL}"
# per-role http server — remote GitHub MCP - name: github transport: http # stdio | http (default: stdio) shared: false # per-role: Authorization header from role.mcp_env.github url: "https://api.githubcopilot.com/mcp/" tool_prefix: "" # optional — prefix tool names tool_annotations: {} # optional — behavioural-hint overrides (see Tool Capabilities) startup_timeout_seconds: 120 # optional — connect + handshake + discovery request_timeout_seconds: 300 # optional — one tool callFull field reference: name (required; crewlet is reserved — it is the name every agent-mode box sees the seat’s own tool bridge under, and a server called the same would be replaced by it there), transport (stdio/http), shared (default true), command/args/env (stdio), url/headers (http), tool_prefix, tool_annotations, startup_timeout_seconds, request_timeout_seconds.
Timeouts
Section titled “Timeouts”An MCP server is another program, and the failure that matters is not an error but a silence — a server that launches and never completes the handshake, or answers discovery and then never returns from a tool call, raises nothing at all. The engine starts MCP servers on the seat-acquisition path, so a silent one does not merely lose its own tools: it holds up every seat behind it, for the life of the process. Both deadlines therefore always apply.
startup_timeout_seconds(default120) bounds launching the process (or opening the HTTP session), the protocol handshake, and the firsttools/list. The default suits auvx/npxserver whose package is not yet in the local cache — the slow case for a healthy server. Lower it for a server you launch from a local checkout.request_timeout_seconds(default300) bounds one tool call. It matches the MCP SDK’s own SSE-friendly HTTP read default, so a tool behaves the same over stdio and over HTTP. Raise it for a server whose tools genuinely run long (a large code search, a slow report); lower it for one that should always answer quickly, so a wedged call reaches the agent as a failed tool result it can react to instead of a turn that never ends.
A server that exceeds either deadline is logged and skipped; the rest of the company still starts.
Environment Variable References
Section titled “Environment Variable References”String values support ${ENV_VAR} syntax, keeping secrets out of config files. An unanswered reference resolves to the empty string. The two tiers resolve references differently, and it decides where one works:
- Tier B (the company) keeps every reference verbatim — in the store, in a backup, on
GET /config— and resolves it where the provider, transport or integration that uses it is constructed: from the secret store first (an encrypted table the provisioning CLIs can write into directly; inert until you store something), then the environment. It validates what is written, not what a reference resolves to, so a field with a closed set of values or a pattern — an enum, a seat handle, a unit id — takes a reference only where that field’s own entry says so. Three fields are resolved only when the whole value is one reference — a seat’s Slackbot_token, its Mattermostbot_tokenand its Mattermostusername— because their transports take anything else as the literal it looks like: write one${VAR}, or the value itself, and never a reference inside other text (bot-${SUFFIX}), which validation refuses rather than letting it reach the server braces and all. A Mattermostusernamewritten as a whole reference is the one patterned field that takes one. - Tier A (
crewlet.yaml) resolves from the environment only — it holds the keys to the secret store, so it can never read a value out of it — and does so at startup, before the file is decoded. What a reference becomes depends on the field it lands in. In a text field it is substituted whole or embedded (edge-${ZONE}) and the result is that text, character for character — so it works in any text field, one with a pattern or a closed set included (node.id: "${HOSTNAME}",logging.level: "${LOG_LEVEL}"), which is then judged on the value it resolved to. One nothing answers leaves the text empty, which a text field reads exactly as unset — its default where it has one (stream.typeandcoordination.typefall back toembeddedandlocal), and a refusal where a value is required — with one exception:logging.file.path, where empty is the setting “write no file”, so a path whose variable nothing answered stops the boot rather than silently running with no durable log. In a number or boolean field — a port, a replica count, anything*_seconds,*_bytesor*_hours, a count, a switch — a reference must be the whole value, and what it resolves to is read exactly as the same characters written there would be, trimmed of the space around it as every Tier A string is:api.port: "${API_PORT}"withAPI_PORT=8080isapi.port: 8080, withAPI_PORTread from a file ending in a newline too, andstream.debug: "${DEBUG}"withDEBUG=trueisstream.debug: true(a switch also takesyes,no,onandoff, as it does written out). Two shapes are refused there by name rather than guessed at: a reference with other text around it ("80${N}"), and one that resolves to nothing, since a number that quietly fell back to its default because a variable was missing would run with a value nobody chose.crewlet schema bootstrapencodes exactly that.
Only the braced identifier form is substituted — ${NAME} where NAME matches [A-Za-z_][A-Za-z0-9_]*. Bare $NAME and shell parameter expansions (${1:-x}, ${line#host=}) pass through untouched, so config-authored script content — a sandbox setup step’s helper script, say — survives intact.
providers: llm: default: api_keys: # resolved at startup (one or many) - "${LLM_API_KEY}"See Environment Variables for a full list.
Full Example
Section titled “Full Example”See the quickstart for a complete working config.
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.