Subscription LLM Backends
Run agents on a coding CLI you already pay a subscription for — Claude Code, Codex, Gemini CLI, Qwen Code, OpenCode, Cursor, Copilot, Grok, Muse Code, Kimi Code, Hermes or Pi — instead of a metered API key. Supported CLIs below is the full list.
The cli-agent provider type drives the vendor’s own command-line tool
as a headless text model. The CLI holds the operator’s OAuth login;
Crewlet never sees a password and never re-implements a vendor’s auth.
providers: llm: default: type: cli-agent model: sonnet # whatever the CLI's --model accepts cli: agent: claude-code# Already have the CLI logged in on this machine? Adopt that login:crewlet llm login default -from-host# Otherwise log in (or mint a headless token) inside Crewlet's own dir:crewlet llm login default -capture-tokencrewlet llm doctor default # verify before the first turnThe trade-off up front. A subscription CLI is a process, not an HTTP endpoint. It is slower to start, its tool calls ride a JSON envelope rather than a native tool-call channel, and most vendors’ terms are written for interactive use. It is an excellent fit for development, evaluation, and a small company you run yourself; a metered key remains the better fit for a large, latency-sensitive fleet. The two compose — see Falling back to a metered key. There is also a second way to spend a subscription that is not this page’s backend at all — an OAuth proxy behind an ordinary HTTP entry, with a different set of trade-offs you own rather than Crewlet.
Why this needs more than “shell out to a CLI”
Section titled “Why this needs more than “shell out to a CLI””Three problems have to be solved before a coding CLI can sit behind
LLMProvider, and each one is a section
below.
| Problem | Why it bites | Where it’s solved |
|---|---|---|
| Shared memory | A CLI keeps sessions, history, todos, and project notes under one home. Seven seats on one subscription would read each other’s transcripts. | Isolation |
| One model per entry | A CLI takes --model, so per-phase models mean several entries — which must not mean several logins. | Per-phase models |
| No tool channel | The tool loop needs tool_calls back. A CLI prints prose. | Tool calls |
| Browser-only auth | Vendor logins are OAuth (PKCE) with MFA — no password grant to script. | Authentication |
Isolation: the part that actually matters
Section titled “Isolation: the part that actually matters”One provider instance serves every seat in the org. Each call gets its own place to run:
<state_dir>/├── credentials/ # the subscription login — ONE per provider└── seats/ ├── sarah-chen/ │ ├── cache/ # XDG_CACHE_HOME — warm, holds no conversation │ ├── home/ # HOME + XDG config/data/state + vendor dirs │ └── work/<call-id>/ # cwd for one call, then deleted └── marcus-rivera/ └── …Between seats. Every seat gets its own home. HOME,
XDG_CONFIG_HOME, XDG_DATA_HOME, XDG_STATE_HOME, TMPDIR and the
vendor’s own relocation variable (CLAUDE_CONFIG_DIR, CODEX_HOME, …)
all point inside it. Nothing in a CLI’s state layout is reachable across
that boundary.
Between turns. Each profile declares its volatile_paths —
sessions, transcripts, history, todo state. They are deleted before and
after every call. Crewlet’s memory model is the agent
diary and the episode store; a second, invisible
memory inside the CLI would make turns non-reproducible and would carry
one task’s context into the next.
Between the seat and the host. The child process gets an
allowlisted environment — PATH, locale, TLS trust, proxy settings,
plus whatever the profile and your cli.env declare — never
the process environment. Inheriting the engine’s environment would hand every seat
the org’s SLACK_BOT_TOKEN and database DSN. It would also, for a
subscription backend, silently bill a metered ANTHROPIC_API_KEY that
happened to be exported.
Working directory. Each call runs in an empty, per-call scratch
directory that is removed afterwards — so a CLI that reads AGENTS.md /
CLAUDE.md from cwd, or writes scratch files, finds nothing from
anyone else.
Concurrency within a seat
Section titled “Concurrency within a seat”Delegated workers run in parallel and belong to the same agent, so they share that seat’s home — sharing memory between an agent and its own workers is harmless by definition. Pruning is keyed to the seat’s in-flight count crossing zero: the first concurrent call wipes and seeds, the last one to finish wipes again. Parallelism inside a seat is preserved; nothing crosses a seat or a turn.
Two modes: a text model, or the agent itself
Section titled “Two modes: a text model, or the agent itself”A cli-agent entry runs one of two ways, named on the entry with
cli.mode. There is deliberately no default beyond text and no
inference, because both are defensible for the same CLI on the same
seat.
mode: text (default) | mode: agent | |
|---|---|---|
| Who drives the loop | Crewlet’s own tool loop | the CLI’s |
| The CLI’s shell and editor | denied | enabled — that is the point |
| Crewlet’s tools | ride the prompt envelope; the engine executes them | reach the run over the MCP bridge |
| Where it runs | a subprocess of this engine | a sandbox box (cli.run_in) |
| Lifetime | one call, inside the phase | detached — outlives the turn, resumes it later |
| Needs | a login on this host | a login plus providers.sandbox and a reachable bridge URL |
Text mode is predictable: the tool log is the engine’s own, every call goes through the permission model and redaction, and it works with no reachable API. Agent mode is the vendor’s own harness: a real shell, a real editor and a real checkout, which is what makes it worth having for code work.
Only the executor branches. Every other phase — the reviewer, a
delegated worker, the summariser, the round-cap judge — is a text call on
the same entry, and a seat pointing llm at an agent-mode entry keeps
all of them. The reviewer in particular stays native and stays a separate
model call: the point of a reviewer is that it is not the thing being
reviewed.
Agent mode
Section titled “Agent mode”providers: llm: subscription: type: cli-agent model: sonnet cli: agent: claude-code mode: agent # text (default) | agent run_in: direct # direct | container | e2brun_in names a cell of providers.sandbox, and it
sits on the entry rather than on the seat because it is a property of
this runtime: the CLI’s subscription login lives on the engine host, so
direct and container reach it directly while a remote cell needs the
headless token instead. Want both? Make two entries and point each seat
at the one that is right for it — the same way you already choose between
two models. Empty takes providers.sandbox.default_run_in.
The cell is checked like a seat’s, at validation rather than at the
seat’s first turn: it must be one the catalogue configures, an empty one
needs a default to fall to, and agent mode in a company with no
providers.sandbox at all is refused outright. The backend behind the
cell is built for it, so run_in: container needs local.image exactly
as a seat’s would. self is not accepted here — it is a seat’s answer,
meaning “my code work rides my executor’s run”, and an agent-mode entry
is that run. Only entries some seat’s executor actually resolves to
are checked and built for; an entry nobody runs on is checked the day a
seat points at it. A seat that names no llm resolves the company-wide
fallback (the entry called default, else the first declared), so an
agent-mode entry can be reached without any seat naming it — but a
human seat never resolves one at all: it is addressable and never
spawned, so it runs no executor and reaches no entry.
The credential guard that refuses a remote run whose login cannot follow
it (see Code Sandbox) reads this
entry for an agent-mode run — the run is the executor — and the seat’s
llm_sandbox only for run_sandbox work.
An agent-mode run is a detached coding run and reuses that machinery whole: the executor phase suspends, the run’s state goes on a durable row in the coordination store, the completion poll collects it, and the same turn resumes — possibly in another process on another node, days later. Nothing about that is new for agent mode; see Code Sandbox.
The engine’s correctives do not reach an agent-mode run. The rounds
are the CLI’s own, so the engine’s tool loop never sees one end in prose
and cannot ask the model again: a run whose model writes its report as
text instead of calling submit_work over the bridge comes back with no
submission, and the phase is rescued as incomplete straight away. In
text mode the same reply is a round the engine’s loop re-prompts, naming
submit_work, up to twice (see Turn Engine).
What an agent-mode run has instead is its brief, which tells it the run
ends with that call.
The tool bridge
Section titled “The tool bridge”The seat’s tools cannot be shipped into the box: most are MCP children
holding the seat’s credentials, several are engine control, and the
whole point of a sandbox is that its credentials are not the company’s.
So the box gets exactly one MCP server — on the engine, named crewlet —
and every call comes back out through the same tools.Surface a native
loop would call. A tool denied natively is denied there; the skill guard,
the recording and the failure shape are the ones already tested. That
name is reserved: an mcp_servers entry may not use it, because the
bridge is written into every agent-mode box’s server list under it and
would replace the entry there. The bridge advertises the seat’s live
tool set — a tool the coding agent activates mid-run with activate_tool
is listed and callable on its next request, over the connection it already
holds — and every MCP session the box opened is closed the moment the run
ends, whatever ended it.
The endpoint is a per-run URL carrying a signed token that expires with
the run, and the session is closed the moment the run ends, whatever
ended it. Set CREWLET_MCP_BRIDGE_URL to a URL a sandbox can reach;
without it agent mode is refused at launch rather than started — a
coding agent with none of the seat’s tools cannot answer anybody, cannot
touch a ticket and cannot submit its work. The URL has to reach a
listener, so the node also needs api.port: a node that binds none
(api.port: 0) refuses agent mode by naming api.port, rather than
handing a box an endpoint nothing answers. With api.public.port set the
bridge is served on that public listener and nowhere else, so the URL
names it.
In a fleet, that URL must address the node itself, not a load
balancer in front of several. A
session is a live tool surface: the seat’s MCP children, its skill
guard, its per-turn recording, all objects in the process that claimed
the seat. Signing shares authentication across a fleet; it does not
and could not share the surface. Each node mints its endpoint from its
own value, so a per-node-addressable one is correct and a shared one
sends calls to peers that never held the session. Those answer 401
forever, and the response deliberately cannot say why — but the log
can, and does: mcp_bridge_unresolved names this setting when the
token is one the fleet signed.
The run ends by calling submit_work over that bridge, exactly as a
native loop ends by calling it locally, so the outcome vocabulary and the
rescue path are shared. A run that stops without submitting is rescued as
incomplete and judged on its record — the engine never reads the prose
a CLI happened to end with as a delivery.
Every bridged call is appended to the run’s own durable row, bounded at 200 with the middle dropped, because that log is the whole record a resume has: the process collecting a run may not be the one that launched it, and without it a restart mid-run would leave the reviewer judging a turn whose entire tool log is gone.
Each append also carries what the run’s bridged calls have cost the engine so far — the auxiliary rewrites the seat’s tools asked for and the workers the CLI delegated to — because those calls happen after the part of the turn that launched the run has ended and charged what it spent. The part of the turn that resumes from the run pays them to the turn’s work item, beside the run’s own tokens. Each call is recorded under the run it was made for — the bridge session is given the run’s name before its box exists — so a call still in flight when the turn launches its next run (a delegated worker outliving the CLI that asked for it) is dropped rather than landing on the next run’s record, where that run’s resume would pay for it as its own. A late call that finishes after its run was collected but before any next run is recorded on its own run and is not paid: the resume paid what the run’s record held when it was collected, so the task can fall short of the turn’s cost and never exceed it.
Code work inside the run
Section titled “Code work inside the run”A seat whose executor already holds a shell has no use for a second box
beside it — two filesystems, with the work in the one the turn cannot
see. That is what role.sandbox.run_in: self
names: code work rides the executor’s own run, run_sandbox refuses with
a message saying to use the shell it already has, and no second box is
provisioned. self is refused on any other runtime, and is not offerable
as a company-wide default.
The system prompt in text mode
Section titled “The system prompt in text mode”A coding CLI takes a prompt, not a conversation, so the transcript is
flattened into one text: ## system, ## user, ## assistant,
## tool result: name (call id). Everything below is about the one
section that does not belong there.
Folded into that text, a system prompt arrives as user content — the
model is asked to treat an ordinary message as its standing instructions,
underneath a vendor default that keeps saying what it is. So where a CLI
has a channel of its own for it, the profile names it in
system_prompt_args and the text travels there instead. Claude Code’s is
declared as:
system_prompt_args: ["--system-prompt-file", "{file}"]Two decisions in that one line, both measured against Claude Code 2.1.263 rather than assumed:
- Replace, not append. Asked “who are you?” as a PM seat, the same
model answers “I’m Agent PM at Nimbus, an AI assistant helping with
software engineering tasks and project work through Claude Code” with
the default prompt in force, and “I’m Agent PM at Nimbus, your AI
assistant for project management and technical collaboration” with it
replaced. A coding-agent identity over the top of whatever seat is
actually being served is not a cosmetic problem: it is the seat’s
standing instructions arguing with themselves. The default prose also
costs ~3.5k input tokens on every round of every phase, describing
tools this backend denies. Replacing it leaves the web tools working —
a
WebFetchprobe still fetches, which is whatcrewlet llm doctormeasures on every run. - The file variant, not the inline one. A seat’s system prompt
carries the org chart, the company’s policies, that seat’s backstory
and roster, and its
## Personal memoryand## Relevant knowledgeprefetches. On argv all of that is readable by every account on the machine through/proc/<pid>/cmdline, and bounded byARG_MAX(256 KB on macOS — the limit thecopilotprofile’s argv prompt already lives under). The text is written0600into the per-call working directory, which is created empty for one call and removed on release, so it cannot outlive the call or reach the next one.
{file} substitutes that path; {system} substitutes the text straight
into argv, for a CLI that offers no file variant. A profile that declares
neither leaves the system prompt in the transcript, which is what a CLI
with no such flag can take.
Which CLIs actually have one
Section titled “Which CLIs actually have one”Checked by running each CLI’s own --help, at the version named — not
from vendor documentation, which lags. That distinction is not academic:
every one of the three most recent rows was drafted from a vendor’s published
reference and then corrected by the installed binary. hermes’s docs
list --toolsets under a subcommand it also accepts at the top level;
pi’s README lists eight built-in tools where the Linux build registers
four; kimi-code’s tool table lists a WebSearch its binary only registers
once a search service is configured, and omits a whole Goal family it
ships. crewlet llm doctor prints written for beside the version you
actually have, and every field is a
config edit away:
cli.agent | version checked | flag | in the profile |
|---|---|---|---|
claude-code | 2.1.263 | flag, file: --system-prompt-file (--append-…-file appends) | system_prompt_args: ["--system-prompt-file", "{file}"] |
gemini-cli | 0.58.0 | env var, file: GEMINI_SYSTEM_MD — no flag exists | system_prompt_env: GEMINI_SYSTEM_MD |
qwen-code | 0.23.0 | both: QWEN_SYSTEM_MD (file) and --system-prompt (string) | system_prompt_env: QWEN_SYSTEM_MD |
grok | 1.0.13 | flag, string: --system-prompt-override (--rules appends) | system_prompt_args: ["--system-prompt-override", "{system}"] |
codex | 0.153.4 | none, on codex or codex exec alike | — |
opencode | 1.18.29 | none (--agent names a persona from its own config, not a per-call prompt) | — |
copilot | 1.0.83 | none (--no-custom-instructions only disables its own) | — |
cursor-agent | 2026.09.02 | none | — |
muse-code | 1.0.3 | none (AGENTS.md / CLAUDE.md only, and only in a trusted workspace) | — |
kimi-code | 0.42.x | flag, file: --agent-file — but an agent definition, not a bare prompt | system_prompt_args: ["--agent-file", "{file}"] + system_prompt_file |
hermes | 0.21.2 | none (SOUL.md in its own home, which --ignore-rules switches off) | — |
pi | 0.85.x | flag, string: --system-prompt (--append-system-prompt appends) | system_prompt_args: ["--system-prompt", "{system}"] |
There are two channels, and --help only shows one of them. A CLI may
take the prompt as an argument (system_prompt_args, with {file}
substituting a path and {system} the text itself) or from a file named by
an environment variable (system_prompt_env). Gemini CLI and its Qwen fork
have no flag at all and are configured entirely through the second — which is
why both were once recorded here as having no system-prompt channel, on the
strength of reading --help. A profile declares one or the other; naming
both is refused at load, because which copy a CLI honours when handed the
same prompt twice is the vendor’s business.
That exclusivity runs one level down as well. A single
system_prompt_args naming both {file} and {system} is refused for a
sharper reason than ambiguity: the renderer writes the private file and
substitutes the text into argv on the same pass, so
["--agent-file", "{file}", "--system-prompt", "{system}"] produced a 0600
file and put every byte of the seat’s identity in /proc/<pid>/cmdline.
Overrides are where a hand-written argv actually appears, so that is where
the check matters.
Prefer the file wherever both exist. {system} puts the seat’s system
prompt — the org chart, the policies, that seat’s own memory — into argv,
where /proc/<pid>/cmdline makes it readable by every account on the
machine. qwen-code is the one CLI offering both, and this profile takes the
variable for exactly that reason. grok has only the string form, but its
prompt already travels on argv (prompt_mode: argv) so nothing changes
there; on a shared host, cli.overrides.system_prompt_args: [] puts the
prompt back in the transcript.
A vendor’s project file is not a system-prompt channel. muse-code
reads AGENTS.md and CLAUDE.md, but only once a workspace has been
trusted — and Crewlet runs it in a per-call directory created empty and
never trusted, precisely so that no rule, skill or hook from a checkout the
engine does not control is admitted. Writing the seat’s identity there
would mean trusting that directory, which trades the whole guard for a
channel the transcript already provides.
Some channels take a FILE WITH A SHAPE, not a bare prompt. Kimi Code’s
only per-call prompt channel is --agent-file, which takes an agent
definition: YAML frontmatter naming the agent, describing it (required) and
listing its tools, with the body as the system prompt. Handed a bare
prompt it does not degrade — a file passed explicitly “must be valid,
otherwise the CLI reports the error and exits” — so the profile declares the
envelope alongside the channel:
system_prompt_args: ["--agent-file", "{file}"]system_prompt_file: name: crewlet-seat.md # the file to write; some vendors key on the extension template: | # {system} is the seat's own prompt --- name: crewlet-seat description: The Crewlet seat this run serves. tools: - WebSearch - FetchURL --- {system}system_prompt_file only means something on a channel that writes a file
({file} or system_prompt_env); on a {system} channel it is refused,
because the text would go to argv bare and the envelope would configure
nothing. A template with no {system} is refused for the same kind of
reason: every call would hand the CLI the same fixed file and no seat’s
identity.
A profile with a template passes the channel on every call, including a
request that carries no system prompt at all. That is not a detail: on the one
CLI that needs the envelope, the same frontmatter is also the only per-call
place its tools can be denied — so skipping it when there is nothing to put in
it would hand the vendor’s default agent, and every tool it has, to exactly
the calls that exist to prove the tools are off (crewlet llm doctor’s two
isolation probes send no system prompt).
Check the CLI you actually have. grok is the trap: xAI’s own CLI
(x.ai/cli, xai-org/grok-build)
and a same-named community package on npm both put a grok on PATH, and
they are different programs — the official one has the flag, the npm one
has none of this profile’s flags at all. If grok --version prints a
0.0.x, you have the other one.
Qwen Code is where the Gemini fork has diverged. It renamed its parent’s
variable (QWEN_SYSTEM_MD, and this build reads neither the other’s) and
added two flags its parent does not have, so “same shape as gemini-cli” no
longer holds here.
kimi-code is not the only trap of its kind either: the unscoped kimi-code
package on npm is a Claude Code wrapper that installs its own kimi, and
MoonshotAI’s is @moonshot-ai/kimi-code. A kimi --version printing a
1.0.x is the other one.
For the six with no channel at all, cli.overrides.system_prompt_args and
cli.overrides.system_prompt_env are how you adopt one the day its vendor
ships it — no engine release needed.
How the prompt itself travels
Section titled “How the prompt itself travels”The rendered prompt is the largest thing this backend hands a CLI and,
after the system prompt, the most sensitive: the flattened transcript, the
tool catalogue, the conversation and every tool result in it. prompt_mode
says which channel carries it, and the three are not equivalent:
prompt_mode | How | Ceiling | On /proc/<pid>/cmdline? |
|---|---|---|---|
stdin (default) | written to the child’s stdin | none | no |
file | written 0600 into the per-call working directory; prompt_args carries the path through {file} | none | no — only the path |
argv | appended as the last argument (or as prompt_args’ value) | ARG_MAX — ~2 MB on Linux, 256 KB on macOS | yes, in full |
argv is a last resort, taken only where a vendor offers nothing else —
copilot, grok, kimi-code, hermes and pi today. It has both failure
modes: a long transcript fails at exec rather than at the model, and every
account on the machine can read the conversation out of the process table
while the call runs. On pi the system prompt is on argv too, since its
only file channel is a static per-home SYSTEM.md; that is the one entry
where both halves of a turn are visible in the process table, and
cli.overrides.system_prompt_args: [] puts the identity back in the
transcript if that matters on your host.
Three of those five take the prompt as a FLAG’S VALUE rather than as a
positional argument, so the flag sits in prompt_args and stays adjacent to
the prompt: -p on grok, --prompt on kimi-code, -z on hermes. Put
one in complete_args instead and the next flag becomes the prompt. pi
takes it positionally and uses prompt_args: ["--"] for a different reason —
it is the one vendor here that offers an end-of-options separator, so a
transcript beginning with a dash is a prompt rather than an unknown flag.
-p does not mean the same thing on every CLI. It is print mode on
claude-code, cursor-agent, copilot and pi, the prompt’s value on
grok and kimi-code, and on hermes it selects a profile — a different
Hermes instance entirely. Borrowing the reflex from one profile to another is
how a seat ends up running against a profile named after its own transcript.
file is the same trade this backend already makes for the system prompt,
in the same directory, at the same mode, and for the same reasons. It needs
a vendor flag that takes a path; muse-code is the first built-in profile
whose CLI has one (muse exec --prompt-file), and its profile is
prompt_mode: fileprompt_args: ["--prompt-file", "{file}"]A file profile whose prompt_args contains no {file} is refused at
load: without it the CLI is run with no prompt at all, which a vendor
answers by opening an interactive session or printing usage — neither of
which looks like the configuration error it is.
Tool calls in text mode
Section titled “Tool calls in text mode”Every one of these CLIs has its own tools — file edits, shell, web fetch. In text mode Crewlet does not use them: they run in the CLI’s sandbox, invisible to the tool registry, the permission model, secret redaction, and the event stream. Routing agent work through them would fork the engine’s tool surface in two.
So every profile denies the CLI’s shell and file tools wherever the
vendor offers a way to, and each says how: a flag on the command line
(Claude Code’s --disallowedTools, Copilot’s --deny-tool, grok’s
--disallowed-tools, Codex’s read-only sandbox) or a settings file the
engine writes into the seat’s own home or the per-call working directory
before every call (Gemini’s settings.json, OpenCode’s opencode.json,
Cursor’s .cursor/cli.json, Muse Code’s run.toolset). The shell is the one that matters: the
seat’s home and environment are isolated, but the filesystem is not, and
a CLI with a shell on the engine host reads whatever the engine user can
read. A vendor with no such switch is declared as
local_tools: vendor-default with a note saying which switch is missing
— and crewlet llm doctor measures the stance rather than trusting
it (see Operating it).
A deny list, not an allow list, where a vendor offers both. grok has
both and the profile takes --disallowed-tools, which reads backwards
until you look at how each is applied. Its allowlist is honoured only if
every entry resolves: one name the build does not recognise and the whole
filter is skipped with a warning, leaving every tool enabled. The deny list
always applies and only warns about the entry that matched nothing. So a
name this profile gets wrong costs one tool on a deny list and costs
everything on an allow list — and a vendor renaming a tool is exactly the
drift these profiles are built to expect.
A refusal and a removal are not the same guard, and only one of them
is a denial. muse-code is the profile that makes the difference
concrete. Its --disable-shell and --disable-write flags read like tool
denials and are not: measured against 1.0.3, bash, bash_input,
write_file and edit_file are still advertised to the model,
byte-identical to the baseline surface — the flags refuse the call when it
comes. A model that can see a shell will try to use it, and every such
attempt is a wasted round inside a CLI whose tool log the engine never
sees. What actually removes them is run.toolset in the seeded
settings.json: an allowlist of exact tool names, validated against the
CLI’s own registry at startup, which replaces the surface outright. The
profile ships both — the allowlist because it is the denial, the flags
because a settings file that failed to apply should still refuse the call.
codex sits at the other end of the same distinction: its --sandbox read-only contains the shell rather than removing it, and reads stay.
That residual is why local_tools: denied is a claim crewlet llm doctor
measures rather than one you take on trust.
But an allow list is right where a stale name costs one tool. grok’s
rule is about grok’s implementation, not about allow lists — and
kimi-code inverts it. There, an entry the CLI does not recognise “is
reported with a warning” and simply matches nothing, so a name that went
stale costs that one tool rather than lifting the restriction. Read how each
vendor applies the list before choosing the shape; the flag’s name says
nothing about which way it fails.
The denial does not always live in a flag. kimi-code’s tool policy is
[tools] and [[permission.rules]] in config.toml — and config.toml is
where kimi login writes the OAuth reference and the model catalogue, so it
is a credential here and seeding a policy into it would destroy the login
it carries. What is left is the agent file, the same --agent-file that
carries the system prompt: its frontmatter allowlist is the denial. That is
the concrete reason a profile with a system_prompt_file passes its channel
on every call — see Which CLIs actually have one.
An approval prompt is a wedge in a headless run. A CLI that stops to
ask sits on the seat’s concurrency slot until timeout_seconds fires,
because there is nobody to answer. Most profiles therefore remove the
asking rather than the guard: OpenCode’s seeded policy is all allow and
deny and denies the tool that asks a person, and muse-code passes
--disable-approval, which is the posture its own vendor’s headless
guidance asks for — approval prompts off, the OS sandbox still on.
hermes needs neither, and reading its parser is what settled that. Its
one-shot flag documents its own posture — “approvals are auto-bypassed” — so
there is no prompt to wedge on and --yolo would widen what a run may do for
nothing. And --toolsets, which the vendor’s prose lists under hermes chat,
is a top-level flag whose own help says “Applies to -z/—oneshot”: passing
it replaces the enabled set for the invocation and the config file is not
consulted. So the denial is --toolsets web on argv — web_search and
web_extract and nothing else — rather than a seeded settings file. That is
the better shape wherever a vendor offers it: a flag fails at argument
parsing when it is renamed, where a settings key a vendor renamed is
ignored and the run quietly keeps every tool.
Web is the one local tool that stays on. A subscription seat must
not have less reach than the same CLI at a terminal, and a fetch is a
read — it never gates a delivery. Where a vendor gates its web tools
behind an approval a headless run cannot answer, the profile allows them
explicitly (--allowedTools WebFetch WebSearch, Copilot’s
--allow-tool); where its default web search answers from an offline
index, the profile switches it live (Codex’s web_search="live"). What
the CLI reads on the web is not in the engine’s event stream — the cost
of an unrecorded read, accepted. Seats on API models reach the web the
way they reach everything external, through the MCP servers you configure.
pi is the one profile that denies every tool outright (--no-tools),
and it is honest there for a reason no other vendor gives it: measured off
the request its CLI actually sent, the baseline is bash, edit, read
and write — there is no web tool in it to keep, and --no-tools sends the
tool array empty. Adopting the same flag on a CLI that has one would cut
the web silently, so a test refuses it for every other profile.
Both stances are profile fields, so an operator can override them like
any other — cli.overrides.local_tools, cli.overrides.local_tools_note,
and cli.overrides.seed_files (a list of {path, in: home|work, content};
lists replace wholesale).
Instead the CLI is used strictly as a text model, and the tool channel rides in the prompt:
-
The phase’s messages flatten into a labelled transcript.
-
The request’s tool definitions (
llm.Request.Tools) render as a JSON catalogue (name, description, JSON Schema). -
A response contract asks for one fenced JSON block:
{"message": "Short note to the operator, or an empty string.","tool_calls": [{ "name": "tool_name", "arguments": { "arg": "value" } }]} -
The reply is parsed back into
Completion.content+Completion.ToolCalls.
The parser is deliberately forgiving — it accepts the last fenced block,
a bare object, arguments as a JSON string, one call object written
without its list, null for no calls, and message / content /
text / response as synonyms. It is strict in the other direction: a
call list it can read no call from — strings for entries, nameless
objects, a string or a number where the list belongs — makes the reply
not an envelope at all, because reading it as one would report a model
that asked for no tools when it asked for some. When nothing parses, the whole reply
becomes assistant content with no tool calls, and in every phase that
has to end in a call the tool loop’s corrective re-prompt takes over —
the finishing corrective naming the phase’s submission. A request never
forces a call, so there is one contract and it is the permissive one —
“use an empty tool_calls list when no tool is needed” — on every phase
and on crewlet llm doctor’s smoke test alike, which therefore certifies
the shape a seat actually sends: the tool offered, the instruction naming
it, and nothing demanding the call. A malformed reply costs a round; it
never crashes a turn.
A call with no tools gets no contract. Auxiliary work (summarisation, the relevance filter) sends a plain prompt and reads a plain answer, with no envelope to get wrong.
Finding the answer in the CLI’s output
Section titled “Finding the answer in the CLI’s output”Before any of that, something has to decide which part of what the CLI
printed is the model’s reply. That is output plus text_paths on the
profile: text takes the whole of stdout, json reads one document and
jsonl concatenates every event that carries a text path, in stream
order. text_paths is a list so a vendor that moved the field
between releases needs no override — the first path that resolves to a
non-empty string wins, and an empty one falls through to the next.
An enveloped stream needs one more thing than paths. Muse Code wraps
every event in a single envelope shape and puts the kind in
payload_type, so payload.text is a token fragment of the reply on a
run.output.delta, a tool’s output on a tool.result, and the
assembled reply on run.terminal.completed. A path walk cannot tell the
three apart: it would splice the tool output into the answer and then
repeat the answer. event_type_path names where the kind lives and
text_events says which kinds carry the reply:
output: jsonlevent_type_path: ["payload_type"]text_events: ["run.terminal.completed"]text_paths: [["payload", "text"]]kimi-code needs the same pair for a different stream shape: its stdout is
an OpenAI-style chat stream discriminated by role, so a profile reading
every line would splice the user’s own prompt and any tool result into the
reply.
output: jsonlevent_type_path: ["role"]text_events: ["assistant"]text_paths: [["content"], ["content", "0", "text"]]Both or neither — one without the other configures nothing and is refused
at load, as is either on a profile that is not jsonl. They scope text
only: usage and error paths are still read across the whole stream,
because a stream reports those wherever it likes and the last value wins.
A profile that names an event filter also streams through it, so a
jsonl profile taking its answer from one terminal event delivers that
answer in a single delta at the end rather than pushing a tool’s output
through as though the model had said it.
Four outcomes, kept apart on purpose, because three of them used to be one — and only two of them are failures:
| What happened | What the engine does |
|---|---|
| A text path resolved to text | That text is the reply. |
| A text path resolved and every one was empty | The CLI answered with nothing. An answer, not a failure: a completion with empty content and the round’s real token usage attached, which is exactly what the openai and anthropic backends return for a model that spends its whole budget thinking. The tool loop corrects it — see When the CLI answers with nothing. |
| No text path resolved at all | The profile no longer matches the installed CLI. A retryable server failure that names text_paths, points at crewlet llm doctor, and prints the tail of what the CLI output so you can write the override. |
| The CLI printed nothing at all on a zero exit | Its own message, because neither of the two above can say anything true about output that does not exist. A retryable server failure carrying whatever it wrote on stderr, which is the only clue there is. |
Output that is not JSON at all is still an answer: a CLI that printed a
banner, a warning, or the vendor’s own sentence about a spent plan is
read as prose rather than refused, which is what lets the
limit sentinels be recognised on a
zero exit. Those sentinels are matched against the CLI’s whole stdout
and stderr, so a drifted profile still yields a real rate_limit with
the vendor’s own reset instant rather than a server fault.
Why the last two are failures rather than answers. They used to be
one case with the empty one, and the answer handed back was the CLI’s raw
stdout — on the reasoning that an operator would then see the shape and
write an override. They would, but only after it had been spoken as an
agent first: an empty result on a Claude Code envelope meant the
seat’s reply became
{"duration_api_ms":11377,…,"result":"","type":"result"}, the tool loop
appended that to the conversation and re-sent it every round, the
reviewer judged the turn on it, and the dashboard printed it as the
sentence the agent had said. The shape belongs in the error message,
where the only person who can act on it is the only one reading. Both
remaining failures are about this build not being able to read the CLI,
which is a fact about your machine — so the chain walking to another
entry is the right move.
When the CLI answers with nothing
Section titled “When the CLI answers with nothing”A model that spends its whole output budget on hidden reasoning exits 0, reports success, bills hundreds of output tokens and leaves the answer field empty. That is a model outcome, so the backend hands it back as an answer of nothing rather than dressing it as an outage:
- The round is charged. An empty answer costs tokens, and it used to be the one outcome that spent them without ever reaching a budget.
- The tool loop asks again. In every phase that finishes by a call — the executor, the reviewer, onboarding, a worker — the round gets the finishing corrective naming the phase’s submission, up to twice in a row on one allowance with a round that answered in prose: there, a round without the call ends the phase only into its rescue, so a second identical send is worth its round. A loop that does not finish by a call gets a single corrective naming what went wrong instead: there a prose answer is a legitimate finish, whatever the model writes next is the result, and there is no rescue for a second nudge to beat.
- It is counted.
empty_answer_roundson the phase record is the number of rounds that reached nobody. A seat whose model habitually answers nothing shows up there, and increwlet llm doctor, which names an empty answer as such rather than reportingit said: "".
If you see it repeatedly, the entry’s model is the field to change.
reasoning_effort and reasoning_budget_tokens are refused on a
cli-agent entry precisely so nobody spends an afternoon on them: they are
per-call API parameters and a headless coding CLI takes neither.
Token accounting
Section titled “Token accounting”Completion.InputTokens / output_tokens come from the CLI’s own
usage report where the profile can find one (Claude Code and Codex
report it; Gemini CLI’s shape varies by version). Where it can’t, the
counts are estimated at four characters per token — an approximation, but
budgets must
keep moving or a seat on this backend would run with no ceiling.
crewlet llm doctor tells you which of the two you are getting.
A CLI can report its tokens honestly and still not put them in the
answer. Hermes’s one-shot entry point is hermes -z, whose whole contract
is “single prompt in, final response text out, nothing else on stdout or
stderr” — so there is no envelope for a usage path to walk, and the figures
ride --usage-file instead. usage_file_args carries that path through a
{usage_file} placeholder; the file lands in the per-call working directory
and is read back through the same usage paths every other profile uses:
usage_file_args: ["--usage-file", "{usage_file}"]usage: input: [["input_tokens"]] output: [["output_tokens"]]Three rules travel with it. A report that never arrived is not a call that
cost nothing — the counts fall back to the estimate, because the vendor writes
the file “even when the run fails”, so its absence means the run did not get
that far and failing a completion the model answered over a count would throw
away work you paid for. A report whose keys this profile cannot read is
drift, not a zero-token turn, and falls back the same way rather than
charging zero. And both prompt counts come from the file or neither does: a
partial overlay would pair one source’s input count with another’s output
count, and the sum is what a budget is charged. (The two cache figures are
not part of that test — a provider that caches nothing reports neither, and
zero is the true answer there.) For the same reason a profile setting
usage_file_args must declare both usage.input and usage.output:
declaring one means every call quietly falls back to an estimate while
crewlet llm doctor reports the vendor’s own figures, which is precisely
what that line exists to settle.
Not every vendor’s richer channel is worth taking. pi has one —
--mode json streams every session event and carries real counts — and this
build deliberately uses print mode instead. Its answer lives at
message.content[N].text, an array of text, thinking and tool-call blocks
whose index moves with whether the model reasoned, and a model that
interleaves thinking with prose spreads one reply across several of them.
Reading it by index would be a guess about which part of the output is the
model’s reply, which is the one thing this backend must never get wrong. The
counts are estimated instead, and the profile’s comment carries the override
that takes the stream if you want it.
Authentication
Section titled “Authentication”Vendor subscription logins are browser OAuth with PKCE, often with SSO, MFA, or a one-time code. There is no username/password grant to script, and driving a headless browser to type into one would break on the vendor’s next login-page change. Crewlet does not pretend otherwise. What it does instead covers every deployment shape:
0. Already logged in on this machine? Adopt it
Section titled “0. Already logged in on this machine? Adopt it”crewlet llm login default -from-hostThe usual starting point: you have been running claude on this box
yourself for months. Crewlet does not use that login on its own —
the child process is given its own HOME, so your ~/.claude is
invisible to it, which is exactly the isolation the rest of this page
depends on. -from-host copies the CLI’s credential files out of your
home directory into Crewlet’s, once, on request.
It is a copy, not a redirect: agents never write into your personal
credential file, so a fleet refreshing a token mid-session is not a
surprise you get handed. The cost is that both copies then descend from
one refresh token, and a vendor that rotates refresh tokens can log out
whichever side refreshes second. Where the CLI mints a headless token
(option 2 below), that is the better answer and avoids the fork
entirely — crewlet llm login -from-host says so after it runs.
-home PATH reads from somewhere other than the engine user’s own home,
for a deployment where the engine runs as a different user than the one
that logged the CLI in.
crewlet llm doctor looks for a host login too, so “no sign-in” on a
machine where the CLI plainly works explains itself:
credentials : none on diskhost login : /home/you/.claude/.credentials.json (not adopted)token env : unsetsign-in : noneproblems: - no sign-in for "default": no credential files in /var/lib/crewlet/llm-cli/default/credentials, and nothing in the environment the CLI is given authenticates it — this machine has a login at /home/you/.claude/.credentials.json: adopt it with `crewlet llm login default -from-host`, or mint a headless CLAUDE_CODE_OAUTH_TOKEN with `-capture-token` (preferred: no shared refresh token)How doctor decides what the sign-in line says is under Operating
it.
1. Broker the vendor’s own login (any CLI)
Section titled “1. Broker the vendor’s own login (any CLI)”crewlet llm login defaultRuns the real claude auth login / codex login / opencode auth login
attached to your terminal — follow its prompts exactly as you would by
hand. The only thing Crewlet controls is where the credential lands:
in the provider’s isolated credentials/ directory, separate from your
personal CLI login on the same machine.
Each profile names its vendor’s own one-shot auth subcommand, which
prints its OAuth URL and returns once you have signed in. A profile that
named an in-session slash command instead would open an interactive
session rather than run a login: the session asks you to sign in itself,
then replays the slash command and asks a second time, and leaves you
in a REPL you have to interrupt — after a login that had already
succeeded. crewlet llm login returning you to your shell is the
signal that it worked; crewlet llm doctor <KEY> confirms it.
pi is the one CLI where the REPL is the login, and its profile says
so rather than pretending otherwise: /login there is a slash command and
the vendor ships no login subcommand at all. So the broker starts pi’s own
interactive session — ephemeral, with --no-session, in the isolated
credential directory — and you type /login, complete the browser flow, then
/exit. Declaring no login instead would have crewlet llm login tell you,
wrongly, that the CLI “authenticates on first use”.
Not every CLI has all three commands. kimi-code has kimi login (a
device-code flow) but no logout and no status subcommand — both are its
TUI’s slash commands, which a headless run cannot reach. hermes has
hermes auth and hermes status, but hermes auth logout requires a
provider name, which is yours to know rather than the profile’s to guess.
A profile declares only the commands its vendor actually has, and
crewlet llm logout / status say so plainly for the ones that do not.
2. Capture a headless token (best where it exists)
Section titled “2. Capture a headless token (best where it exists)”crewlet llm login default -capture-tokenRuns the vendor’s token-minting command (claude setup-token) and puts
the result in the encrypted secret store under the
profile’s token variable — CLAUDE_CODE_OAUTH_TOKEN for Claude Code.
Prefer this whenever the CLI offers it: no credential files to sync,
no refresh-token rotation, and it survives an ephemeral container with
no persistent volume.
Minting is interactive — the CLI opens the same browser sign-in as
option 1 — so its prompts and its sign-in URL are shown on your terminal
while the token itself is captured. The token never touches stdout, which
is what leaves -print-token free to pipe cleanly into your own secret
manager.
Already have a token from elsewhere?
pass show anthropic/crewlet-oauth | crewlet llm login default -token-stdin3. Username / password, where the CLI genuinely has one
Section titled “3. Username / password, where the CLI genuinely has one”vault read -field=password secret/gateway |Available for a profile that declares stdin_login — the built-in
opencode profile, an operator’s own wrapper, or a self-hosted gateway
CLI. The password is read from stdin or a declared environment variable,
never from argv (which is visible in ps and lands in shell history).
The Claude, Codex, and Gemini profiles deliberately leave stdin_login
unset, and the command says so rather than failing obscurely:
Error: the 'claude-code' CLI authenticates through the vendor's browserOAuth flow — there is no username/password login to drive. Run`crewlet llm login` (which brokers that flow), or`crewlet llm login -capture-token` where the vendor mints a headlesstoken. If your build of this CLI does accept a credential, declare itunder providers.llm.<key>.cli.overrides.stdin_login.If your CLI does accept a credential, wire it yourself — no Crewlet change needed:
cli: agent: custom overrides: binary: my-gateway-llm complete_args: ["--json"] stdin_login: args: ["login", "--user", "{username}"] stdin_template: "{password}\n"4. Move a login onto another host
Section titled “4. Move a login onto another host”The engine may run in a container that is rebuilt on every deploy, or on several hosts. Export the credential directory as one blob into the encrypted secret store:
crewlet llm export default -secret-storeThat engine restores it at boot when its own credentials/ directory is
empty, so a fresh container on the same store comes up already authenticated.
It is that node’s store and nothing else’s — the rows do not travel, and a
second host needs its own crewlet llm login, or the same bundle handed to it
through providers.llm[].auth.credential_bundle. The blob is validated on the way back in — only the
profile’s own credential paths, files only, size-capped — because an
archive is an execution surface if it is unpacked on trust.
Only the credential files travel. Sessions, history, and caches never go into a bundle.
Between two hosts that share no database, pipe it instead:
crewlet llm export default | ssh other-host crewlet llm import defaultimport reads the bundle from stdin — a credential on argv is visible in
ps and lands in shell history — and refuses to overwrite a login the target
already has. A host that has been running holds the fresher refresh token, and
restoring a boot-time blob over it is how a fleet logs itself out; crewlet llm logout <KEY> first if you mean to replace it.
5. A provider key in cli.env (hermes, pi, opencode)
Section titled “5. A provider key in cli.env (hermes, pi, opencode)”hermes, pi and opencode each front many model providers and read each
provider’s key from its own variable — OPENROUTER_API_KEY,
ANTHROPIC_API_KEY, NVIDIA_API_KEY and so on — so their profiles name no
single api_key_env. Their key goes in cli.env instead, as a ${VAR},
with auth.mode left at subscription:
providers: llm: herm: type: cli-agent model: anthropic/claude-sonnet-4 cli: agent: hermes env: OPENROUTER_API_KEY: "${OPENROUTER_API_KEY}"Their profiles declare env_sign_in, which is what makes doctor and
crewlet llm list count a credential-named variable in cli.env as a
sign-in (sign-in : cli.env sets OPENROUTER_API_KEY) and report one whose
${VAR} resolved to nothing. A name is credential-named when one of its
parts — the name cut at every character that is not a letter or a digit,
compared ignoring case — is KEY, KEYS, APIKEY, ACCESSKEY,
SECRETKEY, TOKEN, SECRET, PASSWORD, PASSWD, CREDENTIAL,
CREDENTIALS, AUTH, OAUTH or PAT. So OPENROUTER_API_KEY and
GH_TOKEN are, and HERMES_MAX_TOKENS and OPENAI_BASE_URL are not: a
count of model tokens is configuration. It is the same rule that refuses a
credential in a profile’s passthrough_env or env
(below). doctor counts the name, not whether
the provider accepts the key: the smoke test is what proves that — and for a
model served by an endpoint that takes no key, an answered smoke test is what
clears the “no sign-in” problem.
The key also goes into a coding box, the way a headless token does, so
it works in a remote cell; the rest of cli.env is not carried there.
Agent mode is opencode’s alone — hermes and pi have no
coding-agent runner — and a code-sandbox run on any of the three exports the
key into the box, where the seat’s coding agent reads it if it reads that
variable: opencode reads every provider’s own, Claude Code only
Anthropic’s. A key given in the seat’s role.sandbox.env instead is counted
by the remote-box check as well. A CLI that does not read its key from the
environment — kimi-code reads its metered key only from config.toml in
the credential directory — does not declare env_sign_in, and a key in its
cli.env is not counted.
Token refresh across seats
Section titled “Token refresh across seats”OAuth access tokens expire in hours, and the CLI refreshes them mid-run. Most vendors rotate the refresh token at the same time, so Crewlet syncs a changed credential file back to the shared directory when a seat’s generation closes — otherwise the whole fleet would be logged out at the next expiry. Two seats refreshing at the same instant can still race, exactly as two terminals running the vendor’s CLI would. A headless token (option 2) has no refresh file and sidesteps this entirely.
Supported CLIs
Section titled “Supported CLIs”cli.agent | Binary | Subscription | Notes |
|---|---|---|---|
claude-code | claude | Claude Pro / Max | claude auth login (and auth status / auth logout). claude setup-token gives a headless CLAUDE_CODE_OAUTH_TOKEN. Reports full usage incl. cache tokens. |
codex | codex | ChatGPT Plus / Pro | codex login. Streams JSONL events; runs --sandbox read-only. |
gemini-cli | gemini | Google AI Pro / free tier | First run starts the auth picker. GOOGLE_CLOUD_PROJECT passes through. |
qwen-code | qwen | Qwen OAuth | Gemini CLI fork; same shape. |
opencode | opencode | Anthropic / Copilot / any | opencode auth login; the one built-in profile with a credential login. A provider key goes in cli.env under that provider’s own variable, and doctor counts it — see Signing in with a provider key. |
cursor-agent | cursor-agent | Cursor seat | cursor-agent login. A Cursor API key (CURSOR_API_KEY, its api_key_env) is the headless alternative, reached through auth.mode: api-key. |
copilot | copilot | GitHub Copilot seat | Prompt goes on argv, so very long transcripts are bounded by ARG_MAX. Authenticates with a GitHub token, so GITHUB_TOKEN is its api_key_env — reached via auth.mode: api-key or inherit-env, never forwarded silently. |
grok | grok | xAI | xAI’s own CLI from x.ai/cli, not the same-named npm package. Accepts XAI_API_KEY (the variable its own signed-out message names) through auth.mode: api-key. |
muse-code | muse | Muse Code subscription (Everyday / High / Power Usage), or pay-as-you-go | muse login / muse logout; the browser sign-in stores ~/.config/muse/auth.json, which -from-host adopts. No status command — this CLI has none. Mints no headless token: META_API_KEY is a metered Model API key, reached through auth.mode: api-key. Runs muse exec --json, denies its tools through a seeded run.toolset, and puts the prompt in a file rather than on argv. Reports no token counts anywhere on its stream, so they are estimated. |
kimi-code | kimi | Kimi Code OAuth (Moonshot) | MoonshotAI’s own CLI, @moonshot-ai/kimi-code — not the unscoped kimi-code package on npm, which wraps Claude Code behind a proxy and installs its own kimi (a 1.0.x version is the other one). kimi login runs a device-code flow; there is no logout or status subcommand. Runs --output-format stream-json, because the default text output prefixes every line with • and re-wraps it, which destroys the tool envelope. Its tools are denied in the agent file that also carries the system prompt — measured, that is 25 tools down to 1, and an 8.8 KB vendor system prompt replaced by the seat’s own. Reports no token counts on its stream, so they are estimated. |
hermes | hermes | Nous Portal, or any of 30+ providers it fronts | hermes auth for the credential wizard, hermes status for state; hermes auth logout needs a provider name, so no logout is declared. Runs hermes -z, whose contract is the final response text and nothing else — so the tokens come from --usage-file instead. Tools are denied with --toolsets web on argv, which replaces the run’s enabled set; -p selects a profile on this CLI, not a prompt. A provider key (OPENROUTER_API_KEY, …) goes in cli.env, and doctor counts it — see Signing in with a provider key. |
pi | pi | Claude Pro/Max, ChatGPT Plus/Pro, GitHub Copilot — whichever it is logged into | @earendil-works/pi-coding-agent. No headless login: /login is a slash command, so crewlet llm login starts its TUI in the credential directory and you type it there. Denies every tool with --no-tools, which is honest here because its built-in set ships no web tool at all. Prompt and system prompt both on argv. Tokens estimated — see Token accounting. A provider key goes in cli.env under that provider’s own variable, and doctor counts it. |
custom | — | — | Ships nothing; declare everything under overrides. |
CLI flags drift — and that’s a config edit, not a release
Section titled “CLI flags drift — and that’s a config edit, not a release”Every field of every profile is replaceable from YAML. When a vendor renames a flag or changes its JSON shape, fix it in place:
cli: agent: codex overrides: binary: /opt/homebrew/bin/codex complete_args: ["exec", "--json", "--skip-git-repo-check", "-"] text_paths: [["item", "text"], ["msg", "message"]]Lists replace wholesale (position matters in an argv). Overrides are
validated against the profile model, so a typo is refused by crewlet validate and by every API write of the configuration — PUT and PATCH /config, a per-entity write, a revert — rather than by an agent’s first
turn or by every node’s apply after the API had already activated it. crewlet llm doctor prints the CLI
version the built-in profile was written against next to the version you
actually have.
A drift that only shows up at runtime — the flags still work, the JSON
still parses, and the answer field moved — names itself: the completion
fails with the text_paths this profile looked in and the output the CLI
actually produced. See
Finding the answer in the CLI’s output.
limit_markers and auth_markers drift the most quietly. Every other
field fails visibly when it goes stale — a renamed flag is a non-zero exit
doctor reports on the spot. A sentinel is matched verbatim against the
CLI’s own prose, so one the vendor has reworded simply never fires: a spent
plan then classifies as a fatal error instead of rate_limit, the
fallback chain never carries the seat onto
a metered key, and nothing says so until somebody hits their cap. If your
CLI’s wording differs from the built-in profile’s, override it:
cli: agent: claude-code overrides: limit_markers: - sentinel: "Usage limit reached" auth_markers: - sentinel: "Please run /login"Take the sentinel from what your CLI actually prints, not from what it used to print.
And say where it prints it, because the model’s own words are a haystack.
Sentinels are matched against the extracted answer as well as stderr, and
they have to be: Claude Code reports a spent plan on a zero exit with the
vendor’s sentence standing where the answer should be, so a marker confined to
stderr would never fire and the fallback chain would never carry the seat onto
a metered key. The cost is that a seat asked about rate limits can answer in
prose — “our quota resets hourly” — that trips a generic sentinel and benches
a perfectly good credential. This page’s own history has that bug: a bare
429 sentinel was dropped for exactly it.
So a CLI whose failures reach stderr and nowhere else declares that, and then nothing the model writes classifies anything:
marker_scope: stderr # or answer-and-stderr (the default)Three profiles narrow it, each on a measurement rather than an assumption:
kimi-code puts its classification on stderr with the answer stream carrying
only role: meta lines, pi leaves stdout empty and writes
<status>: <the provider's JSON> to stderr, and hermes writes
hermes -z: agent failed: … there while -z guarantees stdout is the reply
and nothing else. Narrow it only where the vendor’s behaviour makes the
answer an impossible place for the report — a profile that narrows it
wrongly stops recognising spent plans, which is silent until somebody hits
their cap.
And check where it prints it. A sentinel can only match what the CLI
puts on stdout or stderr, and one vendor puts the failure nowhere a
plain run would show it: muse exec writes the fixed string run ended with Failed to stderr and carries the real reason only in its
run.terminal.failed event. That is why the muse-code profile runs
with --json even though the event stream buys it no token counts — the
answer would read fine without it, and a spent plan would arrive as a
bare exit 1 that no marker could classify, so the seat would never fall
through to its metered key.
Profile fields
Section titled “Profile fields”Every field a profile has, and therefore every key cli.overrides accepts.
A key not in this table is refused by name. Every mapping field
(config_env, env, usage, stdin_login, system_prompt_file) merges
key by key; lists and single values replace wholesale.
| Field | What it is |
|---|---|
binary | The executable, looked up on PATH unless it is a path. |
vendor | The model family the CLI addresses (anthropic, openai, google, meta, …), for a coding agent that resolves <family>/<model>. |
written_for | The CLI version the profile was written against, printed by doctor beside the one installed. |
version_args | The argv of the version probe. |
complete_args | The argv of one completion, before the model and the prompt. |
model_args | The model flag, with {model} substituted. Required in effect: every entry names a model, so a merged profile with none is refused. |
prompt_mode | How the prompt travels: stdin (default), argv or file. |
system_prompt_args | The flag carrying the system prompt, with {file} (preferred) or {system}. Empty leaves it in the transcript. |
system_prompt_env | A variable naming a file the CLI reads its system prompt from; exclusive with system_prompt_args. |
system_prompt_file | {name, template}: the file a {file} channel writes, and the text around {system}. |
prompt_args | The flag introducing the prompt in argv mode, or carrying {file} in file mode. |
output | How stdout is encoded: json (default), jsonl or text. |
text_paths | Where the answer is in a json/jsonl document. |
event_type_path | Where a jsonl line names its event kind. |
text_events | The event kinds whose text is the answer. |
error_paths | A boolean the CLI sets when it failed despite exiting zero. |
usage | {input, output, cache_read, cache_write}: where the token counts are. |
usage_file_args | The flag asking the CLI to write its counts to a file, with {usage_file}. |
config_env | A vendor’s own relocation variable, mapped to a directory under the seat home. |
env | Fixed child environment. Never a credential, and never token_env or api_key_env. |
passthrough_env | Engine variables forwarded to the child. Never a credential. |
token_env | The variable a headless subscription token goes in (cli.auth.token). |
api_key_env | The variable a metered key goes in (api_keys under auth.mode: api-key). |
env_sign_in | The CLI reads its model provider’s key from its own environment, so a credential-named variable in cli.env signs it in. |
credential_paths | The login files, relative to the seat home. |
volatile_paths | Sessions, transcripts and history, deleted before and after every call. |
login_args | The vendor’s own interactive login, for crewlet llm login. |
capture_token_args | The command that mints a headless token on stdout (-capture-token). |
status_args | The command reporting who the CLI is logged in as. |
logout_args | The command revoking the login. |
stdin_login | {args, stdin_template, password_env}: a real credential login, where the CLI has one. |
limit_markers | Sentences recognising a spent plan — see above. |
auth_markers | Sentences recognising an expired login. |
marker_scope | Where markers match: answer-and-stderr (default) or stderr. |
host_credential_paths | Where the CLI keeps its login in a person’s own home, for -from-host. |
local_tools | The profile’s stance on the CLI’s own tools: denied or vendor-default. |
local_tools_note | Why a vendor-default stance is one. |
seed_files | {path, in, content} (in is home or work): settings files written before a call. |
Configuration reference
Section titled “Configuration reference”providers: llm: subscription: type: cli-agent model: sonnet # passed to the CLI's --model cli: agent: claude-code # or codex | gemini-cli | opencode # | muse-code | kimi-code # | hermes | pi | … mode: text # text (default) | agent — see above run_in: "" # agent mode only: direct | container | e2b
state_dir: /var/lib/crewlet/llm-cli/claude # Where credentials and per-seat homes live. Empty uses # $CREWLET_LLM_CLI_HOME/<key>, falling back to # ~/.crewlet/llm-cli/<key>. Point at a persistent volume when # the engine runs in an ephemeral container. A LITERAL PATH — # unlike the credential fields here it is not ${VAR}-expanded, # because it names where the engine keeps files rather than a # secret (the same reason the store path is a Tier A field).
timeout_seconds: 300 # one CLI invocation, wall clock max_concurrent: 4 # CLI processes at once
env: # extra child env, ${VAR}-resolved ANTHROPIC_SMALL_FAST_MODEL: haiku
auth: mode: subscription # subscription | api-key | inherit-env token: "${MY_OAUTH_TOKEN}" # else the profile's own token var credential_bundle: "${MY_BUNDLE}" # else CREWLET_LLM_CLI_<KEY>_CREDENTIALS
overrides: {} # any field under Profile fieldstimeout_seconds is separate from the entry’s own
timeout_seconds because the transports are not comparable: that one
bounds an HTTP attempt (default 600 s; a streamed call’s silence rather
than its length), while this covers a process
launch — a Node runtime costs seconds before the first byte — plus the
model call and the CLI’s internal retries. On breach the process group
is terminated (so the runtime’s helpers go too) and the call is reported
as timeout, which the role’s fallback chain retries.
max_concurrent: 4 keeps peak memory near 1.5 GB: each CLI is a
full Node or Rust runtime at roughly 200–400 MB resident, and an
unbounded fleet of seats starting turns together can exhaust a small
engine host. Subscription plans also throttle concurrency well below
what an API key allows, so a much higher number mostly buys rate-limit
errors. Raise it on a large host with a plan that permits it.
auth.mode defaults to subscription, not inherit-env, on
purpose: a backend that silently picked up a stray ANTHROPIC_API_KEY
would bill the metered account while you believed you were on a flat-rate
plan.
Every credential an entry configures has to reach the CLI, or a write
keeping it is refused. The engine puts a credential only into a variable
the profile names for it — api_key_env for a key, token_env for a
headless token — so each of these would run with the credential unused:
| Written | Refused because | Instead |
|---|---|---|
auth.mode: api-key on a CLI whose profile names no api_key_env (hermes, pi, opencode, kimi-code) | the key has no variable to go in | hermes, pi, opencode: leave auth.mode at subscription and set the key in cli.env under its provider’s own variable (above). kimi-code: put the key in the credential directory’s config.toml — crewlet llm login, or a credential bundle. A build of the CLI that does read a key variable: name it with cli.overrides.api_key_env |
auth.mode: api-key with no api_keys | there is no key to put in the variable | add one, e.g. api_keys: ["${ANTHROPIC_API_KEY}"] |
api_keys under subscription or inherit-env | only api-key mode reads it | set auth.mode: api-key, or drop it |
more than one api_keys value | an entry holds one login and rotates nothing | one key per entry; another key on an entry of its own in the seat’s fallback chain |
auth.token on a CLI whose profile names no token_env (every built-in profile but claude-code) | the CLI mints no headless token | remove it |
auth.token under api-key or inherit-env | api-key removes the token variable and inherit-env forwards the engine’s own | auth.mode: subscription |
auth.mode: inherit-env on a CLI that names neither variable | nothing would be forwarded | subscription, with the key in cli.env or a login |
the profile’s token_env or api_key_env in cli.env | cli.auth owns those: whether a value there reaches the CLI depends on auth.mode and on whether a credential resolves from the secret store or the engine’s environment, which the document cannot show | auth.token, or api_keys with auth.mode: api-key |
Each is checked against the profile with cli.overrides merged in, and each
refusal names the field and the route to take. They are admission rules:
every write path — crewlet validate, a PUT or PATCH of the config, a
setup write, a revert — refuses a document that breaks one, while a revision
being applied that breaks one is applied and warned about, because the
entry still runs (signed in by whatever the CLI does read) and refusing it
there would take a node off the fleet’s configuration during a rolling
upgrade. A profile that cannot drive its CLI at all — an override typo, a
missing binary, no model_args — is different: no node can build it, so it
is refused at apply too.
A profile’s passthrough_env may not name a
credential — an admission
rule like the ones above. Everything listed there is forwarded from
the engine’s own environment before auth.mode is consulted, so a key
named there would reach every seat whatever the mode says — the same
metered-bill-on-a-flat-rate-plan failure the mode exists to prevent. Use
it for genuine non-secret configuration (GOOGLE_CLOUD_PROJECT, a
region); a CLI’s key belongs in api_key_env or token_env, and
auth.mode: inherit-env is the deliberate way to let the host’s value
through.
Nor may a profile’s env, for the same reason and one more:
cli.overrides is neither ${VAR}-resolved nor redacted, so a key written
there would sit in the stored revision in plain text and come back on every
config read. A credential the CLI needs goes in cli.env, which is both, or
through cli.auth. env may not name the profile’s token_env or
api_key_env either, whatever they are called: cli.auth owns those, and
cli.env and cli.auth are both layered over env, so whether a value
there reached the CLI would turn on the mode and the secret store rather
than on anything the profile says.
Per-phase models
Section titled “Per-phase models”Nothing changes. Phase selection resolves by providers.llm key,
and the resolver never looks at a provider’s type — so llm,
llm_review, llm_subagent, llm_auxiliary,
llm_judge and llm_sandbox all behave exactly as they do for API
entries, including mixing the two kinds in one role and including
list-form fallback chains. See
Turn Engine — per-phase LLM models.
The one difference is where the model string goes: an API entry sends
it as a request field, a cli-agent entry passes it as --model. One
entry is still one model, so per-phase models mean one entry per model:
providers: llm: opus-sub: type: cli-agent model: opus cli: { agent: claude-code, state_dir: /var/lib/crewlet/llm-cli/claude } sonnet-sub: type: cli-agent model: sonnet cli: { agent: claude-code, state_dir: /var/lib/crewlet/llm-cli/claude } cheap: type: openai model: gpt-4o-mini api_keys: ["${OPENAI_API_KEY}"]
roles: - name: Engineer llm: [opus-sub, cheap] # the executor: subscription first, key when spent llm_review: sonnet-sub # the reviewer, on a cheaper subscription model llm_auxiliary: cheap # see the latency note belowPoint them at the same state_dir and they share one login. Both
entries above then use the credential directory a single crewlet llm login wrote, instead of needing one login per entry — the default
state_dir is per provider key precisely so unrelated providers do
not collide, which means entries that should share must say so. They
also share one set of per-seat homes and one generation, so a call on
one entry never wipes a live call on the other.
Entries sharing a state_dir must drive the same CLI: two different
CLIs disagree about which files are credentials and which are
conversation memory, so each would prune the other’s state.
crewlet validate rejects that combination by name.
Concurrency is per entry. max_concurrent caps one provider’s
processes, so two entries at the default of 4 can run 8 CLI processes at
once. Size them together against the engine host’s memory.
Auxiliary work is the one phase to think twice about. Every
reflection, summarisation and the turn-start relevance prefetch goes through
llm_auxiliary, and each one pays a process launch on this backend.
Point it at a cheap API model unless you have no key at all. (Crewlet
does handle the latency: the auxiliary call’s 60-second deadline is
widened to the provider’s own cli.timeout_seconds, so a subscription
aux provider is not cut off mid-call — it is simply slower than it needs
to be.)
Falling back to a metered key
Section titled “Falling back to a metered key”A spent subscription window arrives as prose on a successful exit
(“Usage limit reached · continuing automatically”). Crewlet matches that
wording — and, where the CLI relays the API’s own error instead, the
"type":"rate_limit_error" in it — and
reports it as rate_limit, which is retryable, so the ordinary
provider chain carries the role
onto a metered key for the rest of the window and back again afterwards,
with no operator intervention:
providers: llm: subscription: type: cli-agent model: sonnet cli: { agent: claude-code } metered: type: anthropic model: claude-sonnet-5-5 api_keys: ["${ANTHROPIC_API_KEY}"]
roles: - name: Engineer llm: [subscription, metered] # subscription first, key as backstopAn expired login classifies as auth, which is also retryable, so the
chain keeps the seat working while you re-run crewlet llm login.
What a sentinel may be
Section titled “What a sentinel may be”Both recognitions are limit_markers / auth_markers on the profile: a
literal substring the vendor emits, plus (where it carries one) the field
holding the reset instant, so the retry-after is a datum rather than a
guess. They are matched against whatever the CLI printed — which on a
healthy call is the model’s own answer — and a match benches the
credential for a cooldown and hands the seat to the next entry in the
chain.
So a sentinel has to be the vendor’s wording, and a profile is refused
at crewlet validate if one contains no letters. The rule exists because
a shipped profile carried sentinel: "429", and three digits matched as
a substring is not a rate limit — it is a model quoting an HTTP status, a
stack trace’s line number, a token count, or any ten-digit epoch. Every
one of those took a working subscription out of service.
The other way a sentinel stops working is quieter — the vendor reworded
it, so it simply never fires. Both are fixed the same way, with
cli.overrides.limit_markers; see
CLI flags drift
for the shape, and take the wording from the sentence your CLI actually
printed, which a fatal failure carries verbatim so that you can.
The other shape: an OAuth proxy in front of an HTTP entry
Section titled “The other shape: an OAuth proxy in front of an HTTP entry”Everything above drives the vendor’s CLI as a process. There is a second way to spend a subscription, which Crewlet supports without knowing anything about it: run a proxy that holds the OAuth login itself and re-exposes it as an ordinary Anthropic- or OpenAI-shaped HTTP endpoint, then point a normal provider entry at it.
Crewlet needs no cli-agent block for this. It is an HTTP entry like any
other, and base_url is all that changes:
providers: llm: # The proxy speaks the Anthropic Messages API. subscription-proxy: type: anthropic model: claude-sonnet-5-5 base_url: "${LLM_PROXY_URL}" # e.g. http://127.0.0.1:8317 api_keys: ["${LLM_PROXY_KEY}"] # the proxy's OWN inbound key
# Or it speaks the OpenAI wire format. subscription-proxy-oai: type: openai-compatible model: gpt-5 base_url: "${LLM_PROXY_URL}/v1" api_keys: ["${LLM_PROXY_KEY}"]base_url is not an openai-compatible field. It is honoured on
anthropic and openai entries too — it is only required for
openai-compatible, which has no vendor default to fall back to. The
same field is what points an entry at a corporate egress proxy or an
Anthropic-API gateway, and the code sandbox forwards
an anthropic entry’s value to Claude Code as ANTHROPIC_BASE_URL.
Which header your proxy will be handed depends on the entry’s type, because each backend sends its vendor’s native one:
| Entry type | Credential arrives as |
|---|---|
anthropic | x-api-key — and only that. The backend builds its client with WithoutEnvironmentDefaults, which deliberately disables the SDK’s own bearer-token path so an ambient ANTHROPIC_AUTH_TOKEN cannot redirect a company’s auth |
openai, openai-compatible | Authorization: Bearer — and nothing else from the engine’s environment. The SDK would add OpenAI-Organization, OpenAI-Project and every OPENAI_CUSTOM_HEADERS line from the process it runs in; the backend undoes each, so a proxy never receives headers an operator exported for some other tool. The OpenAI embeddings provider builds its client the same way |
The request is shaped for the model the entry names. An anthropic
entry sends what its Claude model accepts — adaptive thinking and an effort
level on the current generation, never a temperature there — read from the
Claude model table.
A proxy that answers to its own alias (model: sonnet) rather than a Claude
id is shaped as the current generation; if that alias is an older model,
name it with claude_model: claude-haiku-4-5 (or whichever it is), or the
fields only newer models take will be refused through the proxy.
The api_keys value is the credential for the proxy, not for the
vendor: the vendor login lives inside the proxy. Rotation, cooldowns and
the fleet-shared credential bench all apply to that inbound key as they
would to any other.
What you gain, and what becomes yours
Section titled “What you gain, and what becomes yours”Against the cli-agent backend you get a real HTTP provider back: native
tool calls instead of the in-prompt JSON envelope, no
process launch per call, and whatever token accounting the endpoint
reports. Against a metered key you get flat-rate cost.
What you take on is everything this page’s design otherwise handles for you:
- The isolation guarantees do not apply. Per-seat homes, volatile path pruning and the allowlisted child environment exist because a CLI keeps conversation state under one home. A proxy is one process serving every seat, so whatever session, cache or history it keeps is shared across your whole company — that is the proxy’s design to answer, not Crewlet’s.
crewlet llmsees only the HTTP half of it.list,loginand the rest buildcli-agentproviders only, so there is no login state to report and keeping the proxy authenticated is a separate operational job.doctordoes examine ananthropicentry pointed at a proxy — it sends one real round and certifies a tool call comes back — but a proxy rarely serves/v1/models, so the model check reads not served there, and anopenaientry is not examined at all.- A spent window is not translated. The prose sentinel
that turns “Usage limit reached” into a retryable
rate_limitis the CLI backend’s. Over HTTP you get whatever status the proxy returns, and only a 429 / 401 / 403 / 402 / 408 / 5xx is retryable; anything else is fatal and the role’s fallback chain will not walk to the next provider. Check what your proxy returns on an exhausted plan before you rely onllm: [proxy, metered].
Before you choose this
Section titled “Before you choose this”A proxy that spends a subscription rather than an API key has to present itself to the vendor as the vendor’s own client. In practice that means reproducing a specific client build’s headers, its beta flags, sometimes its TLS fingerprint, and often injecting that client’s system prompt ahead of yours — which quietly changes what your prompts say and where prompt-cache breakpoints land.
Vendor terms generally do not permit a third-party client to route
requests through consumer subscription credentials, and vendors have
enforced that. Crewlet’s cli-agent backend is on the other side of
that line by construction: it runs the vendor’s own unmodified CLI,
logged in by you, as a child process — Crewlet never sees a password,
never re-implements an auth flow, and never impersonates a client.
Pointing base_url at a proxy is a supported configuration and a
decision you are making, exactly as the note at the end of this page
says about plan terms generally.
None of this applies to an ordinary gateway — LiteLLM, a corporate egress proxy, a self-hosted vLLM — reached through the same field with a key you were issued. That is just an endpoint.
Operating it
Section titled “Operating it”crewlet llm list # providers, agent, model, sign-increwlet llm doctor # verify them all, anthropic entries toocrewlet llm doctor default -no-smoke # skip the real completionscrewlet llm status default # ask the CLI who it's logged in ascrewlet llm logout default # revoke locally + delete credentialsdoctor is the command that matters. It checks the binary is on PATH,
runs its version probe, reports how it is signed in, says whether
token counts will be real or estimated — and then runs three real
completions: a smoke test with a real tool, because a profile can look
perfect and still not produce a parseable tool call; a shell probe,
which asks the CLI to run date +%s with its own shell and believes it
only if the answer is within minutes of the engine’s clock (a model can
write a token it was asked to echo, but it cannot guess the current
epoch); and a web probe, which asks the CLI to fetch a public
endpoint that reports its own clock and applies the same test:
provider : subscriptioncli agent : claude-codemode : text (a model behind the engine's tool loop)binary : /usr/local/bin/claudeversion : 2.0.31 (Claude Code)written for : Claude Code CLI 2.x (`claude --version`)state dir : /var/lib/crewlet/llm-cli/subscriptioncredentials : presenttoken env : setsign-in : credential files in /var/lib/crewlet/llm-cli/subscription/credentials; headless token in CLAUDE_CODE_OAUTH_TOKENtoken usage : reported by CLIsmoke test : ok — 812 in / 34 outlocal tools : denied by profile — probe: refusedweb : ok — fetched https://www.cloudflare.com/cdn-cgi/traceproblems : noneOn an agent-mode entry the report carries two more lines, and both check something that fails at a seat’s first turn and nowhere earlier:
mode : agent (the CLI runs the executor)agent runtime : runner: "claude-code" is registered : tool bridge: https://engine.example.comThe runner line is whether this build can actually drive that CLI as
a coding agent. Agent mode reuses the coding-agent runners rather than
growing a second way to invoke the same binary, and there are two of
them — claude-code and opencode. An entry naming any other CLI in
agent mode validates cleanly, appears in the schema and reports a
configured provider, then refuses the moment a seat has work. The
tool bridge line is CREWLET_MCP_BRIDGE_URL: without it every
agent-mode launch is refused, because a coding agent with none of the
seat’s tools cannot answer anybody, touch a ticket or submit its work.
A text-mode entry reports neither, rather than reporting that it would not work in a mode it is not in.
How doctor judges a sign-in
Section titled “How doctor judges a sign-in”The sign-in line names every route that authenticates the CLI, read off
the environment the CLI is actually given, after auth.mode has set and
removed what it does — never off the configuration beside it. So a token
that auth.mode: api-key removes is not a sign-in, and crewlet llm list’s
SIGN-IN column is the first route the line names, so the two commands
cannot disagree. The routes are credential files in the entry’s directory,
a headless token, an API key under api-key, a token or key forwarded under
inherit-env, and a credential-named
variable in cli.env on a CLI that reads its key from its environment.
${VAR} references, and the variables inherit-env forwards, are resolved
in the process running crewlet llm — the secret store first, then that
shell’s environment — so run it with the engine’s environment to get the
engine’s answer.
When none of those authenticates the CLI, doctor reports “no sign-in” as a
problem, naming first a route you configured whose ${VAR} resolved to
nothing. One case is not a fault: a CLI pointed at an endpoint that takes no
key (a hermes or pi entry aimed at a local server through cli.env), or
one holding a key in a configuration file of its own. The configuration
cannot tell that from a missing key, so the smoke test decides: when it is
answered, the problem is dropped and the sign-in line says the CLI
authenticates some way the engine does not hand it.
A profile that says denied while the shell ran is a problem naming the
installed version, because the vendor’s switch is not taking effect on
it; a vendor-default profile whose shell ran is a problem stating the
trust you are taking on; a web tool that could not fetch is a problem
pointing at the vendor’s sandbox flags and the egress proxy the child
environment was told about.
One caveat worth stating plainly: doctor spends three real completions.
On a subscription that is a few thousand tokens of your plan’s allowance,
which is why -no-smoke exists for a scripted health check that runs
often — it skips all three and says so on each line.
The metered key behind it is examined too. A role written
llm: [subscription, default] falls through to an anthropic entry
exactly when the plan is spent, which is the worst moment to learn that
entry is refused on every call. So doctor reports on every anthropic
entry beside the cli-agent ones:
provider : defaulttype : anthropicmodel : claude-sonnet-5-5profile : claude-sonnet-5-5endpoint : https://api.anthropic.comkeys : 1 (3f9a0c1b2d4e)request : thinking adaptive (summarized), effort high, max_tokens 128000, never a temperature, reasoning before a tool change shedmodels api : served — Claude Sonnet 5.5 (claude-sonnet-5-5), max_tokens 128000, input …smoke test : ok — called crewlet_smoke (streamed), 1204 in / 61 outproblems : noneThe request line is what a phase sends, read from the Claude model
table
— on a model that binds its thinking to the tools it was written under, that
includes shedding the reasoning from before a tool change;
the models api line is the vendor’s own record of the model, and any
disagreement between the two is listed under drift — a problem when
it puts a field the model refuses on every call, a note when the table is
only more cautious than the model. The smoke test is one round in a
phase’s shape (no tool choice forced, effort low), billed to the entry’s
key, and -no-smoke skips it; the Models API read bills nothing and runs
either way. The full rules are in the CLI
reference.
Limits and caveats
Section titled “Limits and caveats”- The CLI runs on the engine host. It must be installed there, and the engine process must be able to execute it. This is not a remote service.
- Agent mode needs more than a login. A CLI with a coding-agent
runner, a
providers.sandboxcatalogue to place the run in, and a bridge URL a box can dial.crewlet llm doctorchecks all of it — see Operating it. - Code work needs one more decision. A subscription can back the
code sandbox. On any backend including remote E2B,
what travels into the box is the environment that signs the CLI in: a
headless token (
crewlet llm login <key> -capture-token, and Claude Code in the box bills your plan), anapi-keyentry’s key, whatinherit-envforwards, and thecli.envkey of a CLI that reads one (opencode,hermes,pi). The credential files never travel to a remote box: they carry a refresh token whose rotation is shared fleet state. So a CLI that signs in only through its files — Codex, Gemini CLI or Kimi Code with no key — needs a local cell (providers.sandbox.localplusrun_in: directorcontainer), where the coding agent runs on the engine host and reads the login directly. - Latency. Process launch plus model call. Point
llm_auxiliaryat a cheap API-key model rather than paying process startup for every summarisation. - No streaming.
stream()completes and yields one chunk. Crewlet’s phases usecomplete(), so nothing in the engine is affected. - No
reasoningswitch. The CLI’s own plan carries its reasoning configuration and exposes no per-call setting; pick a reasoning model viamodelinstead. Settingreasoning: trueon acli-agententry is rejected at validation. - Check the vendor’s terms. Subscription plans are generally written for interactive use by the subscriber. Running a fleet of agents on one may not be permitted by your plan — that is a decision for you, not something Crewlet can decide for you. It is the sharper question for the proxy shape, where a third-party client is presenting itself as the vendor’s own.
See also
Section titled “See also”- Overview — Provider Layer
- Turn Engine — per-phase LLM models
- Secret Store — where tokens and credential bundles live
- Code Sandbox — the other place Crewlet runs a coding agent
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.