Skip to content
You are reading documentation for unreleased main. This page is not in 0.1 yet.

Subscription LLM Backends

Run agents on a coding CLI you already pay a subscription for — Claude Code, Codex, Gemini CLI, Qwen Code, OpenCode, Cursor, Copilot, Grok, Muse Code, Kimi Code, Hermes or Pi — instead of a metered API key. Supported CLIs below is the full list.

The cli-agent provider type drives the vendor’s own command-line tool as a headless text model. The CLI holds the operator’s OAuth login; Crewlet never sees a password and never re-implements a vendor’s auth.

providers:
llm:
default:
type: cli-agent
model: sonnet # whatever the CLI's --model accepts
cli:
agent: claude-code
Terminal window
# Already have the CLI logged in on this machine? Adopt that login:
crewlet llm login default -from-host
# Otherwise log in (or mint a headless token) inside Crewlet's own dir:
crewlet llm login default -capture-token
crewlet llm doctor default # verify before the first turn

The trade-off up front. A subscription CLI is a process, not an HTTP endpoint. It is slower to start, its tool calls ride a JSON envelope rather than a native tool-call channel, and most vendors’ terms are written for interactive use. It is an excellent fit for development, evaluation, and a small company you run yourself; a metered key remains the better fit for a large, latency-sensitive fleet. The two compose — see Falling back to a metered key. There is also a second way to spend a subscription that is not this page’s backend at all — an OAuth proxy behind an ordinary HTTP entry, with a different set of trade-offs you own rather than Crewlet.


Why this needs more than “shell out to a CLI”

Section titled “Why this needs more than “shell out to a CLI””

Three problems have to be solved before a coding CLI can sit behind LLMProvider, and each one is a section below.

ProblemWhy it bitesWhere it’s solved
Shared memoryA CLI keeps sessions, history, todos, and project notes under one home. Seven seats on one subscription would read each other’s transcripts.Isolation
One model per entryA CLI takes --model, so per-phase models mean several entries — which must not mean several logins.Per-phase models
No tool channelThe tool loop needs tool_calls back. A CLI prints prose.Tool calls
Browser-only authVendor logins are OAuth (PKCE) with MFA — no password grant to script.Authentication

One provider instance serves every seat in the org. Each call gets its own place to run:

<state_dir>/
├── credentials/ # the subscription login — ONE per provider
└── seats/
├── sarah-chen/
│ ├── cache/ # XDG_CACHE_HOME — warm, holds no conversation
│ ├── home/ # HOME + XDG config/data/state + vendor dirs
│ └── work/<call-id>/ # cwd for one call, then deleted
└── marcus-rivera/
└── …

Between seats. Every seat gets its own home. HOME, XDG_CONFIG_HOME, XDG_DATA_HOME, XDG_STATE_HOME, TMPDIR and the vendor’s own relocation variable (CLAUDE_CONFIG_DIR, CODEX_HOME, …) all point inside it. Nothing in a CLI’s state layout is reachable across that boundary.

Between turns. Each profile declares its volatile_paths — sessions, transcripts, history, todo state. They are deleted before and after every call. Crewlet’s memory model is the agent diary and the episode store; a second, invisible memory inside the CLI would make turns non-reproducible and would carry one task’s context into the next.

Between the seat and the host. The child process gets an allowlisted environment — PATH, locale, TLS trust, proxy settings, plus whatever the profile and your cli.env declare — never the process environment. Inheriting the engine’s environment would hand every seat the org’s SLACK_BOT_TOKEN and database DSN. It would also, for a subscription backend, silently bill a metered ANTHROPIC_API_KEY that happened to be exported.

Working directory. Each call runs in an empty, per-call scratch directory that is removed afterwards — so a CLI that reads AGENTS.md / CLAUDE.md from cwd, or writes scratch files, finds nothing from anyone else.

Delegated workers run in parallel and belong to the same agent, so they share that seat’s home — sharing memory between an agent and its own workers is harmless by definition. Pruning is keyed to the seat’s in-flight count crossing zero: the first concurrent call wipes and seeds, the last one to finish wipes again. Parallelism inside a seat is preserved; nothing crosses a seat or a turn.

Executor (seat B)Worker (seat A)WorkspaceExecutor (seat A)Executor (seat B)Worker (seat A)WorkspaceExecutor (seat A)acquire(A) — in-flight 0→1prune A/home, seed settings + credentialsacquire(A) — in-flight 1→2 (same home, no prune)acquire(B) — separate home, own prunerelease — in-flight 2→1 (no prune)release — in-flight 1→0sync refreshed credential out, prune A/home

Two modes: a text model, or the agent itself

Section titled “Two modes: a text model, or the agent itself”

A cli-agent entry runs one of two ways, named on the entry with cli.mode. There is deliberately no default beyond text and no inference, because both are defensible for the same CLI on the same seat.

mode: text (default)mode: agent
Who drives the loopCrewlet’s own tool loopthe CLI’s
The CLI’s shell and editordeniedenabled — that is the point
Crewlet’s toolsride the prompt envelope; the engine executes themreach the run over the MCP bridge
Where it runsa subprocess of this enginea sandbox box (cli.run_in)
Lifetimeone call, inside the phasedetached — outlives the turn, resumes it later
Needsa login on this hosta login plus providers.sandbox and a reachable bridge URL

Text mode is predictable: the tool log is the engine’s own, every call goes through the permission model and redaction, and it works with no reachable API. Agent mode is the vendor’s own harness: a real shell, a real editor and a real checkout, which is what makes it worth having for code work.

Only the executor branches. Every other phase — the reviewer, a delegated worker, the summariser, the round-cap judge — is a text call on the same entry, and a seat pointing llm at an agent-mode entry keeps all of them. The reviewer in particular stays native and stays a separate model call: the point of a reviewer is that it is not the thing being reviewed.

providers:
llm:
subscription:
type: cli-agent
model: sonnet
cli:
agent: claude-code
mode: agent # text (default) | agent
run_in: direct # direct | container | e2b

run_in names a cell of providers.sandbox, and it sits on the entry rather than on the seat because it is a property of this runtime: the CLI’s subscription login lives on the engine host, so direct and container reach it directly while a remote cell needs the headless token instead. Want both? Make two entries and point each seat at the one that is right for it — the same way you already choose between two models. Empty takes providers.sandbox.default_run_in.

The cell is checked like a seat’s, at validation rather than at the seat’s first turn: it must be one the catalogue configures, an empty one needs a default to fall to, and agent mode in a company with no providers.sandbox at all is refused outright. The backend behind the cell is built for it, so run_in: container needs local.image exactly as a seat’s would. self is not accepted here — it is a seat’s answer, meaning “my code work rides my executor’s run”, and an agent-mode entry is that run. Only entries some seat’s executor actually resolves to are checked and built for; an entry nobody runs on is checked the day a seat points at it. A seat that names no llm resolves the company-wide fallback (the entry called default, else the first declared), so an agent-mode entry can be reached without any seat naming it — but a human seat never resolves one at all: it is addressable and never spawned, so it runs no executor and reaches no entry.

The credential guard that refuses a remote run whose login cannot follow it (see Code Sandbox) reads this entry for an agent-mode run — the run is the executor — and the seat’s llm_sandbox only for run_sandbox work.

An agent-mode run is a detached coding run and reuses that machinery whole: the executor phase suspends, the run’s state goes on a durable row in the coordination store, the completion poll collects it, and the same turn resumes — possibly in another process on another node, days later. Nothing about that is new for agent mode; see Code Sandbox.

The engine’s correctives do not reach an agent-mode run. The rounds are the CLI’s own, so the engine’s tool loop never sees one end in prose and cannot ask the model again: a run whose model writes its report as text instead of calling submit_work over the bridge comes back with no submission, and the phase is rescued as incomplete straight away. In text mode the same reply is a round the engine’s loop re-prompts, naming submit_work, up to twice (see Turn Engine). What an agent-mode run has instead is its brief, which tells it the run ends with that call.

The seat’s tools cannot be shipped into the box: most are MCP children holding the seat’s credentials, several are engine control, and the whole point of a sandbox is that its credentials are not the company’s. So the box gets exactly one MCP server — on the engine, named crewlet — and every call comes back out through the same tools.Surface a native loop would call. A tool denied natively is denied there; the skill guard, the recording and the failure shape are the ones already tested. That name is reserved: an mcp_servers entry may not use it, because the bridge is written into every agent-mode box’s server list under it and would replace the entry there. The bridge advertises the seat’s live tool set — a tool the coding agent activates mid-run with activate_tool is listed and callable on its next request, over the connection it already holds — and every MCP session the box opened is closed the moment the run ends, whatever ended it.

The endpoint is a per-run URL carrying a signed token that expires with the run, and the session is closed the moment the run ends, whatever ended it. Set CREWLET_MCP_BRIDGE_URL to a URL a sandbox can reach; without it agent mode is refused at launch rather than started — a coding agent with none of the seat’s tools cannot answer anybody, cannot touch a ticket and cannot submit its work. The URL has to reach a listener, so the node also needs api.port: a node that binds none (api.port: 0) refuses agent mode by naming api.port, rather than handing a box an endpoint nothing answers. With api.public.port set the bridge is served on that public listener and nowhere else, so the URL names it.

In a fleet, that URL must address the node itself, not a load balancer in front of several. A session is a live tool surface: the seat’s MCP children, its skill guard, its per-turn recording, all objects in the process that claimed the seat. Signing shares authentication across a fleet; it does not and could not share the surface. Each node mints its endpoint from its own value, so a per-node-addressable one is correct and a shared one sends calls to peers that never held the session. Those answer 401 forever, and the response deliberately cannot say why — but the log can, and does: mcp_bridge_unresolved names this setting when the token is one the fleet signed.

The run ends by calling submit_work over that bridge, exactly as a native loop ends by calling it locally, so the outcome vocabulary and the rescue path are shared. A run that stops without submitting is rescued as incomplete and judged on its record — the engine never reads the prose a CLI happened to end with as a delivery.

Every bridged call is appended to the run’s own durable row, bounded at 200 with the middle dropped, because that log is the whole record a resume has: the process collecting a run may not be the one that launched it, and without it a restart mid-run would leave the reviewer judging a turn whose entire tool log is gone.

Each append also carries what the run’s bridged calls have cost the engine so far — the auxiliary rewrites the seat’s tools asked for and the workers the CLI delegated to — because those calls happen after the part of the turn that launched the run has ended and charged what it spent. The part of the turn that resumes from the run pays them to the turn’s work item, beside the run’s own tokens. Each call is recorded under the run it was made for — the bridge session is given the run’s name before its box exists — so a call still in flight when the turn launches its next run (a delegated worker outliving the CLI that asked for it) is dropped rather than landing on the next run’s record, where that run’s resume would pay for it as its own. A late call that finishes after its run was collected but before any next run is recorded on its own run and is not paid: the resume paid what the run’s record held when it was collected, so the task can fall short of the turn’s cost and never exceed it.

A seat whose executor already holds a shell has no use for a second box beside it — two filesystems, with the work in the one the turn cannot see. That is what role.sandbox.run_in: self names: code work rides the executor’s own run, run_sandbox refuses with a message saying to use the shell it already has, and no second box is provisioned. self is refused on any other runtime, and is not offerable as a company-wide default.

A coding CLI takes a prompt, not a conversation, so the transcript is flattened into one text: ## system, ## user, ## assistant, ## tool result: name (call id). Everything below is about the one section that does not belong there.

Folded into that text, a system prompt arrives as user content — the model is asked to treat an ordinary message as its standing instructions, underneath a vendor default that keeps saying what it is. So where a CLI has a channel of its own for it, the profile names it in system_prompt_args and the text travels there instead. Claude Code’s is declared as:

system_prompt_args: ["--system-prompt-file", "{file}"]

Two decisions in that one line, both measured against Claude Code 2.1.263 rather than assumed:

  • Replace, not append. Asked “who are you?” as a PM seat, the same model answers “I’m Agent PM at Nimbus, an AI assistant helping with software engineering tasks and project work through Claude Code” with the default prompt in force, and “I’m Agent PM at Nimbus, your AI assistant for project management and technical collaboration” with it replaced. A coding-agent identity over the top of whatever seat is actually being served is not a cosmetic problem: it is the seat’s standing instructions arguing with themselves. The default prose also costs ~3.5k input tokens on every round of every phase, describing tools this backend denies. Replacing it leaves the web tools working — a WebFetch probe still fetches, which is what crewlet llm doctor measures on every run.
  • The file variant, not the inline one. A seat’s system prompt carries the org chart, the company’s policies, that seat’s backstory and roster, and its ## Personal memory and ## Relevant knowledge prefetches. On argv all of that is readable by every account on the machine through /proc/<pid>/cmdline, and bounded by ARG_MAX (256 KB on macOS — the limit the copilot profile’s argv prompt already lives under). The text is written 0600 into the per-call working directory, which is created empty for one call and removed on release, so it cannot outlive the call or reach the next one.

{file} substitutes that path; {system} substitutes the text straight into argv, for a CLI that offers no file variant. A profile that declares neither leaves the system prompt in the transcript, which is what a CLI with no such flag can take.

Checked by running each CLI’s own --help, at the version named — not from vendor documentation, which lags. That distinction is not academic: every one of the three most recent rows was drafted from a vendor’s published reference and then corrected by the installed binary. hermes’s docs list --toolsets under a subcommand it also accepts at the top level; pi’s README lists eight built-in tools where the Linux build registers four; kimi-code’s tool table lists a WebSearch its binary only registers once a search service is configured, and omits a whole Goal family it ships. crewlet llm doctor prints written for beside the version you actually have, and every field is a config edit away:

cli.agentversion checkedflagin the profile
claude-code2.1.263flag, file: --system-prompt-file (--append-…-file appends)system_prompt_args: ["--system-prompt-file", "{file}"]
gemini-cli0.58.0env var, file: GEMINI_SYSTEM_MD — no flag existssystem_prompt_env: GEMINI_SYSTEM_MD
qwen-code0.23.0both: QWEN_SYSTEM_MD (file) and --system-prompt (string)system_prompt_env: QWEN_SYSTEM_MD
grok1.0.13flag, string: --system-prompt-override (--rules appends)system_prompt_args: ["--system-prompt-override", "{system}"]
codex0.153.4none, on codex or codex exec alike—
opencode1.18.29none (--agent names a persona from its own config, not a per-call prompt)—
copilot1.0.83none (--no-custom-instructions only disables its own)—
cursor-agent2026.09.02none—
muse-code1.0.3none (AGENTS.md / CLAUDE.md only, and only in a trusted workspace)—
kimi-code0.42.xflag, file: --agent-file — but an agent definition, not a bare promptsystem_prompt_args: ["--agent-file", "{file}"] + system_prompt_file
hermes0.21.2none (SOUL.md in its own home, which --ignore-rules switches off)—
pi0.85.xflag, string: --system-prompt (--append-system-prompt appends)system_prompt_args: ["--system-prompt", "{system}"]

There are two channels, and --help only shows one of them. A CLI may take the prompt as an argument (system_prompt_args, with {file} substituting a path and {system} the text itself) or from a file named by an environment variable (system_prompt_env). Gemini CLI and its Qwen fork have no flag at all and are configured entirely through the second — which is why both were once recorded here as having no system-prompt channel, on the strength of reading --help. A profile declares one or the other; naming both is refused at load, because which copy a CLI honours when handed the same prompt twice is the vendor’s business.

That exclusivity runs one level down as well. A single system_prompt_args naming both {file} and {system} is refused for a sharper reason than ambiguity: the renderer writes the private file and substitutes the text into argv on the same pass, so ["--agent-file", "{file}", "--system-prompt", "{system}"] produced a 0600 file and put every byte of the seat’s identity in /proc/<pid>/cmdline. Overrides are where a hand-written argv actually appears, so that is where the check matters.

Prefer the file wherever both exist. {system} puts the seat’s system prompt — the org chart, the policies, that seat’s own memory — into argv, where /proc/<pid>/cmdline makes it readable by every account on the machine. qwen-code is the one CLI offering both, and this profile takes the variable for exactly that reason. grok has only the string form, but its prompt already travels on argv (prompt_mode: argv) so nothing changes there; on a shared host, cli.overrides.system_prompt_args: [] puts the prompt back in the transcript.

A vendor’s project file is not a system-prompt channel. muse-code reads AGENTS.md and CLAUDE.md, but only once a workspace has been trusted — and Crewlet runs it in a per-call directory created empty and never trusted, precisely so that no rule, skill or hook from a checkout the engine does not control is admitted. Writing the seat’s identity there would mean trusting that directory, which trades the whole guard for a channel the transcript already provides.

Some channels take a FILE WITH A SHAPE, not a bare prompt. Kimi Code’s only per-call prompt channel is --agent-file, which takes an agent definition: YAML frontmatter naming the agent, describing it (required) and listing its tools, with the body as the system prompt. Handed a bare prompt it does not degrade — a file passed explicitly “must be valid, otherwise the CLI reports the error and exits” — so the profile declares the envelope alongside the channel:

system_prompt_args: ["--agent-file", "{file}"]
system_prompt_file:
name: crewlet-seat.md # the file to write; some vendors key on the extension
template: | # {system} is the seat's own prompt
---
name: crewlet-seat
description: The Crewlet seat this run serves.
tools:
- WebSearch
- FetchURL
---
{system}

system_prompt_file only means something on a channel that writes a file ({file} or system_prompt_env); on a {system} channel it is refused, because the text would go to argv bare and the envelope would configure nothing. A template with no {system} is refused for the same kind of reason: every call would hand the CLI the same fixed file and no seat’s identity.

A profile with a template passes the channel on every call, including a request that carries no system prompt at all. That is not a detail: on the one CLI that needs the envelope, the same frontmatter is also the only per-call place its tools can be denied — so skipping it when there is nothing to put in it would hand the vendor’s default agent, and every tool it has, to exactly the calls that exist to prove the tools are off (crewlet llm doctor’s two isolation probes send no system prompt).

Check the CLI you actually have. grok is the trap: xAI’s own CLI (x.ai/cli, xai-org/grok-build) and a same-named community package on npm both put a grok on PATH, and they are different programs — the official one has the flag, the npm one has none of this profile’s flags at all. If grok --version prints a 0.0.x, you have the other one.

Qwen Code is where the Gemini fork has diverged. It renamed its parent’s variable (QWEN_SYSTEM_MD, and this build reads neither the other’s) and added two flags its parent does not have, so “same shape as gemini-cli” no longer holds here.

kimi-code is not the only trap of its kind either: the unscoped kimi-code package on npm is a Claude Code wrapper that installs its own kimi, and MoonshotAI’s is @moonshot-ai/kimi-code. A kimi --version printing a 1.0.x is the other one.

For the six with no channel at all, cli.overrides.system_prompt_args and cli.overrides.system_prompt_env are how you adopt one the day its vendor ships it — no engine release needed.


The rendered prompt is the largest thing this backend hands a CLI and, after the system prompt, the most sensitive: the flattened transcript, the tool catalogue, the conversation and every tool result in it. prompt_mode says which channel carries it, and the three are not equivalent:

prompt_modeHowCeilingOn /proc/<pid>/cmdline?
stdin (default)written to the child’s stdinnoneno
filewritten 0600 into the per-call working directory; prompt_args carries the path through {file}noneno — only the path
argvappended as the last argument (or as prompt_args’ value)ARG_MAX — ~2 MB on Linux, 256 KB on macOSyes, in full

argv is a last resort, taken only where a vendor offers nothing else — copilot, grok, kimi-code, hermes and pi today. It has both failure modes: a long transcript fails at exec rather than at the model, and every account on the machine can read the conversation out of the process table while the call runs. On pi the system prompt is on argv too, since its only file channel is a static per-home SYSTEM.md; that is the one entry where both halves of a turn are visible in the process table, and cli.overrides.system_prompt_args: [] puts the identity back in the transcript if that matters on your host.

Three of those five take the prompt as a FLAG’S VALUE rather than as a positional argument, so the flag sits in prompt_args and stays adjacent to the prompt: -p on grok, --prompt on kimi-code, -z on hermes. Put one in complete_args instead and the next flag becomes the prompt. pi takes it positionally and uses prompt_args: ["--"] for a different reason — it is the one vendor here that offers an end-of-options separator, so a transcript beginning with a dash is a prompt rather than an unknown flag.

-p does not mean the same thing on every CLI. It is print mode on claude-code, cursor-agent, copilot and pi, the prompt’s value on grok and kimi-code, and on hermes it selects a profile — a different Hermes instance entirely. Borrowing the reflex from one profile to another is how a seat ends up running against a profile named after its own transcript.

file is the same trade this backend already makes for the system prompt, in the same directory, at the same mode, and for the same reasons. It needs a vendor flag that takes a path; muse-code is the first built-in profile whose CLI has one (muse exec --prompt-file), and its profile is

prompt_mode: file
prompt_args: ["--prompt-file", "{file}"]

A file profile whose prompt_args contains no {file} is refused at load: without it the CLI is run with no prompt at all, which a vendor answers by opening an interactive session or printing usage — neither of which looks like the configuration error it is.


Every one of these CLIs has its own tools — file edits, shell, web fetch. In text mode Crewlet does not use them: they run in the CLI’s sandbox, invisible to the tool registry, the permission model, secret redaction, and the event stream. Routing agent work through them would fork the engine’s tool surface in two.

So every profile denies the CLI’s shell and file tools wherever the vendor offers a way to, and each says how: a flag on the command line (Claude Code’s --disallowedTools, Copilot’s --deny-tool, grok’s --disallowed-tools, Codex’s read-only sandbox) or a settings file the engine writes into the seat’s own home or the per-call working directory before every call (Gemini’s settings.json, OpenCode’s opencode.json, Cursor’s .cursor/cli.json, Muse Code’s run.toolset). The shell is the one that matters: the seat’s home and environment are isolated, but the filesystem is not, and a CLI with a shell on the engine host reads whatever the engine user can read. A vendor with no such switch is declared as local_tools: vendor-default with a note saying which switch is missing — and crewlet llm doctor measures the stance rather than trusting it (see Operating it).

A deny list, not an allow list, where a vendor offers both. grok has both and the profile takes --disallowed-tools, which reads backwards until you look at how each is applied. Its allowlist is honoured only if every entry resolves: one name the build does not recognise and the whole filter is skipped with a warning, leaving every tool enabled. The deny list always applies and only warns about the entry that matched nothing. So a name this profile gets wrong costs one tool on a deny list and costs everything on an allow list — and a vendor renaming a tool is exactly the drift these profiles are built to expect.

A refusal and a removal are not the same guard, and only one of them is a denial. muse-code is the profile that makes the difference concrete. Its --disable-shell and --disable-write flags read like tool denials and are not: measured against 1.0.3, bash, bash_input, write_file and edit_file are still advertised to the model, byte-identical to the baseline surface — the flags refuse the call when it comes. A model that can see a shell will try to use it, and every such attempt is a wasted round inside a CLI whose tool log the engine never sees. What actually removes them is run.toolset in the seeded settings.json: an allowlist of exact tool names, validated against the CLI’s own registry at startup, which replaces the surface outright. The profile ships both — the allowlist because it is the denial, the flags because a settings file that failed to apply should still refuse the call. codex sits at the other end of the same distinction: its --sandbox read-only contains the shell rather than removing it, and reads stay. That residual is why local_tools: denied is a claim crewlet llm doctor measures rather than one you take on trust.

But an allow list is right where a stale name costs one tool. grok’s rule is about grok’s implementation, not about allow lists — and kimi-code inverts it. There, an entry the CLI does not recognise “is reported with a warning” and simply matches nothing, so a name that went stale costs that one tool rather than lifting the restriction. Read how each vendor applies the list before choosing the shape; the flag’s name says nothing about which way it fails.

The denial does not always live in a flag. kimi-code’s tool policy is [tools] and [[permission.rules]] in config.toml — and config.toml is where kimi login writes the OAuth reference and the model catalogue, so it is a credential here and seeding a policy into it would destroy the login it carries. What is left is the agent file, the same --agent-file that carries the system prompt: its frontmatter allowlist is the denial. That is the concrete reason a profile with a system_prompt_file passes its channel on every call — see Which CLIs actually have one.

An approval prompt is a wedge in a headless run. A CLI that stops to ask sits on the seat’s concurrency slot until timeout_seconds fires, because there is nobody to answer. Most profiles therefore remove the asking rather than the guard: OpenCode’s seeded policy is all allow and deny and denies the tool that asks a person, and muse-code passes --disable-approval, which is the posture its own vendor’s headless guidance asks for — approval prompts off, the OS sandbox still on.

hermes needs neither, and reading its parser is what settled that. Its one-shot flag documents its own posture — “approvals are auto-bypassed” — so there is no prompt to wedge on and --yolo would widen what a run may do for nothing. And --toolsets, which the vendor’s prose lists under hermes chat, is a top-level flag whose own help says “Applies to -z/—oneshot”: passing it replaces the enabled set for the invocation and the config file is not consulted. So the denial is --toolsets web on argv — web_search and web_extract and nothing else — rather than a seeded settings file. That is the better shape wherever a vendor offers it: a flag fails at argument parsing when it is renamed, where a settings key a vendor renamed is ignored and the run quietly keeps every tool.

Web is the one local tool that stays on. A subscription seat must not have less reach than the same CLI at a terminal, and a fetch is a read — it never gates a delivery. Where a vendor gates its web tools behind an approval a headless run cannot answer, the profile allows them explicitly (--allowedTools WebFetch WebSearch, Copilot’s --allow-tool); where its default web search answers from an offline index, the profile switches it live (Codex’s web_search="live"). What the CLI reads on the web is not in the engine’s event stream — the cost of an unrecorded read, accepted. Seats on API models reach the web the way they reach everything external, through the MCP servers you configure.

pi is the one profile that denies every tool outright (--no-tools), and it is honest there for a reason no other vendor gives it: measured off the request its CLI actually sent, the baseline is bash, edit, read and write — there is no web tool in it to keep, and --no-tools sends the tool array empty. Adopting the same flag on a CLI that has one would cut the web silently, so a test refuses it for every other profile.

Both stances are profile fields, so an operator can override them like any other — cli.overrides.local_tools, cli.overrides.local_tools_note, and cli.overrides.seed_files (a list of {path, in: home|work, content}; lists replace wholesale).

Instead the CLI is used strictly as a text model, and the tool channel rides in the prompt:

  1. The phase’s messages flatten into a labelled transcript.

  2. The request’s tool definitions (llm.Request.Tools) render as a JSON catalogue (name, description, JSON Schema).

  3. A response contract asks for one fenced JSON block:

    {
    "message": "Short note to the operator, or an empty string.",
    "tool_calls": [{ "name": "tool_name", "arguments": { "arg": "value" } }]
    }
  4. The reply is parsed back into Completion.content + Completion.ToolCalls.

The parser is deliberately forgiving — it accepts the last fenced block, a bare object, arguments as a JSON string, one call object written without its list, null for no calls, and message / content / text / response as synonyms. It is strict in the other direction: a call list it can read no call from — strings for entries, nameless objects, a string or a number where the list belongs — makes the reply not an envelope at all, because reading it as one would report a model that asked for no tools when it asked for some. When nothing parses, the whole reply becomes assistant content with no tool calls, and in every phase that has to end in a call the tool loop’s corrective re-prompt takes over — the finishing corrective naming the phase’s submission. A request never forces a call, so there is one contract and it is the permissive one — “use an empty tool_calls list when no tool is needed” — on every phase and on crewlet llm doctor’s smoke test alike, which therefore certifies the shape a seat actually sends: the tool offered, the instruction naming it, and nothing demanding the call. A malformed reply costs a round; it never crashes a turn.

A call with no tools gets no contract. Auxiliary work (summarisation, the relevance filter) sends a plain prompt and reads a plain answer, with no envelope to get wrong.

Before any of that, something has to decide which part of what the CLI printed is the model’s reply. That is output plus text_paths on the profile: text takes the whole of stdout, json reads one document and jsonl concatenates every event that carries a text path, in stream order. text_paths is a list so a vendor that moved the field between releases needs no override — the first path that resolves to a non-empty string wins, and an empty one falls through to the next.

An enveloped stream needs one more thing than paths. Muse Code wraps every event in a single envelope shape and puts the kind in payload_type, so payload.text is a token fragment of the reply on a run.output.delta, a tool’s output on a tool.result, and the assembled reply on run.terminal.completed. A path walk cannot tell the three apart: it would splice the tool output into the answer and then repeat the answer. event_type_path names where the kind lives and text_events says which kinds carry the reply:

output: jsonl
event_type_path: ["payload_type"]
text_events: ["run.terminal.completed"]
text_paths: [["payload", "text"]]

kimi-code needs the same pair for a different stream shape: its stdout is an OpenAI-style chat stream discriminated by role, so a profile reading every line would splice the user’s own prompt and any tool result into the reply.

output: jsonl
event_type_path: ["role"]
text_events: ["assistant"]
text_paths: [["content"], ["content", "0", "text"]]

Both or neither — one without the other configures nothing and is refused at load, as is either on a profile that is not jsonl. They scope text only: usage and error paths are still read across the whole stream, because a stream reports those wherever it likes and the last value wins. A profile that names an event filter also streams through it, so a jsonl profile taking its answer from one terminal event delivers that answer in a single delta at the end rather than pushing a tool’s output through as though the model had said it.

Four outcomes, kept apart on purpose, because three of them used to be one — and only two of them are failures:

What happenedWhat the engine does
A text path resolved to textThat text is the reply.
A text path resolved and every one was emptyThe CLI answered with nothing. An answer, not a failure: a completion with empty content and the round’s real token usage attached, which is exactly what the openai and anthropic backends return for a model that spends its whole budget thinking. The tool loop corrects it — see When the CLI answers with nothing.
No text path resolved at allThe profile no longer matches the installed CLI. A retryable server failure that names text_paths, points at crewlet llm doctor, and prints the tail of what the CLI output so you can write the override.
The CLI printed nothing at all on a zero exitIts own message, because neither of the two above can say anything true about output that does not exist. A retryable server failure carrying whatever it wrote on stderr, which is the only clue there is.

Output that is not JSON at all is still an answer: a CLI that printed a banner, a warning, or the vendor’s own sentence about a spent plan is read as prose rather than refused, which is what lets the limit sentinels be recognised on a zero exit. Those sentinels are matched against the CLI’s whole stdout and stderr, so a drifted profile still yields a real rate_limit with the vendor’s own reset instant rather than a server fault.

Why the last two are failures rather than answers. They used to be one case with the empty one, and the answer handed back was the CLI’s raw stdout — on the reasoning that an operator would then see the shape and write an override. They would, but only after it had been spoken as an agent first: an empty result on a Claude Code envelope meant the seat’s reply became {"duration_api_ms":11377,…,"result":"","type":"result"}, the tool loop appended that to the conversation and re-sent it every round, the reviewer judged the turn on it, and the dashboard printed it as the sentence the agent had said. The shape belongs in the error message, where the only person who can act on it is the only one reading. Both remaining failures are about this build not being able to read the CLI, which is a fact about your machine — so the chain walking to another entry is the right move.

A model that spends its whole output budget on hidden reasoning exits 0, reports success, bills hundreds of output tokens and leaves the answer field empty. That is a model outcome, so the backend hands it back as an answer of nothing rather than dressing it as an outage:

  • The round is charged. An empty answer costs tokens, and it used to be the one outcome that spent them without ever reaching a budget.
  • The tool loop asks again. In every phase that finishes by a call — the executor, the reviewer, onboarding, a worker — the round gets the finishing corrective naming the phase’s submission, up to twice in a row on one allowance with a round that answered in prose: there, a round without the call ends the phase only into its rescue, so a second identical send is worth its round. A loop that does not finish by a call gets a single corrective naming what went wrong instead: there a prose answer is a legitimate finish, whatever the model writes next is the result, and there is no rescue for a second nudge to beat.
  • It is counted. empty_answer_rounds on the phase record is the number of rounds that reached nobody. A seat whose model habitually answers nothing shows up there, and in crewlet llm doctor, which names an empty answer as such rather than reporting it said: "".

If you see it repeatedly, the entry’s model is the field to change. reasoning_effort and reasoning_budget_tokens are refused on a cli-agent entry precisely so nobody spends an afternoon on them: they are per-call API parameters and a headless coding CLI takes neither.

Completion.InputTokens / output_tokens come from the CLI’s own usage report where the profile can find one (Claude Code and Codex report it; Gemini CLI’s shape varies by version). Where it can’t, the counts are estimated at four characters per token — an approximation, but budgets must keep moving or a seat on this backend would run with no ceiling. crewlet llm doctor tells you which of the two you are getting.

A CLI can report its tokens honestly and still not put them in the answer. Hermes’s one-shot entry point is hermes -z, whose whole contract is “single prompt in, final response text out, nothing else on stdout or stderr” — so there is no envelope for a usage path to walk, and the figures ride --usage-file instead. usage_file_args carries that path through a {usage_file} placeholder; the file lands in the per-call working directory and is read back through the same usage paths every other profile uses:

usage_file_args: ["--usage-file", "{usage_file}"]
usage:
input: [["input_tokens"]]
output: [["output_tokens"]]

Three rules travel with it. A report that never arrived is not a call that cost nothing — the counts fall back to the estimate, because the vendor writes the file “even when the run fails”, so its absence means the run did not get that far and failing a completion the model answered over a count would throw away work you paid for. A report whose keys this profile cannot read is drift, not a zero-token turn, and falls back the same way rather than charging zero. And both prompt counts come from the file or neither does: a partial overlay would pair one source’s input count with another’s output count, and the sum is what a budget is charged. (The two cache figures are not part of that test — a provider that caches nothing reports neither, and zero is the true answer there.) For the same reason a profile setting usage_file_args must declare both usage.input and usage.output: declaring one means every call quietly falls back to an estimate while crewlet llm doctor reports the vendor’s own figures, which is precisely what that line exists to settle.

Not every vendor’s richer channel is worth taking. pi has one — --mode json streams every session event and carries real counts — and this build deliberately uses print mode instead. Its answer lives at message.content[N].text, an array of text, thinking and tool-call blocks whose index moves with whether the model reasoned, and a model that interleaves thinking with prose spreads one reply across several of them. Reading it by index would be a guess about which part of the output is the model’s reply, which is the one thing this backend must never get wrong. The counts are estimated instead, and the profile’s comment carries the override that takes the stream if you want it.


Vendor subscription logins are browser OAuth with PKCE, often with SSO, MFA, or a one-time code. There is no username/password grant to script, and driving a headless browser to type into one would break on the vendor’s next login-page change. Crewlet does not pretend otherwise. What it does instead covers every deployment shape:

0. Already logged in on this machine? Adopt it

Section titled “0. Already logged in on this machine? Adopt it”
Terminal window
crewlet llm login default -from-host

The usual starting point: you have been running claude on this box yourself for months. Crewlet does not use that login on its own — the child process is given its own HOME, so your ~/.claude is invisible to it, which is exactly the isolation the rest of this page depends on. -from-host copies the CLI’s credential files out of your home directory into Crewlet’s, once, on request.

It is a copy, not a redirect: agents never write into your personal credential file, so a fleet refreshing a token mid-session is not a surprise you get handed. The cost is that both copies then descend from one refresh token, and a vendor that rotates refresh tokens can log out whichever side refreshes second. Where the CLI mints a headless token (option 2 below), that is the better answer and avoids the fork entirely — crewlet llm login -from-host says so after it runs.

-home PATH reads from somewhere other than the engine user’s own home, for a deployment where the engine runs as a different user than the one that logged the CLI in.

crewlet llm doctor looks for a host login too, so “no sign-in” on a machine where the CLI plainly works explains itself:

credentials : none on disk
host login : /home/you/.claude/.credentials.json (not adopted)
token env : unset
sign-in : none
problems:
- no sign-in for "default": no credential files in
/var/lib/crewlet/llm-cli/default/credentials, and nothing in the
environment the CLI is given authenticates it — this machine has
a login at /home/you/.claude/.credentials.json: adopt it with
`crewlet llm login default -from-host`, or mint a headless
CLAUDE_CODE_OAUTH_TOKEN with `-capture-token` (preferred: no
shared refresh token)

How doctor decides what the sign-in line says is under Operating it.

1. Broker the vendor’s own login (any CLI)

Section titled “1. Broker the vendor’s own login (any CLI)”
Terminal window
crewlet llm login default

Runs the real claude auth login / codex login / opencode auth login attached to your terminal — follow its prompts exactly as you would by hand. The only thing Crewlet controls is where the credential lands: in the provider’s isolated credentials/ directory, separate from your personal CLI login on the same machine.

Each profile names its vendor’s own one-shot auth subcommand, which prints its OAuth URL and returns once you have signed in. A profile that named an in-session slash command instead would open an interactive session rather than run a login: the session asks you to sign in itself, then replays the slash command and asks a second time, and leaves you in a REPL you have to interrupt — after a login that had already succeeded. crewlet llm login returning you to your shell is the signal that it worked; crewlet llm doctor <KEY> confirms it.

pi is the one CLI where the REPL is the login, and its profile says so rather than pretending otherwise: /login there is a slash command and the vendor ships no login subcommand at all. So the broker starts pi’s own interactive session — ephemeral, with --no-session, in the isolated credential directory — and you type /login, complete the browser flow, then /exit. Declaring no login instead would have crewlet llm login tell you, wrongly, that the CLI “authenticates on first use”.

Not every CLI has all three commands. kimi-code has kimi login (a device-code flow) but no logout and no status subcommand — both are its TUI’s slash commands, which a headless run cannot reach. hermes has hermes auth and hermes status, but hermes auth logout requires a provider name, which is yours to know rather than the profile’s to guess. A profile declares only the commands its vendor actually has, and crewlet llm logout / status say so plainly for the ones that do not.

2. Capture a headless token (best where it exists)

Section titled “2. Capture a headless token (best where it exists)”
Terminal window
crewlet llm login default -capture-token

Runs the vendor’s token-minting command (claude setup-token) and puts the result in the encrypted secret store under the profile’s token variable — CLAUDE_CODE_OAUTH_TOKEN for Claude Code. Prefer this whenever the CLI offers it: no credential files to sync, no refresh-token rotation, and it survives an ephemeral container with no persistent volume.

Minting is interactive — the CLI opens the same browser sign-in as option 1 — so its prompts and its sign-in URL are shown on your terminal while the token itself is captured. The token never touches stdout, which is what leaves -print-token free to pipe cleanly into your own secret manager.

Already have a token from elsewhere?

Terminal window
pass show anthropic/crewlet-oauth | crewlet llm login default -token-stdin

3. Username / password, where the CLI genuinely has one

Section titled “3. Username / password, where the CLI genuinely has one”
Terminal window
vault read -field=password secret/gateway |
crewlet llm login default -username [email protected] -password-stdin

Available for a profile that declares stdin_login — the built-in opencode profile, an operator’s own wrapper, or a self-hosted gateway CLI. The password is read from stdin or a declared environment variable, never from argv (which is visible in ps and lands in shell history).

The Claude, Codex, and Gemini profiles deliberately leave stdin_login unset, and the command says so rather than failing obscurely:

Error: the 'claude-code' CLI authenticates through the vendor's browser
OAuth flow — there is no username/password login to drive. Run
`crewlet llm login` (which brokers that flow), or
`crewlet llm login -capture-token` where the vendor mints a headless
token. If your build of this CLI does accept a credential, declare it
under providers.llm.<key>.cli.overrides.stdin_login.

If your CLI does accept a credential, wire it yourself — no Crewlet change needed:

cli:
agent: custom
overrides:
binary: my-gateway-llm
complete_args: ["--json"]
stdin_login:
args: ["login", "--user", "{username}"]
stdin_template: "{password}\n"

The engine may run in a container that is rebuilt on every deploy, or on several hosts. Export the credential directory as one blob into the encrypted secret store:

Terminal window
crewlet llm export default -secret-store

That engine restores it at boot when its own credentials/ directory is empty, so a fresh container on the same store comes up already authenticated. It is that node’s store and nothing else’s — the rows do not travel, and a second host needs its own crewlet llm login, or the same bundle handed to it through providers.llm[].auth.credential_bundle. The blob is validated on the way back in — only the profile’s own credential paths, files only, size-capped — because an archive is an execution surface if it is unpacked on trust.

Only the credential files travel. Sessions, history, and caches never go into a bundle.

Between two hosts that share no database, pipe it instead:

Terminal window
crewlet llm export default | ssh other-host crewlet llm import default

import reads the bundle from stdin — a credential on argv is visible in ps and lands in shell history — and refuses to overwrite a login the target already has. A host that has been running holds the fresher refresh token, and restoring a boot-time blob over it is how a fleet logs itself out; crewlet llm logout <KEY> first if you mean to replace it.

5. A provider key in cli.env (hermes, pi, opencode)

Section titled “5. A provider key in cli.env (hermes, pi, opencode)”

hermes, pi and opencode each front many model providers and read each provider’s key from its own variable — OPENROUTER_API_KEY, ANTHROPIC_API_KEY, NVIDIA_API_KEY and so on — so their profiles name no single api_key_env. Their key goes in cli.env instead, as a ${VAR}, with auth.mode left at subscription:

providers:
llm:
herm:
type: cli-agent
model: anthropic/claude-sonnet-4
cli:
agent: hermes
env:
OPENROUTER_API_KEY: "${OPENROUTER_API_KEY}"

Their profiles declare env_sign_in, which is what makes doctor and crewlet llm list count a credential-named variable in cli.env as a sign-in (sign-in : cli.env sets OPENROUTER_API_KEY) and report one whose ${VAR} resolved to nothing. A name is credential-named when one of its parts — the name cut at every character that is not a letter or a digit, compared ignoring case — is KEY, KEYS, APIKEY, ACCESSKEY, SECRETKEY, TOKEN, SECRET, PASSWORD, PASSWD, CREDENTIAL, CREDENTIALS, AUTH, OAUTH or PAT. So OPENROUTER_API_KEY and GH_TOKEN are, and HERMES_MAX_TOKENS and OPENAI_BASE_URL are not: a count of model tokens is configuration. It is the same rule that refuses a credential in a profile’s passthrough_env or env (below). doctor counts the name, not whether the provider accepts the key: the smoke test is what proves that — and for a model served by an endpoint that takes no key, an answered smoke test is what clears the “no sign-in” problem.

The key also goes into a coding box, the way a headless token does, so it works in a remote cell; the rest of cli.env is not carried there. Agent mode is opencode’s alone — hermes and pi have no coding-agent runner — and a code-sandbox run on any of the three exports the key into the box, where the seat’s coding agent reads it if it reads that variable: opencode reads every provider’s own, Claude Code only Anthropic’s. A key given in the seat’s role.sandbox.env instead is counted by the remote-box check as well. A CLI that does not read its key from the environment — kimi-code reads its metered key only from config.toml in the credential directory — does not declare env_sign_in, and a key in its cli.env is not counted.

OAuth access tokens expire in hours, and the CLI refreshes them mid-run. Most vendors rotate the refresh token at the same time, so Crewlet syncs a changed credential file back to the shared directory when a seat’s generation closes — otherwise the whole fleet would be logged out at the next expiry. Two seats refreshing at the same instant can still race, exactly as two terminals running the vendor’s CLI would. A headless token (option 2) has no refresh file and sidesteps this entirely.


cli.agentBinarySubscriptionNotes
claude-codeclaudeClaude Pro / Maxclaude auth login (and auth status / auth logout). claude setup-token gives a headless CLAUDE_CODE_OAUTH_TOKEN. Reports full usage incl. cache tokens.
codexcodexChatGPT Plus / Procodex login. Streams JSONL events; runs --sandbox read-only.
gemini-cligeminiGoogle AI Pro / free tierFirst run starts the auth picker. GOOGLE_CLOUD_PROJECT passes through.
qwen-codeqwenQwen OAuthGemini CLI fork; same shape.
opencodeopencodeAnthropic / Copilot / anyopencode auth login; the one built-in profile with a credential login. A provider key goes in cli.env under that provider’s own variable, and doctor counts it — see Signing in with a provider key.
cursor-agentcursor-agentCursor seatcursor-agent login. A Cursor API key (CURSOR_API_KEY, its api_key_env) is the headless alternative, reached through auth.mode: api-key.
copilotcopilotGitHub Copilot seatPrompt goes on argv, so very long transcripts are bounded by ARG_MAX. Authenticates with a GitHub token, so GITHUB_TOKEN is its api_key_env — reached via auth.mode: api-key or inherit-env, never forwarded silently.
grokgrokxAIxAI’s own CLI from x.ai/cli, not the same-named npm package. Accepts XAI_API_KEY (the variable its own signed-out message names) through auth.mode: api-key.
muse-codemuseMuse Code subscription (Everyday / High / Power Usage), or pay-as-you-gomuse login / muse logout; the browser sign-in stores ~/.config/muse/auth.json, which -from-host adopts. No status command — this CLI has none. Mints no headless token: META_API_KEY is a metered Model API key, reached through auth.mode: api-key. Runs muse exec --json, denies its tools through a seeded run.toolset, and puts the prompt in a file rather than on argv. Reports no token counts anywhere on its stream, so they are estimated.
kimi-codekimiKimi Code OAuth (Moonshot)MoonshotAI’s own CLI, @moonshot-ai/kimi-code — not the unscoped kimi-code package on npm, which wraps Claude Code behind a proxy and installs its own kimi (a 1.0.x version is the other one). kimi login runs a device-code flow; there is no logout or status subcommand. Runs --output-format stream-json, because the default text output prefixes every line with • and re-wraps it, which destroys the tool envelope. Its tools are denied in the agent file that also carries the system prompt — measured, that is 25 tools down to 1, and an 8.8 KB vendor system prompt replaced by the seat’s own. Reports no token counts on its stream, so they are estimated.
hermeshermesNous Portal, or any of 30+ providers it frontshermes auth for the credential wizard, hermes status for state; hermes auth logout needs a provider name, so no logout is declared. Runs hermes -z, whose contract is the final response text and nothing else — so the tokens come from --usage-file instead. Tools are denied with --toolsets web on argv, which replaces the run’s enabled set; -p selects a profile on this CLI, not a prompt. A provider key (OPENROUTER_API_KEY, …) goes in cli.env, and doctor counts it — see Signing in with a provider key.
pipiClaude Pro/Max, ChatGPT Plus/Pro, GitHub Copilot — whichever it is logged into@earendil-works/pi-coding-agent. No headless login: /login is a slash command, so crewlet llm login starts its TUI in the credential directory and you type it there. Denies every tool with --no-tools, which is honest here because its built-in set ships no web tool at all. Prompt and system prompt both on argv. Tokens estimated — see Token accounting. A provider key goes in cli.env under that provider’s own variable, and doctor counts it.
custom——Ships nothing; declare everything under overrides.

CLI flags drift — and that’s a config edit, not a release

Section titled “CLI flags drift — and that’s a config edit, not a release”

Every field of every profile is replaceable from YAML. When a vendor renames a flag or changes its JSON shape, fix it in place:

cli:
agent: codex
overrides:
binary: /opt/homebrew/bin/codex
complete_args: ["exec", "--json", "--skip-git-repo-check", "-"]
text_paths: [["item", "text"], ["msg", "message"]]

Lists replace wholesale (position matters in an argv). Overrides are validated against the profile model, so a typo is refused by crewlet validate and by every API write of the configuration — PUT and PATCH /config, a per-entity write, a revert — rather than by an agent’s first turn or by every node’s apply after the API had already activated it. crewlet llm doctor prints the CLI version the built-in profile was written against next to the version you actually have.

A drift that only shows up at runtime — the flags still work, the JSON still parses, and the answer field moved — names itself: the completion fails with the text_paths this profile looked in and the output the CLI actually produced. See Finding the answer in the CLI’s output.

limit_markers and auth_markers drift the most quietly. Every other field fails visibly when it goes stale — a renamed flag is a non-zero exit doctor reports on the spot. A sentinel is matched verbatim against the CLI’s own prose, so one the vendor has reworded simply never fires: a spent plan then classifies as a fatal error instead of rate_limit, the fallback chain never carries the seat onto a metered key, and nothing says so until somebody hits their cap. If your CLI’s wording differs from the built-in profile’s, override it:

cli:
agent: claude-code
overrides:
limit_markers:
- sentinel: "Usage limit reached"
auth_markers:
- sentinel: "Please run /login"

Take the sentinel from what your CLI actually prints, not from what it used to print.

And say where it prints it, because the model’s own words are a haystack. Sentinels are matched against the extracted answer as well as stderr, and they have to be: Claude Code reports a spent plan on a zero exit with the vendor’s sentence standing where the answer should be, so a marker confined to stderr would never fire and the fallback chain would never carry the seat onto a metered key. The cost is that a seat asked about rate limits can answer in prose — “our quota resets hourly” — that trips a generic sentinel and benches a perfectly good credential. This page’s own history has that bug: a bare 429 sentinel was dropped for exactly it.

So a CLI whose failures reach stderr and nowhere else declares that, and then nothing the model writes classifies anything:

marker_scope: stderr # or answer-and-stderr (the default)

Three profiles narrow it, each on a measurement rather than an assumption: kimi-code puts its classification on stderr with the answer stream carrying only role: meta lines, pi leaves stdout empty and writes <status>: <the provider's JSON> to stderr, and hermes writes hermes -z: agent failed: … there while -z guarantees stdout is the reply and nothing else. Narrow it only where the vendor’s behaviour makes the answer an impossible place for the report — a profile that narrows it wrongly stops recognising spent plans, which is silent until somebody hits their cap.

And check where it prints it. A sentinel can only match what the CLI puts on stdout or stderr, and one vendor puts the failure nowhere a plain run would show it: muse exec writes the fixed string run ended with Failed to stderr and carries the real reason only in its run.terminal.failed event. That is why the muse-code profile runs with --json even though the event stream buys it no token counts — the answer would read fine without it, and a spent plan would arrive as a bare exit 1 that no marker could classify, so the seat would never fall through to its metered key.

Every field a profile has, and therefore every key cli.overrides accepts. A key not in this table is refused by name. Every mapping field (config_env, env, usage, stdin_login, system_prompt_file) merges key by key; lists and single values replace wholesale.

FieldWhat it is
binaryThe executable, looked up on PATH unless it is a path.
vendorThe model family the CLI addresses (anthropic, openai, google, meta, …), for a coding agent that resolves <family>/<model>.
written_forThe CLI version the profile was written against, printed by doctor beside the one installed.
version_argsThe argv of the version probe.
complete_argsThe argv of one completion, before the model and the prompt.
model_argsThe model flag, with {model} substituted. Required in effect: every entry names a model, so a merged profile with none is refused.
prompt_modeHow the prompt travels: stdin (default), argv or file.
system_prompt_argsThe flag carrying the system prompt, with {file} (preferred) or {system}. Empty leaves it in the transcript.
system_prompt_envA variable naming a file the CLI reads its system prompt from; exclusive with system_prompt_args.
system_prompt_file{name, template}: the file a {file} channel writes, and the text around {system}.
prompt_argsThe flag introducing the prompt in argv mode, or carrying {file} in file mode.
outputHow stdout is encoded: json (default), jsonl or text.
text_pathsWhere the answer is in a json/jsonl document.
event_type_pathWhere a jsonl line names its event kind.
text_eventsThe event kinds whose text is the answer.
error_pathsA boolean the CLI sets when it failed despite exiting zero.
usage{input, output, cache_read, cache_write}: where the token counts are.
usage_file_argsThe flag asking the CLI to write its counts to a file, with {usage_file}.
config_envA vendor’s own relocation variable, mapped to a directory under the seat home.
envFixed child environment. Never a credential, and never token_env or api_key_env.
passthrough_envEngine variables forwarded to the child. Never a credential.
token_envThe variable a headless subscription token goes in (cli.auth.token).
api_key_envThe variable a metered key goes in (api_keys under auth.mode: api-key).
env_sign_inThe CLI reads its model provider’s key from its own environment, so a credential-named variable in cli.env signs it in.
credential_pathsThe login files, relative to the seat home.
volatile_pathsSessions, transcripts and history, deleted before and after every call.
login_argsThe vendor’s own interactive login, for crewlet llm login.
capture_token_argsThe command that mints a headless token on stdout (-capture-token).
status_argsThe command reporting who the CLI is logged in as.
logout_argsThe command revoking the login.
stdin_login{args, stdin_template, password_env}: a real credential login, where the CLI has one.
limit_markersSentences recognising a spent plan — see above.
auth_markersSentences recognising an expired login.
marker_scopeWhere markers match: answer-and-stderr (default) or stderr.
host_credential_pathsWhere the CLI keeps its login in a person’s own home, for -from-host.
local_toolsThe profile’s stance on the CLI’s own tools: denied or vendor-default.
local_tools_noteWhy a vendor-default stance is one.
seed_files{path, in, content} (in is home or work): settings files written before a call.

providers:
llm:
subscription:
type: cli-agent
model: sonnet # passed to the CLI's --model
cli:
agent: claude-code # or codex | gemini-cli | opencode
# | muse-code | kimi-code
# | hermes | pi | …
mode: text # text (default) | agent — see above
run_in: "" # agent mode only: direct | container | e2b
state_dir: /var/lib/crewlet/llm-cli/claude
# Where credentials and per-seat homes live. Empty uses
# $CREWLET_LLM_CLI_HOME/<key>, falling back to
# ~/.crewlet/llm-cli/<key>. Point at a persistent volume when
# the engine runs in an ephemeral container. A LITERAL PATH —
# unlike the credential fields here it is not ${VAR}-expanded,
# because it names where the engine keeps files rather than a
# secret (the same reason the store path is a Tier A field).
timeout_seconds: 300 # one CLI invocation, wall clock
max_concurrent: 4 # CLI processes at once
env: # extra child env, ${VAR}-resolved
ANTHROPIC_SMALL_FAST_MODEL: haiku
auth:
mode: subscription # subscription | api-key | inherit-env
token: "${MY_OAUTH_TOKEN}" # else the profile's own token var
credential_bundle: "${MY_BUNDLE}" # else CREWLET_LLM_CLI_<KEY>_CREDENTIALS
overrides: {} # any field under Profile fields

timeout_seconds is separate from the entry’s own timeout_seconds because the transports are not comparable: that one bounds an HTTP attempt (default 600 s; a streamed call’s silence rather than its length), while this covers a process launch — a Node runtime costs seconds before the first byte — plus the model call and the CLI’s internal retries. On breach the process group is terminated (so the runtime’s helpers go too) and the call is reported as timeout, which the role’s fallback chain retries.

max_concurrent: 4 keeps peak memory near 1.5 GB: each CLI is a full Node or Rust runtime at roughly 200–400 MB resident, and an unbounded fleet of seats starting turns together can exhaust a small engine host. Subscription plans also throttle concurrency well below what an API key allows, so a much higher number mostly buys rate-limit errors. Raise it on a large host with a plan that permits it.

auth.mode defaults to subscription, not inherit-env, on purpose: a backend that silently picked up a stray ANTHROPIC_API_KEY would bill the metered account while you believed you were on a flat-rate plan.

Every credential an entry configures has to reach the CLI, or a write keeping it is refused. The engine puts a credential only into a variable the profile names for it — api_key_env for a key, token_env for a headless token — so each of these would run with the credential unused:

WrittenRefused becauseInstead
auth.mode: api-key on a CLI whose profile names no api_key_env (hermes, pi, opencode, kimi-code)the key has no variable to go inhermes, pi, opencode: leave auth.mode at subscription and set the key in cli.env under its provider’s own variable (above). kimi-code: put the key in the credential directory’s config.toml — crewlet llm login, or a credential bundle. A build of the CLI that does read a key variable: name it with cli.overrides.api_key_env
auth.mode: api-key with no api_keysthere is no key to put in the variableadd one, e.g. api_keys: ["${ANTHROPIC_API_KEY}"]
api_keys under subscription or inherit-envonly api-key mode reads itset auth.mode: api-key, or drop it
more than one api_keys valuean entry holds one login and rotates nothingone key per entry; another key on an entry of its own in the seat’s fallback chain
auth.token on a CLI whose profile names no token_env (every built-in profile but claude-code)the CLI mints no headless tokenremove it
auth.token under api-key or inherit-envapi-key removes the token variable and inherit-env forwards the engine’s ownauth.mode: subscription
auth.mode: inherit-env on a CLI that names neither variablenothing would be forwardedsubscription, with the key in cli.env or a login
the profile’s token_env or api_key_env in cli.envcli.auth owns those: whether a value there reaches the CLI depends on auth.mode and on whether a credential resolves from the secret store or the engine’s environment, which the document cannot showauth.token, or api_keys with auth.mode: api-key

Each is checked against the profile with cli.overrides merged in, and each refusal names the field and the route to take. They are admission rules: every write path — crewlet validate, a PUT or PATCH of the config, a setup write, a revert — refuses a document that breaks one, while a revision being applied that breaks one is applied and warned about, because the entry still runs (signed in by whatever the CLI does read) and refusing it there would take a node off the fleet’s configuration during a rolling upgrade. A profile that cannot drive its CLI at all — an override typo, a missing binary, no model_args — is different: no node can build it, so it is refused at apply too.

A profile’s passthrough_env may not name a credential — an admission rule like the ones above. Everything listed there is forwarded from the engine’s own environment before auth.mode is consulted, so a key named there would reach every seat whatever the mode says — the same metered-bill-on-a-flat-rate-plan failure the mode exists to prevent. Use it for genuine non-secret configuration (GOOGLE_CLOUD_PROJECT, a region); a CLI’s key belongs in api_key_env or token_env, and auth.mode: inherit-env is the deliberate way to let the host’s value through.

Nor may a profile’s env, for the same reason and one more: cli.overrides is neither ${VAR}-resolved nor redacted, so a key written there would sit in the stored revision in plain text and come back on every config read. A credential the CLI needs goes in cli.env, which is both, or through cli.auth. env may not name the profile’s token_env or api_key_env either, whatever they are called: cli.auth owns those, and cli.env and cli.auth are both layered over env, so whether a value there reached the CLI would turn on the mode and the secret store rather than on anything the profile says.


Nothing changes. Phase selection resolves by providers.llm key, and the resolver never looks at a provider’s type — so llm, llm_review, llm_subagent, llm_auxiliary, llm_judge and llm_sandbox all behave exactly as they do for API entries, including mixing the two kinds in one role and including list-form fallback chains. See Turn Engine — per-phase LLM models.

The one difference is where the model string goes: an API entry sends it as a request field, a cli-agent entry passes it as --model. One entry is still one model, so per-phase models mean one entry per model:

providers:
llm:
opus-sub:
type: cli-agent
model: opus
cli: { agent: claude-code, state_dir: /var/lib/crewlet/llm-cli/claude }
sonnet-sub:
type: cli-agent
model: sonnet
cli: { agent: claude-code, state_dir: /var/lib/crewlet/llm-cli/claude }
cheap:
type: openai
model: gpt-4o-mini
api_keys: ["${OPENAI_API_KEY}"]
roles:
- name: Engineer
llm: [opus-sub, cheap] # the executor: subscription first, key when spent
llm_review: sonnet-sub # the reviewer, on a cheaper subscription model
llm_auxiliary: cheap # see the latency note below

Point them at the same state_dir and they share one login. Both entries above then use the credential directory a single crewlet llm login wrote, instead of needing one login per entry — the default state_dir is per provider key precisely so unrelated providers do not collide, which means entries that should share must say so. They also share one set of per-seat homes and one generation, so a call on one entry never wipes a live call on the other.

Entries sharing a state_dir must drive the same CLI: two different CLIs disagree about which files are credentials and which are conversation memory, so each would prune the other’s state. crewlet validate rejects that combination by name.

Concurrency is per entry. max_concurrent caps one provider’s processes, so two entries at the default of 4 can run 8 CLI processes at once. Size them together against the engine host’s memory.

Auxiliary work is the one phase to think twice about. Every reflection, summarisation and the turn-start relevance prefetch goes through llm_auxiliary, and each one pays a process launch on this backend. Point it at a cheap API model unless you have no key at all. (Crewlet does handle the latency: the auxiliary call’s 60-second deadline is widened to the provider’s own cli.timeout_seconds, so a subscription aux provider is not cut off mid-call — it is simply slower than it needs to be.)


A spent subscription window arrives as prose on a successful exit (“Usage limit reached · continuing automatically”). Crewlet matches that wording — and, where the CLI relays the API’s own error instead, the "type":"rate_limit_error" in it — and reports it as rate_limit, which is retryable, so the ordinary provider chain carries the role onto a metered key for the rest of the window and back again afterwards, with no operator intervention:

providers:
llm:
subscription:
type: cli-agent
model: sonnet
cli: { agent: claude-code }
metered:
type: anthropic
model: claude-sonnet-5-5
api_keys: ["${ANTHROPIC_API_KEY}"]
roles:
- name: Engineer
llm: [subscription, metered] # subscription first, key as backstop

An expired login classifies as auth, which is also retryable, so the chain keeps the seat working while you re-run crewlet llm login.

Both recognitions are limit_markers / auth_markers on the profile: a literal substring the vendor emits, plus (where it carries one) the field holding the reset instant, so the retry-after is a datum rather than a guess. They are matched against whatever the CLI printed — which on a healthy call is the model’s own answer — and a match benches the credential for a cooldown and hands the seat to the next entry in the chain.

So a sentinel has to be the vendor’s wording, and a profile is refused at crewlet validate if one contains no letters. The rule exists because a shipped profile carried sentinel: "429", and three digits matched as a substring is not a rate limit — it is a model quoting an HTTP status, a stack trace’s line number, a token count, or any ten-digit epoch. Every one of those took a working subscription out of service.

The other way a sentinel stops working is quieter — the vendor reworded it, so it simply never fires. Both are fixed the same way, with cli.overrides.limit_markers; see CLI flags drift for the shape, and take the wording from the sentence your CLI actually printed, which a fatal failure carries verbatim so that you can.


The other shape: an OAuth proxy in front of an HTTP entry

Section titled “The other shape: an OAuth proxy in front of an HTTP entry”

Everything above drives the vendor’s CLI as a process. There is a second way to spend a subscription, which Crewlet supports without knowing anything about it: run a proxy that holds the OAuth login itself and re-exposes it as an ordinary Anthropic- or OpenAI-shaped HTTP endpoint, then point a normal provider entry at it.

Crewlet needs no cli-agent block for this. It is an HTTP entry like any other, and base_url is all that changes:

providers:
llm:
# The proxy speaks the Anthropic Messages API.
subscription-proxy:
type: anthropic
model: claude-sonnet-5-5
base_url: "${LLM_PROXY_URL}" # e.g. http://127.0.0.1:8317
api_keys: ["${LLM_PROXY_KEY}"] # the proxy's OWN inbound key
# Or it speaks the OpenAI wire format.
subscription-proxy-oai:
type: openai-compatible
model: gpt-5
base_url: "${LLM_PROXY_URL}/v1"
api_keys: ["${LLM_PROXY_KEY}"]

base_url is not an openai-compatible field. It is honoured on anthropic and openai entries too — it is only required for openai-compatible, which has no vendor default to fall back to. The same field is what points an entry at a corporate egress proxy or an Anthropic-API gateway, and the code sandbox forwards an anthropic entry’s value to Claude Code as ANTHROPIC_BASE_URL.

Which header your proxy will be handed depends on the entry’s type, because each backend sends its vendor’s native one:

Entry typeCredential arrives as
anthropicx-api-key — and only that. The backend builds its client with WithoutEnvironmentDefaults, which deliberately disables the SDK’s own bearer-token path so an ambient ANTHROPIC_AUTH_TOKEN cannot redirect a company’s auth
openai, openai-compatibleAuthorization: Bearer — and nothing else from the engine’s environment. The SDK would add OpenAI-Organization, OpenAI-Project and every OPENAI_CUSTOM_HEADERS line from the process it runs in; the backend undoes each, so a proxy never receives headers an operator exported for some other tool. The OpenAI embeddings provider builds its client the same way

The request is shaped for the model the entry names. An anthropic entry sends what its Claude model accepts — adaptive thinking and an effort level on the current generation, never a temperature there — read from the Claude model table. A proxy that answers to its own alias (model: sonnet) rather than a Claude id is shaped as the current generation; if that alias is an older model, name it with claude_model: claude-haiku-4-5 (or whichever it is), or the fields only newer models take will be refused through the proxy.

The api_keys value is the credential for the proxy, not for the vendor: the vendor login lives inside the proxy. Rotation, cooldowns and the fleet-shared credential bench all apply to that inbound key as they would to any other.

Against the cli-agent backend you get a real HTTP provider back: native tool calls instead of the in-prompt JSON envelope, no process launch per call, and whatever token accounting the endpoint reports. Against a metered key you get flat-rate cost.

What you take on is everything this page’s design otherwise handles for you:

  • The isolation guarantees do not apply. Per-seat homes, volatile path pruning and the allowlisted child environment exist because a CLI keeps conversation state under one home. A proxy is one process serving every seat, so whatever session, cache or history it keeps is shared across your whole company — that is the proxy’s design to answer, not Crewlet’s.
  • crewlet llm sees only the HTTP half of it. list, login and the rest build cli-agent providers only, so there is no login state to report and keeping the proxy authenticated is a separate operational job. doctor does examine an anthropic entry pointed at a proxy — it sends one real round and certifies a tool call comes back — but a proxy rarely serves /v1/models, so the model check reads not served there, and an openai entry is not examined at all.
  • A spent window is not translated. The prose sentinel that turns “Usage limit reached” into a retryable rate_limit is the CLI backend’s. Over HTTP you get whatever status the proxy returns, and only a 429 / 401 / 403 / 402 / 408 / 5xx is retryable; anything else is fatal and the role’s fallback chain will not walk to the next provider. Check what your proxy returns on an exhausted plan before you rely on llm: [proxy, metered].

A proxy that spends a subscription rather than an API key has to present itself to the vendor as the vendor’s own client. In practice that means reproducing a specific client build’s headers, its beta flags, sometimes its TLS fingerprint, and often injecting that client’s system prompt ahead of yours — which quietly changes what your prompts say and where prompt-cache breakpoints land.

Vendor terms generally do not permit a third-party client to route requests through consumer subscription credentials, and vendors have enforced that. Crewlet’s cli-agent backend is on the other side of that line by construction: it runs the vendor’s own unmodified CLI, logged in by you, as a child process — Crewlet never sees a password, never re-implements an auth flow, and never impersonates a client. Pointing base_url at a proxy is a supported configuration and a decision you are making, exactly as the note at the end of this page says about plan terms generally.

None of this applies to an ordinary gateway — LiteLLM, a corporate egress proxy, a self-hosted vLLM — reached through the same field with a key you were issued. That is just an endpoint.


Terminal window
crewlet llm list # providers, agent, model, sign-in
crewlet llm doctor # verify them all, anthropic entries too
crewlet llm doctor default -no-smoke # skip the real completions
crewlet llm status default # ask the CLI who it's logged in as
crewlet llm logout default # revoke locally + delete credentials

doctor is the command that matters. It checks the binary is on PATH, runs its version probe, reports how it is signed in, says whether token counts will be real or estimated — and then runs three real completions: a smoke test with a real tool, because a profile can look perfect and still not produce a parseable tool call; a shell probe, which asks the CLI to run date +%s with its own shell and believes it only if the answer is within minutes of the engine’s clock (a model can write a token it was asked to echo, but it cannot guess the current epoch); and a web probe, which asks the CLI to fetch a public endpoint that reports its own clock and applies the same test:

provider : subscription
cli agent : claude-code
mode : text (a model behind the engine's tool loop)
binary : /usr/local/bin/claude
version : 2.0.31 (Claude Code)
written for : Claude Code CLI 2.x (`claude --version`)
state dir : /var/lib/crewlet/llm-cli/subscription
credentials : present
token env : set
sign-in : credential files in /var/lib/crewlet/llm-cli/subscription/credentials; headless token in CLAUDE_CODE_OAUTH_TOKEN
token usage : reported by CLI
smoke test : ok — 812 in / 34 out
local tools : denied by profile — probe: refused
web : ok — fetched https://www.cloudflare.com/cdn-cgi/trace
problems : none

On an agent-mode entry the report carries two more lines, and both check something that fails at a seat’s first turn and nowhere earlier:

mode : agent (the CLI runs the executor)
agent runtime : runner: "claude-code" is registered
: tool bridge: https://engine.example.com

The runner line is whether this build can actually drive that CLI as a coding agent. Agent mode reuses the coding-agent runners rather than growing a second way to invoke the same binary, and there are two of them — claude-code and opencode. An entry naming any other CLI in agent mode validates cleanly, appears in the schema and reports a configured provider, then refuses the moment a seat has work. The tool bridge line is CREWLET_MCP_BRIDGE_URL: without it every agent-mode launch is refused, because a coding agent with none of the seat’s tools cannot answer anybody, touch a ticket or submit its work.

A text-mode entry reports neither, rather than reporting that it would not work in a mode it is not in.

The sign-in line names every route that authenticates the CLI, read off the environment the CLI is actually given, after auth.mode has set and removed what it does — never off the configuration beside it. So a token that auth.mode: api-key removes is not a sign-in, and crewlet llm list’s SIGN-IN column is the first route the line names, so the two commands cannot disagree. The routes are credential files in the entry’s directory, a headless token, an API key under api-key, a token or key forwarded under inherit-env, and a credential-named variable in cli.env on a CLI that reads its key from its environment.

${VAR} references, and the variables inherit-env forwards, are resolved in the process running crewlet llm — the secret store first, then that shell’s environment — so run it with the engine’s environment to get the engine’s answer.

When none of those authenticates the CLI, doctor reports “no sign-in” as a problem, naming first a route you configured whose ${VAR} resolved to nothing. One case is not a fault: a CLI pointed at an endpoint that takes no key (a hermes or pi entry aimed at a local server through cli.env), or one holding a key in a configuration file of its own. The configuration cannot tell that from a missing key, so the smoke test decides: when it is answered, the problem is dropped and the sign-in line says the CLI authenticates some way the engine does not hand it.

A profile that says denied while the shell ran is a problem naming the installed version, because the vendor’s switch is not taking effect on it; a vendor-default profile whose shell ran is a problem stating the trust you are taking on; a web tool that could not fetch is a problem pointing at the vendor’s sandbox flags and the egress proxy the child environment was told about.

One caveat worth stating plainly: doctor spends three real completions. On a subscription that is a few thousand tokens of your plan’s allowance, which is why -no-smoke exists for a scripted health check that runs often — it skips all three and says so on each line.

The metered key behind it is examined too. A role written llm: [subscription, default] falls through to an anthropic entry exactly when the plan is spent, which is the worst moment to learn that entry is refused on every call. So doctor reports on every anthropic entry beside the cli-agent ones:

provider : default
type : anthropic
model : claude-sonnet-5-5
profile : claude-sonnet-5-5
endpoint : https://api.anthropic.com
keys : 1 (3f9a0c1b2d4e)
request : thinking adaptive (summarized), effort high, max_tokens 128000, never a temperature, reasoning before a tool change shed
models api : served — Claude Sonnet 5.5 (claude-sonnet-5-5), max_tokens 128000, input …
smoke test : ok — called crewlet_smoke (streamed), 1204 in / 61 out
problems : none

The request line is what a phase sends, read from the Claude model table — on a model that binds its thinking to the tools it was written under, that includes shedding the reasoning from before a tool change; the models api line is the vendor’s own record of the model, and any disagreement between the two is listed under drift — a problem when it puts a field the model refuses on every call, a note when the table is only more cautious than the model. The smoke test is one round in a phase’s shape (no tool choice forced, effort low), billed to the entry’s key, and -no-smoke skips it; the Models API read bills nothing and runs either way. The full rules are in the CLI reference.


  • The CLI runs on the engine host. It must be installed there, and the engine process must be able to execute it. This is not a remote service.
  • Agent mode needs more than a login. A CLI with a coding-agent runner, a providers.sandbox catalogue to place the run in, and a bridge URL a box can dial. crewlet llm doctor checks all of it — see Operating it.
  • Code work needs one more decision. A subscription can back the code sandbox. On any backend including remote E2B, what travels into the box is the environment that signs the CLI in: a headless token (crewlet llm login <key> -capture-token, and Claude Code in the box bills your plan), an api-key entry’s key, what inherit-env forwards, and the cli.env key of a CLI that reads one (opencode, hermes, pi). The credential files never travel to a remote box: they carry a refresh token whose rotation is shared fleet state. So a CLI that signs in only through its files — Codex, Gemini CLI or Kimi Code with no key — needs a local cell (providers.sandbox.local plus run_in: direct or container), where the coding agent runs on the engine host and reads the login directly.
  • Latency. Process launch plus model call. Point llm_auxiliary at a cheap API-key model rather than paying process startup for every summarisation.
  • No streaming. stream() completes and yields one chunk. Crewlet’s phases use complete(), so nothing in the engine is affected.
  • No reasoning switch. The CLI’s own plan carries its reasoning configuration and exposes no per-call setting; pick a reasoning model via model instead. Setting reasoning: true on a cli-agent entry is rejected at validation.
  • Check the vendor’s terms. Subscription plans are generally written for interactive use by the subscriber. Running a fleet of agents on one may not be permitted by your plan — that is a decision for you, not something Crewlet can decide for you. It is the sharper question for the proxy shape, where a third-party client is presenting itself as the vendor’s own.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.