Skip to content
You are reading documentation for unreleased main. Read the 0.1 version.

Choosing Your Stack

Crewlet is the engine; the surfaces your agents work on — the LLM, the work-item tracker, the knowledge base, the code host, chat, the code sandbox — are services you pick and connect. Every one of them has a hosted and a self-hosted path. This page is the decision guide: what each choice implies, what you must create yourself in the external service, and where the detailed setup steps live.

Only the LLM is required. The tracker and the knowledge base ship with the engine and are on by default; everything else is a surface you add when you want your company working where your people already are. A company with an API key and nothing else runs.

A useful mental model: for each integration there is usually

  1. Something only you can create — an Atlassian site, a Slack workspace, a GitLab group, an E2B account. Crewlet never creates top-level tenancy for you.
  2. Per-agent identities inside it — service accounts, bot apps, tokens. For Mattermost, Slack, and GitLab a provisioning CLI creates these idempotently; for Atlassian and GitHub you create them by hand, because neither third-party app issues a credential on a provisioner’s behalf. Those two still have a CLI — it reports which account each seat’s own credential turned out to be, and registers the webhooks.
  3. A webhook back to the engine — so external activity wakes the right agent. Self-registered where the API allows it, manual where it doesn’t. Mattermost is the exception: it has no usable inbound webhook, so the engine dials out instead and needs no reachable address at all.

OptionConfigNotes
Anthropictype: anthropicOfficial SDK; prompt caching set explicitly by the provider. Each request is shaped for its Claude model — adaptive thinking and an effort level on the current generation, a thinking budget only on the older models that take one (see Claude models). Required if you want Claude Code as the sandbox coding agent.
OpenAItype: openaiOfficial SDK; automatic prefix caching.
Any OpenAI-compatible endpointtype: openai-compatible + base_urlHosted aggregators (OpenRouter, Together, …), cloud gateways, or your own vLLM / LiteLLM deployment. Fully self-hostable. OpenCode (the provider-agnostic sandbox coding agent) can reuse this same provider.
A gateway or proxy in front of a vendorbase_url on anthropic / openaibase_url is not an openai-compatible-only field. It is merely required there. On a vendor entry it redirects that vendor’s own wire format at an egress proxy, an Anthropic-API gateway, or a subscription OAuth proxy. The last of those has terms and trade-offs worth reading first: Subscription LLM Backends § the proxy shape.
A coding CLI you subscribe totype: cli-agent + cli.agentNo API key: drives a coding CLI already installed on the engine host — the claude, codex, gemini, qwen, opencode, cursor-agent, copilot, grok, muse, kimi, hermes or pi binary — on the operator’s own subscription. cli.agent names the profile rather than the binary (claude-code, gemini-cli, kimi-code, …); Supported CLIs pairs them up. The CLI must be installed on the engine host. Flat-rate cost, higher per-call latency, and each seat gets an isolated CLI home so agents never share memory. See Subscription LLM Backends.

You can configure several named providers and pick per role (role.llm), plus a cheap auxiliary model per role (role.llm_auxiliary) for reflection and summarisation work. role.llm also accepts a list, which is how a subscription and a metered key compose: llm: [subscription, default] runs on the flat-rate CLI and falls through to the API key when the subscription window is spent. See Quickstart § LLM options.

Which to pick. A metered key is the default answer for a fleet that has to be fast and always available. A subscription CLI is the better answer when you are developing, evaluating, or running a small company yourself and want predictable cost — check your plan’s terms, which are generally written for interactive use by the subscriber. A subscription can also back the code sandbox — via a headless token on any backend, or via a local cell (providers.sandbox.local plus run_in: direct), which runs the coding agent on the engine host against the same login.

And a subscription CLI can run the turn itself. cli.mode: agent makes the seat’s executor the CLI’s own agentic loop — a real shell, a real editor, a real checkout — with the seat’s tools reaching it over an MCP bridge, as a detached run that outlives the turn. It needs more than a login: a CLI with a coding-agent runner (claude-code or opencode), a providers.sandbox catalogue to place the run in, and CREWLET_MCP_BRIDGE_URL set to something a box can dial. crewlet llm doctor checks all three. Text mode stays the default and is the better answer when you want the tool log to be the engine’s own. See Subscription LLM Backends § Two modes.

Embeddings (providers.embeddings) power the semantic half of a native knowledge search — the company’s own pages and work items, embedded once for the whole fleet, and each search’s query — and the agent-learning subsystem’s similarity recall (personal diary + episode recall). Any OpenAI-compatible embeddings endpoint works via base_url, including a self-hosted one; a model this build does not know states its width and its limits (the tokens one input may hold, and the inputs and tokens one request may carry), and so do the request limits of gemini-embedding-001 and embed-v4.0, whose compatible endpoints document none — see Configuration. Without an embeddings provider the engine still runs: search is keyword only, personal memory is chosen from a seat’s recent notes alone, and similar-prior-work recall renders nothing, because a seat’s recent turns are not similar work.


There is none to choose. The engine is one binary: its event stream is a NATS JetStream server it embeds, and its store is a local file it creates and owns exclusively. A single host runs a whole company with nothing else installed, in development and in production alike.

Two slots change once a deployment outgrows one node:

SlotOptions
Streamembedded (default) — and a fleet is the same embedded server with stream.cluster.name/port/peers naming its peers and stream.replicas: 3 · or nats + stream.url to dial a NATS server or cluster somebody else runs
Coordinationlocal (one node) · embedded-kv (a fleet — one node or three; two has no quorum and is refused by name)

Coordination takes no address of its own, and that is deliberate rather than an omission: the KV holding leases, ledgers and the token counter rides the stream’s own connection. Two connections to one broker fail independently, so a node could keep renewing leases over a link that still works while the one carrying its inbox has dropped — alive to its peers, deaf to its work.

The store is never one of these: it stays one file per node, which is why everything genuinely shared lives in coordination instead. See Running a Fleet and Deployment for sizing and broker authentication.


Agents file and pick up work in a tracker, and search a shared knowledge base.

You do not have to bring either one. The engine ships its own, and they are the default:

tracker:
backend: native # the default
knowledge:
backend: native # the default

That is the whole setup. Items and pages live in the fleet’s own store, there is a board and a page browser on the dashboard, seats get twenty-three tools for them — eighteen over the tracker and five over the pages — and your own AI assistant can reach them over /operator/mcp. Nothing to create, nothing to provision, no per-seat accounts, no webhook.

NativeAtlassian
Setupnonea site, a project, a space, a per-seat account each, webhooks
Where the record livesyour own deploymentAtlassian’s
People can use itthrough the Crewlet dashboardthrough Jira and Confluence, which they may already live in
Workflows, custom fields, sprints, permission schemesnoyes
Existing ticketsnone — it starts emptywhatever you already have

Take the native one unless you have a reason not to. The reasons are real and they are all about the people rather than the agents: a team that already works in Jira should not be asked to watch a second board, and an existing backlog does not migrate (there is deliberately no migration path — see below). If neither applies, the vendor path costs you a provisioning afternoon and buys nothing the agents use.

Either way, the org chart is the same. A unit’s project and space name its project and its container on whichever backend the company runs — which is why the fields are not called jira_project and confluence_space.

No migration between them. Switching tracker.backend does not move anything, in either direction, and the engine does not offer to: a half-migrated tracker where some items answer to one system and some to the other is worse than either, and it is a state nothing can detect from the outside. Choose once, per company.

Atlassian (Jira + Confluence Cloud or Data Center)

Section titled “Atlassian (Jira + Confluence Cloud or Data Center)”

The managed-SaaS path: Atlassian runs the tracker and the wiki, and agents work them through per-agent Atlassian identities. Per-agent setup is manual, since Atlassian exposes no service-account-provisioning API a CLI can drive.

What you do, by hand:

  1. Create the Atlassian site (or use your existing one) — Crewlet never creates the tenancy.
  2. Create per-agent identities: an Atlassian account (or API token identity) per agent seat, so issues can be assigned to agents and comments attribute correctly. Mint one API token per agent and reference them from role.mcp_env (JIRA_API_TOKEN / CONFLUENCE_API_TOKEN for the atlassian MCP server), plus one admin/service token for the engine’s org-level lookups (integrations.jira.token).
  3. Webhooks:
    • Cloud — install the Crewlet Forge app in your site; it forwards Jira + Confluence events to POST /webhooks/forge (signature-verified; needs the forge install extra).
    • Data Center — register webhooks directly against POST /webhooks/jira / POST /webhooks/confluence with an HMAC secret.
  4. Create the spaces/projects your units use (project / space per unit) and an Onboarding page per space.

Details: Jira · Confluence.

One knowledge backend per company, and validation enforces it: a company that sets knowledge.backend: native and declares integrations.confluence is refused. “What do we already know about this” must not depend on which searcher was asked. See Knowledge System.

An empty backend derives: declare integrations.confluence and you get confluence; declare nothing and you get native. So an Atlassian company that has not read this page keeps the backend it had.


Engineer roles read, review, and track code through per-role MCP tools, and author changes through the code sandbox under their own identities — so MRs/PRs come from the agent, not from you.

Option A: GitLab — gitlab.com or self-hosted

Section titled “Option A: GitLab — gitlab.com or self-hosted”

Per-agent identities are API-provisionable end to end, so one CLI run sets up the whole fleet — the same property that makes Mattermost the easiest chat to start on. GitLab Integration walks the bundled Nimbus example onto a local GitLab end to end:

  1. Create the top-level group (you) — e.g. gitlab.com/your-group — or run a self-hosted GitLab (any modern GitLab; a local instance ships as the compose gitlab profile for end-to-end testing). Set integrations.gitlab.url accordingly — the same config shape covers gitlab.com and self-hosted.
  2. Run the provisioner — crewlet gitlab provision company.yaml -public-url https://<engine> creates one service account per engineering seat (mentionable, assignable, reviewer-able), group/project memberships, per-agent PATs minted into your config’s own ${VAR} references, and the webhooks. Requires an admin-capable operator token for the run; see the permission matrix. On gitlab.com, note the prerequisites (service accounts need a paid tier; self-hosted has no such gate).
  3. Agents drive GitLab via the glab CLI’s MCP server (glab mcp serve, spawned per role with that role’s PAT); the sandbox git-auth recipe makes git push + MR creation work headlessly under the agent’s identity.

Details: GitLab integration.

github.com or a GitHub Enterprise Server — leave integrations.github.url unset for the former, name the instance for the latter:

  1. Create the org/repos (you), plus one PAT per engineer seat — GitHub has no API-provisionable service accounts, so per-agent identities are machine users or fine-grained PATs you create by hand (role.mcp_env.github carries Authorization: Bearer ${GITHUB_TOKEN_X}). The engine derives each seat’s login from its own token; nothing is declared.
  2. Register the webhooks with crewlet github provision, which mints the secret and registers one organization hook where the credential may (covering repositories created later) or one per repository where it may not — and reports which account each seat turned out to be, which is the finding that decides whether the integration routes at all.
  3. Agents get the full toolset of the hosted GitHub MCP server per role; the sandbox git-auth recipe has a GitHub form (credential helper on github.com + GITHUB_TOKEN in role.sandbox.env) — see GitHub integration.

A company can run both hosts: they are two hosts with different repositories on them, which is what a migration and an open-source presence both look like.


Chat is the human↔agent conversational surface (DMs, channels, escalations) and is also where the DACI decision framework plays out. Two backends ship.

MattermostSlack
Hostingself-hosted, open sourceSaaS
Credentials per agent1 (bot token)2 (bot token + signing secret)
Manual steps per agentnoneone OAuth Allow click
Engine must be publicly reachableno — the engine dials outyes — the Events API POSTs to it
Working statusfixed “is typing…” (default off)free text, per turn phase

The two are interchangeable as far as the engine is concerned; they differ in what you have to stand up and where your people already are.

Mattermost is the quickest to try. It ships in this repo’s docker-compose.yml, provisioning is one non-interactive command, and nothing has to reach the engine — so you can go from nothing to an agent answering in a channel on a laptop, with no account to create and no tunnel. It is what the bundled example runs on, which makes it the path with the least between you and a working loop.

Slack is where most companies already are. If yours is one of them, that is the answer regardless of anything above: the agents show up where the conversations already happen, under the workspace admin, SSO, retention and compliance setup your organization already runs, with nothing new to host or patch. It also renders the per-phase working indicator as real text, which Mattermost has no API for.

Running both at once is supported — each agent seat carries whichever identities it needs.

  1. Run the Mattermost server yourself (official Docker image; it needs its own PostgreSQL) and create the team agents will live in. Crewlet never creates top-level tenancy.
  2. Declare each agent’s identity in the company YAML — one per-agent bot token as a ${VAR} placeholder under role.integrations.mattermost, and the Mattermost MCP tool server in mcp_servers.
  3. Provision the bots with crewlet mattermost provision, which creates one bot account per agent, adds it to the team and its channels, and mints its access token into the ${VAR} the YAML references. The only thing you do by hand is generate one system-admin token, once.

Nothing needs to reach the engine from outside: it opens outbound websockets per seat rather than receiving webhooks, so no tunnel and no public URL. Details: Mattermost integration.

There is no self-hosted variant:

  1. Create the Slack workspace yourself (or use your company’s).
  2. Declare each agent’s Slack identity in the company YAML — the per-agent bot token + signing secret as ${VAR} placeholders under role.integrations.slack, and the shared Slack MCP tool server in mcp_servers. Each agent is its own bot identity: own token, own DM, own @-mention.
  3. Provision the apps with crewlet slack provision, which creates one app per agent through Slack’s App Manifest APIs, points each app’s event subscriptions at POST /webhooks/slack/{handle}, and writes the obtained secrets back under the exact ${VAR} names the YAML references. Two things stay manual, because Slack has no API for either: generating one app configuration token (once, ever) and clicking Allow on each app’s install. The engine’s API must be reachable from Slack (public URL, or a tunnel during development) for both the events-URL verification and the OAuth landing page.

Clicking through api.slack.com/apps per agent still works if you prefer it — see Manual Setup.

Details (scopes, Events API, thread routing, the working-status indicator): Slack integration.


Section titled “Code sandbox (optional but recommended for engineer roles)”

Lets a role’s Execute phase run a real coding agent (Claude Code or OpenCode) with a shell, a filesystem, and a git checkout, inside an isolated sandbox — see the Code Sandbox concept page.

OptionHowNotes
Noneomit providers.sandboxRoles use the native executor tool-loop; no code authoring.
E2B cloude2b: {api_key: "${E2B_API_KEY}"}Sign up at https://e2b.dev, create an API key in the dashboard. Fastest path.
Self-hosted E2Be2b: {api_key: …, domain: "${E2B_DOMAIN}"}Deploy e2b-dev/infra on your own cloud account, then point the same SDK/code path at it via domain. Your cluster issues its own E2B_API_KEY.
Local — containerlocal: {image: …} + run_in: containerDocker/Podman on the engine host. Real host isolation, no E2B account. You supply an image with the coding CLI installed. Can use a subscription login instead of an API key.
Local — directlocal: {} + run_in: directA process tree on the engine host, using the CLI (and login) already installed there. Fastest to stand up and the natural pair for a subscription backend — but the coding agent runs as the engine user, so it isolates state, not the host. Workstation or dedicated VM only.

providers.sandbox is a catalogue, so these are not exclusive: configure e2b: and local: together and give each seat its own run_in — the seat that needs the host’s subscription login runs direct, the seat whose generated code must never touch that host runs e2b. Where a seat that names none goes is default_run_in, which is required whenever some seat (or agent-mode entry) would fall to it on a catalogue offering more than one cell (and local: alone offers two); a company whose every seat names its own needs none — as long as something names a cell, since a backend is only built for a cell that is reached.

Coding-agent choice: OpenCode is provider-agnostic (reuses any OpenAI-compatible provider you already configured — no extra secret); Claude Code requires an anthropic provider entry, a cli-agent one, or an ANTHROPIC_* credential in role.sandbox.env (select it per role via role.llm_sandbox).

One networking caveat for local development: a cloud E2B sandbox cannot reach services on your laptop (localhost GitLab) — in-sandbox tool access to those needs a reachable deployment, a tunnel, or self-hosted E2B on the same network.


Two bundled Nimbus examples model the same seven-seat company at opposite ends of this page. examples/nimbus.company.yaml + examples/nimbus.config.yaml is the reference: a pick from every row — GitLab, Mattermost, a metered openai-compatible key, a sandbox its engineers push merge requests from, and the engine’s own tracker and knowledge base rather than a vendor’s. Read it to see what a full stack looks like written out, and Jira / Confluence for the blocks that move those two halves to Atlassian.

examples/nimbus-claude-cli.company.yaml + examples/nimbus-claude-cli.config.yaml is the short path: every pick that costs nothing extra to stand up — chat on Mattermost, the model a coding CLI on your own subscription, the tracker and knowledge base the engine’s own, and the code sandbox on the engine host (run_in: direct), which reuses that same subscription login rather than needing an account of its own. Its three engineering seats run that CLI in agent mode, so their executor is the CLI’s own agentic loop with a real shell. The rows that need somebody else’s service are deliberately empty, which is what lets the whole thing run from one compose profile and one crewlet llm login:

Terminal window
docker compose --profile mattermost up -d --wait
COMPANY=examples/nimbus-claude-cli.company.yaml scripts/mattermost-dev-bootstrap.sh
crewlet llm login default -from-host \
-company examples/nimbus-claude-cli.company.yaml \
-config examples/nimbus-claude-cli.config.yaml
crewlet run -config examples/nimbus-claude-cli.config.yaml \
-company examples/nimbus-claude-cli.company.yaml

Reading it top to bottom is the fastest way to see the choices on this page made concretely — each block carries the rationale in comments, including the ones it did not make and what adding them would take. Fill in the other rows one at a time from the integration pages linked above; nothing in the example has to be undone first.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.