Run an AI agent company — not a pile of prompts.
Crewlet is an open-source engine for orchestrating hierarchically organized AI agent companies. It treats the organizational hierarchy as the primary orchestration structure — knowledge, permissions, communication, and decisions are all scoped by position in the org chart.
name: "Acme AI"mission: "Ship AI-powered products fast"
policies: - "All features need PM sign-off before development starts" - "Communicate decisions in writing"
providers: llm: default: type: anthropic model: claude-sonnet-5-5 api_keys: - "${ANTHROPIC_API_KEY}" embeddings: type: openai model: text-embedding-3-large # the model decides the vector width # (3072 here), so there is no # `dimensions` to set api_key: "${OPENAI_API_KEY}" # used by the agent-learning subsystem # (diary vector search + episode recall)
# Org-wide roles — these sit above all teams and manage team leads.roles: # You, in the chart. A `kind: human` seat is addressable but never # spawned (no runtime, no inbox, no LLM) — it gives escalation a person # to stop at, and lets agents recognise your activity on the surfaces # you connect later. Needs at least one `contact` identity; scope # `manages` to the top seat so you aren't copied on everything. - name: Your Name kind: human manages: [CEO] contact: # One identity per surface you connect. Swap this for # `slack_user_id` (a `U…` member ID) if Slack is your chat. mattermost_user_id: "${MATTERMOST_FOUNDER_USERNAME}" # your chat username
- name: CEO handle: ceo # see the note under this block — set these now goal: "Set product vision, prioritize initiatives, and make final calls" backstory: "Experienced founder who balances speed with quality" manages: [CTO, PM] # A zero-integration way to see your first agent turn: a scheduled task. # Delete this once you have real integrations delivering work. schedules: - name: hello-crewlet cron: "*/5 * * * *" task: "Write a short status note on what the company should focus on this week."
# Flexible org structure — use any nesting depth and unit types.units: - name: Product Management type: team lead: PM purpose: "Define what gets built and why" # A unit's identity on the tracker and the knowledge base. Both default # to the engine's own backends, so these two keys are all it takes to # give the team somewhere to file work and write things down — no site, # no project to create, no per-seat account. They name a Jira project and # a Confluence space just as well if you switch the backend later, which # is why they are not called `jira_project` / `confluence_space`. project: PROD space: PROD roles: - name: PM handle: pm goal: "Turn business goals into clear specs and prioritized backlogs" backstory: "Data-driven product manager who writes crisp requirements" manages: [Engineer]
- name: Core Engineering type: team lead: CTO purpose: "Build and ship the product" project: ENG space: ENG goals: - "Ship MVP in 4 weeks" - "Maintain test coverage above 80%" roles: - name: CTO handle: cto goal: "Set technical direction, make architecture decisions, unblock engineers" backstory: "Senior architect with deep distributed systems experience" manages: [Engineer]
- name: Engineer handle: eng goal: "Implement features, write tests, and ship quality code" backstory: "Full-stack engineer who writes clean, tested code"logging: level: debug # every subsystem's DEBUG lines, in colour when you # are watching a terminal. Drop the block (or set # `info`) once the company runs format: console # console (default), text, json file: # optional, and worth it even here: a durable copy path: "./acme-data/crewlet.log" # IN ADDITION to what you are watching. # It rotates itself at 100 MB, keeping five files, # so scrolling past the interesting line costs # nothing. See Deployment → Logging
stream: type: embedded # a JetStream server inside this process: no # listener, no port, no service to operate store_dir: "./acme-data/stream" # leave empty and the stream is # in-memory, and nothing published survives a # restart. This company is on the engine's own # tracker and knowledge base — the defaults — so # every item and every page lives there too, and # the engine REFUSES to boot on an in-memory # stream rather than lose them at the first # restart. Name a directory
store: path: "./acme-data/acme.db" # this node's own database, owned # EXCLUSIVELY by this process. Not a shared # database and no DSN: two engines pointed at one # path corrupt it. A second file is created beside # it for the replicated estate — a snapshot is a # copy of one of the two, which is why they are # not one file. See concepts/architecture.md
coordination: type: local # a single node holding its own seat leases; # a fleet needs embedded-kv (see guides/fleet.md)
api: host: "0.0.0.0" port: 8000 # a port > 0 makes `crewlet run` serve the API EMBEDDED in # the engine process (dashboard + webhooks included) — one # process is the whole stack. (Any free port will do; pick # one nothing else on the host has already taken. Port 80 # is worth the privileged bind only once an external # service registers a webhook URL against this engine — # see guides/deployment.md.) auth: # Needed for WRITES and for /config. Reads — the dashboard, /events, # /agents — serve without one by default; add # `allow_anonymous_read: false` here to guard those too. tokens: - id: founder token: "${CREWLET_API_TOKEN_FOUNDER}"go install github.com/crewlet/crewlet/cmd/crewlet@latestexport CREWLET_API_TOKEN_FOUNDER="$(openssl rand -hex 32)"export ANTHROPIC_API_KEY="sk-ant-..."export OPENAI_API_KEY="sk-..." # embeddingsexport MATTERMOST_FOUNDER_USERNAME="you" # your chat username (the human seat)Getting Started
Section titled “Getting Started”- Installation — Install the binary; why there is no infrastructure to bring up, and what the compose profiles are for
- Quickstart — Build a four-agent company and watch its first turn, with LLM provider options (Anthropic / OpenAI / any OpenAI-compatible)
- Choosing Your Stack: the decision guide for every external dependency (the LLM provider, the tracker and knowledge base, the code host, chat and the code sandbox), what each path sets up for you, and what you must create manually
- Authoring with an AI Assistant — Let an AI write your company config: a step-by-step walkthrough, the
company-architectskill,crewlet schemafor editor autocomplete, and thecrewlet validate -jsonfix loop - Configuration Reference — Full YAML config schema and examples
Core Concepts
Section titled “Core Concepts”How the engine works, one subsystem per page:
- Overview — The org chart as execution graph, design principles, high-level architecture
- Architecture — The whole engine at six zoom levels: what sits outside the process boundary and what each dependency is for, what runs inside one node under
data/ingress/seats/workers, the hop-by-hop path a trigger takes from a third-party app webhook to an agent’s reply across a fleet, what a turn does and how one survives its own process, which of the four estates every table, bucket and stream belongs to, and what changes when a second node appears - Configuration — Two-tier config (ops-owned
crewlet.yaml+ a founder-owned revision in the store), bootstrap sequence, unconfigured state, live propagation, auth, the apply stage by stage, whole-config encryption at rest - Scaling Out — Why one node is the design’s degenerate case rather than a lesser path, what a node is (
data/ingress/seats/workers) the node that holds no data at all and how every node reaches the one replicated estate every data node holds whole, the five kinds of coupling that had to be resolved and which one a lock actually fixes, what the fleet shares in the coordination slot versus what stays per-node, the measured broker numbers the lease TTL and prefetch cap come from, and what the design does not promise - Object Store — Where a company’s files live: the row that names a file replicated like everything else, its bytes stored as one object per upload, under a key minted for that upload and never reused, in one store the whole fleet shares — a JetStream object store bucket on the fleet’s own broker at
stream.replicascopies (nats, the default) or an S3-compatible bucket (s3, with the permissions its identity needs), chosen by every node’sstore.objectsand recorded once so a node configured otherwise refuses to boot. Uploads that store the object before the row and are checked against the row’s SHA-256 on every read, theobject-collectorduty’s hourly collection of objects no row names and of unfinished uploads, its daily audit for missing and damaged files, why a deletion needs no lock, theobjects_missingalarm, choosing a backend, why there is nothing to drain, and how far it scales - Coordination: The fleet’s shared store and the line between it and the node’s own database: the three-valued answer every question here returns, which direction each contract fails in when the store cannot be reached, why a listing is certified against the stream’s own key index and read back from its leader so it is complete or unknown but never short, the slots a fleet is discovered from, why every retention is a bucket’s age rather than a per-write TTL and what sweeps the removal markers a bucket with no age would keep for ever, why fleet duties keep their leases in a bucket apart from the seats’, and what stays node-local or per-process on purpose
- Seat Ownership: How a fleet decides which node runs which seat, and why no two ever run the same one: TTL leases with epoch fencing, fair-share placement with give-back, the two release modes, freshness-based admission, the line between a copy that is behind and one that is wrong (and why lag at any size never moves work), owner-only inbox and sandbox-control attachment, the durable subscription that holds an unowned seat’s mail, the broker settings that must not delete it, how a removed seat’s mailbox is found and retired after a grace period, and why a wedged-but-alive node ends its own process
- Control Plane — How every node converges on one company config: the shared activation pointer whose own revision is the epoch, the reconcile poll, per-node apply status (
ok/error/degraded), the posture a lagging node takes (serve/wait/shed/isolated/stuck), what a running turn sees through a live apply, and what/healthand/readyreport - Secret Store — The company’s encrypted credentials on the coordination KV, read by every node and consulted ahead of the process environment when resolving
${VAR}:crewlet secrets set/list/unset/get/rekey, the/secretsAPI they write through, the-secret-storeprovisioning sink that hands minted credentials straight to the engine, store-wins precedence, and the Tier A root-of-trust boundary - Organization Model — Hierarchy, departments, teams, roles (seats), handles
- Humans in the Org Chart — Human seats (
kind: human): hierarchy membership, contact identities, notify delivery, escalation terminus, prompts and lookup — and the person at the dashboard: an API token bound to their seat, the Inbox and My work that answer for them, and every write made as them - Agent Runtime: how a configured role becomes a running seat, the states the dashboard shows, what each phase’s prompt is built from (including the chat thread a turn is handed rather than told to fetch), the built-in tools, and graceful shutdown
- Turn Engine — Per-agent executor / reviewer loop, how the engine decides a turn delivered, delegating to workers, colleague-surface tools, per-phase LLM models
- Subscription LLM Backends — Run agents on a coding CLI you already subscribe to (Claude Code, Codex, Gemini CLI, OpenCode, …) instead of a metered API key: the two modes (a text model behind the engine’s tool loop, or the CLI running the executor itself as a detached run with the seat’s tools bridged in), per-seat state isolation, the in-prompt tool-call channel, how the answer is found in a CLI’s own output and what happens when the profile no longer matches it,
crewlet llm login, falling back to a key when the window is spent, and the other shape — an OAuth proxy behind an ordinary HTTP entry, and what that moves onto you - Code Sandbox — Sandboxed coding-agent execution: the
run_sandboxtool,providers.sandboxas a catalogue with a per-seatrun_in(includingself, for a seat whose executor already has a shell), E2B cloud/self-hosted, local boxes (direct or containerised, on a subscription login), Claude Code & OpenCode runners, git-auth recipes, mid-run clarifications - Tool Skills: Knowledge-base-sourced prompt fragments (Confluence pages, or the engine’s own) that teach agents how to use each tool / MCP server, and how every node keeps its copy current: a walk at boot, on an apply that moves the skills source and every 10 minutes, a read of the one page a webhook names broadcast to the rest of the fleet, and a bounded retry behind a walk that failed
- Tool Capabilities — How the engine stays tool-stack agnostic: capability prose + MCP annotations, no hardcoded tool names
- Event System: EventQueue, topics, routing, inbox batching, listing the durable subscriptions a broker holds, reading the fleet’s history from every node’s own store (and why a departed node’s detail departs with it), distributed tracing
- Integration Reconcile — What checks that an integration still works: the finding vocabulary every third-party app reports in, why the excess-access advisory is ranked last, a cadence derived from who has to act rather than from what failed, the fleet singleton that runs it, and what the loop deliberately will not do (tear anything down, or rotate a credential that works) — it does register the webhooks and create the accounts, because connecting an integration is the permission
- The Tracker — the two shapes a company’s work can take: the engine’s own tracker (
tracker.backend: native, the default — every change a record on an ordered log the fleet shares, applied into identical SQL copies on every node, with a board, seat tools and an MCP surface), or an external PM tool the engine deliberately mirrors none of - Scheduling — Role/unit-scoped cron-style recurring work (standups, audits, nightly jobs)
- Knowledge System — Shared knowledge behind the
knowledge.Searcherseam — exactly one backend per company, either the engine’s own pages (keyword, semantic or both fused) or a live Confluence search, who read each page and who links to it, edited from the dashboard as the signed-in person — plus the privateagent_diary, andcrewlet search evalfor measuring the semantic half on your own vectors - Agent Learning — Reflection loop, skill induction, episodic memory, counterparty profiles
- Conversation Sessions — What a seat already said in one Slack thread / issue / pull request, carried into that conversation’s next turn: the entry shape and its elision budgets, which turns are recorded, why the key is the conversation and the dedupe is the work key, and why this is a structured ledger rather than a transcript replay
- One-on-Ones — Manager↔report coaching as a usage pattern over the scheduler + A2A channels + learning loop
- Decision Framework — DACI model for multi-agent decisions, and the structured ask a decision is recorded as
Integrations
Section titled “Integrations”Connecting the external surfaces agents work on. The tracker and the knowledge base are the two that are optional: the engine ships its own and runs them by default, so Jira and Confluence are the path you take when your people already live there.
An alternative to the engine’s own tracker. Cloud or Data Center: webhook routing by mention / assignee / watcher / project lead, derived seat account ids, per-role MCP tools, and crewlet jira provision
An alternative to the engine’s own knowledge base, and the one that searches as the asking seat so Confluence’s own page permissions are the boundary. Cloud or Data Center: live CQL search, page-change routing, tool-skill pages, and crewlet confluence import
gitlab.com or self-hosted: API-provisioned per-agent service accounts, crewlet gitlab provision, webhook routing, per-role MCP tools, sandbox code authoring
github.com or Enterprise Server: one GitHub App per agent (create it, then install it) with three access tiers, webhook routing by review request / assignment / review verdict / mention, derived seat logins, participant fan-out, organization or per-repository hooks, and crewlet github provision
Monitor alerts as inbound events through the Webhooks integration: routing by monitor tag with a required fallback seat, a trigger and its recovery as one conversation, and shared-token verification because Datadog cannot sign a body
One app per agent: crewlet slack provision builds and installs them from a manifest, per-seat webhook routing with thread follows, a text-carrying working indicator, and the Slack MCP tool server
Self-hosted open-source chat: one bot account per agent, crewlet mattermost provision, a websocket event fleet instead of webhooks (no public URL needed), and thread routing
Guides
Section titled “Guides”- Tools & MCP — Built-in tools, MCP integration, the two-value tool-origin grammar the dashboard groups on, and how you extend a binary that loads no plugins
- Deployment — The single host, the compose profiles, the stream beyond one host (an embedded cluster or an external NATS one, with its auth and TLS), the event store, tracing, logging (levels, the three formats, colour, a rotating log file beside stderr — with its own shape and level, or taking the stream over entirely — and the embedded broker’s own output, which is its own switch)
- Budgets and Spend — The two numbers that describe what the models consumed and why they are never comparable: the budget counter a ceiling is enforced on, per day, ISO week and month on the company’s clock, and the spend rollup that says where the tokens went — the live 24 hours from each node’s projection, and every named window of up to 90 company days from the replicated usage domain, the same on every node. How far back each reaches (181 days of spend, 30 of turns), the four bands a phase chart draws, the provider split, the cache share, and the seats’ auxiliary model — what it spent by stage and purpose, which of it a turn and its task pay for, a person’s questions as their own row, why the counter and the rollup can still differ, and what is not counted (embeddings)
- Backups & Restore —
crewlet backupagainst a running node, what the directory it writes contains and what the copy is a copy of, the cold runbook behind it, restore ordering and its hazards (epoch rewind, node identity), and which losses are survivable without a backup - Running a Fleet — When to run more than one node, node roles (and nodes that hold no data), seat placement, draining and rolling upgrades, and removing a data node for good without leaving the company’s files a copy short
- Running One Agent Somewhere Else — Put a single seat on a host that can reach what it needs — an internal API, a licensed binary, a GPU, a lab network — without moving the company: what a satellite is, whether it holds data or is a small disposable node that holds none, what moves with the seat (its MCP servers above all), what the node still needs outbound, and what a pin costs when the host is down
- The Org Builder: editing the organization, and creating the company, from Agents › Edit org: what the builder shows and why, arriving from a link, creating a company from a template, reading the organization on the structure chart, the reporting chart and the table (their keys, the lead chip and reordering), how every draft is checked by the engine, what each editor can and cannot change and why, the integration fields a seat carries, what moving, deleting and changing a seat’s kind do, reviewing and saving (including a save whose answer never arrives), following the revision until every node applies it, keeping a
company.yamlin step with-import-companyand Copy as YAML, keeping a draft across a reload and what asks before work is left behind, updating a draft when somebody else saves first, and undo, redo and the keyboard - Configure via API — End-to-end curl recipes for bootstrapping a company through
/config/* - The Work Tracker — The product guide to the engine’s own tracker: projects and keys, everything a task carries, the catalogue of types and fields a company declares, the five view shapes and the six every container has without anybody saving one, what a board groups on and why the due bands are cut on the company’s own day rather than each reader’s, why the trash is a query rather than a shape, the timeline’s bars, bands and dependency arrows and what it refuses to place, who is carrying how much and the two things the workload will not do, manual board order and what a re-spread is, a person’s own inbox, queue and pins and the three different authorities over them, the tags any seat may declare and the rules that keep them one grouping, the eighteen tools a seat has and the five that count as a delivery, the one wake that is about something other than a task, why a handle is checked before it is stored, what a person reaches through the dashboard and their own assistant, and the difference between removing, deleting and purging
- Replication — How the fleet’s copies of the company’s work stay in step: the three numbers a reader must not fold together, the three outcomes a write has and what to do with
pending, what an acknowledged write has actually reached at each topology, the two regimes above and below the trim floor, replication lag as two positions, what a bulk gesture costs every peer, and what the design does not promise - Search — How a knowledge search is answered and what to do when it is slow: the two rankers and the bucket both of them carry, when a fleet divides the scan and the floor below which it does not, why the merge is by score and only then fused, why BM25 stays comparable across a divided scan, and what a partial answer names rather than silently truncates
- Read Consistency — Picking how fresh an answer has to be: the four levels and what each costs, how
linearizableworks as a barrier append rather than a field check, the thirteen refusal codes with a remedy each, completeness as a fact separate from freshness, and the three things no level can strengthen - Retention — What the log keeps and why it sometimes will not shrink: the one number to watch, the six trim terms in plain words, why a company that never backs up never trims, snapshots and the join runbook, eviction and readmission, what a full log refuses, the storage forecast, and the recovery-operations profile
Reference
Section titled “Reference”Command reference
REST API routes and schemas, the WebSocket protocol (the socket reads and pushes; REST writes), the health envelope, the tracker and page surface over the native backends, /operator/mcp — the company’s own tracker and knowledge base served to your AI assistant over MCP — and /operator/act, where the dashboard writes as the person its token is bound to, and how the dashboard’s own files are served, cached and compressed
The dashboard’s information architecture and its visual system: the one sidebar and the three nouns that keep it one (a sidebar row is a workspace or a kept object, a section is a path drawn as a tab, an object tab is a query on an object’s page — and nothing is two of those), the route table and the gate that holds it against the router’s own resolver, the frame every screen wears and its three breakpoints, why Home is the landing screen and what the wake reason on each Inbox row is, the one rule (colour carries state, never identity), a palette whose every contrast and separation claim is recomputed from the shipped stylesheet in both themes, the four rules that keep a running turn’s transcript from moving under the reader, ⌘K, acting as yourself (every write control is shown, and disabled with its reason when you cannot make it), honest empty states, the turn trace, how it is built (one REST transport, one write client, src/contract/ and the gates that hold it), why the built bundle is committed, and the rules a change has to keep — tokens, never money, among them
Every instrument the engine exports, generated from the catalogue: what each one measures, its unit, its attributes, and the failure it makes visible. Plus the OTEL_* variables that switch the export on, and what is deliberately not exported
Every condition the engine raises about itself, generated from the alarm table: what each one means and what to do about it. One table, reaching you as a gauge and as a named log line on entry and exit
All configuration env vars
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.