Skip to content
You are reading documentation for unreleased main. Read the 0.1 version.

Tools & MCP Integration

Agents interact with external systems and internal engine operations through tools. Crewlet supports both built-in tools and dynamically discovered MCP tools.


The engine registers these into each epoch’s tool registry with the origin builtin. A tool whose dependency is absent (no store, no knowledge backend, no sandbox) is omitted rather than registered and broken; Agent Runtime § Built-in Tools says when each one is registered.

ToolDescription
lookup_colleagueResolve any colleague identifier (handle, role name, a human’s contact ID) to one seat, case-insensitively, with partial and fuzzy fallbacks; ambiguous queries return the candidate list so the LLM picks rather than guessing
reflect_and_persistCapture a durable fact in the agent’s private diary (kind: long or short)
refresh_memoryRe-run the personal-memory filter mid-turn after gathering richer context
query_episodesRecall the agent’s own past turns: by meaning, by conversation, or most recent first
use_skillLoad one of the agent’s own synthesized skills on demand
refine_skillReplace a synthesized skill’s body with a corrected procedure; the previous version is kept
mark_onboardedStamp the agent’s onboarding marker after reading the relevant onboarding pages (offered to the onboarding pass)
a2a_askAsk one AI colleague one question over the private A2A channel (see Turn Engine § Colleague-surface tools)
search_knowledgeRe-run the shared-knowledge search mid-turn, once the agent knows what the task actually needs. Registered wherever the company has a knowledge backend at all
load_tool_skillLoad the full body of a Tool Skill by key
run_sandboxHand a code task to a coding agent in a sandbox

delegate, the tool that hands work to short-lived workers, is not registered here: it is built per turn for the executor, because it carries that turn’s grant.

Twenty-three more, registered only where the company runs the engine’s own backends (tracker.backend: native / knowledge.backend: native, which are the defaults). A company on Jira and Confluence gets none of them, and that is the point: a seat offered a tool against a tracker its company does not run would reach for it and fail at the call, and a model shown a tool that always fails learns to distrust the whole catalogue. The seventeen below are the item, file and page tools; the other six read the catalogue, the projects and the activity feed, and The Work Tracker lists the tracker’s eighteen in full.

ToolDescription
list_work_itemsThe board, filtered — what you are assigned, what is open in a project, whether something was already filed
get_work_itemOne item’s description, thread, history and links, by key or id
create_work_itemFile one. project defaults to the seat’s own unit’s, and is required when the unit owns none. status says where it starts (default todo). With no assignee it goes to the project’s default assignee, else to triage for the lead, and the answer’s assignee says which. ask (with an optional decision) files it as a question to that person, in one record
update_work_itemMove it — status, assignee (with an optional reason the new assignee reads), priority, labels, links, one checklist change — with an optional if_match that refuses on a concurrent edit
comment_on_work_itemPost to the thread. Mentions wake the seats they name; the turn’s own key makes a re-run turn post once. ask puts a question to somebody and decision structures it as options; answers with choice answers one
merge_work_itemFold a duplicate into the item that survives — linked, its subtasks re-parented, and closed as cancelled
move_work_itemMove a top-level item and its subtasks to another project, re-keyed there with the old keys still resolving — the project lead’s or a person’s own
search_work_itemsFind an item by what it says, ranked over titles and descriptions
list_project_filesA project’s files — reports, specs, notes — in path order, a page at a time
read_project_fileOne file’s text, a page at a time from an offset; bytes as base64 on request
write_project_filePut a file at a path, creating it or replacing it, with an optional if_version. Writing exactly what the file already holds, with the same type, stores nothing and answers outcome: unchanged
remove_project_fileTake a file out of a project; its content is deleted from storage
list_pagesBrowse the knowledge base by container, parent or title
get_pageOne page’s body, breadcrumb, children and history
write_pageCreate one. Titles are addresses and are unique per container; a parent must be a page in the same container that is not in the trash; a body links another page by id, [its title](#/knowledge/pages/<page id>), which is what “Linked from” reads
save_pageEdit one, stating the version you read — there is no per-field merge that makes overwriting prose safe. A parent moves it, and may not be the page itself or one beneath it; an empty one moves it to the container’s top
comment_on_pageRemark on a page, or replace one of your own with edit

The writes on each side count as a delivery for the turn’s own did-this-reach-anybody gate, and each waits for its own write to reach this node’s projection before answering — so a turn that files an item and then lists the project sees what it just filed.

Every write answers with its outcome — applied, pending or unknown, the three values every write has — and the position it is durable at. A write whose outcome is unknown is never answered with the id, key, version or revision of something that may not exist, on any write tool on either surface: the call fails, saying that the write may have landed and may not, under which operation, and what to do. That includes create_work_item, whose unknown answer names the key this attempt minted, if it minted one, as the key the item has if it was filed — never as a receipt — and a comment, whose id, mentions and ask are not reported beside a remark nobody can say was posted. An update_work_item whose change is unknown writes none of the dependency changes it was also asked for: it stops there, and the same call made again answers the change first and writes them after. And the version an update answers with is always the item’s own and its newest — the one to send back as if_match — never another item’s that the same call wrote, and never one its own dependency change has already moved the item past, when a call changes both; its position is likewise the call’s last write, so waiting for it waits for all of them.

Where this node’s operation ledger cannot vouch for the operation — it was minted before the ledger may have lost rows, a seat woken by a backlog trigger just after its node adopted a snapshot — the answer says this node cannot tell, because the same operation asked here answers the same way until the write reaches this node, and tells the caller to look before writing it again; a gesture that stopped part of the way through at such a step is not told to repeat itself either. A seat’s own write under a plain lost acknowledgement is told to repeat the call with exactly the same arguments, before any different call to the same tool — that is the same operation, and it lands once. Reworded, it is a new one. Where the ledger cannot vouch for it, that same repeat is still safe — it publishes nothing this node cannot vouch for — but it answers the same way until the write reaches this node, so looking says sooner. The two writes that state a whole value rather than change one — write_project and write_work_catalogue — say instead that repeating them is harmless: each call is a new operation stating the same thing.

A create or an update that declares its labels first (labels_create_missing) and cannot tell whether that declaration landed stops there too, and is answered under its own name as the write it did not make — the item not filed, the change not made — because the declaration comes before the item’s own write. The same call made again answers the declaration and then makes the write, once; where this node’s ledger cannot vouch for the declaration, the answer says to declare the tags with write_project first, which is harmless if they already landed, and then to make the same call — and says what that call will do. Its own write dates from the same instant as the declaration, so this node cannot vouch for it either: it answers with the item an earlier attempt filed, or the change this node still holds a record of, and otherwise answers unknown again, which then means looking for the write (list_work_items, get_work_item) rather than making it another way — or, from the operator’s surface, making the same call with its op_id through another node whose ledger reaches back that far.

The same tools are served to your AI assistant over /operator/mcp, with the writes attributed to your token rather than to a seat, and eleven more beside them that no seat is given: the saved views, the catalogue write, a person’s own queue and inbox, the trash, and the board drag. That set is built once per company as ONE operator catalogue, and every operator transport serves it, so what a person can do is the same whichever way they reach the company.

Over MCP, each tracker write’s answer also carries the op_id of the operation the call was, and the write tools take it back as an argument: an assistant has no turn to repeat, so sending the same call with that op_id is how it finishes a write that came back unknown or stopped part of the way through, instead of filing it twice. An op_id is that one call and no other: it carries a digest of the call’s arguments, and brought back with any other argument, or to another tool, it is refused before anything is written. A seat is never offered the argument — its turn is its identity — and a seat’s call that sends one is refused. The dashboard’s transport, /operator/act, takes no op_id either: the operation is derived from the request’s own request_id (a UUIDv7) and the token that sent it, so the retry of a gesture is the same request sent again, and an op_id beside it is refused.

Note the deliberate split between personal and shared writes: reflect_and_persist is personal-only (it writes to the agent’s private agent_diary), while team-shared content is a knowledge-base page — write_page on the native backend, or the vendor’s own MCP tools on Confluence (see Knowledge System). use_skill resolves the agent’s own synthesized skills; shared procedures are knowledge-base pages.

On a vendor tracker there are still no task builtins. create_task, assign_task, update_task and list_tasks are not registered against Jira or GitLab issues — an agent works those through the vendor’s own MCP tools, so the engine mirrors no state it would have to keep in step. See The Tracker.

Roles with GitHub credentials in mcp_env.github get a per-role instance of the remote GitHub MCP server (declared as a shared: false http entry in mcp_servers), giving them the full GitHub toolset for reading/reviewing/tracking code (issues, PRs, repos, code search, actions); code authoring goes through the code sandbox. See GitHub Integration.

Every registered tool records who registered it, and that is recorded at registration because it cannot be recovered afterwards: a tool an MCP server serves is structurally identical to one the engine ships — same name, same schema, same call signature. With nothing recorded, a tool missing because its server failed to start reads as a missing builtin, which sends an operator to debug the wrong subsystem.

GET /tools reports it as each tool’s source, and the dashboard’s Settings › Tools & MCP screen groups on it:

sourceWhere the tool came from
builtinShipped by the engine. The agent-to-agent tools are builtins too — a2a_ask is registered by the same walk, so “a2a” is a capability rather than an origin
mcp:<server>Discovered on an MCP server. <server> is the bare template name, never the per-role instance: two seats’ children of one template are the same integration to a reader grouping the catalogue

Those two are the whole grammar. A server that fails to start is visible as a missing group, rather than its tools quietly going absent from the builtins — and, above the catalogue on the same screen, as a row of its own: every node re-publishes what its MCP starts concluded on its presence heartbeat (per server, its instances counted), so MCP servers shows each server as Running, Partly failing, Failing, Not started or Not reported with one cell per live node and the first failure’s reason and seat (GET /mcp-servers, operator-only; the heartbeat carries each reason bounded, and the whole text is the node’s mcp_server_failed log line). Clicking a server opens its own page: its state, reach and launch, every node’s reason as the node reported it rather than clamped to the grid’s two lines, and the catalogue narrowed to its tools — which, for a server that never started, says it registered none and why.

Add an MCP server on Settings › Tools & MCP (operator token required) writes the same mcp_servers entry you would write in YAML: a name, a command and its arguments (stdio) or an address (http), shared or one per seat, and the environment or headers — values as ${NAME} pointers into the secret store, which the form offers as you type $. It checks the whole company with the server in it before storing, and adds it after every server already declared. A name the configuration already carries is refused; a name only a node still runs, from a revision it has not applied past yet, is free: the create is judged against the configuration alone. Over the API it is PUT /config/mcp-servers/{name} with If-None-Match: * — see configure via the API. A per-seat server still needs a seat to declare credentials for it under mcp_env, which is written in the org builder.

Beside the origin, every catalogue row carries the behavioural hints the tool was registered with — what an MCP server advertised, plus whatever an operator overrode on top — and where calling it lands:

FieldWhat it says
annotations.read_onlyThe tool modifies no state
annotations.destructiveThe tool may perform irreversible updates
annotations.idempotentRepeat calls have no additional effect
annotations.open_worldThe tool reaches entities outside the local system
deliversWhere a call puts something in front of somebody outside the turn — empty for a tool that reaches nobody

Each hint is three-valued — yes, no, unknown — and unknown is a first-class answer, not a soft no: it means the server did not advertise the hint at all. The engine reads them that way everywhere. An unannotated tool is not a known read, because treating unknown as read-only would exempt most of a fresh server from the delivery fence the moment it is added.

delivers is not “was this served by MCP”. A proven read-only MCP tool delivers nowhere, and the engine’s own work-item comment tool delivers although it is a builtin — the surface is a first-party declaration made at registration, and for an MCP tool it is the server, because nothing else about it says where its call lands.

The Tools screen renders all of this: the strongest thing a tool’s hints positively assert, whether it reaches outside the company, where it delivers, and how many arguments its schema declares. A tool whose server advertised nothing shows as unknown rather than as a read — which is the row to check before granting a seat a new server.

A tool that cannot do what it was asked answers with a failed result rather than an error: the turn is fine, this call is not, and the sentence goes back to the model so it can try again with a better argument. That sentence is written for a model and tuned against how models behave, so it is not a contract anybody else can parse.

The engine’s own tools therefore also say what kind of refusal it was — a machine-readable class beside the sentence, which is what a surface that is not a model acts on:

ClassWhat it meansWhat a caller should do
invalidAn argument is wrong, and the sentence names whichSend something different
not_foundThe item, page, skill or colleague the call is about does not existStop; it is gone or was never there
forbiddenThis caller may not do this here — outside a turn, somebody else’s inbox or priorities, a view protected for its owner, a project’s policy without the lead’s authority, a reserved containerAsk whoever the sentence names
stale_versionThe object changed after the caller read itRead it again and decide from what it says now
conflictThe write lost its race to other writers, or the object’s state moved under itRead it again; a retry may land
existsWhat the call would create is already thereEdit the existing one
already_answeredThe question this answers has an answerRead the answer
reassignment_budgetThe item has been handed on as often as it may beDo not reassign it again
inbox_fullA person’s inbox list is at its ceilingMark older entries read
not_runningThe run or turn the call addresses is not running, or not waiting for thisNothing to act on
steer_unsupportedThe running turn’s runtime cannot take a note mid-turnWait for the turn to end
budget_exhaustedThe company’s token budget has no room left in one of its windows, so a call that would spend tokens was not made; nothing was spentWait for the window the sentence names to reset, or raise its ceiling
unavailableThis node cannot serve the call right now, or the company does not run what it needs — a read that failed, a log that refused the append (a maintenance or sealed fleet included), an unconfigured backend. Never “it does not exist”Retry, or use the backend the company does run

An argument that merely names something missing — a parent, a waiting_on item, an assignee nobody has — is invalid, not not_found: the fix is the argument, and the call is not about that object.

invalid is claimed, never assumed. A write the work tracker refuses on what it was asked — a field value it will not round, a tombstoned item, a cap the object would pass — is marked as such where the refusal is written. Any failure that is not marked is read as the node’s: a store read or a log call that failed mid-write is unavailable, so a person is never told to change an input that was never wrong.

A call refused before any tool ran is classed too. A tool name nothing registered is not_found; a real tool this phase was not offered, or one a skill guard holds back until its skill is loaded, is forbidden.

An MCP server’s failure carries no class. Its prose is the server’s, and the engine will not guess a class from text it did not write; a reader that needs one treats an unclassified failure as the server’s own.


There is no plugin API and no runtime loading. Crewlet ships as one binary, and nothing under internal/ is importable from outside the module — so an extension cannot be a library the engine loads.

The extension point is MCP, deliberately. A tool server is a separate process (or a remote URL), it carries its own credentials, it can be written in any language, and a server that crashes takes down a tool group rather than the engine. Everything above about mcp_servers is that surface.

Two things MCP does not cover, and what to do instead:

  • A new chat or tracker third-party app. Routing an inbound delivery to a seat needs a parser, and that is an in-tree Go interface — the notification spine is backend-neutral by design, but a third-party app contributes a client, a parser and a transport as code. That is a pull request, not a config entry. The eight this build serves are Mattermost, Slack, Jira, Confluence, GitLab, GitHub and Datadog, plus Atlassian’s own Forge relay — every one of them routes end to end. A config block the engine cannot honour is refused rather than ignored, because a silently dropped integration block looks exactly like one that is working until somebody notices the messages never arrived.
  • Company-wide periodic work. An MCP server is called by an agent; it does not get a tick of its own. Schedule it as cron work against a seat, which gives it an agent, a turn, and the engine’s own at-most-once delivery across a fleet — rather than a loop that would run once per node.

Instead of building hardcoded API wrappers for external tools (Jira, Slack, Confluence, GitHub, etc.), Crewlet uses the Model Context Protocol (MCP) for dynamic tool discovery. This gives agents access to the full capabilities of external tools — not just a curated subset.

Bridgeowns every MCP server:start, stop, restart on apply

stdio childthe engine spawns the serverand speaks JSON-RPC 2.0 overstdin/stdout

HTTP / SSE clientconnects to an already-runningserver by URL

tool registryeach discovered tool registeredas mcp:<server>, globallyor in a per-role map

the MCP server's own API(a tracker, a code host, a wiki)

Crewlet engine

A stdio server is a process tree, not a process: npx execs a launcher that execs the real server, so the engine puts each child in its own process group and signals the group. Killing only the pid it spawned leaves the grandchild holding the credentials and the port.

  1. Stdio (Crewlet launches the server) — The engine spawns MCP servers as child processes (e.g., npx @anthropic/mcp-atlassian). Crewlet manages the full lifecycle: start on engine boot, stop on shutdown. Communication uses JSON-RPC 2.0 over stdin/stdout.

  2. HTTP/SSE (externally provided) — The engine connects to an already-running MCP server via URL. Supports JSON and SSE response modes with automatic session management and reconnection.

On connect the engine (identifying itself as crewlet in the handshake) negotiates the newest protocol the server speaks: it probes the modern server/discover method first and falls back to the legacy initialize handshake automatically. Servers built on older MCP SDKs may log a one-time “unknown method” warning when they see the probe — harmless, and it stays out of your console because of the rule below.

A stdio server’s stderr is never passed through raw. Every line the child process writes (startup banners, tracebacks, that probe warning) becomes a structured server_stderr DEBUG event attributed to the server, instead of foreign log lines interleaving with the engine’s own stream. When a server fails to start, the last lines it wrote are surfaced with the failure as a single server_stderr_tail ERROR event — that tail usually names the real cause (bad token, missing binary, import error). Run with -debug to watch a server’s full stderr live.


Each Role names its per-server credentials directly in mcp_env, so every agent authenticates as itself in external tools. The engine applies these as env vars for stdio servers and HTTP headers for http servers — it stays tool-agnostic, reading only mcp_env (and, for the Slack transport, integrations.slack):

roles:
- name: Senior Engineer
integrations: # per-agent transport identity (inbound webhook
slack: # verification + the working indicator)
bot_token: "${ALICE_SLACK_BOT}"
signing_secret: "${ALICE_SLACK_SIGNING}"
mcp_env:
atlassian:
JIRA_USERNAME: "${ALICE_JIRA_USER}"
JIRA_API_TOKEN: "${ALICE_JIRA_TOKEN}"
slack:
SLACK_MCP_XOXB_TOKEN: "${ALICE_SLACK_BOT}" # same token, the Slack MCP subprocess
github:
Authorization: "Bearer ${ALICE_GH_TOKEN}"

A Slack-enabled agent names its bot token in both role.integrations.slack.bot_token and role.mcp_env.slack.SLACK_MCP_XOXB_TOKEN: the two are different consumers — the notification transport vs. the Slack MCP subprocess — and both reference the same ${VAR}, so no secret is duplicated. The Atlassian token (JIRA_API_TOKEN / CONFLUENCE_API_TOKEN) and the GitHub PAT (Authorization: Bearer …) likewise live wherever the consuming MCP server reads them.

When the engine launches a per-role MCP server instance, it merges the base server config with the role’s mcp_env (applied as env vars for stdio servers, HTTP headers for http servers).

A unit declares its tracker project and knowledge container identity under project / space (used for inbound webhook routing and as the team’s write home; not a tool credential, and it does not scope knowledge reads). The two keys are vendor-neutral: they name a native project and container, or a Jira project and a Confluence space, depending on which backends the company runs. Real per-agent tool credentials still live in mcp_env, which the unit’s direct agent seats inherit:

units:
- name: Backend
type: team
lead: Tech Lead
project: "BACK" # the unit's tracker project (integration identity)
mcp_env:
atlassian:
JIRA_URL: "${JIRA_URL}" # shared by the whole unit
roles:
- name: Tech Lead
mcp_env:
atlassian: { JIRA_API_TOKEN: "${TL_JIRA_TOKEN}" }
- name: Engineer
mcp_env:
atlassian: { JIRA_API_TOKEN: "${ENG_JIRA_TOKEN}" }

Inheritance: the unit’s mcp_env is the base and a seat’s own values override it variable by variable, so a seat that sets one variable of a server keeps the unit’s other variables for that server. Only the unit’s direct agent seats inherit it. A child unit inherits nothing (it declares its own block), and a human seat inherits nothing, because a human seat runs no tools and may not carry an mcp_env. The unit’s project / space identity is separate from these credentials.


shared: decides which of two quite different things a server is, and the difference is a lifetime as much as a scope.

shared: true (default)shared: false
What it isOne child for the companyA template: one child per role that declares credentials for it
Whose identityNobody’s — it carries no seat’s credentialsThat seat’s, from role.mcp_env[name]
Who can call itEvery seatOnly the seat whose child it is
LifetimeThe config epoch — started on apply, replaced on the next oneThe seat’s lease — spawned when this node claims the seat, killed when it releases it
Use it forA shared knowledge base, a read-only reference serverA tracker, a chat backend, a code host — anywhere the action must be attributable to this agent

A shared mcp_servers edit takes effect on the next turn, not at the next restart. Applying a revision reconciles the shared bridge server by server: an entry that did not change is left alone, and one that was added, removed or re-pointed starts, stops or restarts only that child. A seat mid-turn finishes on the tool surface it started with, and its next turn renders the new one, the same next-turn promise tool skills, embeddings and the org chart make. A per-role (shared: false) child is not on the apply path: it belongs to the seat’s lease, so an apply rebuilds the catalogue each held seat’s turns are built against (its builtins and shared servers) and leaves the running child alone. A change to a per-role template or to a seat’s mcp_env reaches that child when the seat next changes hands.

A per-role child belongs to a seat, not to a node. In a fleet each node claims a slice of the company, and it spawns children only for the seats it holds — so the company’s processes are spread across the fleet rather than run N times over. It also means a seat that moves to a peer takes its identity with it: the credentials in a child are that seat, and one left running after the lease moved would let the old node keep acting as an agent it no longer serves.

Each seat gets its own surface, and that is a correctness property rather than tidiness. Two children of one template publish the same tool names, so a single shared catalogue would keep whichever registered last and hand it to everyone — every seat calling one child, acting under one agent’s identity in the tracker, invisibly, because the call looks identical from the engine’s side. A claimed seat therefore gets its own registry (the company’s catalogue plus its own children’s tools) and its own bridge (holding only its own children).

A seat that declares no mcp_env for a template gets no child. A template with nobody’s identity in it is a server nobody can act through, and offering its tools anyway would put entries in the prompt whose every call fails authentication.

A server that will not start costs its own tools and nothing else. The seat keeps its builtins, the other servers keep working, and the operator sees that server’s group missing from the Tools room — which points at the right subsystem, where builtins quietly shrinking would not. It is logged as mcp_server_failed with the reason. Failing the apply instead would take a working company offline because one vendor’s binary was absent from an image.


All tool servers go in mcp_servers — including the Jira/Confluence (atlassian), Slack, and GitHub servers. The integrations.jira / .confluence / .slack / .github sections carry only non-tool config (admin credentials, webhook secrets) — MCP servers are never declared there. Per-agent identity comes from role.mcp_env[name] — env vars for stdio servers, HTTP headers for http servers:

mcp_servers:
# stdio, shared by all agents
- name: tavily
command: npm
args: ["exec", "--yes", "--", "tavily-mcp@latest"]
env: { TAVILY_API_KEY: "${TAVILY_API_KEY}" }
# stdio, per-role (Jira + Confluence share one mcp-atlassian)
- name: atlassian
shared: false
command: uvx
args: ["mcp-atlassian"]
env: { JIRA_URL: "https://mycompany.atlassian.net" }
# stdio, per-role Slack
- name: slack
shared: false
command: npm
args: ["exec", "--yes", "--", "slack-mcp-server@latest", "--transport", "stdio"]
tool_prefix: "slack_"
# http, per-role remote GitHub MCP (token supplied per agent)
- name: github
transport: http
shared: false
url: "https://api.githubcopilot.com/mcp/"
roles:
- name: Senior Engineer
integrations:
slack: { bot_token: "${ALICE_SLACK_BOT}", signing_secret: "${ALICE_SLACK_SIGNING}" }
mcp_env:
atlassian: { JIRA_API_TOKEN: "${ALICE_JIRA_TOKEN}" }
slack: { SLACK_MCP_XOXB_TOKEN: "${ALICE_SLACK_BOT}" }
github: { Authorization: "Bearer ${ALICE_GH_TOKEN}" }

Environment variables and HTTP header values support ${VAR} references that resolve from the process environment at startup — both whole-value ("${TOKEN}") and embedded ("Bearer ${TOKEN}") — so secrets stay out of config files.

The engine derives some behaviour from a tool’s capabilities — for example, the worker guard denies tools that write to a shared surface, classified from the MCP readOnlyHint / destructiveHint / openWorldHint annotations the server advertises. Most servers advertise these; for one that doesn’t, supply them per server via tool_annotations (keyed by bare tool name):

mcp_servers:
- name: linear
command: uvx
args: ["mcp-linear"]
tool_annotations:
linear_create_comment: { read_only: false, open_world: true }
linear_get_issue: { read_only: true }

Keys accept snake_case (read_only) or the MCP camelCase (readOnlyHint); overrides win over whatever the server advertised. Because every tool server is now an mcp_servers entry — including atlassian, slack, and github — tool_annotations is declared there for all of them. This is the only place tool names appear in config for behaviour purposes — the engine itself never hardcodes them. See Tool Capabilities for the full mechanism.


From the LLM’s perspective, builtin tools and MCP tools are identical — both appear as JSON schema tool definitions. The model does not know whether a tool is a function inside the engine or an MCP server talking to a tracker.

Tool resolution order:

  1. Per-role MCP tools — checked first (role-specific credentials)
  2. Global tools — builtin tools + global MCP tools

Tool output is returned to the LLM in full. Control characters are stripped, secrets redacted and binary rejected, but results are never length-truncated: a truncated result silently hides content the agent reasons over. The same principle applies across the engine: the turn’s own trigger text, the turn-start prefetch blocks (personal memory, similar prior work — but for a recalled turn’s ask and its account of what it did, each past 600 bytes where the bullets are shown or past about 11 KB where the episode summary reads them, below — synthesized skills, counterparty profiles), the tool catalogue and every list_mcp_server_tools listing, the draft handed to Review, the agent diary and the conversation ledger as stored, coalesced notification digests, and the knowledge-base search query all carry their full text.

Where a bound is genuinely unavoidable, the engine never just cuts the text. A cut keeps an opening, and a reader — a model or a person — takes an opening for the whole: a message’s greeting for the message, a report’s summary for the report. So text somebody acts on is carried whole where it is bounded upstream, and where it is not, one of four shapes applies, chosen deliberately:

ShapeWhereWhy
Condensethe prior-work ledger’s payloads past their budget (an argument value, a failed call’s error, a round’s produced text) and a conversation’s older turns; a long chat thread’s middle at turn start; the round-cap judge’s view of the task and the last text; a delegated worker’s answer past its share of the next task; a source document answer_knowledge reads; a turn card’s one-line account; a recalled turn’s ask and its account of what it did, each past 600 bytes (## Similar prior work when its bullets are shown rather than summarised, and query_episodes) or past about 11 KB — a sixth of one auxiliary call’s input, so three turns’ asks and accounts fit one call — before the episode summary reads them (learning.summarize_episodes on), every rewrite of one answer held to one thirty-second deadline; an episode’s long task or outcome before the monthly compaction; a coding run’s report, failure or question past what its record carries (a question none of whose lines fits is not asked at all)The text has to fit, and what fits must still be what the text said. It is rewritten by the seat’s own auxiliary model (llm_auxiliary, falling back to llm) through internal/compact, charged to the seat like any auxiliary call, labelled as condensed wherever it is shown, and cached per text. Where no model can condense it, the fallback is never a cut: whole entries are left out and counted, a size and a digest stand in its place, or it is carried whole
Refusereflect_and_persist and mark_onboarded notes, refine_skill bodies and reasons, a secrets value, a Slack app manifest name, a counterparty trait, a search or answer_knowledge query past 400 bytes, a purge reason too long for its one-line record, a coding run’s report, question or result line read back from its box past 32 MiB (its event and error streams are read as streams and from their end instead, and the refused piece is described by its size); in the dashboard, a title, a view’s name, a question or a save note past the engine’s cap, said under the field before the pressThe value is stored or sent on, so half of it is a lasting half-fact. The caller can shorten it and retry, and the refusal names the field and the limit
Describe, don’t quotea task’s description and a long custom-field value on a history row (N bytes), free text past 600 bytes on a row, a set on a row (what joined and left rather than both sides), a vendor’s refusal body (an HTML page’s title rather than its markup)A history row is a log line kept forever on every node; it says what moved and by how much, and the change record beside it holds the text
Say it, in whole unitsa coding run’s transcript (whole lines from its start and its end, 64 KiB and 192 KiB, its middle counted — the plan opens it and a crash explains itself at the bottom), the end of its error stream (2 MiB, its unread start said by size) and its delivered refs (16 KiB of whole refs, the rest counted); an MCP server’s stderr tail and a control command’s output (whole lines, counted); error text on an event (events.ClipDiagnostic, 64 KiB, head kept); a vendor refusal past 2 KiB (marked); a knowledge snippet or an inbox excerpt (a preview of something one click away, marked)The full text genuinely cannot travel — an unbounded subprocess, or an event the queue would REFUSE, which costs the operator the whole record rather than its tail — or it is a preview of something the reader can open; so it carries a marker, and where a reader can get the rest, says how

The embeddings input — the text a provider turns into a vector — is the one bound on what a machine is given rather than what a reader sees. The bound is the embedding model’s own, carried per model with the vendor’s documented limits (Configuration § Providers), and the provider refuses an input past it before sending anything. What a longer text becomes is each caller’s choice: the knowledge corpus embeds a document’s opening, because the opening is what a search for the document is about, while a turn’s ask, an episode and a diary note are split between words into pieces whose vectors are pooled into one, so all of the text is represented. Either way only the vector’s input is bounded; the stored and displayed text stays complete.

Two rules run through the table. Bound the render, never the record — the stored row is usually the only copy, so a cut at write time is not a shortened rendering. And drop whole units where you must: an entry, a line, a member, never the middle of a sentence.

A cut with none of these properties is a bug, not a budget: a “character budget” on a prompt block is not a ceiling, it is a silent decision about which of the agent’s own memories it is allowed to see.

See Agent Runtime for the full execution loop.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.