Agent Learning
The agent-learning subsystem turns finished turns into durable, retrievable lessons — so the same agent (and its team, and the org) gets better over time without retraining the underlying model.
This page describes the shipped architecture: what runs in-engine, where each piece slots into the Turn Engine and Knowledge System, and the deliberate non-goals.
Provider-agnostic by design. Learning lives in the org/data layer, not in a model checkpoint. Any llm.Provider can back any role. Model fine-tuning is never required and is not part of the in-engine learning loop.
Why tools alone are not enough
Section titled “Why tools alone are not enough”A common misconception is that adding memory/skill tools is sufficient to make an agent “learn.” It is not. For a learning loop to be effective — i.e. for the LLM to reliably produce good lessons and invoke the right tool at the right moment — four layers must line up:
| Layer | What it does | Where Crewlet carries it |
|---|---|---|
| 1. Model training | Weights that know the memory/reflection protocol | Not required. Layers 2–4 do the work; any stock Claude/GPT works. |
| 2. Per-phase contract | Short, per-phase system-prompt rules that remind the LLM when to persist, reflect, recall | The executor and review prompt builders in internal/agent/prompts: guidance blocks injected only when the matching tool is registered. See Prompt scaffolding. |
| 3. Tool descriptions | One-line when to use text on each tool — Crewlet pushes guardrails into descriptions, not prompts | Builtins (query_episodes, reflect_and_persist, refresh_memory, refine_skill, use_skill, mark_onboarded) have precise one-line descriptions. |
| 4. Deterministic harness | Post-turn code that runs reflection regardless of whether the LLM “remembers” to | the reflect engine — the load-bearing piece. LLM cooperation is a bonus, not a dependency. |
Crewlet puts the weight on layers 2–4. Layer 1 is desirable but optional — effectiveness is not gated on any one vendor’s checkpoint.
Subsystems
Section titled “Subsystems”Small components, each with a single responsibility, plus the orchestrator that wires them.
Two ways in, and they are not the same shape. Everything learned from a turn arrives on one event through one dispatcher; the two passes driven by a clock rather than a turn are their own loops, and fleet singletons.
1. PersistDecider (post-turn personal memory)
Section titled “1. PersistDecider (post-turn personal memory)”Replaces “hope the LLM remembers to capture a durable fact” with a deterministic post-Review decision.
- Trigger: after the turn settles
done, orfailed(a reviewer’sfailed, or one the engine set: a fired guard, a spent budget, an exhausted provider chain).self_iterateis a mid-state; reinforcing it would teach the agent from incomplete work. - Decision: small auxiliary-model prompt answering what, if anything, should persist? Defaults to NOOP. The classifier picks a tier:
LONG— durable preference / fact (no TTL).SHORT— situational, with a TTL in days (sprint focus, vacation, delegation context).DOC— would be team-relevant; the decider does not write personal memory but logs the recommendation. Real cross-agent propagation goes through the team knowledge base, not the diary.NOOP— nothing worth persisting.
- Writing-style rule: persisted entries are declarative facts, not instructions.
"User prefers concise responses"✓ —"Always respond concisely"✗. Instructions drift out of date and get re-discovered as contradictions; facts compose cleanly. (Adopted verbatim from Hermes’s memory guidance.) - Effect: writes a row to the agent’s
agent_diaryvialearning.Diary.Write— agent-scope only.
2. Diary and reflect_and_persist (in-flight personal memory)
Section titled “2. Diary and reflect_and_persist (in-flight personal memory)”The agent’s private observation log. Two kinds:
| Kind | TTL | Use |
|---|---|---|
diary_long | None | Durable preferences and facts (Stakeholder X prefers digests). No deadline — a durable fact does not stop being true — but capped per seat at 500 entries, because recall scans and cosines every one of them at every turn start. Past the cap the sweep evicts by use, not age: least-retrieved first, then least-recently-retrieved, then oldest — so a fact the seat actually reaches for survives being old, and what goes is what has never once been recalled. A use is an entry the relevance filter selected, not one it merely considered: the candidate pool is similarity ∪ recency, so counting candidates would move every entry’s counter on every turn and the ordering would mean nothing. Both read paths — the ## Personal memory block and refresh_memory’s hinted re-filter — run through that one filter, so both count. The write is detached from the turn’s context, since a recall the seat already benefited from must still be recorded when the turn is finishing. |
diary_short | Set | Situational state (Sprint freeze runs through 2026-05-10, Opened PR-123 from sandbox run, awaiting review). Excluded by the read’s own SQL predicate once the TTL passes — so an expired row never consumes a slot in the recency window — and physically deleted by the retention sweep, which runs on the same fleet-wide singleton tick as the other short-horizon tables. |
Two writers converge: the post-turn PersistDecider (above) and the in-flight reflect_and_persist LLM-facing builtin. Both go through learning.Diary.Write, which embeds the content in the same INSERT as the row — whole, and tagged with the model it came from — so the row is reachable by vector similarity later and the memory changelog carries the vector with it. The embed has a two-second budget (learning.DiaryEmbedBudget, a note being one short input and reflect_and_persist having a model waiting on it); a note whose vector cannot be had in time, or at all, or comes back with no direction (all zeros, or not finite — a vector no recall could ever return, which stored would read as one the note has), is stored without one, never refused; an episode’s is too. The holder fills what is missing: once a minute the node holding a seat embeds, in batches, its live notes with no vector of the current model and width — any whose embed failed, and every note of a model or width the company has since left. A minute’s fill is one pass over every seat the node holds, held to what the provider answers: a rate limit, a timeout, a refused key or a model the endpoint does not serve ends the minute for every seat rather than for one; a refused batch is split in halves until the note it refuses is alone, and that note is held back for an hour, then offered again alone; a note the provider accepts and answers with a vector that has no direction — all zeros, or not finite — costs only itself: it is held back for an hour too (memory_fill_vector_unusable names it with its size), costing no request, and then offered again with its neighbours — stored, such a vector reads as a note that has one, and forgotten, the note was sent again every minute, for ever; and a refusal that is the minute’s first answer is judged by one plain word sent alone: refused too, the refusal is the configuration’s (a dimensions the model does not take, say) — memory_fill_configuration_refused is logged and nothing is sent for an hour, or until an apply changes providers.embeddings’ model, width, limits or endpoint, and the first refusal after the hour is judged again — and accepted, it was the notes’, which are isolated as above. What the fill holds back, notes and pause alike, is about the provider as configured — the knowledge corpus’s rule, on the same four — so an apply that changes anything else, or re-activates an unchanged revision to rotate the key, keeps it, and one that changes any of the four tries every held note again at once. The word judges it rather than the notes because the notes a fill meets are the ones whose embed already failed: a rule that read two notes refused alone as the configuration’s stopped a node’s fill for an hour over two poison notes after any restart. What a minute may send is bounded by what it sends, refused requests included: sixteen requests a node, and a million bytes of text across the whole company — shared equally among the nodes that run seats, because a provider limits an account rather than a node — which is the million tokens a minute of OpenAI’s lowest paid tier for these models, about 500 notes at the 2,000-byte note bound. A minute that runs out starts the next at the seat it ran out at. Setting a note’s vector stamps it with a fresh change sequence, so the changelog carries it like a new note, and a peer already holding the vectorless copy takes the filled vector over it. The ## Personal memory prefetch and refresh_memory read the diary via hybrid candidate selection: learning.Diary.Recall (vector top-K matches to the trigger) unioned with learning.Diary.Recent (recency top-K), deduped by row id — the two halves are 50 each, so the union is the bound and there is no separate cap over it — then passed to an aux-LLM relevance filter that picks the final digest. The two halves serve different needs: vector search catches topical / semantic matches to the current trigger; recency catches broadly-applicable operational rules that may not be a topical match (e.g. “use semantic commit messages on every PR”). The aux filter judges from the merged pool.
Write-boundary hygiene. learning.Diary.Write has no duplicate guard of its own: what keeps the diary from filling with paraphrases is upstream of the store — the reflect dispatcher’s redelivery guard on the work key, and the PersistDecider’s prompt, which is shown the seat’s recent notes and asked not to repeat one — and a byte-exact guard at the store would catch only the copy those already catch, and none of the paraphrases. Content is stored verbatim — never length-truncated, so the agent reads back exactly what was written. A note past learning.MaxContentChars (2 000) is refused, never trimmed, by both writers: reflect_and_persist refuses with the limit named, so the model can tighten the text and retry, and the post-turn PersistDecider skips the row and logs it, because there is nobody there to ask. One store, one rule — the two used to disagree, the tool refusing while the decider stored whatever the classifier produced. Nothing is sliced for the embeddings provider either: a note longer than the model’s input is split between words and its pieces’ vectors pooled into one. The post-turn PersistDecider is additionally skipped when the turn already self-persisted in-flight (the executor called reflect_and_persist), so the two writers don’t double-write the same fact. Prompt-injection scanning at this boundary is a separate concern, deliberately not bundled into the hygiene pass.
The diary is read by:
- The
## Personal memoryprefetch block (see Personal memory prefetch + refresh). - The mid-turn
refresh_memorybuiltin, which re-runs the same diary query with an enriched context hint.
3. Profiler (entity modeling)
Section titled “3. Profiler (entity modeling)”Crewlet’s multi-party equivalent of Hermes’s “model of who you are.”
- Input: observed interactions per counterparty (colleague, stakeholder, external human) — the inbound messages a turn’s notification trigger carries, from a chat surface, the tracker or a code host. An agent-to-agent ask carries none (it is an internal trigger, like a schedule’s fire or a sandbox completion), so the colleague who asked is not profiled from it. A coalesced trigger runs one observation pass per distinct sender (
Profiler.subjectsOfgroups a sender’s messages in order first), so a thread where one human sent four messages is one counterparty and a multi-human thread is genuinely several. A seat never profiles itself. - Output: one
learning.Profilerow per(observer_handle, subject_handle | subject_external_id, subject_platform): preferred communication style, past decisions, sensitivities, topics of interest. Stored in thecounterparty_profilestable (not the diary; not Confluence). - Scope: per-observer always — a fact one agent learns about Bob is private to that agent. Cross-agent propagation goes through humans + the team knowledge base, not auto-merging.
- Retrieval: the turn-start prefetch injects the trigger counterparties’ profiles into the executor’s prompt when the trigger has identifiable senders (one block per distinct sender with a stored profile).
lookup_colleagueresolves who a colleague is and does not return a profile.
4. Episodes and query_episodes (search own past)
Section titled “4. Episodes and query_episodes (search own past)”Agents can search their own prior turns.
- Source: the
episodestable in the node’s own store, replicated onto the memory changelog so it follows the seat across nodes — one row per completed turn (agent_handle,task_summary,ask,plan_summary,tool_sequence,skills_used,review_outcome,started_at,ended_at,duration_ms,embeddingand theembedding_modelit came from), filed against the work item the turn was charged to as its backend-qualified ref (native:<id>,jira:<id>— see Which work a turn is on); a turn on nothing, and every compacted row, is filed against none. The column iswork_itemsince node migration0033renamed thetask_idit was declared as, which never held anything: the turn record it was copied from declared a task id and never set one. - Vector: each turn is embedded as ONE text — the label of the event that woke it (
task_summary, the line a feed shows: “Message from Ana: Slack message”), what it was asked (each interaction’s body; for a wake with no interactions — a colleague’s question, a schedule’s task, the resumed half of a coding turn — theaskfieldturn_completedcarries instead, parked with the conversation across a suspension), and what it did (plan_summary: the last review’s account of what had landed, or the turn’s final answer) — and embedded whole: a text past the model’s input is split between words and its pieces’ vectors pooled. The label alone, which is what an episode was embedded from before, is the same line for every message on a surface. A scheduled turn’s label names its schedule (swe was assigned scheduled work weekly-report) and never the fire’s run id, which differs on every fire: led by the id, two fires of one schedule were unalike to every similarity search, and every recall of one showed the seat a key that names no item. The ask is stored on the row (ask, node migration0042), whole, so the vector is a function of the row alone and the node holding the seat fills it like a diary note’s — for a row written while the embedder was unreachable, and for every raw row after the company moves to another model or width — under the same once-a-minute pass and the same budget (above). A row whose text is longer than a minute’s pass sends — an ask is bounded only by the event that carried it — is sent a request at a time rather than in one call: the pieces each request embeds are kept on the node,memory_fill_input_continuednames the row and how many of its pieces are in, and the next minute sends the rest, so a long row takes as many minutes as its length needs and never more of the account’s minute than any other row; the vector is still the pool of the whole text. A piece the provider refuses, or answers with no direction, holds the row back for the hour with the pieces before it kept, so its retry resumes at that piece rather than at the first and the rows and seats behind it are filled meanwhile — forgotten, a row whose prefix cost a minute took every minute; and a failure that is not about the row — a rate limit, a timeout, a server error — ends that minute wherever it falls, and the next minute sends that request again. Until a row is filled it is found by recency and conversation only. A turn that was told nothing stores no ask, and is filled from its label and what it did. What storing the ask costs is not bounded by a count of rows: the lifecycle’s threshold only makes a pass due, and a pass keeps every settled turn of the last thirty days raw, the two exemplars of every cluster it folds, and — with no horizon at all — every settled turn that called tools and never joined a cluster of three; a compacted summary keeps no ask. That is about thirty days of a seat’s asks on each node holding a copy (1.5 MB at fifty turns a day of a kilobyte each) plus what never folds; and on the memory changelog, which keeps one message per row and carries no deletes, every ask the seat was ever asked — about 18 MB a year at the same rate, beside a vector that is already 6 KiB a row at 1 536 dimensions. - Builtin:
query_episodes(query?, conversation?, outcome_filter?, limit?): similarity to each past turn’s text (above) whenqueryis given, the seat’s own turns in one conversation whenconversationis, recency otherwise; scoped to the calling agent’s handle, available to the executor. Each turn is answered with what woke it — the waking event’s label, markedwoken by:, because written unmarked after the date a model read “Message from Ana: Slack message” as the question it had been asked — what it was asked (asked:, the ask the row stores; a turn that was told nothing has none and shows none), how it ended, and what it did (what it did:, its account,plan_summary). The ask and the account are each on one line, whole up to 600 bytes (learning.EpisodeAccountBytes, the three or four sentences a review’s account runs to) and past that condensed by the seat’s auxiliary model and marked as a rewrite, or named by its size where no rewrite can be had in time: every rewrite one answer makes, four at a time, shares one thirty-second deadline (learning.EpisodeRewriteTimeout), so an answer of any number of turns waits at most thirty seconds for them, where a deadline per rewrite was one per wave of four, never cut. A compacted row is answered as its pattern: the dates it spans, how many turns it stands for, how many of them endeddone(counted from the members), and what varied. Aqueryanswer is headed as most similar first, and says how many of the seat’s turns its search could not reach — rows with no vector of the current model yet, which the holder is filling — namingconversationand recency as the way to read them; an empty answer is “this is new work” only when there were none. Aquerythat cannot run says what stopped it, because each cause sends the seat somewhere different and only one is worth asking again: a company with no embeddings (stop asking by meaning); an embedder that refuses its configuration — a rejected key, a missing model or endpoint, a vector of a width the store was not sized for (embeddings.ErrConfiguration) — which no query will get past until an operator fixesproviders.embeddings, so the seat is told to stop asking by meaning until then; an embedder that refuses this query (embeddings.ErrRefused), which only a different query may get past; an embedder that could not answer now — a 429, a 5xx, a network failure, a timeout — which may answer if called again; anything else, such as an episode store that could not be read, said as what it was with no promise either way; and a registry with no search wired. Each namesconversationand recency as the paths that still work. - Auxiliary summarization: the turn-start block’s hits are passed through the role’s
llm_auxiliarymodel (a cheap one), which writes one short briefing of them in their place, keeping the system prompt small (learning.summarize_episodes, capped bysummarize_max_tokens). Falls back to the raw entries when it is off or no aux model answers.query_episodesis never summarised: a seat that asked for its turns gets them. - Frozen-at-turn-start: the
## Similar prior workprefetch resolves once per turn and bakes the summary into the system prompt. Aself_iterateround (review, then the executor again) reuses the same prefix so the LLM provider’s prompt cache keeps working.
5. Synthesizer (skill induction)
Section titled “5. Synthesizer (skill induction)”Mines recurring successful trajectories and drafts a new procedural skill.
- Single-turn induction runs inline in the reflect engine: a settled turn (
doneorfailed) that used ≥min_tool_callstools is offered to the auxiliary model, which drafts a skill or declines. Declining is the ordinary answer — most turns are not procedures — so a turn with no reusable shape costs one cheap call and writes nothing. Every post-turn worker’s prompt describes the turn under the names of what its fields are — woken by (the label of the waking event, which says what kind of event it was and nothing of what it said), asked (what the turn was asked; the PersistDecider renders each interaction with its sender instead), and what it did (the review’s account of what landed, or the final answer) — never as a “task” and a “plan”, which a model reads as the whole of what was asked and what was intended. - Three gates, all before the model call, because a draft made only to be discarded is money spent for nothing: the per-seat cap (
max_skills_per_agent), and a duplicate check comparing the turn’s tool set against the seat’s existing skills by Jaccard similarity (duplicate_jaccard_threshold). The comparison is over the tool set, not the ordered run — two turns calling the same four tools in a different order are the same procedure, and treating order as identity is how a seat ends up with a skill per permutation. A draft the model returns without a name, summary or body is dropped rather than written with the gap. - Clustered synthesis (
scheduler_enabled,cluster_*) is the other half, and it catches what single-turn induction cannot: the shape a seat arrives at over a fortnight — three tools, unremarkable on any one turn, run the same way eleven times. Repetition is evidence a single turn cannot offer, and it is invisible from inside any one of them. A daily singleton pass reads each seat’s lastepisode_fetch_limit(default 200) turns, greedy-clusters them by tool-sequence Jaccard atcluster_jaccard_threshold(default 0.6), and drafts from the largest cluster of size ≥cluster_min_size(default 3). It is off by default —scheduler_enabled: false— because the pass costs an auxiliary call per seat per day and a young company has nothing to cluster yet.- One draft per seat per pass, largest cluster first. Not every qualifying cluster: each draft is a completion, and a seat with three real patterns learns them over three days with the strongest evidence going first. A pass that drafted everything could also fill the per-seat cap in a single tick.
- A cluster the seat has already learned is skipped, not a stop. The next pattern down may be one it has not — the same
duplicate_jaccard_thresholdthe inline path uses, which is stricter than the pooling threshold on purpose: pooling asks “is this the same kind of work”, rejecting a draft asks “is this the same skill”. - Only raw, settled turns with ≥
min_tool_callstools are evidence. A compacted row is already a summary of a cluster and would count a fold as one turn; aself_iterateround is work the agent judged incomplete. - The stored
tool_sequenceis a run that actually happened — the cluster’s representative — rather than a union of its members, because the duplicate check compares stored sequences against new turns and a union nobody performed matches everything loosely. - The
skill_synthesizedevent carriestrigger: clusteredand thecluster_size, and noturn_id: the draft came from a group, and naming any single member would put a trace on the event that explains none of the others.
- Output: a row in
synthesized_skillskeyed by(agent_handle, name)— agent-scope only. The body is stored in the familiar SKILL.md Markdown shape, whichuse_skillreturns verbatim. - Cross-agent promotion is the third path, and its output is deliberately not a skill row. Every other skill here is agent-scope (one seat’s row, in one seat’s prompt) because a skill is a procedure a particular seat follows. A procedure four seats independently arrived at is something the team has, which makes it documentation. So when ≥
min_sibling_countdistinct seats in one unit converge on a similar tool run, a daily singleton pass distils the cluster into a draft page in the team knowledge base under the unit’sAuto-Drafted Skillsparent, for a lead to review.- Distinct seats, not skills. One seat that drafted four near-identical skills is a catalogue that needs curating, not a team convergence — counting rows rather than owners would promote it and present one agent’s habit as the unit’s practice.
- Direct members only. A parent unit does not pool its children’s catalogues: it would find the convergence the child already promoted and draft it again one level up, on a page naming a team that never converged on anything.
- The draft is hidden until a person publishes it. The
## Relevant knowledgesearch excludes the auto-drafted subtree (and, as a fail-closed backstop where a backend has no parent chain, the[Auto-draft]title prefix), so an unvetted draft never reaches another agent. A lead adopts one by moving it out of that parent; once published it is an ordinary knowledge-base page reachable through the query-time search. Rejecting one is a delete — it is re-drafted only if the team converges again. - One backend, matched. The pass writes through a small
learning.PromotionWriterseam:confluence.PromotionWriterposts rendered storage-format XHTML under the unit’sspace, creating theAuto-Drafted Skillsparent if the space has none, and refusing the draft rather than filing it at the space root if that parent cannot be created, because a page outside the subtree is one every agent can read. The container is the unit’s wiki space, never its tracker project: a unit carries both identities, and filing a draft under the tracker’s key would create a page in whatever space happened to share the name, or fail against nothing at all. - Cross-tick dedup is the writer’s job, because the pass re-clusters the same persisted rows every tick and would otherwise yield one draft a day forever. Confluence keys on the title, which is unique within a space. One converging cluster yields one page, and a tick that finds the existing draft stays quiet rather than re-announcing the promotion.
- A unit with no container is soft-skipped with the field to set in the log; a company that configured knowledge for one team and not another is supported, and failing would stop the configured team’s promotions too. A write failure announces nothing, and the next tick retries.
- Success publishes
skill_promotedcarrying the unit, thecontainer_key, thepage_id/page_title, and bothsibling_countanddistinct_agents, because one agent repeating itself and five agents converging are different findings. - The engine carries no unit-scope skill rows of its own.
- Collision guard: the synthesizer rejects names that already exist in the agent’s own
synthesized_skillstable. There’s no global skill registry to guard against — synthesized skills are per-agent, and shared procedures live in the team knowledge base rather than in an engine-side registry.
6. Refiner and refine_skill (improve skills during use)
Section titled “6. Refiner and refine_skill (improve skills during use)”When a synthesized skill was central to a successful turn, append an observed-in-practice bullet; when it contributed to a failed turn, append a counter-example.
- Auto path: the reflect engine dispatches the refiner after every settled turn (
doneorfailed) whoseskills_usedis non-empty. Aself_iterateround is not refined — it is work the agent itself judged incomplete, so a lesson drawn from it is one the next round may contradict, and the turn will emit anotherturn_completedwhen it does settle. The auxiliary model picks one observation (or NOOP); successful turns produceObserved in practice: …, failures produceCounter-example: …. - One call, one bullet, at most one skill. The turn’s whole offered catalogue goes into a single prompt and the model chooses which skill — if any — learned something. Per-skill calls would cost a completion per skill per turn for answers that are almost always NOOP, and a turn rarely teaches two procedures something new at once. A name the model invents is dropped rather than matched onto the nearest candidate: a bullet appended to the wrong procedure is worse than no bullet.
- NOOP is the expected answer, and it is not an error. A model asked what a turn taught will produce something for any turn at all, and a skill that grows a bullet per turn stops being a procedure and becomes a diary of the turns that read it. The prompt says twice that answering nothing is correct.
- Only skills that still exist. The turn’s
skills_usedis a list of ids captured when its prompt was built; the refiner intersects it with the live catalogue, so a skill the curator archived in the meantime is not resurrected. A turn whose skills have all been archived costs no model call at all. - Bullets collect under one
## What practice addedheading at the end of the body, rather than scattering a new section through the steps on every refinement — a reader sees the procedure first and what practice added to it second. - Manual path: the LLM-facing
refine_skillbuiltin lets the executor correct its own skills mid-turn. It takesskill_name, the full correctedcontentand an optionalreason; the new text replaces the body in its entirety. A whole-body replacement rather than a patch, because a model asked for a diff produces something diff-shaped that does not apply, and a half-applied edit leaves a procedure that is neither the old one nor the new — with nothing to compare against, since the prior body is already archived by then. Patch-on-encounter: a seat that finds a skill outdated corrects it immediately rather than waiting for a separate consolidation pass. A seat may only refine its own skills. - Versioning: every refinement archives the prior state to
synthesized_skill_versionsand bumps the live row’sversion. Rollback is a forward-step operation (the archived body becomes the body of a new version), so rewinding then un-rewinding works without losing history. History is bounded bymax_versions_kept(default 10) per skill. - Body cap:
max_body_chars(default 20 000) — a refinement that would breach the cap is refused, never truncated, so a runaway loop can’t blow up a skill body. A clip lands mid-step and the model reads the remainder as the whole procedure. The auto path skips silently and logs it; the manual tool refuses with the field name, because there a model can tighten the text and retry. enabledgates both halves.learning.skill_refinement.enabled: falsewithdraws therefine_skilltool and leaves the post-turn refiner unwired — they write the same rows through the same version archive, so a company that turned refinement off and still had the tool would watch its skills change under a knob it had set to false.use_skillis unaffected: reading a skill is not changing one. Setting bothauto_refine_on_successandauto_refine_on_failureto false is refinement-off spelled the long way, and the engine leaves the worker unbuilt rather than skipping every turn.
7. Reflector (the orchestrator)
Section titled “7. Reflector (the orchestrator)”The deterministic harness. Owns when reflection runs and coordinates the workers above.
- Hooks: reflection runs on the node holding the seat. When a turn closes, the node that ran it publishes
turn_completed— the record — and areflection_duewake carrying that same turn onto the seat’s own reflection subject (crewlet.agent.<handle>.reflect, groupagent-<handle>-reflect). The node that acquires a seat attaches that subject after hydrating the seat’s memory and lets it go on release, before the release’s last memory flush — the same way it treats the seat’s sandbox control subject — so a wake published while the seat is moving waits on the subject for whoever holds it next. A seat’s own subject rather than one fleet-wide group, because reflection writes the seat’s memory into the store of the node that runs it, and only the node holding the seat carries that memory to the next holder and answers reads of it: on a fleet of N nodes a fleet-wide group put all but one in N of a seat’s reflections where nothing of the seat was held, writing diary rows, episodes and skills that no holder ever read. A node withoutseatsholds no seat and so reflects on nothing, and a node withoutdatareflects on the seats it holds exactly as a data node does — memory is node-local on both. - What a move costs. A release lets the subject go without waiting for a pass already running (the fenced release must not wait on an auxiliary model), so a pass that finishes after the flush writes rows this node no longer carries — the same bounded loss a crash costs between flushes. A wake still waiting when a seat is removed from the company is deleted with the seat’s other subscriptions when its mailbox retires.
- One wake, not one per writer. Everything a seat learns is learned from one turn, and every writer is gated on the same questions about it. Three subscriptions would mean three redelivery windows over one turn, three places to discover a company that learns nothing, and three chances for one of them to be quietly unwired.
- One dispatcher per process, not per config revision. A config apply swaps the org and the worker set behind it; the seats’ subscriptions and the redelivery guard stay put. Rebuilding it per revision would empty that guard, so a redelivery landing either side of an apply would be classified twice: two auxiliary calls, two differently-worded rows for one fact. A refused worker set leaves the previous one serving: reflecting against a stale org is a far smaller wrong than not reflecting at all. A node that boots with no company has no org to build the dispatcher over, so its first apply builds it; a wake that reaches a node before then is acknowledged and logged, and none can carry a turn, since no seat has run one.
- Gates, in order. Each is reported by name, per worker, per turn, because “this company never learns anything” needs an answer that says which gate closed:
- No workers — nothing is wired, so there is no question to ask about this turn.
- No role — the turn came from a seat this revision no longer has. Learning about a renamed or removed role writes memory under an identity nothing can read back.
- Per-role opt-out —
learning_enabled: falseopts a noisy or sensitive role out without disabling the subsystem globally. Unset inherits the company-wide setting. - No budget — the company or the seat is already at its
token_budgetceiling, so the pass declines to start. Reflection is best effort and this is the skip that costs nothing: a pass that runs and then discovers it is over budget has already made its auxiliary calls. It sits before the duplicate ring on purpose — an exhausted budget is the one transient refusal in this list, so a turn skipped for it stays reflectable when the ceiling moves, while every other gate would still hold on a redelivery. An unreachable counter reflects anyway: unknown is not “no”, and a coordination blip must not silently stop a company learning. - Duplicate — a bounded ring of recently-processed RUN ids (not work keys: a redelivered trigger genuinely runs again, and its reflection is a second pass over a second execution — what collapses THAT is the episode row’s own work-key index, not this ring). Reflection is not idempotent: each pass is a fresh auxiliary call that can write a second, differently-worded row for the same fact. The ring is per-process and deliberately not durable — a second node reflecting the same turn writes a second diary row, which is the bounded duplication the engine promises rather than exactly-once. It evicts rather than growing: past the bound, a redelivery is far outside any backend’s redelivery window.
- No engagement — the turn was skipped (
review_outcome: skipped), or finisheddonehaving called no tool at all. Either way the agent processed nothing externally observable, and a fact read off the trigger would teach it a directive it never received. - Per-worker skip — each worker states its own applicability, because they genuinely differ: the persist decider must not run on an unsettled turn, the counterparty profiler must.
- Failure mode: every dispatch is best-effort and each worker runs under its own panic recovery, so one worker’s bug costs neither the pass’s remaining workers nor its sentinel. A failed worker logs and the next one still runs; a failed reflect never fails the parent turn. The handler always acks: reflection is work about a turn that is already over, so a nak would redeliver it to spend another round of auxiliary tokens reaching the same conclusion.
What the turn event has to carry
Section titled “What the turn event has to carry”The dispatcher reads a wake, and the node that runs it is the seat’s holder when the wake is taken — which, after a seat moved, is not the node that ran the turn and never saw its trigger. Everything the gates read therefore rides on the turn_completed payload itself — the tool sequences, the outcome and the opt-out decision, the skills the prompt offered, the inbound interactions with their senders resolved, and, for a wake with no interactions, what the turn was asked (ask). A field left off that payload is a fact no worker can consult, and the gates fail open-looking: an absent tool sequence reads as “the agent engaged with nothing”, which silently skips every worker on exactly the successful turns worth learning from, while the dispatcher reports a clean pass.
Prompt scaffolding
Section titled “Prompt scaffolding”Short, conditional guidance fragments are appended to the executor’s system prompt, injected only when the matching tool is registered for the role. This scaffolding is sourced from the Tool Skills registry — knowledge-base pages (Confluence) operators can edit at runtime — rather than being hardcoded in engine prose. The bundled examples/tool-skills/ files ship ready-made versions:
| Bundled skill | Trigger | What it teaches |
|---|---|---|
examples/tool-skills/reflect-and-persist.md | tool: reflect_and_persist | Persist declarative facts, not instructions to yourself. |
examples/tool-skills/refine-skill.md | tool: refine_skill | Patch a loaded skill when it goes stale; don’t wait to be asked. |
examples/tool-skills/retrieval-research.md | any_of of query_episodes / the atlassian MCP server / refresh_memory | The consolidated retrieval re-search rule — see below. |
examples/tool-skills/observed-directives.md | tool: mattermost_post_message | Share team-relevant directives via the agent’s broadcast surface. |
examples/tool-skills/getting-unstuck.md | any_of of colleague-surface tools (mattermost_post_message, the atlassian MCP server, a2a_ask) | Manager-handoff conventions: when stuck, mention manager on the surface where the problem lives. |
examples/tool-skills/channel-discovery.md | any_of of the Mattermost channel, post and user-search tools | How to choose the right Mattermost channel or surface, and how to fall back when membership is missing. |
The retrieval-research skill carries the consolidated retrieval re-search block. The three relevance prefetches — ## Similar prior work, ## Relevant knowledge, ## Personal memory — are all derived from the triggering message as it stood at turn start, before any recon, so they share one rule: after recon has given the seat a richer query, re-query the corresponding tool — even when the initial block already had entries. Rather than repeat that rule in three near-identical blocks, the shared preamble states it once and one terse per-tool line (query_episodes / the knowledge backend’s page-search tools / refresh_memory) is appended for each re-query tool the role actually has. On a thin trigger the turn-start message genuinely is a bare pointer (the thin-trigger gate skips the prefetch entirely); on a substantive trigger it is the whole message but still pre-recon. Either way the guidance makes the assumption legible to the LLM so the re-query pattern does not rest on the model guessing.
Plus four always-on prefetch blocks (rendered when the data exists):
## Similar prior work— top 3 episode-search hits, each as what woke it (Woken by:), its outcome and tools, what it was asked (Asked:) and what it did (What it did:), summarised by the aux model whenlearning.summarize_episodesis on. The summary reads each ask and account whole — up to a sixth of one auxiliary call’s input, about 11 KB, so three turns’ asks and accounts fit one call, past which a text is condensed first — because its briefing replaces the bullets; the bullets condense an ask or an account past 600 bytes, asquery_episodesdoes, only where they are what the seat is shown: with the summary off, or when it does not answer. Both render through one function (learning.PastTurns), so at most four rewrites run at once and all of one render’s share one thirty-second deadline (learning.EpisodeRewriteTimeout), here and inquery_episodesalike, and a text no rewrite reached by then is named by its size. The block holds everything it asks a model — its rewrites and its summary — to one auxiliary deadline (prefetch.AuxTimeout, thirty seconds) past the recall, as when the summary was all it called: with the summary off it waits at most that for its rewrites, and with it on the bullets a summary that did not answer falls back to get what is left of the same deadline rather than thirty seconds more from the auxiliary chain that just did not answer. Its worst case is therefore the turn’s embed (two seconds) plus thirty seconds, the memory block’s own bound.## Personal memory— diary entries selected via hybrid vector ∪ recency candidate selection, then filtered for relevance to the current task / trigger by the aux model.## Synthesized skills you've learned— names + descriptions of the agent’s own synthesized skills, loadable viause_skill.## Relevant knowledge— knowledge-base pages from a live query-time search: the aux LLM generates a short search query from the trigger, and theknowledge.Searcherruns it scoped to the role’s accessible containers. See Relevant-knowledge prefetch below.
Plus the conditional prefetches:
## The thread so far— the chat thread this turn was woken in, read at turn start on the seat’s own chat credential and handed over rather than left for the agent to fetch. Rendered only for a chat thread reply; a top-level message has no earlier conversation. Bounded by whole messages (8000 bytes; past it the messages between the root and the newest are condensed by the seat’s auxiliary model rather than dropped, and left out with the count reported only where no rewrite can be had; the root and the newest of what was read always survive — a root with nothing readable in it, such as an alert app’s attachment-only post or a deleted opening, keeps its place as a line saying so rather than letting the oldest reply take it), senders resolved through the party registry, and the seat’s own replies marked. A thread that could not be read renders a different sentence from one that was read and was empty, and a thread too long for the backend to read to its end renders a third that says the newest messages are missing. It is not a learning prefetch and is not gated by the thin-trigger gate — it is the thing that makes a thin trigger thick. See Agent Runtime.## First-turn onboarding— rendered until the agent callsmark_onboarded. Lists the relevantOnboardingknowledge-base pages on the agent’s unit chain. Stored markers live inagent_onboarding_markers, keyed byagent_idand stamped with achain_hash; an org-chain change invalidates the marker so the hint re-fires for the new structure.
These blocks are layer 2 from the four-layer table above. The reflect engine (layer 4) runs regardless of whether the LLM follows them — the scaffolding is an optimization that lets well-behaved models cooperate, not a dependency.
Personal memory prefetch + refresh
Section titled “Personal memory prefetch + refresh”The ## Personal memory block runs once at turn start (prefetch.Fetcher.personalMemory): it assembles a candidate pool of the agent’s diary rows, filters them by relevance to the trigger via the aux model, and renders a digest into the system prompt, which is not re-queried mid-turn.
Hybrid candidate selection
Section titled “Hybrid candidate selection”Fetcher.memoryCandidates builds the candidate pool that feeds the aux filter. Given a trigger query, it returns the union of two top-K reads against the agent’s diary:
- Vector top-K —
learning.Diary.Recallranks by cosine distance against the diary’s embedding column, scoped to the agent’s id, to unexpired rows and to vectors of the query’s own model (embedding_model): two models of one width are two spaces, and a row from another is no match at all. The distance is computed by the database (vector_distance_cos), so only the rows above the relevance floor cross the driver boundary rather than every embedded row the seat owns. This catches topical / semantic matches to the trigger. - Recency top-K —
learning.Diary.Recentreads the most-recent unexpired rows, again scoped to the agent’s id. This catches broadly-applicable operational rules that may not be a topical match to this particular trigger — “use semantic commit messages on every PR,” “always tag the security channel before merging auth changes” — which the vector half would miss when the trigger is unrelated to the rule’s topic but the rule still applies.
The two sets are deduped by row id. Their sizes (50 each, memoryVectorLimit and memoryRecencyLimit) are what bounds the pool; a further cap over the union used to sit here and was unreachable arithmetic, since a dedup of two 50-row halves cannot exceed 100. The aux-LLM relevance filter then judges from this merged pool: the same filter as before, over a better-recall candidate pool. The hybrid is not pure vector (which would miss the broadly-applicable rules) and not pure recency (which falls off for long-lived agents with >100 LONG entries, where old-but-relevant rows would drop off the window and never reach the filter).
PersistDecider’s write-side dedup reads the diary’s recency list alone (Diary.Recent, up to 50 rows), which is the correct shape for the “is this paraphrase already in the diary?” check.
Failure modes
Section titled “Failure modes”The prefetch filters against the salient inbound message — the raw message, not the notification builder’s enriched task description (see Salient-body sourcing). That leaves two failure modes:
- Context-thin triggers (“yes”, “+1”, a thread reply with little semantic content) — the salient message itself is thin, so the filter has nothing to match on and the block ends up empty. When that happens and the agent has memory rows, the block renders the gate-path hint line nudging the seat to refresh after recon.
- Richer triggers can produce a non-empty block, but the entries the trigger-time filter chose may not be the most relevant once the executor has read the thread / fetched the ticket / queried knowledge and learned what the conversation is actually about.
The refresh_memory(context_hint=…) builtin fixes both. The refresh_memory line of the bundled retrieval-research Tool Skill (examples/tool-skills/retrieval-research.md) tells the executor to call refresh after any tool call that materially changed its understanding of the conversation — even when the initial block already had entries, not only as an escape hatch from the empty case. The tool re-runs the filter — and the similarity half of its pool — against the executor’s context_hint alone, which is its own account of what the task is about once recon has made it real, tells it who triggered the turn exactly as the turn-start filter was told (the filter’s per-subject rule has nothing to judge “party to the task” by otherwise) — on the resumed half of a coding turn too, which re-reads no trigger and so is told the senders its turn parked with beside the ask — and returns the freshly-rendered digest as the tool result. Bounded by:
- Per-turn cap —
learning.personal_memory.max_refreshes_per_turn(default 3). A hint beyond the cap is refused with the count spent, so the model learns the shape of the limit and stops trying instead of silently no-op’ing. A hint whose filter call failed still spends its slot — otherwise a failing call is retryable without bound, which is the same unbounded spend the cap exists to stop — but retrying that same hint is allowed and does re-run the filter. - Idempotency cache — a repeat of a hint already used this turn (case- and whitespace-normalised) is answered from the ledger without a fresh auxiliary call, including when the answer was “nothing bears on this”. A repeat is free because it is answered from here, not merely uncharged: re-running the filter for free would leave the cap bounding nothing, since a model alternating two hints could spend a completion per round forever. What is cached is the filtered rows, not the rendered text, so a repeat asking for a larger
limitgets the extra notes rather than the first call’s rendering. - Per-turn isolation — state keyed by the RUN id (
turn_id, which names one execution — see a turn’s two identities) and bounded to the most recent 256 turns, far more than a node runs at once. Per run rather than per unit of work, deliberately: a redelivered trigger runs again with a fresh context and must not inherit the spend or the answers of the attempt it is repeating; the bound is what makes it a cache rather than a leak, since nothing tells the tool when a turn ended. State from one turn never leaks into another. - Frozen-prefix-cache safe — refresh output lands as a tool-result message, not a system-prompt rewrite. The LLM provider’s prompt cache stays valid across iterations.
Relevant-knowledge prefetch
Section titled “Relevant-knowledge prefetch”The ## Relevant knowledge block surfaces team-published documents — playbooks, runbooks, ADRs, conventions, design docs, anything in the agent’s accessible knowledge-base containers — without forcing the seat to discover them by guessing names against use_skill or by remembering to call the knowledge-search tool first. It runs a live knowledge-base search once per turn through the knowledge.Searcher seam (Confluence CQL — one backend per org); the Personal memory prefetch is the closest sibling in spirit, though that one reads the private diary via hybrid vector ∪ recency candidate selection filtered by an aux-LLM relevance pass.
Why “knowledge” and not “skills”
Section titled “Why “knowledge” and not “skills””An alternative design would carve out a special “team-skill” label so operators could mark certain pages as procedures meant for agents. That reintroduces an operator-curated/synthesized skill split — two parallel skill surfaces to maintain — which the project deliberately avoids.
The shipped design takes the opposite stance: a knowledge-base page is a knowledge-base page. The search runs against the agent’s accessible containers and the backend’s own relevance ranking decides which pages come back; there is no engine-side “skill” label or parallel surface to maintain.
Source: query-time knowledge-base search
Section titled “Source: query-time knowledge-base search”For each turn:
- The searcher gate runs:
Searcher.CanSearch(seat, org), a cheap, no-I/O check that a search could return anything (the role has accessible containers, or its own backend credentials for an unscoped search). When it says no, the aux-LLM query-generation call is skipped entirely. - The role’s auxiliary model (
role.llm_auxiliary) turns what the turn was asked — its ask, not the integration’s wrapping, whose worked examples would otherwise become search terms — into a short plain-text keyword query (the user prompt endsKnowledge-base search query:). Scope is not the aux model’s job — the searcher derives it internally from the org-wideknowledge.*list via accessible containers. There is no per-unit/role union: a unit’sspaceis integration identity (webhook routing + write home), not read scope. - The searcher runs the query as the agent’s own backend user (the seat’s own Confluence credential from its
mcp_env, falling back to the org-level token) as a CQLtext ~ "..."clause narrowed byspace IN (...). The backend enforces page permissions natively, so restricted pages the agent cannot see never appear; unreviewed auto-drafts are excluded by the query’s default ancestor exclusion (knowledge.AutoDraftedParent, “Auto-Drafted Skills”).
Loading full bodies
Section titled “Loading full bodies”The bullets render title + snippet — enough for the executor to decide which pages to open. To pull a full body or run a fresh search, it calls search_knowledge or the backend’s own MCP tools — confluence_get_page / confluence_search. The block prose describes the capability, never a hardcoded tool name.
Hardening
Section titled “Hardening”- Show-nothing on a failed query. When query generation fails there is nothing to search with, and the block renders nothing rather than erroring the turn.
- A search that did not run says so. When the search never ran — the backend could not be reached, or no copy of the knowledge base answered it, so the outcome served no mode (
ServedModeempty) — the block renders the could-not-be-searched hint (UnsearchedKnowledgeHint), neverEmptyKnowledgeHint. “Nothing surfaced” is a claim about what the company has written down, and a search that never ran makes none: a seat told it concludes the page does not exist and writes a duplicate of one that does. Neither path errors the turn —Searchis best-effort by the seam’s contract — andsearch_knowledgeanswers a search that served no mode the same way. - The gate-path hints rendered when the block would otherwise go silently empty, each pointing the agent at
search_knowledgeas the mid-turn escape hatch.EmptyKnowledgeHintwhen the thin-trigger gate skipped the search or the search ran and returned nothing — it mirrorspersonal_memory’s hint — andUnsearchedKnowledgeHintbeside it when the search did not run at all, which says nothing about whether a page exists. - Frozen at turn start. The block is part of the system-prompt prefix, so a
self_iterateround reuses the same prefix and the LLM provider’s prompt cache stays valid. - Once per turn. The query is generated and the search runs once; anything more the agent needs it asks for.
The executor asks instead (thin triggers)
Section titled “The executor asks instead (thin triggers)”On a thin-trigger turn the turn-start ## Relevant knowledge prefetch is gated off — generating a search query from a bare pointer is noise. That leaves a gap: the agent does its recon inside the turn, and until it has, it has no relevant-knowledge block at all.
The search_knowledge builtin closes it. The gated block renders a hint saying to search again once the task’s real shape is known, and the executor calls the tool with a query it writes itself — over the same knowledge.Searcher seam, with the same auto-draft exclusion, authenticating as the same seat. The result comes back as an ordinary tool result, spliced into the conversation the agent is already in.
The three-phase engine had a push here instead: a second search the engine ran between the phases, keyed on the plan summary, because the actor could not ask for itself — it was a different conversation. With one loop there is nothing between the phases to hang a push on, and there no longer needs to be: the frame that just did the recon is the frame that searches.
Being a tool rather than a seam, it is also cheap to be honest about. A backend that is unreachable, unconfigured, or scoped to nothing answers with a sentence saying which of those it is, rather than an empty block the agent has to interpret.
Telemetry
Section titled “Telemetry”The turn-start prefetch_summary event’s relevant_knowledge_hit, relevant_knowledge_bytes and relevant_knowledge_selection_count are recorded alongside the other prefetch blocks. The selection count distinguishes the two paths where the hit is true: a non-zero count means real pages were rendered; zero with a true hit means a hint was rendered instead: the gate-path hint (the thin-trigger gate skipped the search, or the search ran and returned nothing), the could-not-be-searched hint (the search did not run), or the still-indexing one. Operators investigating low effectiveness pivot on this field to tell “no signal” from “hint nudge only.”
A block stuck at 0% hit rate over a representative window is almost always one of:
- No
knowledge.scopeconfigured and the agent has no per-agent backend credentials, so it can’t search unscoped (a credential-less / fallback-token agent with no containers searches nothing, andCanSearchgates the whole prefetch off). - No
integrations.confluenceconfigured, so no searcher is wired, or no pages in the read scope match. Also check the seat’s own page permissions: the search runs as that account, so a space it cannot read silently contributes nothing (see Confluence § Knowledge search). - Aux LLM unavailable (
llm_auxiliarynot configured and the role’s primaryllmdoesn’t resolve as an aux provider), so query generation cannot run.
Thin-trigger gate
Section titled “Thin-trigger gate”All three relevance-driven turn-start prefetches — ## Personal memory, ## Relevant knowledge, and ## Similar prior work (episode recall) — run an aux-LLM call against the bare trigger at turn start, before the agent has done any recon. For a self-contained trigger (a full task assignment, a detailed issue body) that’s high-value: the agent gets relevant memory / docs / episodes baked into the system prompt for free.
But for an event-driven turn the trigger is a pointer, not the context. A Jira webhook says “POC-518 got a comment”; a Slack thread reply says “+1”. The real context only exists after the agent fetches the issue / reads the thread. Running the aux filter against the bare pointer is near-guaranteed low-value — it has nothing substantive to match against — and we’d also spend prompt space rendering “nothing matched, go look later”. So on the common webhook turn we’d pay twice (a wasted aux call + prompt clutter) for a result the agent has to redo via tools anyway.
The gate skips the aux call when the trigger is a pointer. It is pure logic — no LLM call: the decision is read from notification metadata (issue_key / thread_ts / event_type), which is exactly why it’s cheap enough to gate on.
| Stage | Carries the signal |
|---|---|
| Notification builder | notify.Prompt.RequiresRecon: true when the builder emitted a “go fetch the real thing” directive. Jira and Confluence page events (## Get Full Context), GitHub review_requested (“read the diff”), Slack thread replies (read-the-thread). The generic builder returns false, because its body is the message. |
| the notification service | Carries the builder’s answer onto the notification it publishes, so nothing downstream has to re-derive it. |
| The inbound interaction | The flag is read off the trigger event into the interaction: the one normalized, platform-agnostic property workers may branch on (it is not an event-type check). A coalesced trigger yields one interaction per constituent message, all carrying the event-level merged flag, and the whole-trigger predicate is true when any of them is. A2A and internal task_assigned triggers carry their own context, so always false. |
| Prefetch | All three relevance prefetches read it: personal memory, relevant knowledge and episode recall. When set: skip the aux call (for personal memory the relevance filter, for relevant knowledge the query generation and live knowledge-base search, for episode recall the vector query). All three then render a gate-path hint so the block stays visible and self-explanatory rather than vanishing — EmptyMemoryHint, EmptyKnowledgeHint and EmptyRecallHint respectively — and the matching per-tool line in the retrieval re-search guidance carries the same nudge. |
The signal lives at the notification builder because the builder decides whether to emit a recon directive — classifying from event.type downstream would duplicate that decision and let the two drift. A raw token-count heuristic doesn’t work here: a webhook task_description is long (title + event metadata + multi-step “How to Handle This” boilerplate) but thin on substance — length would wrongly classify it as rich.
Personal memory still does its cheap diary recency list on a thin trigger (a DB read, no LLM) so it can render the hint only when the agent actually has memory rows to refresh — the vector half of the hybrid would key on a bare pointer that has nothing substantive to match, so it’s skipped alongside the aux filter, and nothing is embedded. A trigger whose ask is empty is gated the same way, because there is equally nothing to judge relevance against; trigger_requires_recon stays the builder’s flag and reads false for it. That is rare by construction — a chat parser drops a message with no text before it is routed, and an attachment arrives as a line naming the file — so what reaches the gate is a wake carrying nothing the engine can read as an ask: a source that sends neither a subject nor a message, or an event whose payload holds no text. Relevant knowledge skips the query generation and live knowledge-base search entirely — it only needs CanSearch to confirm a search could return anything (so the search-tool nudge is actionable) before rendering the hint. Episode recall skips the vector query outright and renders its hint unconditionally: unlike a diary list or an accessible-spaces check, the only way to know whether an agent has matching past episodes is the vector query the gate exists to skip — so the hint is phrased conditionally (“if this task resembles something you have done before…”) to read correctly even for an agent with no episodes.
Observability. The summary’s trigger_requires_recon records the gate decision once per turn. Without it, a gated prefetch and a filter that ran-and-found-nothing look identical in telemetry (both report a false *_hit and a zero selection count); with it, an operator seeing an empty ## Relevant knowledge block can tell the prefetch was gated (the trigger was a pointer) rather than broken. The event’s summary line surfaces it in the trace view: the count of blocks that hit out of seven (prefetch: N/7 hits), marked as a thin trigger with its filters gated when the gate fired. The denominator is the length of the engine’s own block list rather than a literal, so a block added without one reads as “7/6” rather than going unnoticed.
This makes the prefetch honest about its role: it’s an optimization for rich triggers, and for event-driven turns the tool-call path is the primary retrieval path — re-query-after-recon is the expected pattern, not a fallback. The agent pulls mid-turn via refresh_memory / search_knowledge / query_episodes, guided by the retrieval re-search guidance — one loop, so the frame that discovers it needs something is the frame that asks for it.
Two blocks are not gated, deliberately. The first is ## The thread so far: on a chat thread reply the gate fires precisely because the trigger is a pointer, and the thread is what the pointer points AT — so gating it would withhold the context on exactly the turns that need it most. Reading it is one HTTP call on a credential the node already holds, with no embedding and no aux LLM, so there is nothing for the gate to save. The flag stays true regardless: it describes the trigger body, and the three filters behind it still have only “+1” to judge relevance against.
The second is the conversation ledger. A pointer-shaped trigger is very often a follow-up on a conversation this seat already worked — the second comment on POC-518, the reply in a thread it answered yesterday. Conversation sessions render what the seat itself already said there, and that lookup is a keyed read on the trigger’s conversation identity rather than a similarity match: no embedding, no aux LLM, nothing for the gate to save. So on exactly the turns where all three prefetches above go quiet, the seat still arrives knowing what it last said and did in this conversation — which is what stops it answering the same question twice while it goes off to re-read the thread.
Salient-body sourcing
Section titled “Salient-body sourcing”The relevance prefetches, the counterparty profiler, the PersistDecider, and refresh_memory all reason about what the sender said. None of them want the notification builder’s scaffolding.
A source’s notify.Prompt builds the enriched body: for a Slack message, about 1.5k characters of ## Triage instructions front-loaded before the actual message. That enriched body becomes the turn’s task text (the executor needs the triage contract). But a relevance judgement made against it is made mostly against boilerplate that is byte-identical on every Slack turn — embedded, it dominates the vector, so every chat turn looks alike; and its worked examples (“@PM open a ticket for @SWE”) read to a model as roles and people the task involves, and as search terms.
So the raw message rides separately. The notification’s SalientBody carries the inbound body verbatim — the message, no scaffolding — alongside the enriched body. InboundInteraction.body is sourced from it (falling back to the enriched body for events that carry no salient_body); a coalesced trigger sources one interaction body per constituent message, and the merged notification’s own salient_body is the same messages joined chronologically with sender attribution (Alice: …).
The turn’s ask. At turn start the engine derives, from the trigger, what the turn was asked (prefetch.Request.Ask) beside the task the executor is handed: each notification’s salient body, led by its subject wherever the source says the subject is part of what was sent (an issue’s key and title, a page’s title, a monitor’s alert, and on a tracker comment the only place the topic is named) — a chat message’s subject is only the surface’s name (“Slack message”), the same on every message, so a chat turn’s ask is what was said (notify.Prompt.SubjectIsLabel, stamped on the wake as subject_is_label) — a coalesced burst’s merged salient body, a colleague’s question with who asked, and a schedule’s name and task. Every relevance judgement is made against it:
| Surface | Reads |
|---|---|
| Counterparty profiler / PersistDecider | InboundInteraction.body — the salient text, one entry per constituent; the PersistDecider is also shown turn_completed.ask for a wake with no interactions (a colleague’s question, a schedule’s task) |
| Synthesizer / Refiner | the turn’s ask as turn_completed records it — the interactions’ bodies, or its ask |
| The episode’s vector | the turn’s label, its ask and what it did, as one text — see Episodes |
## Personal memory prefetch | the turn’s ask → the similarity half’s vector, and the aux filter prompt |
## Relevant knowledge prefetch | the turn’s ask → aux-LLM query generation + knowledge-base search |
## Similar prior work (episode recall) | the turn’s ask → the vector query, and the episode summary prompt |
refresh_memory | its context_hint, which is the executor’s own account of the task after recon |
The ask is embedded once a turn, shared by the memory and episode searches, under a two-second budget (prefetch.EmbedBudget, the knowledge search’s own QueryEmbedBudget, stated for “a turn starting”), and whole: an ask longer than the embedding model’s input — a long task description, a busy thread’s digest — is split between words and its pieces’ vectors pooled into one, in one request. Nothing is embedded on a thin trigger, or for an ask with nothing in it, which the searches treat as a thin trigger. The executor’s own task text is unchanged.
Episode lifecycle
Section titled “Episode lifecycle”The episodes table is the raw substrate of agent learning. Without lifecycle management it grows forever. The episode lifecycle worker drains it on a threshold-gated pass, walking each seat and doing nothing for the ones that are not due.
Episodes have no plain retention sweep, deliberately. Every other short-horizon table gets a range delete on a single horizon. Episodes cannot: their retention is this pass, which applies four different horizons to four different row states. A single DELETE WHERE ended_at < cutoff would collapse all four, and the row it would take first is the compacted summary — the only record of a whole era of a seat’s work, standing in for hundreds of turns that are already gone.
Trigger: threshold-gated, on a slow loop
Section titled “Trigger: threshold-gated, on a slow loop”The worker ticks hourly and, for each seat, runs one indexed count(*). A seat under max_raw_episodes_per_agent (default 500) is skipped without touching the pass; a seat over it gets the full lifecycle run. The count is the gate rather than the pass’s own early return, because “not yet” is the overwhelmingly common answer and it must cost one query rather than a walk.
The cadence is far shorter than the skill curator’s because what it watches is a count, not a clock: a busy seat crosses its threshold in a burst, and every turn past that point pays the recall scan over rows that should already have been folded. An hour bounds that overshoot to one hour of one seat’s traffic.
It is a fleet singleton, claimed per tick under the node’s own incarnation: two nodes compacting one seat’s episodes would summarise the same cluster twice and pay for it twice. The claim fails closed: not knowing whether a peer holds the duty is exactly the case where running anyway produces the double write. No background pass fires on start, because every node in a fleet starts within seconds of a rolling restart: firing on start means every node races for the duty at once, and a crash-looping node spends the company’s tokens on every restart.
The loops are the process’s, and the passes are the revision’s. A node arms the four background loops (episode lifecycle, skill curator, clustered synthesis, promotion) once, whatever its company configures and with no company at all, and every config apply hands them the passes that revision turns on, built from its models, credentials and knobs. A loop keeps its clock across an apply: a pass the revision turns on runs at that loop’s next tick, a cadence the revision changes (skill_curator.interval_hours, skill_synthesis.scheduler_interval_seconds) starts over from the apply, and an apply that leaves a cadence alone leaves the next tick where it was, so a company edited more often than its curator ticks still curates. A loop whose pass is off claims no duty. That is what lets the first company a fresh node is handed (every company created from the dashboard) run its passes without a restart, and what starts compaction, clustering and promotion, the three that call a model, on the apply that gives a company with no providers.llm its first provider.
Compaction needs a summarizer. The pass folds a cluster by asking the seat’s own llm_auxiliary chain to describe what its members had in common — per seat, so a company whose seats run on different models has each one’s memory compacted by the model that seat is configured with. With no auxiliary model configured anywhere, the pass does not run at all: what it could still do is delete, and deleting is the half an operator least wants unsupervised. The rows stay raw and readable instead.
One worker, four actions
Section titled “One worker, four actions”For each seat over its threshold the worker runs the full lifecycle pass:
- Drop non-terminal episodes older than
non_terminal_max_age_days(default 14).self_iterateis a mid-state — the reflect engine’s terminal-outcome gate already excludes it from skill synthesis, and it only feedsquery_episodesrecall as noise. Cheap SQL DELETE; no LLM. - Drop tool-free turns older than
tool_free_max_age_days(default 90). A turn that called no tools cannot be compacted — clustering pools turns by tool-sequence overlap, and there is no overlap to measure — so without this the raw rows of a chat-only seat grow for the life of the deployment, and every one of them is scanned and cosined at the start of every turn. The horizon is far longer than the two either side of it because this sweep drops the only record of work that really happened: a fact worth keeping past a quarter is one the seat should have written to its diary withreflect_and_persist. Cheap SQL DELETE; no LLM. This is the one sweep whose deletions are irrecoverable, so its volume is reported on its own:episode_lifecycle_passcarriestool_free_droppedbeside the other per-pass counts. Watch it before shortening the horizon — theCompactionCompletedevent folds the same number intonon_terminal_dropped, so the log is the only place it is visible alone. - Drop skill-consolidated episodes older than
consolidated_grace_days(default 30). When the synthesizer drafts a skill from a cluster of episodes it stampsconsolidated_into_skill_idon each source row; the lifecycle worker drops them after grace because the skill itself now carries the learning forward. The grace gives operators a chance to audit / detect bad consolidations before the source disappears. - Compact the rest, the centerpiece. Pulls the oldest
compaction_batch_size(default 200) remaining raw episodes older thancompaction_min_age_days(default 30) that a fold could take — never a turn that called no tools, which nothing can pool, nor a summary’s retired exemplar, which is kept raw on purpose: both only collect at the old end of the batch, and counted in it they took its slots until a pass found nothing to fold — greedy-clusters them by tool-sequence Jaccard, and for each cluster of size ≥compaction_min_cluster_size(default 3) calls the role’sllm_auxiliaryto summarise into a compacted row (common_task_pattern,common_outcome,success_rate,subjects_involved,notable_patterns). Writes onekind='compacted'row, deletes the cluster’s originals (except 2-3 exemplars retained as raw rows for drill-down, referenced by the new compacted row’sexemplar_turn_ids). Each member reaches that call as its task, its outcome, its whole tool sequence and its date. A task or outcome past 280 bytes is first condensed by the same seat’s auxiliary model and labelled(condensed), never cut — a pattern inferred from a stack trace’s first 280 bytes is the trace’s opening and none of what the turn was for — and one whose rewrite fails is carried whole. A long-lived seat’s outcomes are mostly over that line, so a full cluster is a few hundred small rewrites, run four at a time and charged to the seat like any auxiliary call. - Optional: evict ancient compacted entries older than
compacted_max_age_days(default 0 = disabled). Hard long-tail storage cap for orgs that need years-out limits; off by default since compacted summaries are 10-100× smaller than the raw rows they replaced.
Two physical row shapes share the same table
Section titled “Two physical row shapes share the same table”After the migration episodes rows distinguish on kind:
| Field | kind='raw' | kind='compacted' |
|---|---|---|
count | always 1 | N original episodes collapsed |
task_summary / plan_summary / tool_sequence / review_outcome | per-turn detail | the cluster’s common values |
started_at / ended_at | one turn’s timestamps | the cluster’s window |
common_task_pattern / common_outcome / success_rate / subjects_involved / notable_patterns | unused | LLM-summarised aggregate |
exemplar_turn_ids | empty | 2-3 raw rows kept as drill-down anchors |
consolidated_into_skill_id | set when a skill drafted from this row | always NULL |
Similarity recall reads raw rows only. A compacted row summarises a cluster of turns and reads in a prompt like one turn that did all of them, so it is never embedded and learning.Episodes.Recall filters on kind = 'raw'. The time-window reads return both kinds, and their callers branch on kind:
query_episodesbuiltin: itsquerypath is similarity, so raw turns only; its recency andconversationpaths return both kinds, and render a compacted row as its pattern, count, outcome tally and variations rather than as a turn.## Similar prior workprefetch block: similarity, so raw turns only.Synthesizer: raw rows only, on both paths. Compacted aggregates are too coarse to draft a clean skill body from, and the clustered pass would count one fold as one turn.Refiner: reads no episodes at all. It is shown the skills the turn was offered and the turn itself (what woke it, what it was asked, what it did, its tool sequence and outcome), which is the whole question it answers.
What this protects
Section titled “What this protects”- Storage growth —
max_raw_episodes_per_agentis when a pass runs, not a cap. A pass removes mid-state turns after 14 days, tool-free ones after 90 and absorbed ones after the grace, and folds settled tool-using turns past 30 days into summaries ~10-100× smaller per unit of original work. What it does not bound: a settled tool-using turn that never joins a cluster of three has no horizon and stays raw for the life of the seat, as do the two exemplars of every fold unlesscompacted_max_age_daysevicts them; and a pass clusters only inside its batch of the oldest candidates, so turns of one kind spread thinner than three to a batch are never folded, and once the batch’s oldest rows are all such turns it folds nothing at all. - Recall pollution — non-terminal noise drops fast; old patterns become aggregate summaries instead of crowding similarity hits.
- Learning drift — when a skill captures a workflow, the source episodes get out of the seat’s view (after grace), so the agent stops being shown stale per-turn detail of work the skill now represents abstractly.
- Long-tail signal preservation — routine work that never qualifies as a skill (most agent turns) survives as a compacted aggregate rather than getting dropped wholesale. The seat can still answer “you’ve done this kind of work N times” via the compacted entry.
What does NOT happen
Section titled “What does NOT happen”- No work on the caller’s path — nothing about compaction runs inside a turn. Reads pay no latency cost for it, and neither does the write that crossed the threshold.
- Compaction never feeds skill synthesis — the consolidation hierarchy is one-directional: raw → skill, raw → compacted. A compacted entry doesn’t get re-promoted to a skill; if the same pattern recurs after compaction, the new raw episodes form a fresh cluster the synthesizer can pick up.
- No work for idle seats — a seat under its threshold costs one indexed count per tick and nothing else. The loop wakes; the seat does not.
Skill curator
Section titled “Skill curator”A synthesized skill that nothing uses any more should leave the catalogue, and one that is used again should come back. That is a clock, not an event, so it is the second background pass — a fleet singleton on the same claim discipline as the episode lifecycle, ticking daily. A day, because the transitions it makes are measured in tens of days: a pass an hour would scan the whole catalogue 24 times to make the same zero transitions, and the one it eventually makes would land at most an hour earlier, against a threshold nobody set to the hour.
The state machine is active → stale → archived on disuse, and stale → active on use:
| Transition | When | Effect |
|---|---|---|
active → stale | unused for stale_after_days (default 30) | Still listed and still loadable — the prefetch renders it with an ageing marker, so the agent knows. |
stale → archived | unused for archive_after_days (default 90) | Listings hide it and the loader refuses it. Archived is not deleted: the row stays readable, so restoring one is an operator edit rather than a re-synthesis. |
stale → active | the skill is used again | Revival happens in the same transaction as the use, so the skill is back in the very next turn-start prefetch rather than after the curator’s next tick, which on the default schedule is up to a day later. |
An archive window inside the stale window is a misconfiguration, and taken literally it archives rows the same policy calls fresh. It is widened to the stale window instead, which is the reading both halves agree on.
Skills the operator has pinned are exempt from every automatic transition. Nothing promotes a skill to pinned.
Being offered is being used
Section titled “Being offered is being used”A skill’s staleness clock is its last-used stamp, and the thing that moves it is the turn-start prefetch offering the skill — not the model then loading its body.
That is the honest reading of what the stamp answers. A skill rendered into the prompt is in the catalogue and is what the seat is being asked to work from; whether the model loaded the body is a question about that turn, not about the skill’s currency. Keying on the load would age out every skill whose menu line was enough — which is the well-written ones.
The ids follow the prompt’s own character budget: a skill whose menu line did not fit was never offered, so its clock does not move. A stamp that cannot be written is announced as a telemetry failure rather than swallowed, because an operator has to see a clock that stopped before the curator archives a hot skill.
Without this the whole catalogue ages out over a quarter while the prefetch is putting it in front of a model the entire time — and not as a slow degradation anyone notices. The menu simply gets shorter.
Telemetry harness
Section titled “Telemetry harness”The learning loop produces durable artefacts (synthesized skills, diary entries, counterparty profiles, episodes) and the surfaces that read them. Without per-surface measurement an operator cannot answer two basic questions:
- Are skills being used? Berlot-Attwell et al. (2024) showed that in some library-learning systems the apparent gain from skill induction comes from extra LLM sampling rather than skill reuse. Crewlet’s induction pipeline does real work; whether the resulting skills earn their keep is an empirical question that requires telemetry.
- Are the turn-start prefetches actually firing? A block stuck at 0% hit rate (e.g.
episode_recallreturning empty for every turn) is almost always a configuration / data problem, not a turn problem — but only visible if hit / miss is recorded.
The harness lives in:
| Surface | What’s tracked | Where it lands |
|---|---|---|
synthesized_skills.use_count / last_used_at | Per-skill use count + most-recent-use timestamp | Bumped by learning.Skills.MarkUsed, called from the use_skill builtin after a successful resolution and from the SkillUse reflection worker for every skill a turn was offered. |
skill_used event | One per use_skill(name) resolution (source_kind: synthesized) and one per load_tool_skill(key) (source_kind: registry, naming the tool skill’s page in source_page_id / source_container) | Published on crewlet.events.skill_used; correlated to the host turn via trace_id / span_id. |
knowledge_read event | One per read of the knowledge base — a get_page, a search_knowledge, the turn-start knowledge block, a tool-skill load, a phase’s tool-skill catalogue — naming the pages it reached. See Knowledge System § What agents read. | Published on crewlet.events.knowledge_read. |
skill_synthesized / skill_refined / skill_promoted events | Lifecycle markers: induction, refinement, cross-agent promotion | Published on crewlet.events.skill_*; the dashboard groups them by trace. |
prefetch_summary event | One per turn after the seven context prefetches resolve, recording per-block hit (bool) + bytes (rendered size), the knowledge selection_count, the chat thread’s thread_context_posts / _read / _stopped_short, the trigger_requires_recon gate decision, and turn_embedding — what became of the one vector the memory and episode searches rank by: embedded, failed (the embedder refused or did not answer inside the two-second budget, so those searches did not run — an empty ## Similar prior work is then not “nothing similar”), unconfigured, or absent where no search asked for one. Every block but the knowledge and chat-thread ones — which say in words when their source could not be read — degrades silently rather than failing, so this is the only signal that tells an unreachable store from one with nothing to say. | Published on crewlet.events.prefetch_summary once per turn. |
persist_decider_completed classification / ttl_until | Tier label (LONG / SHORT / DOC / NOOP) + TTL on SHORT writes | Existing event extended so dashboards can plot the per-agent tier distribution. |
learning_health SQL view | Per-agent rollup: total_skills, skills_used_at_least_once, total_skill_uses, most_recent_skill_use, avg_uses_per_skill, avg_skill_age_days | Created by the store’s 0002_learning.sql migration; query it directly from the store. |
Berlot-Attwell threshold
Section titled “Berlot-Attwell threshold”The single load-bearing metric is avg_uses_per_skill from learning_health. The literature’s working threshold:
avg_uses_per_skill < 0.1 → the library isn't doing what it claims; investigate retrieval, granularity, or whether the gain is just from extra samplingA new agent will sit at zero until it has been alive long enough to retrieve. Combine with avg_skill_age_days to discount young rows.
Best-effort rule
Section titled “Best-effort rule”Every telemetry write (Skills.MarkUsed, the skill_used and knowledge_read publishes, the prefetch_summary publish) is best-effort: a failure is logged once and swallowed so the host path (skill load, turn) is never broken by measurement. Test mode (no event queue / no DB) is a silent no-op.
What the auxiliary calls cost
Section titled “What the auxiliary calls cost”Every completion made on a seat’s behalf, the auxiliary calls on this page
included, is charged to the seat’s token_budget
counter and the company’s, in the windows current when it returns. For the
auxiliary calls ONE SEAM does it: the learning workers, the turn-start prefetch,
every compaction and a person’s answered question resolve their model through
the engine’s auxiliary seam, stating whose cost the call is — its stage
(turn, reflection, background, operator), its purpose and the turn it
serves — and the seam both charges the call and records it as an
auxiliary_spend event, so a worker added later is charged and counted in every
spend figure without anyone wiring either. A call whose attribution names no
stage or purpose is refused before a model is resolved. Because an auxiliary
call’s size is known only from its answer, each one is recorded after it
returns, past the ceiling included; where a gate reads the room left first, it
decides whether the work starts at all. See
Budgets and spend § Auxiliary spend
for where each stage’s spend is drawn:
| Path | Calls | Gate before it starts |
|---|---|---|
| Turn loop | every executor and reviewer round | the round’s own meter |
| Round-cap extension judge | one call per exhausted phase | the turn’s meter, asked before the judge is consulted and again once its evidence is condensed; the call is charged on it after it answers |
| Coding sandbox | the box’s whole run | the seat’s headroom, read before the run launches; the spend is recorded when the run is collected |
| Turn-start prefetch | memory filter, knowledge query, episode summary | the budget park: a seat with no room has its delivery parked, so no prefetch runs — and then the turn’s meter, as for every in-turn call below |
| In-turn rewrites | the conversation block’s condensation, and every rewrite the turn’s ledgers, judge, tools and workers need | the turn’s meter: every in-turn auxiliary call is charged through it, so a window one fills is held before the turn’s next round is sent, and none is made while it holds a full window — except the task card’s, written after the last round as the turn’s record |
| Reflection pass | persist decider, profiler, refiner, single-turn induction | the pass’s no-budget skip |
| Conversation entry | the rewrite of the last round’s long tool payloads before the thread’s entry is written, after the turn | the same gate as the reflection pass, since the entry is what the seat remembers of the turn: with no room left on the seat or the company — a turn the budget ended among them — no rewrite is made, and each long payload is named by size and digest instead |
| Background learning | episode compaction, clustered synthesis, skill promotion | none: each call is recorded, but these passes do not read the room left before they start |
Integration points
Section titled “Integration points”| Touchpoint | Role |
|---|---|
internal/engine (the turn’s telemetry) | Emits turn_completed when a turn closes, carrying everything the reflection gates read: what the turn did, what it was asked where no interaction says (ask), its outcome, the final round’s tool sequence and every tool name the turn called, the review outcome, the skills the prompt offered, and the inbound interactions with their senders resolved. |
internal/agent/prompts | The executor’s prompt builder injects conditional guidance blocks gated on tool availability. |
internal/knowledge, internal/confluence | The knowledge-search seam and its one backend, the Confluence searcher (CQL), backing the ## Relevant knowledge prefetch and search_knowledge; the org-wide read scope narrows it by space. See Knowledge System. |
internal/agent/builtin | Builtins, registered into internal/tools: query_episodes, reflect_and_persist, refresh_memory, refine_skill, use_skill, mark_onboarded. |
internal/events | turn_completed, episode_written, persist_decider_completed, counterparty_profile_updated, reflection_completed, skill_synthesized, skill_refined, skill_promoted, skill_used, skill_staled, skill_archived, skill_revived, skill_telemetry_write_failed, prefetch_summary, knowledge_read, compaction_requested, compaction_completed. |
internal/store | Holds episodes, agent_diary and the dashboard’s event log, in the node’s own file. |
internal/learning/memsync | Makes that file a cache rather than the only copy: every memory row is published to a compacted changelog on the stream, and a node acquiring a seat replays it into its own store before the mailbox attaches. Without it a seat that moved node would run its next turn having forgotten everything. See A seat’s memory follows it. |
internal/learning/memread | Reads a seat’s memory — and its conversation ledger — from the node that HOLDS the seat, which is the one copy memsync keeps current: the reader follows the seat’s lease, answers here, asks the holding incarnation on an ephemeral scatter, or answers empty for a seat no node holds, and every answer names who gave it (held_by). Every node answers for the seats it holds. The OVERVIEW (memory_overview) is the same rule for every agent at once, in one round: one listing of the seat leases, one scatter naming each holder’s seats, one reply per holder, and a coverage naming any holder that did not answer. See A seat’s memory follows it. |
internal/learning | The reflect dispatcher (Reflector) and its per-turn workers (PersistDecider, Episodist, Profiler, SkillUse, Synthesizer, Refiner), the background passes behind Background (episode Lifecycle, the skill curator, clustered synthesis, cross-agent Promoter), Skills for synthesis and refinement, Diary, and the onboarding marker store. The turn-start prefetches are internal/agent/prefetch. |
internal/config | learning: block — per-role enable flag, reflection budget, promotion thresholds, lifecycle knobs. See Configuration. |
internal/api | GET /agents/{id}/memory (and the agent_memory query) serves the diary, the episodes, the synthesized skills and the counterparty profiles — each a page with its counted total — plus the latest reflection and when the seat onboarded, for the seat page’s Memory tab, its Overview and one agent’s page under Knowledge › Agent diaries; the memory_overview query serves that screen’s list — every agent seat’s totals and newest note, each counted by its holder; the rows are projected by memread rather than marshalled from the domain types. See API endpoints. |
Data model summary
Section titled “Data model summary”| Table | What it holds | Keyed by |
|---|---|---|
episodes | One row per completed turn (raw) or per cluster (compacted) | id; indexed by started_at |
agent_diary | The agent’s private observation log; rows carry an embedding and its embedding_model for the vector half of the ## Personal memory prefetch’s hybrid candidate selection | id; indexed by agent_id, created_at; no vector index — recall is a per-agent scan the database ranks |
synthesized_skills | Auto-drafted skills, agent-scope | id; unique on (agent_handle, name) |
synthesized_skill_versions | Refinement history | id; references skill_id |
counterparty_profiles | One row per (observer, subject, platform) | composite |
agent_onboarding_markers | mark_onboarded bookkeeping | agent_id (PK) |
Shared knowledge has no table here: natively it is rows in the REPLICATED estate — the pages themselves, and the vectors derived from them — and on Confluence there is no local copy at all, only a live query (see Knowledge System).
Deliberate non-goals
Section titled “Deliberate non-goals”- Single-user persona model. Crewlet is multi-party; a counterparty profile is per-identity and observer-scoped.
- Model-level fine-tuning as a core feature. Optional, downstream of a stable trajectory dataset. No role is required to use a learning-aware model.
- Cross-org knowledge leakage. Synthesized skills are agent-scope only; cross-agent promotion lands as a knowledge-base draft for human review, not as an engine-side row.
- Black-box self-modification. Every synthesized skill edit is versioned and rollback-able. Counterparty profiles are written through a single observer, never auto-merged.
- Auto-promotion of casual remarks to team rules. A directive issued in Slack to one agent reaches another only when (a) a human or authorized agent updates the relevant knowledge-base page, (b) the receiving agent broadcasts to the team, or (c) someone with structural authority decides to formalise. The system does not auto-promote personal counterparty-profile observations to unit-shared knowledge.
- A monolithic “learning agent.” Six small, independently testable components beat a single reflective super-loop.
Prior art: Hermes Agent
Section titled “Prior art: Hermes Agent”The learning subsystem was designed with Nous Research’s Hermes Agent as a reference point — reimplemented rather than taken as a dependency. Hermes is a vertically-integrated single-user CLI agent, not a library — its memory manager, skill tools, and session search are threaded through a 600k-line monolith with assumptions (home-directory storage, single user, single agent, no hierarchy) that are incompatible with Crewlet’s org model.
That said, several Hermes design choices are directly useful and adopted above:
| Hermes pattern | Adopted where |
|---|---|
| Conditional prompt-guidance blocks injected only when the matching tool is registered | Prompt scaffolding |
| “Declarative facts, not instructions to yourself” memory-writing rule | PersistDecider writing-style rule |
| “Patch skills on encounter; don’t wait to be asked” | Refiner patch-on-encounter norm |
| 5-tool-call default threshold for treating a turn as skill-worthy | Synthesizer default trigger |
| Cheap auxiliary model for summarizing session/episode-search hits | query_episodes + ## Similar prior work prefetch |
| Frozen memory snapshot at session start for prefix-cache stability | Context prefetches frozen at turn start |
Pluggable MemoryProvider interface (mem0, honcho, supermemory, …) | Validates the agent_diary store shape |
Explicitly rejected:
- Monolithic CLI coupling — Hermes’s learning loop is threaded through its agent entry point; ours sits behind the
EventQueueas its own package. - LLM-nudge-only triggers: Hermes’s pipeline fires only if the model invokes the tool. Ours pairs nudges with the deterministic reflect dispatcher.
- Single-user
USER.mdpersona: replaced by multi-party counterparty profiles keyed by(observer, subject, platform). - Home-dir file storage — replaced by tables in the engine’s own store, each row carrying its embedding and the model it came from (no vector index — recall is a per-seat scan the database ranks).
- Unversioned skill overwrites — Crewlet keeps prior revisions for rollback.
- No model fine-tuning requirement — notably, Hermes itself also runs on stock models; Crewlet’s in-engine learning never touches weights.
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.