Skip to content
You are reading documentation for unreleased main. This page is not in 0.1 yet.

Conversation Sessions

A seat that answered a Jira comment on Monday and is woken by a follow-up on Wednesday used to arrive with no memory of Monday. Everything the engine carried across turns was either agent-scoped and similarity-addressed (episodes, the diary) or content-free (thread follows, the completion ledger) — nothing was keyed to the conversation itself.

The conversation session ledger closes that. Every completed turn appends a structured entry to its conversation, and the next turn of that same conversation gets those entries back as an ## Earlier in this conversation block on the executor’s user message.

It is the cross-turn counterpart of the prior-work ledger, and it deliberately follows the same doctrine one scope wider.


Not a transcript replay. The engine can already round-trip a whole LLM conversation — the detached run’s execute_state persists the full message list, each assistant turn kept as the vendor’s own content blocks (signed thinking included, in the order the model wrote them) beside the backend and model that wrote it and a digest of the tools its reasoning was written under, and splices it back into a running loop, which replays those turns exactly as they were written. That is right for a turn parked on a question whose dangling tool call is waiting for one answer. It is wrong here: a conversation’s next turn arrives against a thread that has moved, and replaying raw prior context invites acting on state that is no longer true.

Not episodic memory. Episodes answer “have I done something like this before?” by embedding similarity across every conversation. This answers “what did I already say here?” by identity. They share a write site and nothing else:

EpisodesConversation sessions
Keyed byagent + vector similarityagent + conversation_key (the conversation identity)
Holdstwo ≤2000-char summaries, tool namesplan, reasoning, calls, the reply, the verdict
Retrievalcosine top-3, done only, no recencythe newest N of this conversation
On thin triggersprefetch gated offalways rendered
Compactionclusters collapse per-turn detailnone; trimmed by count and age

The last row of that table is the sharpest difference. A Slack thread reply or a Jira comment webhook is a pointer, and the engine’s thin-trigger gate skips all three aux-LLM prefetches on exactly those turns — the ones that continue an existing conversation. The session block is deterministic (no embedding, no aux LLM), so it renders there regardless.

It is not the same thing as the thread block, and neither replaces the other. ## The thread so far is what everyone said on the chat surface, read back live from the vendor. A conversation session is what this seat said and did, across turns, recorded by the engine — the plan, the tool calls, the reply and the reviewer’s verdict, none of which is visible in a chat thread. A seat reading only the thread cannot tell which of its own replies it has already reasoned through; a seat reading only the ledger does not know what the other five people in the thread have said since. Chat is also only one surface: a Jira comment, a GitHub review and a scheduled fire all have a conversation session and no thread to read.

Not a second, invisible memory. The cli-agent workspace deletes a coding CLI’s own sessions before and after every call, precisely so that one task’s context cannot leak into the next through a channel nobody can see. This ledger is the inverse shape of what that rule rejects: engine-owned rather than tool-private, scoped to one conversation rather than leaking across tasks, and rendered into the prompt as a visible block — so the context it adds is stated in the turn that uses it, where a person reading that turn can see it, rather than applied out of sight.


Built at turn end from data already in hand — there is no summarisation call, because a summariser that drops the line naming the reply the seat already sent re-creates the duplicate-answer bug in a place nothing else can catch.

  • Triggered by — the ask the turn was given, verbatim. For a coalesced conversation that is the whole merged digest — its header, each earlier message with its sender, and the latest message in full — because the entry is the store’s only record of what the turn was answering, and keeping only the last constituent would make a five-message thread read back as one
  • You set out to — the executor’s own summary of what it did
  • You called — the tool-call lines, writes always recorded, reads marked (read)
  • You replied — the turn’s final text, and only when the turn actually reached the party waiting on it
  • You did NOT reply — the same text when nothing did. A turn can end with real work behind it and no way to say so: the round budget ran out, the loop broke, the reviewer closed it anyway. The text is kept because the conclusion is real context for the next turn; what it must never do is read as a reply. Filed as one, it made the failure seal itself — the seat’s next turn on the thread read back that it had already announced work nobody had been told about, and answered the follow-up against that. The line is spelled out rather than left implicit in a missing You replied, because “no reply line” is something the reader has to notice while “nobody received this” is something it has to act on
  • Reviewer — completed_work, the prose on what already landed
  • You ended that turn blocked — what stopped it, taken from the executor’s own evidence on a blocked round and empty on every other outcome. A turn that put a question to somebody and ended there records done exactly like one that finished the work, so without this the seat’s next turn on the thread cannot tell “I answered this” from “I asked about this and I am waiting” — and read as the first, a question the seat asked becomes work it believes it delivered. It is also precisely the turn that reads the line: the reply to that question is what wakes the seat, so the message in front of it is usually the answer
  • Turn ended — only when the decision was not done

The reviewer’s other field, turn.Review.Notes, is deliberately not carried. It is what the executor’s next round is shown when the decision is self_iterate, and is written as an instruction to that round: “the next round should retry posting X”. Replayed into a later turn it stops being history and becomes a standing order the reviewer never issued, aimed at a round that already came and went. Nothing is lost by dropping it: the calls that failed are in the tool lines and the verdict is in Turn ended.

Every field is written verbatim, apart from a tool argument or a failed call’s error past the ledger budgets, which is rewritten to fit by the seat’s auxiliary model before the row is written — never cut. Where no rewrite can be had it is named by its size and a digest instead, and the call line beside it still says which tool ran and whether it worked: the model failed, or the seat or the company has no token budget left. The rewrite is reflection-stage spend, so it waits on the gate the reflection pass waits on, and a turn the budget ended pays for none. Nothing else is touched at write time, and that is deliberate: this row is the store’s only record of the turn, so a trigger trimmed on the way in is not a shortened entry, it is the only copy. How much of it a later turn is shown is a read-side decision — see Cost below — and answering that display question at write time destroyed the data.

Inherited wholesale from the within-turn ledger, and stronger here: a read from last Tuesday is stale by construction. Tool results are never carried, reads render with a (read) marker, and the block’s header tells the model it may re-run exactly those. Telling it not to repeat a jira_get_issue would push it to invent the data instead.


A turn is appended when all of these hold:

  • it completed — a crashed turn has nothing coherent to record;
  • it was not a detached-sandbox suspend — the resumed turn records once, for real;
  • its trigger has a reproducible conversation identity (see below);
  • it has a work key — the constituent-trigger identity used for dedupe.

A turn that ended failed on a guard breach is recorded. “I tried this and it did not land” is exactly what the conversation’s next turn must not rediscover the hard way.

The row is keyed on (agent_handle, conversation_key) and deduped on work_key — never turn_id. A turn id names ONE RUN (two identities), and one trigger legitimately runs more than once: a turn that broke before reaching outside the engine is redelivered, and two nodes completing one trigger mint two runs besides. A turn-keyed row would record each of those instead of collapsing them, and the next turn would read its own reply twice.

The row still carries the run id beside the key, because that is what it renders: the seat reads “(turn a1b2c3d4)” back on its next turn of the thread, and that has to name an execution somebody can open.

The dedupe index is partial, over work_key <> '': an empty work key is the documented “a turn with no ledgerable trigger”, and those turns are legitimately distinct rows that must never collide onto one.

Separately, each row carries an entry_id its writer mints — the row’s name across the fleet, so memory replication can carry the ledger with the seat. It is not a second dedupe: the table’s id is an AUTOINCREMENT that starts at 1 on every node and means nothing off the one that wrote it, and entry_id says nothing about what a row means — two unkeyed turns get two ids and stay two rows.

conversation_key holds the conversation identity: the {source}:{local} grammar of the event system — jira:POC-7, slack:C9:1718.001, github:acme/api#42. It is the durable thread, NOT the partition key the seat’s inbox coalesces on. The two are the same string for every source but chat, and for chat they differ on exactly one surface: a direct message is one conversation however it is threaded, so it keys on the bare channel (slack:D1) while a reply inside it still partitions on its thread. Keying this row on the partition instead is what made the ledger silently empty on the surface it was written for — a DM’s first turn was filed under the channel, its thread reply looked up the thread, and the seat re-read its own 1:1 line as a first turn.

A trigger with no derivable conversation — a scheduled fire, a task assignment, an A2A wake — keys as event:{uuid}, which no later message can reproduce; those are not recorded, because the row could never be read back.

Two seats legitimately serve one conversation (a lead and its report on one ticket) and each keeps its own ledger.


The block is resolved once at turn start and frozen for the whole turn — the same rule the turn-start prefetches follow, so a self_iterate loop cannot invalidate the provider prompt cache. It rides the user message, never the frozen system prefix, alongside the prior-work ledger:

## Earlier in this conversation
…your prior turns, oldest first…
## Task
…the newest thing said…
## Already done earlier in this turn
…the within-turn ledger, on iterations after the first…

All three are headings at the same level, because they are peers: the ask was a bare Task: label once, and under a structural reading of the message that filed the newest thing anybody said inside the conversation history above it.

Oldest to newest top to bottom — the ask is the newest thing said, and the within-turn ledger below it is the newest thing done, so the most recent context sits nearest the model’s answer. The header states the contract explicitly: do not repeat a reply already given, (read) calls may be stale, everything else already took effect, and where the history and the task disagree the task wins.

Review does not receive the block. It judges this turn’s record — what the executor said it set out to do against what the tool log says ran — and its duplicate-delivery rule is already served by the within-turn ledger; feeding it prior turns invites judging work this turn never promised.


The block is re-sent on every executor and reviewer round, and Anthropic bills cache reads at full token value, so its cost multiplies by rounds used. Two things bound that product, and which one applies where is the whole design:

  • max_entries, at write time — how many turns of a conversation are kept at all. A chat DM is one conversation for the whole channel, so its ledger never stops receiving entries.
  • ledger.InjectedMaxChars (24 000 bytes), at render time — how much reaches one prompt. Past it the newest entries render whole in three quarters of the room, and the older ones are condensed by the seat’s auxiliary model into one account in the last quarter, under a heading that says how many turns it stands for and that it is a rewrite. Only where no rewrite can be had are the older entries left out — and the block then says how many. The newest always survives however long it is. An entry carries the seat’s own reply, which is unbounded — a turn that produced a document puts that document in the row — so without this a busy conversation eventually exceeds the model’s context and the turn cannot run at all. Dropping was the old behaviour, and it told a seat on a long thread that it had delivered nothing in the turns it could no longer see.

Never a cut inside an entry, and never at write time. The stored row is the only copy of that turn, and a half-recorded reply reads as the whole of what the seat said — which is how a seat repeats a reply it cannot see it already gave. Two config knobs (injected_max_entries, injected_max_chars) used to be documented here; neither was ever threaded to a caller, so both validated, defaulted and described a truncation that did not happen. The prompt.size telemetry event is where the delta shows up fleet-wide — read its user_bytes, which is where this block lands, rather than its approximation: that figure also carries the tool-definition array, which is usually larger than the ledger and moves for reasons of its own.

Against that: the re-recon it displaces costs a list_mcp_server_tools round, an activate_tool round and the read itself, on every turn of the conversation — and recovers only what was posted, never the agent’s own intent or the results it gathered.


turn_engine:
conversation_session:
enabled: true # the feature gate — a live kill switch
max_entries: 20 # kept per conversation, trimmed at write time.
# What ONE PROMPT shows is bounded separately,
# at render, by ledger.InjectedMaxChars — see
# Cost above
retention_days: 30 # matches the event store's own horizon

Nested under turn_engine rather than beside it, which is load-bearing: it rides the live the turn-engine settings cell, so it hot-reloads through the existing turn-engine diff handler with no extra apply-config wiring. Setting enabled: false restores the previous prompt exactly — which is why it is safe to leave on.

retention_days is the one retention an operator sets rather than the engine (a company running quarter-long tickets has a real reason to keep more). It is read when the maintenance worker is built, so a change lands at the next process start, and it is floored at the sweep interval so a hostile value cannot undercut the worker’s own invariant.


One row per recorded turn in conversation_sessions, deduped by a plain partial unique index over work_key <> '' and an ordinary ON CONFLICT DO NOTHING (internal/store/schema/node/0005_turn_ledgers.sql).

episodes reaches the same guarantee by the other route — a NULLable work_key under a plain unique index, so ON CONFLICT can name it (0002_learning.sql). Two shapes for one rule, and the difference is only what ON CONFLICT can target: aiming it at a partial index is a parse error unless the predicate is repeated verbatim.

Bounded twice: max_entries trims on write (a chat DM keys on the whole channel rather than a thread, so its ledger never stops receiving entries), and the maintenance worker sweeps past retention_days.

Every field of a recorded entry is stored verbatim apart from the tool arguments and errors past the ledger budgets, which are rewritten to fit — never cut — or, where no rewrite can be had, named by size and digest. This row is the store’s only record of the turn, so a field cut at write time would not be a shortened rendering — it would be the only copy. Bounding what a prompt shows is a separate decision, made at render time; see Cost above.

Failure never stops a turn. A write that fails is swallowed — it happens on a completed turn’s tail, where there is nothing left to tell. A read that fails raises, and the caller decides: the turn engine renders no history (exactly the pre-ledger prompt) and logs that it could not read it. Swallowing it in the store would make “unreadable” and “nothing said yet” one answer, and a seat would run without its history with nothing anywhere to say why.

The ledger lives in the store of the node running the seat, which every node opens, so it is never replicated: a seat that moves between nodes simply arrives with no history, which is the same fail-open answer.


There is no read surface, deliberately. The ledger is prompt context: the engine renders it into the next turn of the same conversation and nothing else consumes it.

A dashboard tab and a /conversations endpoint did exist, and both were removed. They were a viewer for somebody ELSE’s threads — a Slack channel, a Jira issue — reconstructed from what the engine happened to record about them, always a worse version of the thread than the surface it lives on, and a conversations screen in this product is meant to be Crewlet’s own messaging when there is one to show. Keeping a half-view until then would have promised a chat system the engine does not have.

The entries reach a person through the prompt they shape, and through the conversation_key shown on a phase, which names the external thread a turn served so a reader can go to it.


  • Turn Engine — the phases, and the within-turn prior-work ledger
  • Agent Learning — episodes, the diary, counterparty profiles
  • Event System — the two keys, and inbox coalescing
  • Scaling Out — why shared state lives in the coordination slot

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.