Knowledge System
The knowledge system (internal/knowledge) is the read path agents use to find context they don’t already have in their system prompt. It is two purpose-specific reads composed into the agent runtime:
- Shared knowledge — the team knowledge base. Exactly one backend per company, chosen by
knowledge.backend, behind aknowledge.Searcherseam that every consumer reads through. ASearchertakes plain text — never a backend fragment, never a space key — and answers ranked hits; the turn-start prefetch translates the trigger into that plain text once per turn with the auxiliary LLM, and the executor can re-run the same search itself withsearch_knowledge. agent_diary(a vector column, scanned per agent — there is no vector index) — the agent’s private observation log. One row per declarative fact the agent captured for itself viareflect_and_persist(or that the post-turnPersistDecidersaved on its behalf), scoped to the agent’s id. Rows are embedded on write, and the node holding the seat fills any note left without a vector of the current model; the## Personal memoryprefetch picks candidates via a hybrid selection — the union of a vector top-K (semantic matches to the trigger) and a recency top-K (broadly-applicable operational rules that may not be a topical match), deduped by row id (the two halves are 50 each, so the union is the bound), then handed to an aux-LLM relevance filter.
A diary is read where it is kept current. An agent’s diary is written to
the store of the node running the agent and follows the agent when placement
moves it, so the dashboard’s Knowledge › Agent diaries reads every diary
from the node HOLDING its agent — the list (memory_overview) counts every
agent in the chart at its holder in one round, and one agent’s page
(agent_memory) is its holder’s answer, naming that node. An agent no node
holds shows nothing rather than a copy of unknown age. See
seat ownership.
One backend, and that is a rule rather than a limitation. “What do we already know about this” must not depend on which searcher was asked, so the config refuses a company that wires two.
The two backends
Section titled “The two backends”knowledge: backend: native # the default — the engine's own pages# backend: confluence # a live search against Confluence at query time# backend: none # no knowledge base; every turn gets an empty blocknative | confluence | |
|---|---|---|
| Where pages live | the fleet’s own ordered log, applied into every node’s database | a Confluence site |
| How search works | keyword (BM25 over the node’s own lexical index), semantic (two-stage 1-bit retrieval — an index over the codes, then an exact rerank), or hybrid — both, fused | CQL against the site’s search API, live at query time |
| Who it searches as | the engine — every seat reads every page, so there is no per-seat credential to be missing | the agent’s own user, so Confluence enforces its page permissions natively |
| Staleness | the index is built behind the node’s own applied rows; a node still indexing SAYS SO rather than answering empty | none — there is no local copy at all |
| What it costs to set up | nothing | a site, a space, and a per-seat account |
Native: an index, and what that means
Section titled “Native: an index, and what that means”There is a local copy, and being honest about it is the whole design. Every page change is one record on an ordered log, every node applies it into its own database, and a lexical index is built behind those rows asynchronously — tokenising a large wiki takes minutes, and doing it inline would put the whole index build inside the apply transaction that holds the node’s only writer.
So a search on a freshly joined node can be against an index that is still building, and that is a different fact from an empty company. Both the dashboard and a seat’s own prompt say which: a seat is told “the knowledge base is not searchable from this node yet — ask a colleague rather than concluding nothing has been written down”, because a seat that read an empty result would act on it by writing a page that already exists.
Ranking is BM25 with term-frequency saturation and length normalisation — the part that stops a 20 KB runbook outranking the one-paragraph page that is actually the answer. There is no phrase query, no proximity and no query language, because the seam deliberately does not have one: an agent writes a keyword line and a person types into a box.
Semantic search: two stages, the first one indexed, no new dependency
Section titled “Semantic search: two stages, the first one indexed, no new dependency”Keyword search finds what shares words. Semantic search finds what shares meaning — the question “how do we handle rate limits” against the page titled “429 backoff in the GitLab client”, which share no term at all. That is the class the semantic half exists for, and it is worth naming precisely because it is also the class that is hardest to keep: a document only the semantic half found leaves the fused answer entirely if the semantic half drops it, where a document both halves found merely slides down.
The driver this engine ships has no approximate-nearest-neighbour index —
internal/store/caps.go measures for one once per process and reports what it
found on every open — and the alternatives to building one were a full exact
scan of every vector on every query, or embedding a search library with its
own index format, file and backup story on every node. Instead the search is
two stages, which is the shape every production vector engine uses anyway:
- Stage one ranks a narrow table of 1-bit sign codes — one bit per dimension, 387 bytes at 3 072 dimensions against 12 KB for the vector — and keeps the nearest 1 200 by Hamming distance.
- Stage two reranks exactly those candidates against their full vectors, by primary key, and returns 150.
The narrow sibling table is the load-bearing part rather than a compression detail. A row is stored contiguously, so reading any column of a 12 KB row costs traversing that row’s overflow pages: the identical 1-bit column measures 7.31 µs/row inside the wide row and 0.99 µs/row in a narrow one. A generated column, an expression index and a second column on the vector table all buy the compression; none of them buys the speed.
The score you see is always the exact one. A sign code decides which documents are looked at and never how they are ordered.
Three modes, and the query’s own vector
Section titled “Three modes, and the query’s own vector”Every ranked search takes a mode: hybrid (the default — both halves,
fused), keyword (BM25 alone) or semantic (meaning alone; a screen labels it
“Meaning”). One vocabulary covers the knowledge search and the tracker’s own
item search, on the knowledge and work_search queries alike.
A semantic ranking needs the query in the documents’ embedding space, so
the asking node embeds it once through providers.embeddings — at the model
and width the corpus is embedded at — and sends the vector with the request.
It keeps the last 1 024 query vectors per node, emptied when the model
changes, and bounds one query embedding at two seconds.
The answer says what it actually served. With nothing to rank by meaning —
no provider, knowledge.vectors: false, or a provider that did not answer in
time — a hybrid search serves its keyword half and says served_mode: keyword; a semantic search serves nothing and says why, because a keyword
ranking is exactly the one that cannot find a page sharing no word with the
question. The full table of reasons is in the search guide.
The first stage is an index over the codes
Section titled “The first stage is an index over the codes”Stage one used to read every sign code on every search, which is what capped the corpus one node could search inside the one-second budget. It is now an inverted file built in-tree over the same codes (ADR-0028): the codes are filed in lists by k-means in Hamming space, and a search ranks the lists by its own code and reads only the nearest ones. The rerank above it is unchanged.
- It is replicated state, like the vectors. The embedding duty trains it from the corpus’s own codes, publishes it as a record on the vector log, and every node applies that record — so every node holds the same centroids and files every document in the same list, and a node that adopts a snapshot adopts the index inside it. The arithmetic is integer Hamming distance with a fixed tie break, so two CPUs never disagree about a list.
- How many lists a search reads is measured, never configured. Every training measures recall against the exact scan on 25 documents sampled from the corpus — the evaluation’s own method, with the sampled documents kept out of the training and never counted as their own answer — in every shape a search is issued in: unfiltered, narrowed to each source, and narrowed to one container. It installs the smallest number of lists at which every shape meets the recall floor with no document missing from the top ten. An index that would have to read more than half its lists to do that is not installed: the full scan stays the first stage, because an index reading most of the table costs the scan’s time for the scan’s answer.
- A narrowed search reads more lists, or scans. A search narrowed to one source or to some containers keeps only the rows its filter matches, so its answer is spread over lists an unfiltered search never reads: on the test corpus, reading the unfiltered count of lists recalled 0.949 with three top-ten misses for a search narrowed to the tenth of the corpus that is pages, and under 0.75 for one narrowed to a single container. So a narrowed search reads lists, nearest first, until it has seen as many of its own rows as an unfiltered search reads rows — and runs the full scan instead when that would take more than half the lists, which is what a filter matching few rows near the query needs anyway.
- It follows the corpus. It is retrained when the corpus has doubled or halved since training, and when its fullest list has grown past four times the mean and doubled its share since the training filed it (a corpus full of identical documents piles them into one list whatever the training, and retraining it would change nothing). The duty re-measures it every day against the corpus as it is then, adjusts how many lists a search reads, and retrains it on the spot if no count within half the lists meets the floor any more. A corpus below 1 024 sources has no index at all.
- A new index never hides a document. It is installed by one record and its rows are re-filed by batches of a thousand whose ranges its training fixed, published on the duty’s next tick; until the last batch has applied, a search reads the rows not yet re-filed in full beside the lists it probes.
- It is decided only from a node that has applied the whole log. A node
still catching up would see the index the log has already replaced, or
none, and train over a healthy one, so the duty takes no step there and logs
search_index_behind. - It costs a second copy of the narrow table — a covering index, ≈ 450 bytes a source, about 3 % of what the vectors themselves hold — which is what makes a list one sequential read.
What it buys depends on your corpus, and the benchmark says so rather than promising a factor. Measured at 3 072 dimensions on four cores, p95, on the test fixture’s topical corpus (documents clustered by subject, which is what an index can find):
| 40 000 sources | Full scan | Index |
|---|---|---|
| 1 search in flight | 116 ms | 73 ms (half the lists) |
| 8 in flight | 294 ms | 219 ms |
| Recall against the exact scan | 0.994 | 0.992 |
Half its lists is the most an installed index may read, so this table is the least an index buys. How many a training needs is a property of the corpus rather than of the index, and the fixtures disagree at every size:
| Sources | Topical corpus | Corpus with no topics |
|---|---|---|
| 20 000 | half the lists | every list — no index installed |
| 120 000 | an eighth (recall 0.985) | half (0.966) |
| 500 000 | half (0.993) | every list — no index installed |
Where the training installs nothing the scan answers, exactly as before. The
rule that holds the count highest at scale is no document missing from the
top ten: at 500 000 topical sources the recall floor alone is met by 8 of
2 048 lists, and one top-ten miss among 25 held-out queries keeps the count
at half. That rule is the evaluation’s own definition of passing, so an index
never installs a first stage crewlet search eval would report as failing.
These figures are for a search that is not narrowed — or narrowed to a source that is all of its corpus. A search narrowed to a small share of the corpus reads proportionally more lists, and past half of them it is on the full scan’s figures.
What the quality of this can and cannot be promised
Section titled “What the quality of this can and cannot be promised”A sign code keeps only each vector’s orthant, and how much an orthant says about cosine rank is a property of your corpus’s distribution and of nothing else. Over a family of embedding-shaped generators at one corpus size, recall at the shipped over-fetch spans 0.29 to 0.98. So the engine’s own gate measures the arithmetic — that an exact rerank over a 1-bit candidate pool recovers the exact ranking at sufficient depth, and that an index trained by the engine meets the floor on queries its training never saw — and deliberately makes no claim about recall on your documents.
crewlet search eval is what answers that, against your own vectors. The
ground truth is the exact scan’s own top-K, so nobody authors a judgement.
When the corpus has an index the search is measured through it — what
searches actually run — and again with the full scan as its first stage, so
the report says what the index costs. Each query document is measured
unfiltered, narrowed to each source and narrowed to its own container, each
against its own floor, and never counted as its own answer:
$ crewlet search eval -store /var/lib/crewlet/crewlet-replicated.dbcorpus 118432 sources, text-embedding-3-large at 3072 dimensionsmeasured 25 queries at depth 150 from 1200 candidatesfirst stage the semantic index: 128 of 1024 lists probed (generation 1099511744562)index measured on 118432 sources: 0.9812 against a 0.9800 floor in its worst shape (container:task), 0 head miss(es)recall 0.9761 (floor 0.9312 for this corpus size)worst query 0.9467head misses 0 (documents dropped from the exact top ten)scan recall 0.9803 with 0 head miss(es) — the same search with the full scan as its first stagenarrowed source:page recall 0.9950 floor 0.9800 head misses 0 (25 of 25 scanned) — scan 0.9950, 0 head miss(es)narrowed source:task recall 0.9772 floor 0.9368 head misses 0 — scan 0.9810, 0 head miss(es)narrowed container:task recall 0.9967 floor 0.9800 head misses 0 (19 of 21 scanned) — scan 0.9967, 0 head miss(es)narrowed container:page recall 1.0000 floor 0.9800 head misses 0 (4 of 4 scanned) — scan 1.0000, 0 head miss(es)verdict the two-stage search recovers the exact ranking at the shipped depth, in every shapewindow 8192 bytes a source (the corpus's opening): a semantic search sees each source's title and body up to it, a keyword search the whole bodypast window task 2110 of 104208 sources (2.0%), 7.4 MiB of 196.0 MiB of text (3.8%)past window page 9874 of 14224 sources (69.4%), 402.1 MiB of 518.6 MiB of text (77.5%)It exits non-zero when any shape’s recall is below its floor, or any drops a
top-ten document, so it can go in a schedule. Run it monthly, and after any
change to providers.embeddings.model or .dimensions — those are the two
inputs that move the answer. It reads a file rather than a running node:
point it at the copy inside a backup, which needs nothing stopped and measures
the same rows. -probes N measures what reading N lists would recall
instead of the count the index reads now (a narrowed search still reads more
from there).
The last lines say what a search by meaning cannot see. A source is
embedded as its opening (below), so for each
corpus the report counts the sources whose text runs past that window and how
much of the corpus’s text lies past it: each source’s title and whole body
with the whitespace collapsed, against the opening the duty really sends —
formed, as the duty forms it, from the first 16 384 characters of the body, so
a body that opens on a long run of whitespace (deeply indented code, a padded
table) counts what follows that run as past the window, because the duty never
read it. That text is still
found by the words it uses; it is never found by its meaning, so a semantic
search cannot reach it and a hybrid one reaches it only through its keyword
half. It is the number that decides whether a source needs more than one
vector: a tracker of short items loses almost nothing, a wiki of long runbooks
may lose most of each page. The window is the corpus’s 8 KiB, cut to the
model’s own per-input bound where this build knows a smaller one; the store
does not carry the company’s configuration, so where
providers.embeddings.max_input_tokens or max_batch_tokens lowers the bound
(one input must fit one request, so a request total below the window lowers
the per-input bound too), pass the bound the duty runs at as -window BYTES. This part of the report reads every source’s
whole body, which is why it is a command an operator runs and not a gauge the
engine evaluates.
The floor is a curve rather than a number, because recall from a sign code
decreases as the corpus grows — 0.98 at twenty thousand sources, 0.93 at a
hundred and twenty thousand, 0.88 at half a million. A single threshold would
certify the smallest deployment and say nothing about the largest. A narrowed
shape is judged at the size of the corpus it searches. The index’s own
training and its daily re-measurement are judged against the same curve, and
the ivf_recall_below_floor alarm fires when the
latest one found recall below it even reading every list — which is the full
scan’s own candidate pool, so it means the codes are failing this corpus rather
than the index.
If a run comes back below the floor, the remedy is decided in advance:
- If only the index is below it and the scan recall beside it is not, the
corpus has moved since the index was last measured. The duty re-measures
it every day and, when no count of lists within half of them meets the
floor, retrains it in the same tick;
-probesshows what reading more lists recovers in the meantime. - If the scan is below it too, raise the quantization over-fetch. It measured free in latency, because the rerank’s cost does not depend on how many candidates stage one keeps.
- Failing that, an 8-bit first stage, which is a code change shipped in the same release that moves the model default.
Where the vectors come from
Section titled “Where the vectors come from”Embedding a document costs a provider call, and it is the one thing in the search path a node cannot recompute on its own. So it is not done per node and not on the write path — a page save would otherwise carry a third-party HTTP round trip inside its own transaction. One fleet-singleton duty embeds each source once and publishes a record; every node applies it. The company pays the bill once and holds the answer everywhere.
The duty ticks every minute and embeds at most 1 024 sources in at most 32 provider requests per tick, 128 sources a request at most — so a tick on a caught-up company reads the semantic index’s head and its per-list counts, counts the embedding space and runs one indexed anti-join that returns nothing, and stops; a tick on one that is behind cannot monopolise the provider budget. The one long tick is a training of the semantic index: about 120 µs a source to read every code and make one exact pass, plus a k-means and a filing of every code that run on half the node’s cores, and never fewer than two — so the seats and the searches on a node of four cores or more keep the other half — which comes to a little over three minutes at the largest corpus an index serves. Two is the floor because on a two- or three-core node one worker bought the searches nothing measurable and nearly doubled the training. A node with a single core trains on it, as it always did, and with searches always in flight on that core the same training is about three and a half minutes of reading and seven and a half of arithmetic. The training renews the duty’s lease as it runs, and stops publishing — handing its cores back within a fraction of a second — the moment it cannot. And a tick is cut off once it has gone five minutes without progress, never after five minutes of running: every request answered, every vector published or withdrawn, every 1 024 rows a training reads and the end of its k-means and filing are progress, so a wedged tick never holds the duty while a slow node still finishes its training once rather than starting it again every tick.
What is embedded is a source’s opening. The title, one space, then the
body, every run of whitespace collapsed — and of that, the first 8 KiB, or
less where the model’s own per-input bound is smaller. Collapsed first and cut
second, so indentation and blank lines spend none of the 8 KiB. The figure is a
representation rather than a limit somebody hit: one vector stands for one
source, and 8 KiB of prose is about 2 000 tokens — what a page or a task is
about — while a vector over the whole of a long page about several things
matches none of them well. It is also the most that is provably inside
OpenAI’s 8 192-token window without counting tokens, since no tokenizer emits
more than one token a byte. The keyword half indexes the whole body, so a
hybrid search still finds a passage deep in a long runbook by the words it
uses; a semantic search cannot see past the window at all. A selection reads
only what the opening can need — the first 16 384 characters of a body — rather
than the whole of every page it considers, and the digest each vector carries
is of exactly the bytes the provider received.
A change that leaves the text alone costs no provider call. Every change to a work item moves its version — a status, an assignee, a label, a move to another project — and the duty selects it; but where the vector it already has was computed from exactly the text it would send now, at the same model and width, it republishes that vector under the item’s new version and container rather than asking the provider again. A status change costs one replicated record, and a project move takes the container a scoped search filters on with it. Only a changed title or body is embedded again.
Both source kinds are covered: the tracker’s work items and the knowledge base’s published pages. A page is selected again exactly when what its vector was computed from moved: its body (the page’s own edit number, which a save moves), its title (which a rename or a retitle moves, and which is the first thing the vector embeds), or its container (which a rename across containers moves, and which every scoped semantic search filters on). A renamed or retitled page is embedded again; a page moved to another container under the same title keeps its text, so its vector is restamped under the new container with no provider call — and a scoped search finds it where it now lives, as the keyword half already did. A comment, a watcher change or a re-parent moves none of the three and costs nothing, which is why the selection does not key on the page’s log version, which all of them stamp.
Those requests are the whole company’s, not each corpus’s, and they are handed out round robin between the two, one request at a time. With both behind, each gets half a tick’s requests and holds half its sources in reserve while it has work; with one caught up, the other takes them all. So a tracker being cold-filled — or written to faster than 1 024 items a minute — cannot stop the wiki being embedded, which is the failure the division exists to prevent: a corpus that is never reached is not slow, it is permanently unsearchable by meaning, and the coverage figure below sums both corpora and would report it as merely behind. Equal shares rather than shares weighted by backlog, so how stale a corpus gets depends on its own size rather than on the size of the biggest corpus in the company.
A cold fill of 110 000 sources is roughly 108 minutes and 860 requests —
again across every corpus together, and more requests but the same minutes
where sources run long — and those numbers do not move with the configured
width: providers bill per input token, and dimensions is a truncation
parameter the request already carries wherever the endpoint takes one.
Every request is one the model accepts for its size. The duty forms each request itself, through the provider’s own packing rule: at most 128 sources, and no more than the model’s own limits admit — inputs a request and tokens a request, counted in bytes so no tokenizer is needed (see Configuration). On OpenAI, whose request total is 300 000 tokens, 128 sources of the full 8 KiB are four requests; sent as one, a corpus of code, markup or a script that is not Latin ran past the total and was refused on every tick. Every request is held to a one-minute ceiling of its own — a single embedding is held to fifteen seconds, and a search’s query vector to its own two-second budget.
A source the provider refuses costs only itself. A refusal (HTTP 400, 413
or 422) says the request is unacceptable and not which input, so a refused
request is split in halves, sent ahead of everything else, until the input it
refuses is alone — at most fifteen requests for one input among 128 — and every
half it accepts on the way is embedded as it goes. With the two corpora, each
is guaranteed sixteen of a tick’s requests and half of its sources, so a
neighbour embedding a backlog of short items cannot spend the tick out from
under the isolation and it finishes inside the tick it is met in. Where it
cannot — more corpora, a model that takes few inputs a request, or a rate limit
that ends the tick’s requests partway — the halves a tick did not reach are
kept, the one the failure met in flight among them, and the next tick resumes
the isolation where it stopped rather than starting again from the whole
request. The input refused alone is logged as
search_embed_input_refused, naming the source, the model, the bytes it was
sent and the per-input bound the model’s limits assume (a refusal inside that
bound means max_input_tokens is declared wider than the endpoint enforces, or
the endpoint refuses the text for what it says — and the error says which,
because it carries the endpoint’s own message and codes, such as
code context_length_exceeded, redacted). It is then held back for an
hour: the selection passes over it, so it costs no request and takes no
place in the tick’s 1 024 — a thousand refused sources at the front of the
oldest-first order do not stop anything written after them being embedded.
What a held source still costs is the coverage figure, which counts it as
behind for as long as the provider refuses it. When its hour is up it is
offered again alone, one request and one warning, at most two of a
corpus’s a tick, the longest refused first, so refusals that fall due
together cannot take a corpus’s requests either; a corpus holding more than
the 120 an hour that reaches offers each of them less often than hourly. A
rewritten source is a new text and is offered at once, with its neighbours; a
change that leaves its text alone (a status, a move) keeps it held. The
memory follows the provider’s configuration — the model, the width, the
limits and the endpoint — so changing any of those (a lowered
max_input_tokens, another gateway) forgets every refusal and the fix is tried
at once, while an apply that changes something else, or rotates the key, keeps
it; a duty that moves to another node isolates each one again once. A source
the provider accepts but answers with a vector the duty will not publish —
a component that is not finite, which every search would score a perfect
match, or every component zero, which no search would ever find — costs only
itself too: its neighbours are embedded, search_embed_vector_refused names
it with its bytes, and it is held back for the hour and offered again alone
exactly as a refused one is, where it was left to the next selection and so
sent, and discarded, on every tick. Any other failure — a rate limit, a timeout, a server down, a
credential refused — is about the provider rather than an input, so it ends the
tick’s requests and the next tick asks again; nothing is lost, because the
selection is derived from the rows. When the provider has accepted no
request since the node began embedding with it as configured, and a tick sees
it refuse at least two different inputs sent alone for the first time with
nothing else failing, the duty says so once that tick
(search_embed_every_request_refused): that is the configuration being refused
— a parameter the endpoint does not take, a model it does not serve at that
width — not any document. One refused document, the hourly retry of one already
held, and refusals beside requests the provider accepted are each a document’s
refusal and never blame the configuration. A node that restarts, or takes the
duty over, with two refused documents and nothing else to embed cannot tell
those apart from a refused configuration, and says so as well.
A batch response has to say which input each vector answers. A request
carries many texts, and the API allows the results back in any order — so
each one is filed by the index it carries rather than by where it arrived.
The engine accepts only a response that maps onto the batch exactly once: as
many results as inputs, every index inside the batch, no index twice, and
either every result indexed or none of them. A response carrying no indices at
all is read in arrival order, which is what keeps a compatible server that
omits the field working; one that indexes only some of its results is refused,
because position and index are two different claims about the same result and
nothing in the response says which to believe.
“No index” means the field is absent, and only that. An index that comes
back null, or holding anything that is not a position in the batch, is
refused rather than read as silence — it is a claim the server did make and
the engine could not parse, and taking the arrival-order fallback there would
file a batch somebody deliberately ordered onto whatever turned up first, which
is the one outcome the index exists to prevent.
That is a requirement on a self-hosted endpoint or a gateway rather than on OpenAI itself, and the refusal is why it is stated: a response that repeated an index would store one document’s vector against another — a wrong search answer with nothing left to trace it to, since a stored vector carries no evidence of the text it came from — and leave a third document with no vector at all, which the duty reads as “nothing to embed” and re-selects on every pass for ever. So a non-conformant endpoint surfaces as coverage that never rises, the refusal named in the log, and eventually the alarm below; it never surfaces as a corpus that is quietly wrong.
How much of the corpus is covered is published, as
crewlet.tracker.vector.coverage — the fraction of sources carrying a current
vector, summed across both corpora rather than averaged, so a small fully
embedded corpus cannot mask a large uncovered one. The
recall_below_floor alarm fires below 95 %, which is
how a stalled backlog is reported: it never drops a seat, and semantic recall
answering from a corpus it does not cover has no other symptom. A company with
no embeddings configured measures nothing rather than zero.
A model change at the same width is the case to know about. Until the
refill finishes, the corpus holds two incompatible embedding spaces, and a
search filters on the pair — so documents still on the old model are not
ranked badly, they are simply not in the candidate pool. crewlet search eval
names every space it finds, which is how you see a refill in progress.
Every document carries a bucket
Section titled “Every document carries a bucket”Both halves of the native backend stamp each document with a search shard — a stable hash of its own identity into 64 fixed buckets, written beside the row by the same function in both estates.
It is unrelated to everything else the engine partitions on. Not a stream, not a project, not a container, not a log position, not the node holding the row. That independence is the point:
- A source that moves keeps its bucket. Bucketing on the filed project would re-bucket every task a re-file touches, and a search over the old bucket would miss it — silently, because a result set that is one document short looks exactly like a corpus that is one document short.
- The division follows the corpus, not the company’s shape. Bucketing on a project puts the busiest project in one bucket.
- A document with no embedding is still in a bucket, because the bucket is a function of the id rather than of anything derived from it.
A single node reads every bucket, exactly as it did before the column existed. A fleet above 10 000 documents divides them: each live node scans a contiguous range, returns its best candidates with scores, and the asking node merges. There is no routing plan to compute and nothing to configure — the division is a sorted roster and a remainder, and every node computes the same one. See Search for the merge order, why BM25 stays comparable across a divided scan, and what a partial answer names.
The count is fixed for the life of a deployment and is deliberately not configurable. Changing it re-buckets every document, which costs a full index rebuild rather than a rebalance — a knob that can be turned exactly once, at the price of the whole index, is one that gets turned by somebody who did not know the price.
Confluence: no local copy at all
Section titled “Confluence: no local copy at all”Shared knowledge is read straight from the backend on demand, so there is no sync worker to run, no index to keep fresh, and no staleness window. It authenticates as the agent’s own user, which is what makes the backend’s own permissions the ones that apply.
There is no shared vector index and no scope ladder for shared docs on this backend.
Data flow
Section titled “Data flow”The two reads are independent: the diary is read by hybrid candidate selection (vector top-K ∪ recency top-K → aux-LLM relevance filter), scoped to the calling agent; the knowledge base is read through whichever backend the company wired — this node’s own applied rows natively, a live query at the site on Confluence — scoped to the role’s accessible containers. Neither depends on the other, and each renders into its own block of the executor’s prompt.
The knowledge.Searcher seam
Section titled “The knowledge.Searcher seam”internal/knowledge defines the one seam between the agent runtime and the knowledge backend:
type Hit struct { Title string URL string // shareable human link; "" when unbuildable Container string // Confluence space key PageID string Snippet string // plain text, <= 200 chars; may be "" Ancestors []string // ancestor page titles, outermost first}
type Query struct { Text string // plain language, never a CQL fragment Seat *org.Role // whose credential the search runs as Org *org.Organization // supplies the read scope, per call Limit int // 0 takes DefaultLimit (8) ExcludeAncestors []string // nil takes the auto-draft default}
type Searcher interface { // Backend names the integration answering, for logs and for the // operator surface that reports which one a company wired. Backend() string
// CanSearch is the cheap, no-I/O pre-gate. CanSearch(seat *org.Role, o *org.Organization) bool
// Search returns up to Query.Limit ranked hits, and what the search // did: the mode it served, the modes this backend can serve, what part // of the fleet it covered, and why it served less than was asked. It // never reports an error: every failure path is an answer with no hits. Search(ctx context.Context, q Query) Result}A search that did not run says so. The native knowledge base is searched on
a data node — this one when it has the data role, another data node
otherwise — and a search no data node could run is answered with no hits, no
mode served and none of the corpus covered, never as a plain empty list:
“nothing matched” and “the knowledge base could not be searched” send a seat to
different places. The turn-start block and search_knowledge say the second in
words.
Contract semantics every backend honors:
- Scope lives behind the seam.
Searchderives its container scope from the organization (knowledge.scope); callers pass a role, a plain-text query, and ancestor-title exclusions — never CQL fragments, space keys, or project lists. Because the organization is a per-call parameter, live config edits toknowledge.scopeflow through with no engine refresh hook. - Unscoped-vs-nothing is enforced inside
Search: empty scope + a self-authenticating role ⇒ unscoped search (the backend’s own ACLs bound the hits); empty scope + a credential-less role ⇒ no results. CanSearchis a cheap, no-I/O pre-gate — “could a search possibly hit anything?” Its only job is letting the relevant-knowledge prefetch skip the aux-LLM query-generation call when the search is a guaranteed no-op.- Best-effort, never silent:
Searchnever reports an error; every failure path returns no hits, and a search that did not run serves no mode — so the turn-start block (UnsearchedKnowledgeHint) andsearch_knowledgeboth say the knowledge base “could not be searched” rather than rendering an empty block or “nothing surfaced”. What it does not do is fail quietly: theResultcarries anOutcome—ServedMode,Modes,Coverage{nodes, complete, buckets_missing}and aDegradedreason — so a caller can tell “nothing matched” from “part of the corpus was not scanned” and from “this could not rank the way it was asked”.search_knowledgesays the second to the seat in words, and says “the knowledge base could not be searched just now” for a search that never ran rather than “no team documents match”. Query.Modeishybrid(the zero value),keywordorsemantic— see modes. A backend that cannot rank that way says so in the outcome; Confluence answersmodes: [keyword].Query.ExcludeAncestorsdrops hits whose ancestor/parent chain matches any listed title. Left nil it takes the default,"Auto-Drafted Skills"(knowledge.AutoDraftedParent), so unreviewed promotion drafts never surface before a lead publishes them; an empty, non-nil list disables the exclusion. Every draft title also carries the[Auto-draft]prefix (knowledge.AutoDraftTitlePrefix) as a fail-closed backstop for a backend whose parent lookup fails.
Selection is by knowledge.backend, and single-homed. Engine start constructs exactly one searcher: the native one over this node’s own page index, or the Confluence one, or none. One knowledge home is what makes the turn-start prefetch, the search_knowledge builtin, onboarding hints and skill promotion agree about what the company knows — two searchers would make an agent’s answer depend on which was asked, and neither would be wrong. With backend: none, the searcher stays unwired and the ## Relevant knowledge block renders empty. A live config change re-points the running turn engine at the new searcher (or at none).
An empty backend derives rather than defaulting blindly: a company that declares integrations.confluence gets confluence, and one that declares nothing gets native. That is the compatible half of the rename — an Atlassian company that has not read this page keeps the backend it had, and a quickstart company gets a wiki without asking for one.
The seam has two implementations, and it was written for the second. knowledge.Searcher is declared by its consumers — the prefetch, the onboarding hint, the promotion pass — so the native backend arrived as a new implementation rather than a rewrite of everything that searches. A seam collapsed into its last backend is what makes the next one a rewrite.
Native backend — the engine’s own pages
Section titled “Native backend — the engine’s own pages”internal/pages + internal/search. The knowledge base is a state-log domain: every change is one record on CREWLET_PAGES_LOG, a deterministic applier writes it into every node’s replicated database, and a lexical index is built behind those rows. The log’s byte ceiling is stream.pages_log_max_bytes, reserved beside the tracker’s and the vector index’s inside one budget (see how the byte ceilings are sized). A search is BM25 over that index: term-frequency saturation and length normalisation, so a long runbook that mentions a word thirty times does not outrank the short page that is about it.
A search here is hybrid by default: the query is embedded once through the company’s embeddings provider, and the BM25 ranking and the two-stage semantic scan are fused by reciprocal rank fusion. With no provider, or with knowledge.vectors: false, it is BM25 alone — and the answer says served_mode: keyword with degraded: no_embeddings rather than presenting the words as the whole of it.
Two properties differ from the vendor path and both are visible:
-
Every seat reads every page. There is no per-seat credential, so
CanSearchreduces to “is there an index at all” — the credential-less case below does not arise. -
An index that is still building says so. It is a different fact from an empty company, and a seat is told which: “the knowledge base is not searchable from this node yet — ask a colleague rather than concluding nothing has been written down”. A seat that read an empty result would act on it, by writing a page that already exists. The gate is this node’s FIRST BUILD — one lap over every corpus — and not “nothing is waiting to be indexed”: a page saved a moment ago is ordinary staleness, and reading the gate off a pending count made every empty search on a company with people in it answer “still building” instead. After the first lap a search is a true answer over slightly older rows, which is what a search always is.
-
A CONTAINER IS A DOCUMENT, and the engine writes one for every
space:the org chart names — a unit’s, a seat’s own — plus the two reserved ones, on every config apply and every boot. It is idempotent: a container whose row already says what the chart says is not written again, so the log grows with edits rather than with restarts.Its name and purpose follow the configuration activated last, whichever node applies what when. The row carries the instant its configuration was activated (
chart_epoch, in Unix milliseconds), and a configuration activated earlier never overwrites one activated later — so a node restarting on a revision the fleet has since replaced leaves the newer names alone, exactly as the chart’s projects do (see the work tracker). A later activation over unchanged settings is still written — a re-stamp — because a row left at the older stamp is open to any activation between the two.A page merely names its container, so a page can exist in a container with no document — it is reachable by address and by search, and it is missing from
GET /containersand from the Knowledge tree. That is what a space nobody declared looks like. -
THE DASHBOARD BROWSES IT AS A TREE, one level at a time: a space’s top (
pages{container, roots: true}) and one page’s children (pages{parent}), each read in windows of 500 with the listing’stotaland anaftercursor, and each page saying how manychildrenthe same listing would show under it — so an expander never opens onto a folder whose pages are all in the trash. A search there reads THIS NODE’S OWN COPY (the rows, the lexical index and the replicated vectors), and the screen says so: it is as current as this node’s place on the pages log. See the dashboard’s Knowledge. -
A PAGE KNOWS WHO READ IT AND WHO LINKS TO IT. Every read a seat makes is a recorded
knowledge_read, and the replicatedusagedomain keeps it per company day, sopage_readsnames each seat, how the page reached it and in which turn — the same answer on every node, a departed node’s reads included. Backlinks are this node’s: the lexical indexer extracts every page id a body links to (pages.Links—/pages/<id>or the dashboard’s#/knowledge/pages/<id>, outside code) intopage_linksin the node estate, beside the index it is derived with, and a task links a page either as a page relation or by an address in its description. A link is BY ID, never by title — a title is an address a rename moves — andwrite_pageandsave_pagetell a seat so on theirbodyparameter, with an example a test holds againstpages.Links:[its title](#/knowledge/pages/<page id>). A[[CONTAINER/Title]]wiki link is not a grammar the engine reads; it is drawn as the brackets it is and counts for nothing in “Linked from”. A build that changes what the index derives re-derives every older row once, on its own (search.IndexDerivation), with no rebuild command to remember. -
THE DASHBOARD EDITS AS THE PERSON. A save, a new page and a comment go through
/operator/actbound to the signed-in seat, so the page’s history names who wrote it — never “the dashboard”. A save states the revision it edited; one against a stale revision is refused rather than overwriting prose somebody else just wrote. -
A BODY HAS A HISTORY, and revision N is the body at version N. The dashboard reads any one of them back and shows what a save changed against the version before it, by line.
-
A body is MARKDOWN, and the only format — see
internal/pages. The dashboard renders it (headings, lists, tables, code, links), with raw HTML shown as its own text and a link’s scheme restricted tohttp(s):,mailto:and the app’s own routes: a page is written by an agent acting on content it read somewhere else, so a href in one is untrusted input. -
A title is an ADDRESS. It is unique within its container, claimed first-writer-wins on the fleet, and a page is fetched by
CONTAINER/Titleas readily as by its id. The address is the title NORMALISED — lowercased and with runs of whitespace collapsed — soENG/deploy runbookreaches a page called “Deploy Runbook”, and two people cannot create pages whose titles differ only in spacing or case. -
The address and the displayed title are two things, and a rename can move either. The page stores both: the address it is claimed at, and the title as its author capitalised it — a link is resolved by the first and rendered by the second. So renaming “Deploy Runbook” to “Deploy Guide” moves the address (the old one is freed and can be taken again), while renaming it to “DEPLOY RUNBOOK” leaves the address exactly where it is and changes only what every reader sees. Both are real changes: each writes a revision to the page’s history and tells its watchers. Only a rename to the title the page already displays does nothing — and it reports success, because it has already happened.
-
A container is a TREE, and a parent has to keep it one. A page’s
parentmust be a page that exists, in the same container, not in the trash, and neither the page itself nor one beneath it. Anything else is refused when it is written, namingparent_idand what to do instead — because the tree is read from the top down, and a page filed under a page that is not there, or two pages filed under each other, are pages no walk of the container ever reaches. An empty parent is the container’s top. Two moves that are each sound can still race — A under B on one node, B under A on another — and so can a move and its parent’s purge. Every node then files the page where it can, identically: a move whose parent it can no longer hold leaves the page where it was, a create lands at the top, and apages_parent_salvagedwarning names the page. Purging a page re-files its children under its own parent (or the top) rather than leaving them pointing at nothing. -
A save’s note is the one its writer gave. The one line
save_pagetakes asmessageis what the page’s activity and each watcher’s wake show under the saver’s name; a save without one shows none. It never falls back to a line of the page — the opening heading is the same before and after an edit anywhere below it, and shown there it reads as a note nobody wrote. What changed is the revision’s to say, and History shows it by line.
The tool-skills container is excluded from every result. A tool skill is machinery the engine injects into a phase, and a seat told to read one as knowledge would follow it as an instruction. The exclusion costs a result that page and never a place: the search walks its fused ranking past every page it does not return, so a company whose skills lead a ranking still gets a full answer.
Confluence backend — the Confluence searcher
Section titled “Confluence backend — the Confluence searcher”internal/confluence (confluence.Searcher). The query text is wrapped into a Confluence CQL text ~ "..." clause (confluence.BuildCQL), optionally narrowed by space IN (...) from the read scope, and run against the Confluence REST API (/rest/api/content/search), so Confluence’s own search backend does the matching and the relevance ranking. Authentication is as the agent’s own Atlassian user, using the seat’s Confluence credential from its mcp_env (the atlassian or confluence server entry, read by atlassian.CredentialOf, which accepts CONFLUENCE_API_TOKEN, CONFLUENCE_PERSONAL_TOKEN, CONFLUENCE_TOKEN or ATLASSIAN_API_TOKEN, and a JIRA_API_TOKEN on the shared atlassian entry). Confluence enforces its page permissions natively: a restricted page the agent’s user cannot see simply doesn’t come back, and there is no engine-side restricted-page handling. Seats without their own credential fall back to the org token (integrations.confluence.token); an agent on the org token sees whatever that account sees, subject to the empty-scope rule below. Hits carry the full ancestor-title chain, so the auto-draft exclusion filters on any depth.
Accessible containers
Section titled “Accessible containers”The search scope is set by one thing: the org-wide knowledge.scope list, normalised once by internal/knowledge —
knowledge.scope: ["HANDBOOK"] # scoped to these containersknowledge.scope: [] # empty ⇒ unscoped / ACL-bound for # self-authenticating agentsIt is role- and unit-independent — every agent has the same read scope.
Read scope ≠ team identity. A unit’s own container — space (runtime org.Unit.Space) — is integration identity: it decides webhook routing (page activity → the unit lead) and is the team’s write / skill-promotion home. It deliberately does not narrow reads. An Engineering agent isn’t limited to the ENG space when searching; it searches across everything its own account can read. (See Confluence § integration identity.)
The list is optional — and empty is the useful default. When the scope list is empty, behaviour depends on how the search authenticates (per-agent token vs. engine/admin fallback):
- A role with its own backend credentials searches unscoped: the container clause is dropped and the backend’s own ACLs bound the results — the agent finds anything its account can read that matches the query.
- A credential-less role (engine/admin-token fallback) searches nothing: an unscoped query would read the shared account’s entire view, so the empty list means “no search” rather than “everything”.
So set knowledge.scope only to narrow reads to a curated floor (e.g. a company handbook); a fully per-agent-credentialled org leaves it unset and lets the backend’s ACLs do the scoping. The backend’s own permissions remain the hard boundary regardless.
Single-homed, and validation enforces it. A company that sets knowledge.backend: native and declares integrations.confluence is refused, because “what do we already know about this” would depend on which searcher was asked. Also refused: a read scope with no backend behind it — knowledge.scope under backend: none reads as a working narrowing and narrows nothing.
Where content comes from
Section titled “Where content comes from”Shared knowledge is the backend — there is no separate engine-managed store to populate. The writers feeding the two reads:
| Writer | Read by | Reach |
|---|---|---|
Humans + agents via the backend’s MCP tools (confluence_create_page / confluence_update_page / confluence_add_comment) | knowledge.Searcher (live query) | Whoever the page’s backend permissions allow |
crewlet confluence import (below) | knowledge.Searcher (live query) | Same |
Agents via reflect_and_persist (in-flight) and PersistDecider (post-turn) | agent_diary (hybrid vector ∪ recency selection → aux-LLM filter) | The writing agent only |
Static org configuration (mission, vision, policies, role profile, team roster, unit context, integration hints) is a third source, but it is not “knowledge” in the read-path sense: it renders straight into the executor’s system prompt via the section builders in internal/agent/prompts. There is no startup seed step and no reconcile pass, because the prompt is the configuration. Documents that change frequently (procedures, ADRs, runbooks) live in the knowledge base, where humans and agents already author them.
Answering a question
Section titled “Answering a question”The same search also answers a person’s question: the dashboard’s ⌘K answer
block calls the operator tool
answer_knowledge,
which writes a short Markdown answer from what the company has written down and
lists what it was written from.
- From the sources and nothing else. The model is told to say only what the
numbered sources say, to cite each claim as
[n], and to say so in one sentence when they do not answer the question. Source[n]is the n-th entry of the answer’ssources, so a screen links every citation. A question nothing matches is answered without a model at all, and costs nothing. - A person’s, not a seat’s. Only a token bound to a person may ask — the
spend is on somebody’s behalf — and no seat is given the tool: a seat has
search_knowledgeand its own model, and a second model’s summary in its context would be one it could not check. - Charged to the company. It runs on the asker’s own seat’s auxiliary model and is judged against the company’s day, week and month — a person has no seat budget. A company window with no room refuses it before the call; after the call, exactly what the reply spent is recorded on the company’s counter — a reply the model refused included, since its prompt was billed even though the answer is refused rather than shown.
- Cached at a corpus position. Each node keeps 256 answers, keyed on the
question (case and spacing folded) and where its tracker, pages and vector
logs are applied through. Any write that could change an answer moves the
position and retires every answer cached at the old one; a repeat spends
nothing and says
cached: true. A company on Confluence has no position to key on — the wiki changes without the node hearing — so its answers are never cached.
Publishing knowledge docs
Section titled “Publishing knowledge docs”Most shared knowledge is authored directly by humans and agents — a seat calls
write_page (native) or the vendor’s own MCP tools (Confluence). For docs an
operator wants to keep in version control — onboarding pages, runbooks,
playbooks — there are two paths, one per backend.
On the native backend: your own assistant
Section titled “On the native backend: your own assistant”There is no import CLI, and there does not need to be one. Point any MCP client
at /operator/mcp
with your API token and tell it what to publish:
Publish everything under examples/nimbus-docs/ — one container per
directory, the page title from each file’s first # H1.
It calls write_page per file, with your token’s own name on each page as the
author. That handles the parts a flag-driven CLI handles badly: the parent
chain, a title that already exists (save_page with the version it read), and
a file that turns out to be a tool skill rather than prose.
The reserved containers are refused to it, exactly as they are to a seat — a page written into the tool-skills container would be injected into a phase as an instruction rather than read as knowledge.
On Confluence: the import CLI
Section titled “On Confluence: the import CLI”crewlet confluence import <company.yaml> <directory>The first positional argument is the Tier B company YAML (the importer reads the backend credentials from its integrations.confluence block); the second is the directory to publish, walked recursively, with dot-directories skipped. The importer routes each .md file by frontmatter: a file with a trigger: is a Tool Skill (published to the Tool Skills container); every other file is a knowledge doc.
Knowledge docs follow a directory-based convention — the files are pure prose, no frontmatter required:
- Container = the file’s immediate parent directory name. A file at
<root>/ENG/onboarding.mdpublishes to the containerENG— a native container, or a Confluence space of that key. - Title = the file’s first
# H1heading. That H1 line is stripped from the published body (the backend shows the page title separately, so leaving it would duplicate the title on the page).
examples/nimbus-docs/ is a worked set of these: the pages the Nimbus
example company publishes. Both bundled examples run the native
knowledge base, so neither is an argument to this CLI: they publish
through the assistant path above. A company that moves to Confluence adds
the confluence: block per Confluence,
and its company YAML is then the positional argument the importer reads
those credentials from.
examples/nimbus-docs/├── HOME/│ └── Onboarding.md → the ROOT container: the page every seat│ reads before its own team's├── ENG/│ └── Onboarding.md → space ENG, title "Onboarding"├── LEAD/│ ├── Onboarding.md → space LEAD, title "Onboarding"│ ├── Repo Ownership.md → space LEAD, title "Repo Ownership"│ └── Manager 1-1.md → space LEAD, title "Manager 1:1"└── PROD/ └── Onboarding.md → space PROD, title "Onboarding"HOME is knowledge.root_space’s default. It holds what is true of the
whole company, which is what makes it worth a container of its own: the
alternative is the same four paragraphs in every team’s Onboarding, where
three of the four copies go stale and nobody can tell which.
# Onboarding
## Who is here...Optional frontmatter is supported only for overrides — a plain-prose doc needs none of it:
| Field | Required | Description |
|---|---|---|
title | no | Overrides the H1 as the page title. (Onboarding pages must be titled exactly Onboarding — see the onboarding convention; name the file Onboarding.md and that falls out of the H1 automatically.) |
space | no | Overrides the container for a doc that lives somewhere the tree cannot express. |
parent | no | The title of a page in the same space to nest this one under. The plan is ordered parents-first, so a parent published by the same run resolves; a parent: cycle stops the walk naming the files. A parent nobody publishes is a note and a page at the space root — a doc nobody can read is worse than a doc in the wrong place. An existing page is never re-parented: where a page sits is something people move deliberately, and a run that dragged it back every time would be fighting them with no way to say so. |
labels | no | The author’s own page labels, lower-cased and de-duplicated because that is what Confluence stores and answers with. Attached on every run, not only on create; a label that will not attach is a note, not a page failure. |
A doc with neither a frontmatter title: nor an # H1 has no determinable title and stops the walk naming the fix. So do two files that would publish as the same page. Both are things an operator corrects in their editor, and a run that skipped them would report success with a doc silently unpublished.
Key properties:
- Clean prose. Knowledge-doc pages render the markdown body straight to the backend’s page format (Confluence storage XHTML) — no YAML metadata box on the page (unlike skill pages, which carry a binding-metadata code block the engine parses back out).
- Idempotent, and always a write. A page that exists is updated in place. This is a publisher, and skipping existing pages would mean an edited file never reaching the backend.
-dry-runpreviews without page writes. The match key is(space, title). The importer never creates its container: a missing space fails before a single page is written, naming the space to create. - Searched live, not registered. Knowledge docs are read on demand through the query-time
## Relevant knowledgesearch — they are not loaded into any in-memory registry, so there is nothing to resync (crewlet confluence resyncis skills-only).
Wiring it up
Section titled “Wiring it up”There is no orchestrator object to construct. The two reads are wired independently by engine start:
- The
knowledge.Searcheris constructed from whichever backendknowledge.backendnames (see the seam): the Confluence searcher, which needs the site connection and nothing local; or the native one, which needs this node’s own store and its lexical index. The native one fuses the semantic half wheneverknowledge.vectorsis on, embedding each query throughproviders.embeddings— a model that produces vectors, not an LLM. Neither takes an LLM — writing the query text is the prefetch’s job, on the seat’s auxiliary model, andsearch_knowledgehas the executor’s own words to search with. Withbackend: none, or aconfluencecompany whose integration is missing, no searcher is wired and the## Relevant knowledgeblock stays empty. learning.Diaryis built over the node’s store (learning.NewDiary), so a node with no store has no diary and the## Personal memoryblock stays empty without error. Writes are embedded whenproviders.embeddingsis configured, and the node holding a seat fills the notes written without a vector or under another model; without it the diary degrades to a pure recency list (vector candidate selection becomes a no-op) but writes and recency reads still work.
The two are independent: an org can have knowledge search without reflection, or reflection without knowledge search.
Relevant-knowledge prefetch
Section titled “Relevant-knowledge prefetch”Beyond agents calling the backend’s search tools directly, the executor’s prompt carries a ## Relevant knowledge block that pre-runs a knowledge-base search for the seat. Once per turn, the auxiliary LLM generates a short plain-text search query from the trigger context, the searcher runs it live (scoped to the role’s accessible containers), and the executor sees title + snippet bullets without having to think to call a tool. When the block is thin — a pointer trigger gates the search off — the executor searches the same seam itself with the search_knowledge builtin, once it knows what the task actually needs. The CanSearch pre-gate skips the aux-LLM call entirely when a search could not return anything. Because the search runs as the agent’s own backend user, restricted pages the agent cannot see never appear — there is no draft-page or restriction filter to apply engine-side.
Full page bodies open via the backend’s page-read MCP tool; further searches via its search MCP tool (confluence_get_page / confluence_search).
See Agent Learning § Relevant-knowledge prefetch for the design rationale and failure modes.
What agents read
Section titled “What agents read”Every time a seat reads from the knowledge base the engine records it, as one knowledge_read event per act. A page’s history says who wrote it and a search’s ranking says what it matched; neither says whether anybody opened it, and “which runbooks does this company’s staff actually read” is the question a knowledge base is curated against — a page nobody has read in a quarter is one to retire, and a page every turn is given is one whose mistakes are expensive.
A read names the pages it reached (pages[{id, container, title, rank}]), the backend their ids are addresses in, the seat, the turn and the phase, and how the pages arrived (via):
via | What happened | Phase |
|---|---|---|
get_page | The seat read one page in full with get_page (native backend). | the phase that called it |
search | The seat ran search_knowledge; pages are the hits it was shown, ranked, and query is what it searched for. | the phase that called it |
prefetch | The turn-start ## Relevant knowledge block put these pages in front of the seat; query is the auxiliary model’s. | none — it runs before the first phase |
skill_loaded | The seat loaded a tool skill’s body with load_tool_skill; the page is the one the skill was read from. | the phase that loaded it |
skill_injected | A phase’s prompt carried the tool-skill catalogue; pages are the pages behind each skill it listed. One read per rendered catalogue — the executor’s, the reviewer’s, and each delegate worker’s. | the phase whose prompt it was |
Three rules keep the record honest:
- One event per act, listing its pages. A search that showed six pages is one read with six pages, because a page’s rank is a fact about that search and six separate rows could not say what else was on the list.
- Nothing is recorded for nothing. A search with no hits, a page that was not found and a catalogue that listed no skill publish no event — a read of nothing is not a read, and a row with an empty page list would count toward every “read today” total.
- Only a seat reads. The operator surface serves
get_pageandsearch_knowledgeto people too; their calls are audited by the operator surface’s ownoperator_actedrecord and are never aknowledge_read.
A search’s query is clipped to 200 bytes on a character boundary (with a trailing … when it was cut): it is a label a reader recognises a search by, not an input anything re-runs. The event is stored under the learning category, beside skill_used; see Event System for the catalogue entry.
Onboarding markers
Section titled “Onboarding markers”The mark_onboarded builtin records that an agent has read its team’s Onboarding pages so the onboarding hint stops re-rendering on every turn. Markers live in their own small table, agent_onboarding_markers, keyed by agent_id with UPSERT semantics (so re-onboarding never accumulates stale rows). The marker carries a chain_hash over the agent’s org chain (learning.ChainHash); a chain change (role moved units, ancestor renamed, new unit inserted) silently invalidates the marker, and the dedicated onboarding pass runs again until the agent re-reads and re-marks. The same row holds the cross-process pass lease that stops two turns onboarding one seat at once.
A dedicated table, because learning.Onboarding.Onboarded answers with one indexed equality lookup instead of a per-agent metadata-filter scan.
Configuration
Section titled “Configuration”The knowledge system has a block of its own — knowledge.backend, knowledge.scope, knowledge.skills_container, knowledge.root_space and knowledge.vectors, field by field in Configuration. Two upstream configs determine the rest:
integrations.confluence— required bybackend: confluence, and refused besidebackend: nativebecause pages would then live in two places with nothing keeping them in step. The query-time search authenticates with each role’s per-agent token (mcp_env.atlassian), falling back to the org-level token (confluence.token); aconfluencecompany missing it has no searcher at all, so the## Relevant knowledgeblock stays empty and only the agent’s diary contributes. The native backend needs none of it — it searches as the engine, over this node’s own applied rows.providers.embeddings— required for the diary’s vector candidate path (the vector half of the## Personal memoryprefetch’s hybrid selection, plus the diary write-side embedding step), forepisodesvector recall in the learning subsystem (query_episodesand the## Similar prior workprefetch), and for the semantic half of the native knowledge search.knowledge.vectorsis the switch that fuses that half into the query, and it derives from whether this block is configured — so a company already paying for embeddings for its diary gets the better search, and an explicitvectors: truewith no provider is refused at validation rather than degrading quietly. An explicitvectors: falseturns off everything that embeds the corpus at once: the embedding duty stops (no provider bill for documents), the coverage gauge measures nothing, and every search is keyword and says so — the diary and episode recall keep using the provider regardless. Without an embeddings provider the native search is lexical only (BM25 over this node’s own index), the diary degrades to its recency-only path (still functional, just without semantic candidate matching), and episodic recall is disabled. Onbackend: confluencethe question does not arise: that search is a live CQL query against the site, which embeds nothing either way.
See Configuration for the full YAML shape, Confluence integration for setup, and Agent Learning for diary mechanics.
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.