Skip to content
You are reading documentation for unreleased main. This page is not in 0.1 yet.

Search

How the knowledge search is answered, what it costs, and what a fleet can do about it when the corpus outgrows one node’s CPU.

Search is answered from this node’s own tables. Every node holds the whole corpus — see Scaling — so there is no lookup that has to leave the machine and no answer that depends on a peer being up. What a fleet can divide is the work, not the data.


The two halves, and the bucket both of them carry

Section titled “The two halves, and the bucket both of them carry”

A hybrid search runs two rankers and fuses them:

  • Lexical — BM25 over the engine’s own inverted list, in this node’s database.
  • Semantic — a two-stage vector search over the replicated vectors, whose first stage probes the corpus’s semantic index when it has one.

Every indexed document carries a search shard: a stable hash of its own identity into 64 fixed buckets, written beside the row in both estates by the same function. See Every document carries a bucket for what the bucket is a function of and — just as important — what it is deliberately not.

A search can be told to read only part of that range. On a single node nothing tells it one: it holds every bucket and reads every bucket.


Three modes, and what an answer says it served

Section titled “Three modes, and what an answer says it served”

Every ranked search — the knowledge search and the tracker’s own item search alike — takes one of three modes, with one vocabulary on every surface:

ModeRanks byLabel on a screen
hybrid (the default)both halves, fused by reciprocal rank fusionHybrid
keywordthe words the query used — BM25 aloneKeyword
semanticwhat the query means — the vector scan aloneMeaning

The two halves do not read the same text. The keyword half indexes a source’s title and its whole body. The semantic half ranks one vector per source, computed from its opening — the title and then the body, whitespace collapsed, the first 8 KiB of the two together, or less where the model’s own per-input bound is smaller (see Where the vectors come from). So a passage deep in a long page is found by hybrid and keyword through the words it uses, and never by semantic, which cannot see past the window. crewlet search eval reports, per corpus, how many sources and how much of their text lie past it (see What the quality of this can and cannot be promised).

A semantic ranking needs the query in the same embedding space as the documents. The asking node computes that vector once through the company’s embeddings provider — the same model and width the corpus is embedded at — and sends it with the request, so no participant makes a provider call of its own. It caches the result: 1 024 query vectors per node (about 12 MiB at 3 072 dimensions), least recently used first out, emptied the moment providers.embeddings.model changes so a vector from the retired model is never ranked against rows the refill is replacing. A phrase the ⌘K palette embedded is a cache hit for the same phrase searched as knowledge, and for the turn-start prefetch. Computing a query’s vector is bounded at two seconds — twice the one-second budget of the scan itself — so a slow provider costs a search its meaning half rather than holding the person who asked.

A query is at most 400 bytes, on every surface that takes one: a seat’s search_knowledge, search_work_items and answer_knowledge, the operator’s tools, the API’s knowledge and work_search (bad_params, naming the size and the limit) and the dashboard, which sends nothing past it and says why. A longer one is refused, never cut — a query is a question, and a search on its first part answers another one; past four hundred bytes it is a pasted thread, which no ranker turns into a better search. This is the opposite of what the corpus does with a long page, and deliberately: a source is embedded as its opening because one vector stands for one source, while a query’s vector is always computed from the whole query. Every model this build knows takes at least 2 032 bytes an input; on a model stated with a narrower window (max_input_tokens), a query longer than one input is embedded in pieces at the model’s bound, in one request, and the pieces’ vectors pooled.

A mode asked for is not always a mode served, and every answer says which. It carries served_mode (the ranking the hits actually came from), modes (what this backend can serve as asked, right now), degraded (why those differ from what was asked) and coverage (below):

degradedWhat happenedHybrid servesSemantic serves
no_embeddingsNo embeddings provider, or knowledge.vectors: false.its keyword halfnothing
embedding_failedThe provider did not produce the query’s vector in time. Nothing to configure; the next search asks again, and a failure is never cached.its keyword halfnothing
semantic_partialPart of the fleet ran without its vector scan.both, with meaning over less of the corpusmeaning over less of the corpus
unsupportedThe backend has no such ranker — Confluence, whose own CQL search is keyword.the keyword answernothing

Semantic never falls back to keyword. Somebody asking for meaning is asking for the pages that share no word with the query, and a keyword ranking is the one ranking guaranteed not to find them — so it answers no hits and says why, rather than a keyword answer labelled as meaning. Hybrid does fall back: the words are half of what was asked for.

The modes on offer are known before anybody types. A search with an empty phrase is the seam’s PROBE: it runs nothing and reads nothing, and answers modes and — for the mode asked — the degraded the CONFIGURATION decides (no_embeddings, unsupported; never a transient embedding_failed, which is a property of one search). The dashboard’s Knowledge tree asks it for semantic and disables a mode the answer does not list, with the reason written under the control, rather than letting a reader pick Meaning and find out from a degraded answer. The same answer sets the dashboard’s DEFAULT: Hybrid where it is served, otherwise the first mode listed — Keyword on a company with no embeddings provider — so a search nobody chose a mode for runs in a mode the engine serves as asked instead of reporting a degradation of a choice nobody made. The wire default is unchanged: a request that names no mode is still hybrid.

A keyword search sends no vector at all, so a participant runs no vector scan for it; the rankers a query needs travel with it, and the asking node fuses only the ones it asked for.

The semantic scan has a one-second budget, and how much fits inside it depends on how many searches the node is running at the same moment — seats taking turns, the dashboard, a person searching — and on which first stage answers: the full scan of every sign code, or the corpus’s semantic index, which reads only the lists nearest the query (see the knowledge system). Measured side by side in one run, at 3 072 dimensions on four cores, p95, over 40 000 sources of the test fixture’s topical corpus — where the index’s training chose to read half its lists:

Searches in flightFull scan, per documentInside one secondIndex, per documentInside one second
1 (an idle node)2.90 µs≈ 345 0001.84 µs≈ 545 000
87.34 µs≈ 136 0005.48 µs≈ 183 000

The second row is the one for a node that is running a company. The metric crewlet.tracker.search.concurrency says which row your node is actually on; see Metrics. The index’s column is per probe share: an index that reads an eighth of its lists — which the same corpus’s training chooses at 120 000 sources — reads a quarter of the rows this one does, and a corpus with no index is on the scan’s column. So is a search narrowed to a small share of the corpus — to the pages in a corpus that is mostly tasks, or to one container: it reads lists until it has seen as many of its own rows as an unfiltered search reads, and past half the lists it runs the scan instead, which reads every row its filter keeps. crewlet search eval names the first stage and the share. Every figure here is a benchmark on one machine — the scan’s own benchmark measured 2.58 µs and 6.14 µs on a quieter host — so treat them as a projection until the benchmarks have run on yours; Replication says the same.

Past the row your node is on, a single node’s search is slower than its budget and search_slow fires. A fleet divides the scan, below.


Above 10 000 documents, a company running more than one node divides the buckets between them. Each node scans its own contiguous range, returns its best candidates with their scores, and the node that asked merges them.

The division is computed, not configured, and every node computes the same one: sort the live node ids, give each a contiguous range, hand the remainder to the first few. 64 buckets over three nodes is 22, 21, 21.

Below that floor a search is answered by the asking node alone. The floor is not a preference — a broker round trip is about a millisecond, and a semantic scan of 10 000 documents measures 49 ms at p95 on an idle node and 125 ms with eight searches in flight. Under it the scan’s fixed cost dominates (its per-document cost at 10 000 is nearly twice its cost at 40 000), and a fan-out divides the documents but makes every node pay that fixed cost again, so it spends more wall clock arranging the work than it saves.

More nodes is not always faster, and the limit is CPU rather than count. Measured on four cores over 4 000 documents, with every participant in one process: 89 ms at one replica, 64 ms at two, 70 ms at four and 98 ms at eight. A real fleet spreads those scans across separate machines, so the turn is further out — but the shape is the same, and it is why the floor is priced in scan time rather than in node count. Recall against the single-scan answer was exactly 1.000 at every width.

node-bnode-aSeat (asking node)node-bnode-aSeat (asking node)scan buckets 0-21slice request (assignment table)slice request (assignment table)top candidates for 22-42, with scorestop candidates for 43-63, with scoresmerge by score per method, then fuse once

The request carries the query and the whole assignment table; each node answers only for its own row. Nothing about it is durable: there is no stream, no consumer, no acknowledgement and no record afterwards. A request nobody serves is a request that never existed. That is deliberate — a search that wrote two records and an audit row per keystroke would make the audit log a function of how often somebody typed.


Why the merge is by score, and then fused once

Section titled “Why the merge is by score, and then fused once”

Each node returns its top candidates per method, with scores. The asking node merges each method’s candidates into one global list by score, and only then fuses the two global lists by reciprocal rank fusion at k = 60.

The order matters and it is not a style choice. Reciprocal rank fusion combines different rankers over one corpus. Run over one ranker across disjoint slices it ranks by placement: a document ranked first on a weak slice contributes 1/61 = 0.0164 and a far stronger document ranked second on a strong slice contributes 1/62 = 0.0161 — so a 0.10 outranks a 0.98.

Merging by score first removes the problem rather than tuning around it. It is also exact: the slices are disjoint, so the global top-N is contained in the union of the per-slice top-N, and a search fanned out four ways returns the identical order a search fanned out no ways does.

The scores are comparable, for a different reason per method. Semantic similarity is comparable by construction — one model, one metric, one vector space. BM25 would not be comparable if each node computed its statistics from its own slice, and it does not: every node holds the whole corpus, so document frequency and the document count come from that node’s complete tables whatever buckets it scanned. The scan is bucket-limited; the statistics are global.


A node that is restarting, overloaded or gone does not answer its assignment. The asking node holds the whole corpus and could have answered alone — which is exactly why a short answer must say so, because a search that returns one fewer result looks identical to a corpus with one fewer document.

There is a second way not to cover an assignment, and it is the quieter one. Every node holds the whole corpus, but the lexical index over it is each node’s own — built by that node’s own walk, in its own database, on its own schedule. A node that joined a few minutes ago therefore holds every document and can find none of them. Such a node answers its range and says so, and the coordinator counts its buckets missing rather than merging an almost-empty answer under complete coverage. It clears itself when that node finishes its first lap, and it applies to the asking node too: a freshly booted node reports its own range missing rather than reporting an empty corpus.

Either way the answer is labelled partial, in the answer itself rather than in a log line on the asking node: every answer carries coverage{nodes:[{id,answered,error}], complete, buckets_missing} — every participant, whether it covered its range and, if not, why (silent inside the budget, still building its index, or the fleet could not be asked). A seat’s own search_knowledge and search_work_items answers say so in words, because a seat reading five results out of what should have been eight otherwise concludes the other three do not exist. The search_scoped alarm reports the fraction of searches answered that way; see Alarms.

The three search alarms are kept apart because they cost different things:

AlarmWhat happenedWhat the answer lost
search_scopedA node did not cover its assignment — it was silent, or its own lexical index has not finished its first lap.A range of the corpus went unscanned.
search_degradedA semantic ranking was asked for and did not run — a node’s vector scan failed, or the provider could not compute the query’s vector.Over the range it did scan, only what shares words with the query was found.
search_slowInteractive search is over its p95 target.Nothing — yet. The corpus has outgrown what one node’s share can scan in the budget.

A company with no embeddings provider is not degraded as far as the alarm is concerned. search_degraded counts semantic rankings that were asked for and failed, never a keyword search and never a company that has nothing to rank by meaning with — an alarm red for the life of a deployment is one nobody reads. The answer itself still says no_embeddings, because the person reading it is the one who can configure a provider.

The asking node always scans its own range itself, never through the broker. A search that returned nothing because the broker hiccupped would be a fleet-wide outage of a read every node can serve alone, so only the peers’ ranges can go missing.

The bucket division above is how the corpus is scanned. On a node without the data role the search itself is asked of a data node — every data node holds the whole estate — and a search none of them could run is answered with no mode served and none of the corpus covered (served_mode empty, coverage.complete false), rather than as an empty list: a seat is told the knowledge base could not be searched, and never “nothing matched”. A work-item search no data node could run is an error the caller is told, for the same reason. The bucket coverage above says something narrower: a search that ran, with a range of the corpus unscanned.


Section titled “The index also derives a page’s backlinks”

The lexical index reads every page body and task description to tokenise it, and the same read extracts the page ids that body links to — /pages/<id> or the dashboard’s #/knowledge/pages/<id>, outside code — into this node’s page_links table. That is where a page’s Linked from comes from (the page answer’s linked_from), so it has the index’s staleness: a body saved a moment ago lists its links once the next lap reads it.

It has the index’s first lap too. Until a node has finished its first lap over both pages and tasks it cannot say what links to a page, and the page answer says so — linked_from_status: building — rather than sending an empty list, which would read as “nothing links here” about a page that is linked.

An upgrade that changes what the index derives needs no rebuild command. Every indexed row records the derivation it was built under (kb_docs.derivation, compared against the build’s search.IndexDerivation), and a row built under an older one is re-derived on the next lap exactly as a row whose source moved. A later build that bumps search.IndexDerivation re-reads each row once, and until that lap finishes the node reports building.


  1. Check search_degraded first. A failing embeddings provider makes every search worse and cheaper at the same time, which does not look like a slowdown.
  2. Check the corpus against the fleet. search_slow fires against the interactive p95 target. Adding a node divides the buckets again with no configuration and no rebuild.
  3. Check recall_below_floor. A corpus whose vectors are behind is answering semantically from a fraction of itself.
  4. Check the first stage. crewlet search eval says whether searches are probing the corpus’s index or scanning, and why — unfiltered, and for each narrowed shape how many of its searches scanned. A corpus scans below 1 024 sources, while a model change is re-embedding it, when its training measured that no index reading half its lists or fewer meets the recall floor in every shape — which is a property of the corpus, and ivf_recall_below_floor says when it is the codes rather than the index that fall short.

There is nothing to tune: the index’s list count follows the corpus, how many lists a search reads is its own training’s measurement, and there is no shard count to set and no routing table to maintain. The bucket count is fixed for the life of a deployment: changing it re-buckets every document, which costs a full index rebuild rather than a rebalance.

Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.