Search
How the knowledge search is answered, what it costs, and what a fleet can do about it when the corpus outgrows one node’s CPU.
Search is answered from this node’s own tables. Every node holds the whole corpus — see Scaling — so there is no lookup that has to leave the machine and no answer that depends on a peer being up. What a fleet can divide is the work, not the data.
The two halves, and the bucket both of them carry
Section titled “The two halves, and the bucket both of them carry”A hybrid search runs two rankers and fuses them:
- Lexical — BM25 over the engine’s own inverted list, in this node’s database.
- Semantic — a two-stage vector search over the replicated vectors, whose first stage probes the corpus’s semantic index when it has one.
Every indexed document carries a search shard: a stable hash of its own identity into 64 fixed buckets, written beside the row in both estates by the same function. See Every document carries a bucket for what the bucket is a function of and — just as important — what it is deliberately not.
A search can be told to read only part of that range. On a single node nothing tells it one: it holds every bucket and reads every bucket.
Three modes, and what an answer says it served
Section titled “Three modes, and what an answer says it served”Every ranked search — the knowledge search and the tracker’s own item search alike — takes one of three modes, with one vocabulary on every surface:
| Mode | Ranks by | Label on a screen |
|---|---|---|
hybrid (the default) | both halves, fused by reciprocal rank fusion | Hybrid |
keyword | the words the query used — BM25 alone | Keyword |
semantic | what the query means — the vector scan alone | Meaning |
The two halves do not read the same text. The keyword half indexes a
source’s title and its whole body. The semantic half ranks one vector per
source, computed from its opening — the title and then the body, whitespace
collapsed, the first 8 KiB of the two together, or less where the model’s own
per-input bound is smaller (see
Where the vectors come from).
So a passage deep in a long page is found by hybrid and keyword through the
words it uses, and never by semantic, which cannot see past the window.
crewlet search eval reports, per corpus, how many sources and how much of
their text lie past it (see
What the quality of this can and cannot be promised).
A semantic ranking needs the query in the same embedding space as the
documents. The asking node computes that vector once through the company’s
embeddings provider — the same model and width the corpus is embedded at — and
sends it with the request, so no participant makes a provider call of its own.
It caches the result: 1 024 query vectors per node (about 12 MiB at 3 072
dimensions), least recently used first out, emptied the moment
providers.embeddings.model changes so a vector from the retired model is
never ranked against rows the refill is replacing. A phrase the ⌘K palette
embedded is a cache hit for the same phrase searched as knowledge, and for the
turn-start prefetch. Computing a query’s vector is bounded at two seconds —
twice the one-second budget of the scan itself — so a slow provider costs a
search its meaning half rather than holding the person who asked.
A query is at most 400 bytes, on every surface that takes one: a seat’s
search_knowledge, search_work_items and answer_knowledge, the operator’s
tools, the API’s knowledge and work_search (bad_params, naming the size
and the limit) and the dashboard, which sends nothing past it and says why. A
longer one is refused, never cut — a query is a question, and a search on
its first part answers another one; past four hundred bytes it is a pasted
thread, which no ranker turns into a better search. This is the opposite of
what the corpus does with a long page, and deliberately: a source is embedded
as its opening because one vector stands for one source, while a query’s
vector is always computed from the whole query. Every model this build
knows takes at least 2 032 bytes an input; on a model stated with a narrower
window (max_input_tokens), a query longer than one input is embedded in
pieces at the model’s bound, in one request, and the pieces’ vectors pooled.
A mode asked for is not always a mode served, and every answer says which.
It carries served_mode (the ranking the hits actually came from), modes
(what this backend can serve as asked, right now), degraded (why those differ
from what was asked) and coverage (below):
degraded | What happened | Hybrid serves | Semantic serves |
|---|---|---|---|
no_embeddings | No embeddings provider, or knowledge.vectors: false. | its keyword half | nothing |
embedding_failed | The provider did not produce the query’s vector in time. Nothing to configure; the next search asks again, and a failure is never cached. | its keyword half | nothing |
semantic_partial | Part of the fleet ran without its vector scan. | both, with meaning over less of the corpus | meaning over less of the corpus |
unsupported | The backend has no such ranker — Confluence, whose own CQL search is keyword. | the keyword answer | nothing |
Semantic never falls back to keyword. Somebody asking for meaning is asking for the pages that share no word with the query, and a keyword ranking is the one ranking guaranteed not to find them — so it answers no hits and says why, rather than a keyword answer labelled as meaning. Hybrid does fall back: the words are half of what was asked for.
The modes on offer are known before anybody types. A search with an empty
phrase is the seam’s PROBE: it runs nothing and reads nothing, and answers
modes and — for the mode asked — the degraded the CONFIGURATION decides
(no_embeddings, unsupported; never a transient embedding_failed, which is
a property of one search). The dashboard’s Knowledge tree asks it for
semantic and disables a mode the answer does not list, with the reason
written under the control, rather than letting a reader pick Meaning and find
out from a degraded answer. The same answer sets the dashboard’s DEFAULT:
Hybrid where it is served, otherwise the first mode listed — Keyword on a
company with no embeddings provider — so a search nobody chose a mode for runs
in a mode the engine serves as asked instead of reporting a degradation of a
choice nobody made. The wire default is unchanged: a request that names no
mode is still hybrid.
A keyword search sends no vector at all, so a participant runs no vector scan for it; the rankers a query needs travel with it, and the asking node fuses only the ones it asked for.
How big a corpus one node scans
Section titled “How big a corpus one node scans”The semantic scan has a one-second budget, and how much fits inside it depends on how many searches the node is running at the same moment — seats taking turns, the dashboard, a person searching — and on which first stage answers: the full scan of every sign code, or the corpus’s semantic index, which reads only the lists nearest the query (see the knowledge system). Measured side by side in one run, at 3 072 dimensions on four cores, p95, over 40 000 sources of the test fixture’s topical corpus — where the index’s training chose to read half its lists:
| Searches in flight | Full scan, per document | Inside one second | Index, per document | Inside one second |
|---|---|---|---|---|
| 1 (an idle node) | 2.90 µs | ≈ 345 000 | 1.84 µs | ≈ 545 000 |
| 8 | 7.34 µs | ≈ 136 000 | 5.48 µs | ≈ 183 000 |
The second row is the one for a node that is running a company. The metric
crewlet.tracker.search.concurrency says which row your node is actually on;
see Metrics. The index’s column is per probe
share: an index that reads an eighth of its lists — which the same corpus’s
training chooses at 120 000 sources — reads a quarter of the rows this one
does, and a corpus with no index is on the scan’s column. So is a search
narrowed to a small share of the corpus — to the pages in a corpus
that is mostly tasks, or to one container: it reads lists until it has seen
as many of its own rows as an unfiltered search reads, and past half the
lists it runs the scan instead, which reads every row its filter keeps. crewlet search eval names the first stage and the share. Every figure here is a benchmark on
one machine — the scan’s own benchmark measured 2.58 µs and 6.14 µs on a
quieter host — so treat them as a projection until the benchmarks have run on
yours; Replication says the
same.
Past the row your node is on, a single node’s search is slower than its budget
and search_slow fires. A fleet divides the scan, below.
When a fleet divides the scan
Section titled “When a fleet divides the scan”Above 10 000 documents, a company running more than one node divides the buckets between them. Each node scans its own contiguous range, returns its best candidates with their scores, and the node that asked merges them.
The division is computed, not configured, and every node computes the same one: sort the live node ids, give each a contiguous range, hand the remainder to the first few. 64 buckets over three nodes is 22, 21, 21.
Below that floor a search is answered by the asking node alone. The floor is not a preference — a broker round trip is about a millisecond, and a semantic scan of 10 000 documents measures 49 ms at p95 on an idle node and 125 ms with eight searches in flight. Under it the scan’s fixed cost dominates (its per-document cost at 10 000 is nearly twice its cost at 40 000), and a fan-out divides the documents but makes every node pay that fixed cost again, so it spends more wall clock arranging the work than it saves.
More nodes is not always faster, and the limit is CPU rather than count. Measured on four cores over 4 000 documents, with every participant in one process: 89 ms at one replica, 64 ms at two, 70 ms at four and 98 ms at eight. A real fleet spreads those scans across separate machines, so the turn is further out — but the shape is the same, and it is why the floor is priced in scan time rather than in node count. Recall against the single-scan answer was exactly 1.000 at every width.
What travels, and what does not
Section titled “What travels, and what does not”The request carries the query and the whole assignment table; each node answers only for its own row. Nothing about it is durable: there is no stream, no consumer, no acknowledgement and no record afterwards. A request nobody serves is a request that never existed. That is deliberate — a search that wrote two records and an audit row per keystroke would make the audit log a function of how often somebody typed.
Why the merge is by score, and then fused once
Section titled “Why the merge is by score, and then fused once”Each node returns its top candidates per method, with scores. The asking node merges each method’s candidates into one global list by score, and only then fuses the two global lists by reciprocal rank fusion at k = 60.
The order matters and it is not a style choice. Reciprocal rank fusion combines different rankers over one corpus. Run over one ranker across disjoint slices it ranks by placement: a document ranked first on a weak slice contributes 1/61 = 0.0164 and a far stronger document ranked second on a strong slice contributes 1/62 = 0.0161 — so a 0.10 outranks a 0.98.
Merging by score first removes the problem rather than tuning around it. It is also exact: the slices are disjoint, so the global top-N is contained in the union of the per-slice top-N, and a search fanned out four ways returns the identical order a search fanned out no ways does.
The scores are comparable, for a different reason per method. Semantic similarity is comparable by construction — one model, one metric, one vector space. BM25 would not be comparable if each node computed its statistics from its own slice, and it does not: every node holds the whole corpus, so document frequency and the document count come from that node’s complete tables whatever buckets it scanned. The scan is bucket-limited; the statistics are global.
When part of the corpus is not scanned
Section titled “When part of the corpus is not scanned”A node that is restarting, overloaded or gone does not answer its assignment. The asking node holds the whole corpus and could have answered alone — which is exactly why a short answer must say so, because a search that returns one fewer result looks identical to a corpus with one fewer document.
There is a second way not to cover an assignment, and it is the quieter one. Every node holds the whole corpus, but the lexical index over it is each node’s own — built by that node’s own walk, in its own database, on its own schedule. A node that joined a few minutes ago therefore holds every document and can find none of them. Such a node answers its range and says so, and the coordinator counts its buckets missing rather than merging an almost-empty answer under complete coverage. It clears itself when that node finishes its first lap, and it applies to the asking node too: a freshly booted node reports its own range missing rather than reporting an empty corpus.
Either way the answer is labelled partial, in the answer itself rather than
in a log line on the asking node: every answer carries
coverage{nodes:[{id,answered,error}], complete, buckets_missing} — every
participant, whether it covered its range and, if not, why (silent inside the
budget, still building its index, or the fleet could not be asked). A seat’s
own search_knowledge and search_work_items answers say so in words, because a seat reading five
results out of what should have been eight otherwise concludes the other three
do not exist. The search_scoped alarm reports the fraction of searches
answered that way; see Alarms.
The three search alarms are kept apart because they cost different things:
| Alarm | What happened | What the answer lost |
|---|---|---|
search_scoped | A node did not cover its assignment — it was silent, or its own lexical index has not finished its first lap. | A range of the corpus went unscanned. |
search_degraded | A semantic ranking was asked for and did not run — a node’s vector scan failed, or the provider could not compute the query’s vector. | Over the range it did scan, only what shares words with the query was found. |
search_slow | Interactive search is over its p95 target. | Nothing — yet. The corpus has outgrown what one node’s share can scan in the budget. |
A company with no embeddings provider is not degraded as far as the alarm
is concerned. search_degraded counts semantic rankings that were asked for and
failed, never a keyword search and never a company that has nothing to rank by
meaning with — an alarm red for the life of a deployment is one nobody reads.
The answer itself still says no_embeddings, because the person reading it is
the one who can configure a provider.
The asking node always scans its own range itself, never through the broker. A search that returned nothing because the broker hiccupped would be a fleet-wide outage of a read every node can serve alone, so only the peers’ ranges can go missing.
When the estate does not answer
Section titled “When the estate does not answer”The bucket division above is how the corpus is scanned. On a node without the
data role the search itself is asked of a data node — every data node holds
the whole estate — and a search none of them could run is answered with no
mode served and none of the corpus covered (served_mode empty,
coverage.complete false), rather than as an empty list: a seat is told the
knowledge base could not be searched, and never “nothing matched”. A work-item
search no data node could run is an error the caller is told, for the same
reason. The bucket coverage above says something narrower: a search that ran,
with a range of the corpus unscanned.
The index also derives a page’s backlinks
Section titled “The index also derives a page’s backlinks”The lexical index reads every page body and task description to tokenise it,
and the same read extracts the page ids that body links to — /pages/<id> or
the dashboard’s #/knowledge/pages/<id>, outside code — into this node’s
page_links table. That is where a page’s Linked from comes from (the
page answer’s linked_from), so it has the index’s staleness: a body saved a
moment ago lists its links once the next lap reads it.
It has the index’s first lap too. Until a node has finished its first lap over
both pages and tasks it cannot say what links to a page, and the page answer
says so — linked_from_status: building — rather than sending an empty list,
which would read as “nothing links here” about a page that is linked.
An upgrade that changes what the index derives needs no rebuild command.
Every indexed row records the derivation it was built under (kb_docs.derivation,
compared against the build’s search.IndexDerivation), and a row built under an
older one is re-derived on the next lap exactly as a row whose source moved. A
later build that bumps search.IndexDerivation re-reads each row once, and
until that lap finishes the node reports building.
If search is slow
Section titled “If search is slow”- Check
search_degradedfirst. A failing embeddings provider makes every search worse and cheaper at the same time, which does not look like a slowdown. - Check the corpus against the fleet.
search_slowfires against the interactive p95 target. Adding a node divides the buckets again with no configuration and no rebuild. - Check
recall_below_floor. A corpus whose vectors are behind is answering semantically from a fraction of itself. - Check the first stage.
crewlet search evalsays whether searches are probing the corpus’s index or scanning, and why — unfiltered, and for each narrowed shape how many of its searches scanned. A corpus scans below 1 024 sources, while a model change is re-embedding it, when its training measured that no index reading half its lists or fewer meets the recall floor in every shape — which is a property of the corpus, andivf_recall_below_floorsays when it is the codes rather than the index that fall short.
There is nothing to tune: the index’s list count follows the corpus, how many lists a search reads is its own training’s measurement, and there is no shard count to set and no routing table to maintain. The bucket count is fixed for the life of a deployment: changing it re-buckets every document, which costs a full index rebuild rather than a rebalance.
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.