Integration Reconcile
Every integration can break quietly. A seat’s tracker token gets revoked, an administrator narrows a group’s access, somebody deletes a webhook while tidying. Nothing about that is visible from the engine: a webhook that stopped arriving looks exactly like a quiet afternoon, and an agent that hears nothing does nothing, which is what an idle agent looks like too.
The reconcile loop (internal/integration) is what looks. It runs each configured surface on a cadence, reports what it found in one vocabulary shared by every third-party app, and records the result where the API and dashboard read it.
What a pass reports
Section titled “What a pass reports”A pass produces findings, and a finding is one observation that is not “fine”. A third-party app says what it found; it does not decide what that means or which of several findings matters most, because those two decisions have to agree across every surface and cannot be checked from inside any one of them.
| Finding | Means |
|---|---|
credential_missing | The block is enabled and the credential its ${VAR} names resolved to nothing. |
credential_rejected | The credential resolved and the third-party app refused it: a revoked token, a rotated key, an account that lost its access. |
credential_expiring | A credential that works today and stops on a date the third-party app has already published, inside the 14-day warning window (integration.ExpiryWarning). The finding carries the date in expires_at. GitLab is the case: every seat’s token is minted and replaced by the pass itself, but the group Owner token the pass runs on was pasted in by a person, nothing rotates it, and the day it lapses every pass is refused. So each pass reads its own token’s expires_at (GET /personal_access_tokens/self, GitLab 15.5+) and warns while there are still two weekly looks left to act on it. An advisory — the integration is ready — owned by the operator, because the replacement is a value in this deployment’s secret store. A GitLab that cannot say (an older instance, a credential that is not an access token) leaves a note and no finding. |
approval_required | A person must install or approve something at the third-party app. |
ingress_blocked | Deliveries cannot reach this engine, and it will not fix itself. |
ingress_pending | The delivery path is not established yet; the next pass tries again. |
identity_missing | A seat has no account at the third-party app yet. |
identity_failed | A seat’s account could not be created or its credential was refused. |
grant_pending | The third-party app accepted access and has not applied it yet. Atlassian is the case to picture, twice: an invitation is accepted immediately and the account appears in the organization’s directory some time later, so the grant this engine just made reads back as no such user; and once the account is granted, Jira and Confluence go on refusing its brand-new credential for about a minute with a 401 that is byte for byte what a wrong credential looks like. Both are a wait, not a failure — reported as one, the first sent operators looking for a broken organization key while the invite was in flight, and the second told them an admin had to act on an agent that worked a minute later. See When a new account is still coming up. |
unknown_tier | The company document names something the third-party app does not have. |
grant_short | A seat holds less access than its role asks for. |
grant_excess | A seat holds more access than its role asks for. |
registration_orphaned | Something this engine registered at the third-party app that it no longer manages, because the name it is held under changed. Datadog’s webhook definition is the case: it is addressed by NAME, that name is also the handle a monitor writes, and Datadog serves no listing — so the previous definition goes on delivering correctly for every monitor still naming it, and nothing could ever find it again. Reported rather than removed, because removing it would silence exactly those monitors. Datadog’s service accounts are the second case and take the same verdict for the same reason: a renamed seat leaves an identity behind, and an account is a colleague with history attached rather than a row a timer deletes. Only the enabled ones are reported — a disabled account has already reached the state the advisory asks for, and reporting it asks for work somebody has done. That is the general shape of an advisory: it names what somebody might still want to act on, never what they already acted on, or it teaches them that the card does not respond to what it asks for. |
coverage_partial | An integration working exactly as it was asked to, over less than the whole of what it could reach — reported so nobody has to infer a decision from silence. GitHub is the case it was added for: a company whose agents each carry their own App hears about the repositories those Apps are installed on and nothing else in the organization, which is the arrangement its setup form recommends. The mildest of the three advisories, and ranked last: the other two are loose ends somebody may want to tidy, this is a decision already taken. Its actor is the operator, not an admin — widening it is a value in the company’s own configuration rather than a grant somebody at the third-party app has to make. It exists because the alternatives were silence or a false alarm: reported as ingress_blocked it read as Action required over agents receiving events perfectly well, and reported as a note it reached nobody, because a pass returns findings and its notes are discarded. |
Those findings fold into one report, which is what an operator reads:
phaseis where the integration got to:disconnecting,unconfigured,awaiting_admin,provisioning,activating,degradedorready.actoris who has to act for the phase to end: nobody, theengine, theprovider, anadmin(a person, at the third-party app), or theoperator(a person, in this deployment’s own config).detailis one sentence naming what is outstanding, andaction_urlis where the person named byactorgoes to do it. Both are filled only when a person owes something.
The advisories are always last
Section titled “The advisories are always last”A finding about many things names three of them, and carries the rest beside the sentence. A finding’s detail is the card’s one-line status, and the engine caps it at 500 characters — not for tidiness but because an oversized status row is refused by the coordination store rather than truncated, which would stop the surface recording anything at all. A finding that listed its subjects inline therefore arrived as a wall cut off mid-item: measured, 36 Datadog service accounts ending …@agents.cr…. So the sentence says the count and up to three examples, and the whole list travels in the finding’s own subjects, which the screen folds away under the sentence. A finding about one thing names it and carries no list.
And a seat that needs a person shows action required, never a finished status. Where an act belongs to somebody at the third-party app and the engine can never perform it — creating a GitHub App, installing one — the finding is approval_required, whose verdict is awaiting_admin, owed to an admin. identity_missing is the engine’s own work and reads as Setting up agents, which over an act nobody is performing is a card waiting for a pass that will never change anything. The rule extends to the tool’s own satisfied: a surface that requires per-seat identities and has no working seat at all is not satisfied, whatever its company block says.
Four findings have a verdict of ready: credential_expiring, grant_excess, registration_orphaned and coverage_partial. credential_expiring ranks first among them because it is the only note with a deadline — it becomes credential_rejected on its own date — and it is the one advisory the tool roll-up reads as needing attention. Everything below is about grant_excess, and applies to the other two.
grant_excess is the older of the two. The engine did not grant that access and cannot revoke it: it comes from the operator’s own scheme, usually inherited from a parent group or a second role. Agents keep working, so the integration is ready with a note rather than blocked.
That makes its rank load-bearing. Anything the advisory outranks disappears from the report entirely, so it is ranked below every real problem. The control plane this was ported from wrote one classifier per integration and three of them returned the advisory early, which hid a short grant, a failed agent, and a webhook that reached nobody. The ordering now lives in one place with a test that pins it.
The cadence follows who has to act
Section titled “The cadence follows who has to act”Not what failed. A third-party app applying a grant it already accepted finishes in seconds; a person told to install an app is usually installing it as they read; a ${VAR} nobody set will still be unset in an hour. One retry interval would be wrong for all three.
| The report says | Next pass |
|---|---|
ready | 10 minutes, or integrations.check_interval_seconds |
the engine or the provider is working | 30 seconds, doubling to 5 minutes |
an admin must act at the third-party app | 15 seconds, doubling to 10 minutes |
the operator must edit config | 1 hour, flat |
The brisk admin cadence is the point of the whole design: install the app, and provisioning continues without you pressing anything. The flat operator cadence is the opposite case, because nothing at the third-party app will ever change a variable this deployment did not set, so backing off buys nothing and asking often only spends requests.
An applied revision ignores all of it and reconciles now. Every interval above is a wait for asking a third-party app again. A configuration change is the answer changing here, so it marks every surface due and brings the tick forward instead of waiting out the cadence: save the setup dialog and the pass runs within about a second, not at the end of whatever wait the last report earned. One operator action is one pass, because the dialog writes one request per surface (saving Atlassian applies three revisions in a row) and the applies inside a short window fold into a single tick.
And so does somebody finishing the thing the card asked for. The admin backoff is right in general and exactly wrong at one instant: it exists to avoid nagging a third-party app while a person gets round to acting there, so the moment they do act is the moment the wait is longest — up to ten minutes — and least deserved. Measured: a GitHub App installed in about eight seconds, followed by minutes of a card still asking for the install, reloaded by hand, read as the install not having worked.
GitHub’s App-install redirect lands back at this engine, and that arrival now brings the surface’s next pass forward. Nothing about the redirect is believed. It carries no installation id into anything and makes no claim; it says look now, and the pass that follows is the ordinary verified one, which lists the app’s own installations and trusts GitHub rather than a query string. That distinction is what makes it safe on a route nobody signed, where adopting the id off the query would not be — the worst an arrival can do is make the engine ask a question it was going to ask anyway.
Because that route is unauthenticated, the ask is rate limited to one every five seconds. The number comes from the two flows it has to tell apart: a GitHub install round trip cannot complete in under five seconds, so consecutive installs — which a multi-agent company does back to back — each get their own look, while a caller hammering the URL cannot drive more than twelve passes a minute against somebody’s GitHub rate limit. Two other bounds sit under it: the loop folds asks that arrive before its next tick into one pass, and a pass takes the surface’s own lease. Losing an ask is not losing the work; the surface falls back to its cadence.
It is deliberately narrower than an applied revision. A config change moves the answer for every surface; a person finishing something at one third-party app moves it for one, and sweeping all eight would spend seven other vendors’ rate limit on a click that said nothing about them.
The surfaces are visited in dependency order, not in reading order. Atlassian is where an agent’s service account is created; Jira and Confluence are products that account then works in, and each of them checks for the account using the credential Atlassian minted. So the loop visits Atlassian first, then its two products, then everything else. Measured over a reconnect without that ordering: the tracker asked first, found the seat mapped to the account a disconnect had deleted, and reported a 401 as Action required — you, at the third-party app for about thirty-five seconds, until Atlassian’s own pass ran and made the new one. Nothing was wrong and nobody had anything to do; the card was simply asking the products about an account that was one pass away from existing. The order is a separate list from the one the screen reads in, because the two answer different questions and would drift the moment either moved for its own reason — a test pins them to the same set, so a surface added to one and forgotten in the other is caught rather than silently never converged.
Slack is not on this cadence, because it is not reconciled at all. Its apps are created from the command line — one per agent, from a manifest — so there is nothing here to converge: it is registered teardown-only, which gives a disconnect somewhere to run without giving the loop a pass to run. Teardown-only is not nothing to tear down, which is how it was read for a long time — see Slack’s teardown is on the document, not at the vendor below. It carries no reconcile report, and its row holds only the address its setup form was saved against (below) — a row that now ends when the slack: block leaves the document. It used to outlive it for the life of the deployment, because the only thing that forgets a departed surface is a pass reporting ErrNotConfigured and Slack has no pass; a company that removed and later re-added the block inherited an address from before. Slack is the one surface the loop asks the document about directly, precisely because it is the one with nothing to ask.
The settled interval is the company’s to choose, because its cost is the company’s own size. It is the only thing that ever finds access somebody revoked by hand at the third-party app — nothing tells this engine — so it decides how long a card can say Connected over an agent that has already been cut off. Measured on a live deployment: an operator deleted an agent’s token and the card stayed green for eight minutes. Against that sits what a converged pass costs, which scales with the company: a pass asks each seat’s own credential who it is and reads the memberships and hooks per seat and per project, so it is O(seats × projects) requests per surface per interval — tens for a small company, a few hundred for a large one on GitLab or Mattermost. Ten minutes is the default because it is affordable at the large end; a five-seat company can set integrations.check_interval_seconds: 60 and be told within the minute, and a two-hundred-seat one should probably lengthen it. The floor is 60 seconds, and a shorter value is refused naming the field rather than clamped — the likeliest way to type one is meaning minutes and writing seconds, which a clamp hides. Zero is the field being unset, never “off”: a settled surface nothing ever reads back is one this engine would report healthy for the life of the deployment. It is read fresh on every pass, so shortening it takes effect on the pass after the edit rather than after the interval it replaced.
A surface can be given a settled interval of its own, for a third-party app whose reads are rate limited hard enough that the shared one would be spent waiting. No surface in this build sets one.
When the deployment’s address moves
Section titled “When the deployment’s address moves”Every registration a third-party app holds points at integrations.public_base_url as it was when the registration was made, and that value moves: a tunnel restarts, a deployment is renamed, a proxy goes in front.
Where a pass registers the hook, the next tick registers it again at the new address and the surface heals itself. Jira and GitLab match their own hooks by name — webhook_name on each block, defaulting to crewlet — so the address is a field they rewrite, and a hook this engine left at an address it no longer uses is removed rather than abandoned. Confluence has no such field and matches on the address instead: its Cloud endpoint takes no name at all (it is the one Atlassian has never documented, and it swallows every field it does not know), so the only thing a hook can be recognised by is where it points. Two deployments watching one Confluence site therefore cannot be told apart there — which is what webhook_name exists for elsewhere — so give them two different public_base_urls, as you would have to anyway for either to receive a delivery. GitLab sweeps both levels while it is there: a run that establishes a group hook removes this engine’s project hooks and a run that registers project hooks removes its group hook, because the level a pass writes at moves with the group’s plan and a hook at the level nobody writes any more delivers everything twice.
GitHub is the exception, and it is the vendor’s: a GitHub webhook has no name, only its delivery URL, so a hook this engine registered at a previous address is indistinguishable from one a second deployment of the same company registered at its own. It is therefore neither removed nor reported — a claim that a live hook is orphaned would send somebody to delete another deployment’s working registration. A GitHub hook left at a moved address is debris to remove by hand, and GitHub disables one after repeated delivery failures.
Where nothing registers the hook, nothing heals. Slack’s request URL lives in each agent’s app at Slack and can only be read back with an app-configuration token an operator may not have, so the engine cannot see that it is stale, cannot fix it, and the app goes on delivering to an address that no longer answers. And nothing reconciles Slack at all, so the surface reports no phase — the dashboard draws that as Connecting, which is what it means for a configured block the loop has not reported on — and the first symptom is an agent that stopped replying.
So the address is recorded and compared. Every pass stamps the base it ran against onto the surface’s status row, and a surface no pass converges is stamped when its setup form is saved. A row whose recorded address is not the one in force is an ingress fault: the card reads Action needed, the surface carries an address moved badge, and the note names both addresses, because the fix is to replace one with the other at the third-party app and a badge cannot say that.
It clears itself where it should. A surface with a pass is re-stamped on the next tick, so the warning appears only where a person really does have to act. A surface nothing has recorded an address for reports null rather than false: “nothing here can say” is not the claim “the address moved”, and a fresh company must not open with a warning on every card.
When a new account is still coming up
Section titled “When a new account is still coming up”Atlassian creates an agent’s service account, grants it access to Jira and Confluence, and mints its credential — and for about a minute after that the products refuse the credential. Measured on a live Cloud site: about seventy seconds, twice in a row. The gateway accepts the token; Jira answers its own 401 Client must be authenticated to access this resource. Nothing is wrong, nobody can do anything, and a reconnect hits this window every time, because a disconnect that removes accounts means the next connect creates new ones.
Three things used to say three different things about that one minute:
- The tracker reported the seat as
identity_failed— degraded, owed by an admin, “sre-lead has no Jira account”. A person was told to act at Atlassian on an account that was working sixty seconds later. - The engine’s own wiring reported the same seat as
identity_missing— provisioning, owed by the engine. Two findings about one fact, and the engine’s outranks the admin’s, so the card put “the engine is working on it” in the headline and “a person must act at the third-party app” directly underneath it. - The roster badged the same agent ready, because a credential was sealed where the app looks for one.
Now there is one answer. The organization surface reports grant_pending for a seat it has just granted, because a grant that has been accepted is not a grant that is in force. The tracker reports grant_pending too — but only for a credential sealed within the last five minutes, which is the only thing that separates a grant still landing from a credential that is simply wrong; outside that window a refusal is the failure it looks like and is owed by an admin, as it always was. And the engine’s wiring no longer adds a second finding about a seat the surface’s own pass has already reported: both halves resolve the same seats with the same credentials against the same instance, so the pass’s answer — the one with the vendor’s own words in it — is the one that stands.
The roster’s badge follows. satisfied answers “is a credential sealed where this app looks for it”, which stays true of a seat the app is refusing, so an agent any of the card’s surfaces has a finding about is shown as not ready rather than ready. The reason is printed once, in the surface’s own band.
The five-minute window is a four-times margin over the measured propagation. It is deliberately not longer: a grant that is genuinely not landing has to become somebody’s work inside one settled interval rather than waiting for ever under “the provider is working on it”.
A fleet singleton
Section titled “A fleet singleton”The loop is a worker duty, claimed per tick like the retention sweep and the sandbox waiter, so exactly one node runs it at a time. Here that is correctness rather than economy: two nodes reconciling one surface at the same moment both read a third-party app that has no account for a seat, and both create one. The third-party app ends up with two identities for one agent, and no later pass can detect or repair that.
Three separate things stop a node running it, and a reader debugging “why is nothing being reconciled” needs all three:
node.rolesexcludesworkers— the node never claims the duty at all.- The coordination store cannot say whether this node holds it — treated as not held, because entering the two-nodes case on a store blip is exactly what the singleton exists to rule out.
- This node’s config posture is
shedorstuck— it declines, and does so before claiming, so its lease lapses and a peer on the current revision takes the loop over. Every reconciler reads the live company document, so a node the fleet has moved past would converge a third-party app to a revision that has been replaced. This is the only one of the three that says so in the log:integration_reconcile_shedgoing in,integration_reconcile_resumedcoming out, once per transition rather than once per tick.
The duty outlives one pass, deliberately. Its TTL is derived from the deadline a single pass may take rather than from the tick interval, because a pass is allowed minutes and a tick is fifteen seconds: a short TTL lapsed mid-sweep, a peer claimed it, and both nodes swept the surfaces the other had not reached. Nothing unsafe followed — the surface’s own lease is what stops two writers at one third-party app — but the sweep stopped being deterministic. The cost is on the other side and is worth knowing: a node that dies holding the duty leaves it unclaimable for that long rather than for three ticks.
The status itself lives on the coordination store, not in the node’s own database, because the node that reads it is usually not the node that wrote it: -roles data,ingress puts the API and the seats on separate hosts on purpose.
What the loop does, and what it leaves alone
Section titled “What the loop does, and what it leaves alone”It does not tear anything down. Removing an integration block from the company document says what the engine should stop talking to. It does not say that fifteen service accounts, and everything attributable to them, should be destroyed. A removed block makes the loop forget the surface’s status and nothing else; decommissioning stays an explicit flag on the third-party app’s own subcommand, where you type it and read what it is about to delete.
A teardown reports what it removed, and the credentials go with it. Each vendor teardown walks its own plan and deletes the accounts it created, so it knows exactly which seats went — and because every planned seat carries the ${VAR} names its credentials live in, exactly which sealed values are now dead. That used to be destroyed at an error-only return boundary, and no teardown or decommission path anywhere deleted a single secret.
Measured: disconnecting Atlassian with remove accounts deleted every agent’s service account and left SRE_ATLASSIAN_TOKEN and SRE_ATLASSIAN_EMAIL sealed and resolving. On reconnect the tracker mapped the seat to the account that no longer existed — its identity cache is keyed on the credential, and the credential had not changed — and showed a 401 as Action required, you, at the third-party app for about thirty-five seconds. The setup roster kept reporting the seat satisfied, beside the removed account’s address, for as long as the value survived.
Two rules travel with the report. It names the end state the teardown established, not the delta this call performed: every step is already “remove this if it is there” and the whole thing is retried, so a delta would come back empty on the retry and let the block drop with the credentials still sealed. And it names only what is genuinely dead — a merely disabled account does not count, because a token on one works again the moment anybody re-enables it. Datadog’s teardown deletes each account’s application key for exactly that reason; it used to disable the account and leave the key live. The test underneath that rule is whether the value can be recovered: “gone at the vendor” is the usual way to be sure and not the only one, which is what lets Slack delete credentials for apps it cannot delete (below).
So every teardown revokes the credentials before it touches the account. A revoked token on a live account is an agent that can do nothing; a live token on a removed account is a credential that works again the moment anybody restores it — and the first is the safer thing to be interrupted at. GitLab was the one that did not, and it is where the rule bites hardest: GitLab’s service-account delete blocks rather than erases. Measured on a live disconnect, crewlet-sre-lead ended up state: blocked, out of the group, and holding one active token, while the engine had already deleted the company’s own copy of that value — so the company lost the credential and GitLab kept a working one, on an account one click restores. A token that cannot be revoked now leaves the account visibly intact and the seat unreported, because an account still listed holding a credential nobody could withdraw is a state an operator can see, and deleting the only copy of a live credential is the one move nothing can undo.
Three apps disable rather than delete, and one rule covers all three
Section titled “Three apps disable rather than delete, and one rule covers all three”Datadog, Mattermost and GitLab have no way to erase an agent’s account. Datadog’s user delete disables; Mattermost’s bot delete deactivates; GitLab’s service-account delete blocks. So a disconnect on any of the three leaves the account there, and the next connect finds it. The question every one of them has to answer is whether to bring it back — and the answer has to be the same, because an operator learns it once.
The rule: an account this engine disabled is re-enabled on connect. One somebody else disabled is reported, never touched. Re-enabling what a person deliberately switched off at the vendor is the engine overruling an administrator with no way for them to make it stick; leaving its own disconnect’s account disabled forever means a reconnect never works and the only remedy is to find and re-enable each one by hand.
Which it was is written on the account, not in the engine’s own records. A surface’s status row is forgotten the moment a disconnect succeeds, does not survive a lost coordination store, and does not come back from a restore — the account does. So a teardown marks the account before disabling it, and a connect reads the marker back:
| App | Marker | Where |
|---|---|---|
| Datadog | crewlet:disconnected | the service account’s title |
| Mattermost | crewlet:disconnected | the bot’s description |
| GitLab | — | nothing to mark; see below |
The marker is written unconditionally and before the disable, which is the ordering H1 of a field report caught: Datadog’s teardown wrote it inside an if !account.Disabled branch, so a second disconnect over an already-disabled account — an ordinary retry, or a disconnect after somebody had disabled it at Datadog — disabled nothing, marked nothing, and left an account no connect would ever touch again. The seat sat at degraded / admin with no remedy but a hand edit at the vendor. Writing the marker first also means a teardown interrupted between the two steps leaves a marked but enabled account, which the next connect simply clears — the harmless direction.
A connect clears the marker once the account is enabled again, so an administrator who disables the agent afterwards is not mistaken for a past disconnect. Clearing is best effort and noted rather than fatal: the account is already working, and failing the pass over a cosmetic field would be worse than the stale marker it is trying to avoid.
GitLab reports instead, and that is the vendor’s doing. Unblocking is POST /users/:id/unblock, an instance-admin route, and the ordinary deployment provisions with a group Owner token — so the engine cannot re-enable a blocked account whoever blocked it, and attempting it would be a 403 on every tick. No marker is written either: a marker is only worth having where the engine could act on it. What a connect does instead is stop: it finds the account blocked, mints nothing into it, and reports the seat as identity_failed naming the account and the state GitLab itself reports. It used to walk straight past — mint a token, record it, have the very next request refused, and mint another one on the next tick, forever, with the card saying degraded and nothing about why. An instance-mode deployment also gets a link to /admin/users?filter=blocked, because there its credential is an instance administrator; a group-mode one gets the state in words and no link, because where a group Owner administers service accounts has moved between GitLab versions and a link that 404s costs an operator the trip.
GitHub is the exception, and it is not an omission. GitHub offers no API at any permission for deleting an App registration, so a disconnect uninstalls each agent’s App — which is what actually revokes its access — and hands over a link to the page a person deletes it from. The private_key and the App’s own webhook secret therefore stay valid for an App that still exists, so they are named to the operator rather than deleted: destroying a company’s only copy of a working key is not something any API can undo. What the disconnect does record is the installation, which it removed, so no reader is left believing an agent is installed. That record used to stay: measured on a live disconnect, the seat still reported satisfied: true with its App slug while every other disconnected surface said it was waiting to be connected, and /query/integrations kept a github row alive on the strength of two sealed values — rendering the card as Connecting with a routes nowhere badge, permanently.
Slack’s teardown is on the document, not at the vendor
Section titled “Slack’s teardown is on the document, not at the vendor”Slack is the one surface with no reconcile pass, and that was read as nothing this engine put anywhere. For the apps it is true: they are created from the command line or by hand, and deleting one needs an app-configuration token Slack issues only by hand — the same credential the whole surface exists because an operator may not have. What it is not true of is everything else the connect wrote down. Every agent’s seat carries an integrations.slack block naming a sealed bot token and signing secret, and a disconnect that dropped the company block alone left both exactly where they were.
Measured on a live company: the block went, the seats kept their credentials, and the card settled on Paused — a word for a state an operator chose, over the residue of a removal they had just asked for. Both attempts logged accounts=[] secrets=[], and there were two attempts, because disconnecting again is what a person does when the first one appears not to have worked.
So the teardown removes the seat blocks and deletes the values they named, on remove accounts exactly like Atlassian, GitLab and Datadog. Deleting a credential for an app that still exists is normally the one move nothing can undo — it is why GitHub’s private key is named to the operator rather than deleted (above) — and Slack is the exception for a concrete reason: it shows a bot token and a signing secret on the app’s own settings page on every visit, unlike Datadog’s application key, which is shown once. Whoever owns the app can read back anything deleted here; what they could not reconstruct is which ${VAR} each seat’s credential lived in, and that is in the report.
A seat is addressed through the config entity route, one revision per seat, because a merge patch replaces a list wholesale and seats nest inside units: to any depth. Left unticked, the apps go on working and every credential stays: the disconnect drops the company block, the transport retires, and nothing is deleted behind an app the operator kept.
And what only a person can do is handed over. Each agent is named with the app id the running transport learned from auth.test — the only thing that identifies one agent’s app, since a Slack app is named nowhere in the company document — and the disconnect dialog links each one to the page it is deleted from.
And a shared one is not reported orphaned while a sibling still reads it. Jira and Confluence normally authenticate with the same seat credential — Atlassian issues one API token per account, and the ordinary place for it is the shared mcp_env.atlassian block — so disconnecting one product alone named a token the other went on using, on the one list an operator reads to decide what to unset. A sibling counts as a user unless it is itself disconnecting, which is what keeps a whole-card disconnect honest: its surfaces are taken in order, each request records the intent before the next is made, and the union the dialog shows names the credential exactly once.
Only a seat’s own credentials are deleted. A company-level one — an admin token, a webhook secret, an organization key — survives by design and is named to the operator instead: it may be shared with another deployment, which is not something a disconnect decides about.
It does not rotate a credential that works. A third-party app serves a token once, so the tempting reading of “reconcile” is to mint every pass, and that is an outage on a timer: the engine is authenticating with the old value, and rotating revokes what every running agent is using. Rotation is a flag on the subcommand.
It registers webhooks, and it creates the accounts. The loop runs each third-party app’s pass with the sink and the public base URL supplied, so a pass mints a signing secret when the config’s ${VAR} resolves to nothing, creates or updates the hook, and creates the service account each seat acts as.
That is a reversal, and the reason is what connecting an integration means. The loop used to withhold both, on the reasoning that a webhook base is permission to register a hook and a sink is permission to mint a credential, and neither is a decision a timer gets to make. What that produced was an integration nobody could finish from the dashboard: a person connected an app, the loop reported it incomplete forever, and finishing it meant pressing a second button whose whole content was “yes, I meant it”. Connecting is the permission. It is an explicit act, by a person, naming one third-party app, and the credential it hands over is an administrator’s, given for exactly this.
It still does not tear anything down on its own, and it still does not rotate a credential that works (above). Provisioning converges towards the company’s seats: an account that should exist is created, and one that should not is left alone until somebody disconnects the integration, which is the explicit act on the other end.
A pass that seals a credential re-activates the revision. ${VAR} resolves from a snapshot taken at apply time — that is what keeps the secret store off the path of every config read — so a pass that mints a seat’s token lands in a company where everything reading that token already resolved it, to nothing, at the last apply. Refreshing the snapshot fixes the next read, and for the seat identities, parsers, transports, provider clients and MCP children there is no next read: each is built by the apply and by nothing else. So the pass makes one, through the control plane’s own rotation gesture — re-activating the current revision unchanged, which every node is already watching.
Measured, before it did: connecting Atlassian created an agent’s Jira account and sealed its API token, the reconcile reported the surface ready from its own check, and Jira’s live routing held zero seat identities — resolved minutes earlier against a token that did not exist yet. Every issue naming that agent fell through to its project’s lead, indefinitely, while the card said the integration was fine. On a node with no config surface — a worker-only one — the values are still sealed and the rebuild is logged as outstanding, naming POST /config/reload.
Re-activations are coalesced into one apply per 15 seconds, and both halves of that matter. Connecting a third-party app is a burst: the setup dialog writes one request per surface, and Atlassian’s alone applies three revisions in a row — so an apply per seal turns one button press into several whole-company rebuilds and several permanent config revisions. The first request in a quiet period still runs immediately, because somebody pressed Connect and is watching; the ones behind it fold into a single run at the end of the window. Nothing is dropped — a skipped rebuild is exactly the state this mechanism exists to prevent.
The bound is not optional, and believing it was is what made an incident worse. This used to argue it could not loop, because only a real write triggers it and every reconciler is certified against “a converged pass writes nothing”. That is an argument about a converged world, and a pass that can never converge writes on every tick by construction. One did: GitLab created a service account with an address that could not be confirmed, GitLab refused its tokens, the pass read the refusal as a stale credential and minted another. At the reconcile cadence that was a slow leak; behind an apply that wakes the pass which caused it, it ran every five seconds and left 144 live year-long api-scoped tokens from a single connect. That vendor fault is fixed where it lives — this is what stops the next one being amplified.
A pass also re-resolves this node’s own seat identities. A tracker or code-host webhook names people by account, and nothing in the org model says which account a seat holds — so the engine asks, with that seat’s own credential, and registers whatever answers. That lookup happens during an apply, and a seat whose lookup failed used to stay unresolved until the next one: nothing schedules an apply, so a failure that cleared itself a second later persisted until somebody edited something unrelated. Measured: a seat’s token was created and sealed, the tracker checked it one second afterwards, the account was not grantable yet and answered 403 — and the surface held zero seat identities with every card reading Connected, still not having asked again two minutes later.
So each visit re-resolves before it reports, and a seat still holding no account is an identity_missing finding. That is one change doing two jobs: the surface stops classifying ready over an agent that receives nothing, and because it is no longer ready the loop keeps visiting it on the engine’s cadence — 30 seconds doubling to 5 minutes, which is exactly the shape of a lookup that will probably succeed shortly. A seat that resolves is registered into the live registry and starts receiving work immediately, with no config change and nothing for an operator to press. A seat holding no credential at all is never reported: it has opted out of the surface, and listing it would put every human seat on the card.
A disconnect refused by a concurrent writer is worth repeating, and says so. The guard below is one surface at a time, so a disconnect pressed while a reconcile tick or an operator’s own pass is running is refused — a 503 carrying surface_busy, which is the one refusal on this route that clears on its own. The request waits the collision out briefly first, because a tick a moment from finishing is the common case; past that the answer names the surface and the operator (or the dashboard) repeats it. Every step is idempotent. It used to answer internal_error, which a caller can only treat as terminal: the disconnect dialog submits one request per surface in order, so a collision on the second of Atlassian’s three left the tool half disconnected with nothing retrying. The dialog now sits a busy surface out rather than skipping it, because the order is load-bearing — the organization’s credential is what removes the accounts.
A person can still ask for a pass. POST /setup/integrations/{kind}/provision runs the same function on demand, under the same guard, and folds its outcome into this same status. That guard is one thing, not two: a surface’s own lease plus an in-process claim, taken by the operator’s pass, by the loop’s tick and by a disconnect’s teardown alike. Both halves are needed — a lease claim by an owner that already holds it doubles as a renew, so two goroutines in one process would both be told yes, and a single-node install has no lease at all — and with them, two writers at one third-party app never overlap.
And it is held across the write that records the pass, not only across the pass. That is the half that was missing, and “taken first” was satisfied without it. Each writer took the guard, ran its pass, released it, and only then folded the outcome into a status row it had read before the pass began — so a disconnect an operator pressed in that window was silently overwritten by the loop’s own tick, the card went from Disconnecting back to Connected, and they pressed the button again. The guard now spans the row re-read, the pass and the write, and the decision of what to do — converge or tear down — is made from the re-read row, so a disconnect arriving a moment earlier is answered by a teardown rather than by a converge pass that finds the block still present, converges the surface and reports it healthy.
The pass is bounded inside its own lease. The lease is taken once and never renewed, so work that outlived it would go on creating accounts with nothing left excluding a peer — the two-identities-for-one-seat case again, arriving by the clock instead of by a missing lock. internal/setup holds the two values together: a pass is cut off strictly before the lease expires, and the margin between them is what the status write runs in. Nothing in the dashboard calls it any more, because there is nothing left for it to grant: it is there for an operator who wants a pass to run now rather than at the next tick. See Setting an integration up.
Reading the status
Section titled “Reading the status”Every row of the integrations question carries a reconcile object, or null.
GET /query/integrations{ "key": "jira", "configured": true, "enabled": true, "routes": true, "reconcile": { "phase": "degraded", "phase_label": "needs attention", "actor": "admin", "detail": "swe has no Jira account, so no issue reaches it: no credential under mcp_env.atlassian", "action_url": "", "outcome": "blocked", "attempts": 3, "last_error": "", "last_attempt_at": "2026-03-01T12:00:00Z", "settled_at": "2026-03-01T09:14:00Z", "next_attempt_at": "2026-03-01T13:00:00Z", "findings": [ { "kind": "identity_failed", "subject": "swe", "detail": "...", "action_url": "" } ] }}reconcile is three-valued, like routes and secret_usable beside it:
nullmeans this node cannot say. A node that could not read the fleet’s rows has nothing to report, and rendering that as “nothing has been reconciled” would put an alarming claim on a screen over a coordination store that was briefly unreachable. A surface the loop has not reached yet is alsonull.- An object is a real finding.
The findings list travels as well as the report, because the two answer different questions. The report says what to do next; the findings say what is actually wrong. A company with a broken webhook and four under-granted seats reports the webhook, and an operator who fixes it should not have to wait a full pass to discover there were four more things behind it.
phase_label is the same phase in the words a person reads, derived once, in the engine, so a client never has to know what a phase value means. The words are the control plane’s, so one company reads the same status whichever console it is looking at: backlet’s ReconcilePhase is this vocabulary under another name, and these are the labels its console renders.
| Phase | Label |
|---|---|
ready | Connected |
degraded | Action required |
awaiting_admin | Action needed |
provisioning | Setting up agents |
activating | Waiting for the provider |
unconfigured | Failed |
disconnecting | Disconnecting |
unconfigured reads as Failed rather than “not connected” because this phase is only ever reached with a block present: an absent one is ErrNotConfigured, and the row is forgotten rather than reported. So what it names is an integration somebody configured whose credential is missing or the third-party app refused, and “not connected” would read as nobody having tried. A phase a newer node wrote is rendered as its own value with the underscores opened up, never guessed at.
One of backlet’s phases has no counterpart here, and it is not an omission:
disconnectedis a tenant who has not connected an integration yet. Here that is a company document with no block, so there is no row and no phase. The screen shows no status badge at all, only a Connect button. A tool nobody has configured has nothing to report.
disconnecting is a phase this engine does have, and it sorts first, so it wins a tool row’s tag over every other surface: disconnect asks the third-party app to remove what the engine registered there before the block leaves the document, so a surface sits in disconnecting for as long as that takes and reports it if it fails. Nothing about a surface that is going away is worth reporting over the fact that it is going away.
One state per tool
Section titled “One state per tool”A company connects tools, not surfaces: Atlassian is one tool over the organization, Confluence, Jira and the Forge relay. The answer carries one roll-up per tool in tools, for every tool this build serves whether or not the company configured it, and it is decided in the engine (integration.Rollup) rather than by whatever screen reads it — the dashboard used to derive it with its own copy of the phase order and its own ingress rules, and the copy had drifted.
{ "key": "atlassian", "surfaces": ["atlassian", "confluence", "jira", "forge"], "state": "attention", "label": "Action required", "reason": "swe has no Jira account", "surface": "jira" }state | Means |
|---|---|
attention | A person has to act: a phase whose actor is an admin or the operator, a ready surface whose deliveries cannot work, or a credential about to expire. Only ever something a person can do, which is what makes it countable. |
not_connected | Configured and not working yet, with nobody owing anything: the engine or the vendor is mid-flight (setting up agents, applying a grant, the window before the loop’s first report, a disconnect being carried out), or this node could not read the fleet’s status at all. Nothing a person does moves it. |
connected | Every configured surface works. |
not_in_use | No block at all (label Not in use), or every block switched off (label Paused). |
Each configured surface is judged on its own, by its report’s outcome — who has to act, not what failed: settled is connected, blocked is attention, waiting is not connected. Over a ready phase two more facts are read, because the loop says nothing about them: a delivery path that cannot work (secret_usable, routes or endpoint_current false) is Action needed, and a credential_expiring finding is Credential expiring. A surface with no report is Paused where its block is off; judged on those same ingress facts where no pass converges it (Slack, whose apps are created by hand, never gets a report — so “connecting” would be permanent); Status unavailable where this node could not read the status; and Connecting otherwise, which is the window of one reconcile interval after somebody connects.
The tool then takes one surface’s verdict: a teardown first, whatever else is wrong, because a tool being taken away is not one anybody should be sent to fix; then the first state in the order above — so a surface a person owes something outranks one the engine is still bringing up; then, within one state, the least ready phase by Phases (a phase a newer node wrote sits between the worst phase this build knows and ready, never presented as ready and never masking a phase it does know); then the catalogue’s order. label is that surface’s phase label or the roll-up’s own word, reason its sentence, and surface names which surface it was.
What the dashboard shows
Section titled “What the dashboard shows”On the dashboard Settings › Integrations draws one tile per tool from that roll-up — its label as the tag and, while something is owed, its reason as the tile’s sentence — and the Settings column counts the tools in attention. Each tile carries one action, from the same inputs: Connect where nothing is configured; Rotate token where a configured surface holds a credential_expiring or credential_rejected finding, which opens the tool’s setup form saying so (a new token typed over the held one is sealed and the revision re-activated); Continue where the form is unfinished; and Manage otherwise — including a tool that is not_connected, since nothing a person writes moves it but where it is going is worth watching. Manage goes to the tool’s own page.
That page holds everything else in the object above: each faulted surface with its detail, named, so “Jira: swe has no Jira account” reads on the Atlassian page; the actor, the link, the fault and the findings the phase was not derived from; each agent’s row with its own step at the vendor; the settings form, Disconnect, the provisioning passes and each surface’s deliveries. What a line is drawn as follows the actor, not the phase: amber only where a person owes the next step (admin or operator) or where ingress is broken, because nothing the engine or a third-party app is doing needs a person told about it.
last_error is a fault, not a finding: the engine or the third-party app failing to look at the world at all, rather than a statement about it. A pass that fails drops the previous pass’s findings rather than leaving them standing under a fresh timestamp, because a pass that failed did not observe anything.
A fault is normally reported as activating, owned by the engine, because almost every fault is a third-party app briefly unreachable and the next pass clears it. One fault is not a wait. When the third-party app answers 401 or 403, it did look, and it refused the credential: every later pass is refused identically until a person changes it. That is reported as credential_rejected, which makes the surface unconfigured and the operator’s to fix, so nobody is left watching a retry that cannot succeed. Rate limits, timeouts, 404s and 5xxes stay waits.
Adding a surface to the loop
Section titled “Adding a surface to the loop”A third-party app contributes one method:
type Reconciler interface { Kind() Kind Reconcile(ctx context.Context) ([]Finding, error)}and is registered in internal/engine/integrations.go. Findings and errors are different answers and must not be collapsed: findings are statements about your world, and an error is a failure to read it.
Every implementation is certified against one suite, internal/integration/integrationtest, in the same tradition as queuetest and coordtest. Each third-party app’s package holds its own harness, which stands its world up already converged — every seat has its account, every credential works, every webhook points at the address in force — and then counts what the pass writes into it.
Its load-bearing case is that a pass over a converged world writes nothing, because that is the clause most likely to be wrong and the one whose failure is a credential rotated every ten minutes for ever. A write here means anything a person would have to undo, which is wider than a request to the third-party app — a pass that re-seals a seat’s credential into this deployment’s own sealed store on every run is writing too — and narrower than a non-GET, because some third-party apps model a listing as a POST and a counter keyed on the method would make the clause impossible to satisfy. Count by route.
That clause is what makes “safe to leave switched on” a fact rather than an intention, and it was not one for a while: the suite existed, was believed, and was pointed only at a stub, while three of the seven reconcilers wrote at their third-party app on every single pass. The fix in each case was to read the current state and write only on a real difference — never to soften the clause, which is the pressure a suite like this has to be able to resist.
Part of Crewlet. Generated from crewlet/crewlet main at f665f5a. This is not the current version — see the latest docs.