Keep the instance credential across a TerdutServer delete, and adopt it on recreate

Deleting a TerdutServer removed the credential Secrets but never touched the
database, so a recreated one found a server that was already bootstrapped and
no key for it: /api/bootstrap answered 403 and the operator stopped at
BootstrapStateLost, whose message and DESIGN.md both said "delete and
recreate". That is how the terdut-demo install on the cluster got stuck on
2026-10-03: Helm's cleanupOnFail deleted its TerdutServer after a failed
upgrade, the recreate found the bootstrapped database, and it sat at Ready:
False for five days until the database was reset by hand. Recreating cannot
fix it, because the finalizer clears Secrets and the database is not its to
reset, so "a fresh create starts clean" was only ever true when the database
went with it.

spec.credentials.deletionPolicy is Retain by default: the finalizer keeps the
instance credential Secret (Delete removes it, as before). The bootstrap
checkpoint is always removed. Before calling /api/bootstrap, reconcile now
looks for the retained Secret and asks the server for the operator's own
service account with its token. Accepted: adopt it and skip bootstrap.
Rejected with 401/403: the Secret outlived a database reset, so ignore it and
bootstrap like a first install, which replaces it. Any other error retries.
terdut-server's own tests already call that endpoint with an instance-scoped
key, so the permission is not new.

BootstrapStateLost is still the answer when the server is bootstrapped and no
credential it accepts survives, but its message now names the Secret to
restore and says that recreating does not clear the database. DESIGN.md §6
says the same, and the chart passes the setting through as
terdutServer.credentials.deletionPolicy.

A retained Secret of a TerdutServer that is gone for good is an orphan to
delete by hand. It is inert: nothing adopts it unless the server accepts the
token.

Checked on the kind demo with a locally built image against the real
terdut-server v0.43.0: deleting the TerdutServer kept the Secret, recreating it
reached Ready with the same credential (identical hash) and both TerdutTeams
came back Ready with their original ids. The controller specs cover adoption,
a rejected token after a reset, the bootstrapped-and-rejected failure, and
both deletion policies.

Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
Niklas Ye
2026-10-08 21:07:38 +02:00
parent 6572f63157
commit 62664c93ff
10 changed files with 361 additions and 50 deletions
+38 -13
View File
@@ -169,6 +169,8 @@ spec:
matchers: "alertname=Watchdog"
timeout: 15m
severity: critical
credentials:
deletionPolicy: Retain # Retain (default) | Delete -- what deleting this CR does to the instance credential Secret; see §6 point 2
notify:
ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
fallbackTopic: ""
@@ -529,6 +531,18 @@ when nothing ever crosses into a tenant namespace in the first place.
the same rigor §5's general idempotent-create rule already applies
elsewhere:
- `status.credentialsSecretRef` already set: done, nothing to do.
- Otherwise, look for the instance credential Secret itself
(`<namespace>.<name>-instance-credentials`, point 2 below): a
`TerdutServer` deleted and recreated under the same name leaves it
behind by default (`spec.credentials.deletionPolicy: Retain`), and the
server it logs in to has not changed, because deleting a
`TerdutServer` never touches its database. If it exists, ask the server
for the operator's own service account with that token
(`GET /api/service-accounts?name=terdut-operator`). Accepted: adopt it,
set `status.credentialsSecretRef`, and skip everything below. Rejected
with a `401`/`403`: it is stale (a database reset since), so ignore it
and carry on; the steps below replace it. Any other failure is a retry,
not a guess.
- Otherwise, check for an intermediate
`<namespace>.<name>-bootstrap-admin` Secret in
the operator's own namespace first. If it exists, its key is a still-
@@ -538,15 +552,21 @@ when nothing ever crosses into a tenant namespace in the first place.
on `201`, immediately checkpoint its response's raw admin key
(`{"user": ..., "api_key": {"key": "<raw>", ...}}`) into that Secret
before doing anything else with it. A `403` with neither
`status.credentialsSecretRef` nor this checkpoint Secret present is
the one genuinely pathological case left (the checkpoint deleted out
from under a reconcile already past this point) — handled the same
way the design already handles unrecoverable server-issued material
elsewhere (§5's webhook-Secret-loss rule): fail closed,
`Ready: False, reason: BootstrapStateLost`, with the same recovery as
that case, delete and recreate the `TerdutServer` (its finalizer tears
down the Deployment/database-backing and server-side rows; a fresh
create starts clean) — not a workaround peculiar to this one path.
`status.credentialsSecretRef`, nor an instance credential the server
accepts, nor this checkpoint Secret present is the one genuinely
pathological case left: the server's database is already bootstrapped
and no credential for it survives here (the checkpoint deleted out from
under a reconcile already past this point, `deletionPolicy: Delete`, or
the Secret removed by hand). Fail closed, `Ready: False, reason:
BootstrapStateLost`, with a message that names the instance credential
Secret to restore. Deleting and recreating the `TerdutServer` does **not**
recover from it: the finalizer removes Secrets only, and the database
(the one thing that still says "already bootstrapped") is not the
operator's to reset. Restore the Secret from a copy, or reset the
server's database and then recreate the `TerdutServer`, which then
bootstraps like a first install. (This section used to say a fresh
create "starts clean"; that was only ever true when the database went
with it.)
- With an admin key in hand (fresh or checkpointed): `POST
/api/service-accounts {name: "terdut-operator", scope: "instance"}`.
A `409` here means a prior attempt got this far before being
@@ -563,10 +583,15 @@ when nothing ever crosses into a tenant namespace in the first place.
under a fixed data key, `token`), referenced back from
`TerdutServer.status.credentialsSecretRef: {name, key}` (§4.1). No
`OwnerReference` (those can't cross namespaces, and this Secret doesn't
share a namespace with the `TerdutServer` that caused it); the `TerdutServer`'s
finalizer deletes this Secret directly as part of its own teardown,
the same way it already has to clean up the server-side resources it
created (§5's general finalizer rule extends naturally to this Secret).
share a namespace with the `TerdutServer` that caused it). What the
`TerdutServer`'s finalizer does with it is `spec.credentials.deletionPolicy`:
`Retain` (the default) leaves it for a recreated `TerdutServer` to adopt
(point 1), `Delete` removes it directly as part of the teardown. There are
no server-side resources to undo either way: the bootstrap user and
service account have no delete verb in terdut-server's API. The bootstrap
checkpoint Secret is always removed. A retained Secret of a `TerdutServer`
that is gone for good is an orphan to delete by hand, and it is inert:
adoption asks the server to accept the token first.
3. When a `TerdutTeam` first becomes `Ready` (its `serverRef` resolved,
`allowedTeams` satisfied if cross-namespace), its controller uses the
`TerdutServer`'s instance-scoped credential (read from the operator's own