Keep the instance credential across a TerdutServer delete, and adopt it on recreate
Deleting a TerdutServer removed the credential Secrets but never touched the database, so a recreated one found a server that was already bootstrapped and no key for it: /api/bootstrap answered 403 and the operator stopped at BootstrapStateLost, whose message and DESIGN.md both said "delete and recreate". That is how the terdut-demo install on the cluster got stuck on 2026-10-03: Helm's cleanupOnFail deleted its TerdutServer after a failed upgrade, the recreate found the bootstrapped database, and it sat at Ready: False for five days until the database was reset by hand. Recreating cannot fix it, because the finalizer clears Secrets and the database is not its to reset, so "a fresh create starts clean" was only ever true when the database went with it. spec.credentials.deletionPolicy is Retain by default: the finalizer keeps the instance credential Secret (Delete removes it, as before). The bootstrap checkpoint is always removed. Before calling /api/bootstrap, reconcile now looks for the retained Secret and asks the server for the operator's own service account with its token. Accepted: adopt it and skip bootstrap. Rejected with 401/403: the Secret outlived a database reset, so ignore it and bootstrap like a first install, which replaces it. Any other error retries. terdut-server's own tests already call that endpoint with an instance-scoped key, so the permission is not new. BootstrapStateLost is still the answer when the server is bootstrapped and no credential it accepts survives, but its message now names the Secret to restore and says that recreating does not clear the database. DESIGN.md §6 says the same, and the chart passes the setting through as terdutServer.credentials.deletionPolicy. A retained Secret of a TerdutServer that is gone for good is an orphan to delete by hand. It is inert: nothing adopts it unless the server accepts the token. Checked on the kind demo with a locally built image against the real terdut-server v0.43.0: deleting the TerdutServer kept the Secret, recreating it reached Ready with the same credential (identical hash) and both TerdutTeams came back Ready with their original ids. The controller specs cover adoption, a rejected token after a reset, the bootstrapped-and-rejected failure, and both deletion policies. Co-authored-by: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -169,6 +169,8 @@ spec:
|
||||
matchers: "alertname=Watchdog"
|
||||
timeout: 15m
|
||||
severity: critical
|
||||
credentials:
|
||||
deletionPolicy: Retain # Retain (default) | Delete -- what deleting this CR does to the instance credential Secret; see §6 point 2
|
||||
notify:
|
||||
ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
|
||||
fallbackTopic: ""
|
||||
@@ -529,6 +531,18 @@ when nothing ever crosses into a tenant namespace in the first place.
|
||||
the same rigor §5's general idempotent-create rule already applies
|
||||
elsewhere:
|
||||
- `status.credentialsSecretRef` already set: done, nothing to do.
|
||||
- Otherwise, look for the instance credential Secret itself
|
||||
(`<namespace>.<name>-instance-credentials`, point 2 below): a
|
||||
`TerdutServer` deleted and recreated under the same name leaves it
|
||||
behind by default (`spec.credentials.deletionPolicy: Retain`), and the
|
||||
server it logs in to has not changed, because deleting a
|
||||
`TerdutServer` never touches its database. If it exists, ask the server
|
||||
for the operator's own service account with that token
|
||||
(`GET /api/service-accounts?name=terdut-operator`). Accepted: adopt it,
|
||||
set `status.credentialsSecretRef`, and skip everything below. Rejected
|
||||
with a `401`/`403`: it is stale (a database reset since), so ignore it
|
||||
and carry on; the steps below replace it. Any other failure is a retry,
|
||||
not a guess.
|
||||
- Otherwise, check for an intermediate
|
||||
`<namespace>.<name>-bootstrap-admin` Secret in
|
||||
the operator's own namespace first. If it exists, its key is a still-
|
||||
@@ -538,15 +552,21 @@ when nothing ever crosses into a tenant namespace in the first place.
|
||||
on `201`, immediately checkpoint its response's raw admin key
|
||||
(`{"user": ..., "api_key": {"key": "<raw>", ...}}`) into that Secret
|
||||
before doing anything else with it. A `403` with neither
|
||||
`status.credentialsSecretRef` nor this checkpoint Secret present is
|
||||
the one genuinely pathological case left (the checkpoint deleted out
|
||||
from under a reconcile already past this point) — handled the same
|
||||
way the design already handles unrecoverable server-issued material
|
||||
elsewhere (§5's webhook-Secret-loss rule): fail closed,
|
||||
`Ready: False, reason: BootstrapStateLost`, with the same recovery as
|
||||
that case, delete and recreate the `TerdutServer` (its finalizer tears
|
||||
down the Deployment/database-backing and server-side rows; a fresh
|
||||
create starts clean) — not a workaround peculiar to this one path.
|
||||
`status.credentialsSecretRef`, nor an instance credential the server
|
||||
accepts, nor this checkpoint Secret present is the one genuinely
|
||||
pathological case left: the server's database is already bootstrapped
|
||||
and no credential for it survives here (the checkpoint deleted out from
|
||||
under a reconcile already past this point, `deletionPolicy: Delete`, or
|
||||
the Secret removed by hand). Fail closed, `Ready: False, reason:
|
||||
BootstrapStateLost`, with a message that names the instance credential
|
||||
Secret to restore. Deleting and recreating the `TerdutServer` does **not**
|
||||
recover from it: the finalizer removes Secrets only, and the database
|
||||
(the one thing that still says "already bootstrapped") is not the
|
||||
operator's to reset. Restore the Secret from a copy, or reset the
|
||||
server's database and then recreate the `TerdutServer`, which then
|
||||
bootstraps like a first install. (This section used to say a fresh
|
||||
create "starts clean"; that was only ever true when the database went
|
||||
with it.)
|
||||
- With an admin key in hand (fresh or checkpointed): `POST
|
||||
/api/service-accounts {name: "terdut-operator", scope: "instance"}`.
|
||||
A `409` here means a prior attempt got this far before being
|
||||
@@ -563,10 +583,15 @@ when nothing ever crosses into a tenant namespace in the first place.
|
||||
under a fixed data key, `token`), referenced back from
|
||||
`TerdutServer.status.credentialsSecretRef: {name, key}` (§4.1). No
|
||||
`OwnerReference` (those can't cross namespaces, and this Secret doesn't
|
||||
share a namespace with the `TerdutServer` that caused it); the `TerdutServer`'s
|
||||
finalizer deletes this Secret directly as part of its own teardown,
|
||||
the same way it already has to clean up the server-side resources it
|
||||
created (§5's general finalizer rule extends naturally to this Secret).
|
||||
share a namespace with the `TerdutServer` that caused it). What the
|
||||
`TerdutServer`'s finalizer does with it is `spec.credentials.deletionPolicy`:
|
||||
`Retain` (the default) leaves it for a recreated `TerdutServer` to adopt
|
||||
(point 1), `Delete` removes it directly as part of the teardown. There are
|
||||
no server-side resources to undo either way: the bootstrap user and
|
||||
service account have no delete verb in terdut-server's API. The bootstrap
|
||||
checkpoint Secret is always removed. A retained Secret of a `TerdutServer`
|
||||
that is gone for good is an orphan to delete by hand, and it is inert:
|
||||
adoption asks the server to accept the token first.
|
||||
3. When a `TerdutTeam` first becomes `Ready` (its `serverRef` resolved,
|
||||
`allowedTeams` satisfied if cross-namespace), its controller uses the
|
||||
`TerdutServer`'s instance-scoped credential (read from the operator's own
|
||||
|
||||
Reference in New Issue
Block a user