Deleting a TerdutServer removed the credential Secrets but never touched the database, so a recreated one found a server that was already bootstrapped and no key for it: /api/bootstrap answered 403 and the operator stopped at BootstrapStateLost, whose message and DESIGN.md both said "delete and recreate". That is how the terdut-demo install on the cluster got stuck on 2026-10-03: Helm's cleanupOnFail deleted its TerdutServer after a failed upgrade, the recreate found the bootstrapped database, and it sat at Ready: False for five days until the database was reset by hand. Recreating cannot fix it, because the finalizer clears Secrets and the database is not its to reset, so "a fresh create starts clean" was only ever true when the database went with it. spec.credentials.deletionPolicy is Retain by default: the finalizer keeps the instance credential Secret (Delete removes it, as before). The bootstrap checkpoint is always removed. Before calling /api/bootstrap, reconcile now looks for the retained Secret and asks the server for the operator's own service account with its token. Accepted: adopt it and skip bootstrap. Rejected with 401/403: the Secret outlived a database reset, so ignore it and bootstrap like a first install, which replaces it. Any other error retries. terdut-server's own tests already call that endpoint with an instance-scoped key, so the permission is not new. BootstrapStateLost is still the answer when the server is bootstrapped and no credential it accepts survives, but its message now names the Secret to restore and says that recreating does not clear the database. DESIGN.md §6 says the same, and the chart passes the setting through as terdutServer.credentials.deletionPolicy. A retained Secret of a TerdutServer that is gone for good is an orphan to delete by hand. It is inert: nothing adopts it unless the server accepts the token. Checked on the kind demo with a locally built image against the real terdut-server v0.43.0: deleting the TerdutServer kept the Secret, recreating it reached Ready with the same credential (identical hash) and both TerdutTeams came back Ready with their original ids. The controller specs cover adoption, a rejected token after a reset, the bootstrapped-and-rejected failure, and both deletion policies. Co-authored-by: Claude <noreply@anthropic.com>
Terdut operator
Aims to expose most config as CRD's, so end users can self-service over gitops.
See DESIGN.md for the full design: CRD catalog and specs,
reconciliation semantics, bootstrap/auth, Postgres integration, RBAC, and the
relationship to charts/terdut-server. This README stays a short pitch; the
open questions it used to carry are now resolved decisions there (§2).
CRD's
terdutServers
Creates a server — Deployment, Service, database wiring, bootstrap, operator
credentials, and allowedTeams consent for cross-namespace teams. See
DESIGN.md §4.1, §4.6.
terdutTeams
- team name
- oidc groups
serverRef— explicit reference to itsTerdutServer, may be in a different namespace (one team owns the server, others self-service a team against it), gated by thatTerdutServer's ownallowedTeamsfield (DESIGN.md §2, §4.1, §4.2, §4.6)
terdutEscalationrules
- rule
teamRef— explicit reference to itsTerdutTeam(DESIGN.md §2, §4.3)
terdutDeadmansswitches
- rule
teamRef(DESIGN.md §4.4)
terdutAlertSources
teamRef(DESIGN.md §4.5)- URL/key are generated by the server at creation and surfaced only via a generated Secret, never set explicitly
Demo
examples/demo wires one of every CRD above together
— two teams, each with an escalation rule, a dead man's switch and an
alert source — plus a script that fires synthetic Alertmanager webhooks
at it, so you can watch real incidents open, escalate and resolve without
a real Alertmanager anywhere in the picture.