BootstrapStateLost's documented recovery ("delete and recreate") doesn't actually work #1

Open
opened 2026-10-01 20:31:15 +00:00 by niklas · 0 comments
Owner

What happened

Standing up terdut-demo for real (Ryuvia/charts#275) against this operator: a NetworkPolicy gap in that chart (since fixed, Ryuvia/charts#276) silently refused every POST /api/bootstrap call the operator's own controller made — connection refused on every single reconcile, even once the TerdutServer's own Deployment was Ready. That policy only ever blocked the operator's traffic, though — Envoy's own route to the pod was allowed from the start, so the app was reachable from outside the whole time the operator couldn't reach it.

Corrected from the first version of this issue: the orphaned user wasn't a bootstrap response that got cut off mid-flight — checked against source and that guess was wrong. The operator's own bootstrap call always creates a user named terdut-operator-bootstrap / bootstrap@terdut-operator.local (terdutserver_controller.go's own constants). The orphaned row was a real human identity instead. terdut-server's OIDC callback (internal/api/oidc.go's createSSOUser) auto-provisions a brand-new user on first sign-in with no check at all for whether /api/bootstrap has ever run — only that the identity is in an allowed group. Someone signed in via OIDC against the already-reachable (via Envoy) but not-yet-bootstrapped server, which auto-created the first user and consumed the single-shot slot /api/bootstrap gates on (SELECT COUNT(*) FROM users). Every subsequent operator bootstrap attempt was doomed to 403 from that point on, independent of and invisible to the NetworkPolicy fix.

That's exactly DESIGN.md §6's own documented fail-closed case:

Ready: False, reason: BootstrapStateLost
"This TerdutServer cannot recover a credential on its own; delete and recreate it"

The real gap(s)

Two, now that the trigger is correctly understood:

A. The recovery instruction doesn't work. Deleting and recreating the TerdutServer CR only re-runs the finalizer's own cleanup, which — per its own documented scope — tears down the generated credentials Secret and the owned Deployment/Service. It has no way to touch the underlying Postgres: that's a separate postgresql.acid.zalan.do CR the operator only ever reads a connection string from, never owns. The occupying user row stays put, so a fresh /api/bootstrap call 403s immediately for the same reason the first one did.

The only way we actually recovered this time was DELETE FROM users WHERE id = 1 run directly against the database — not something this operator, or terdut-server's own API, can do on its own, and not something anyone should need to do by hand against a real install.

B. Nothing stops a human from winning the bootstrap race at all. This is the more fundamental gap, and the NetworkPolicy bug is almost incidental to it: /api/bootstrap's single-shot gate and OIDC auto-provisioning both key off the same users table with zero coordination between them. Even with perfect connectivity, anyone who can reach a freshly-created TerdutServer's public hostname and is in an allowed OIDC group can sign in before the operator's very first reconcile completes and take the bootstrap slot — this operator's whole design (§1, §6) assumes it always wins that race because it's the only caller, and that assumption is false the moment OIDC is configured and the hostname is reachable, which is the normal case, not an edge case.

What to discuss

A few directions, not mutually exclusive:

  1. Fix the instructions, not the mechanism. If "delete and recreate the TerdutServer" is wrong, say what actually works (reset/recreate the Postgres alongside it) — in both the event message and DESIGN.md §6.
  2. Give the operator a real way to self-heal this. terdut-server exposes some admin-only "reset bootstrap state" endpoint the operator can call once it can prove ownership. A real new API surface with its own trust questions (similar in weight to the service-account work) — not small.
  3. Close the actual race (gap B). Some way for the operator to claim the install before any human-facing path (OIDC, password login if ever enabled) can create the first user — e.g. terdut-server refusing OIDC sign-in entirely until /api/bootstrap has run once, or the operator's own networking/HTTPRoute work (still not implemented — DESIGN.md's own note on NetworkingSpec) deliberately not exposing the hostname until bootstrap succeeds.
  4. Document the real recovery as what it actually is. At minimum: "delete the TerdutServer CR and the postgresql CR together" as the sanctioned path, even though that's more destructive than the current message implies.

Probing reachability before calling /api/bootstrap (the direction the first version of this issue proposed) would not have caught this: the operator was perfectly able to reach the server before I fixed the NetworkPolicy at the Gateway/Envoy path, since that's what the human used to sign in. It only helps the connectivity-specific failure mode, not this one. Gap B is the one actually worth settling.

The NetworkPolicy root cause from this specific incident is already fixed (Ryuvia/charts#276). This issue is about BootstrapStateLost staying unrecoverable without direct database access (gap A), and about nothing preventing a human from racing the operator's own bootstrap in the first place (gap B) — found this way, not by design review.

## What happened Standing up `terdut-demo` for real (Ryuvia/charts#275) against this operator: a `NetworkPolicy` gap in that chart (since fixed, Ryuvia/charts#276) silently refused every `POST /api/bootstrap` call the operator's own controller made — `connection refused` on every single reconcile, even once the `TerdutServer`'s own Deployment was `Ready`. That policy only ever blocked the *operator's* traffic, though — Envoy's own route to the pod was allowed from the start, so the app was reachable from outside the whole time the operator couldn't reach it. **Corrected from the first version of this issue**: the orphaned user wasn't a bootstrap response that got cut off mid-flight — checked against source and that guess was wrong. The operator's own bootstrap call always creates a user named `terdut-operator-bootstrap` / `bootstrap@terdut-operator.local` (`terdutserver_controller.go`'s own constants). The orphaned row was a real human identity instead. terdut-server's OIDC callback (`internal/api/oidc.go`'s `createSSOUser`) auto-provisions a brand-new user on first sign-in with **no check at all** for whether `/api/bootstrap` has ever run — only that the identity is in an allowed group. Someone signed in via OIDC against the already-reachable (via Envoy) but not-yet-bootstrapped server, which auto-created the first user and consumed the single-shot slot `/api/bootstrap` gates on (`SELECT COUNT(*) FROM users`). Every subsequent operator bootstrap attempt was doomed to `403` from that point on, independent of and invisible to the NetworkPolicy fix. That's exactly DESIGN.md §6's own documented fail-closed case: ``` Ready: False, reason: BootstrapStateLost "This TerdutServer cannot recover a credential on its own; delete and recreate it" ``` ## The real gap(s) Two, now that the trigger is correctly understood: **A. The recovery instruction doesn't work.** Deleting and recreating the `TerdutServer` CR only re-runs the finalizer's own cleanup, which — per its own documented scope — tears down the generated credentials Secret and the owned Deployment/Service. It has no way to touch the underlying Postgres: that's a separate `postgresql.acid.zalan.do` CR the operator only ever reads a connection string from, never owns. The occupying user row stays put, so a fresh `/api/bootstrap` call `403`s immediately for the same reason the first one did. The only way we actually recovered this time was `DELETE FROM users WHERE id = 1` run directly against the database — not something this operator, or terdut-server's own API, can do on its own, and not something anyone should need to do by hand against a real install. **B. Nothing stops a human from winning the bootstrap race at all.** This is the more fundamental gap, and the NetworkPolicy bug is almost incidental to it: `/api/bootstrap`'s single-shot gate and OIDC auto-provisioning both key off the same `users` table with zero coordination between them. Even with perfect connectivity, anyone who can reach a freshly-created `TerdutServer`'s public hostname and is in an allowed OIDC group can sign in before the operator's very first reconcile completes and take the bootstrap slot — this operator's whole design (§1, §6) assumes it always wins that race because *it's the only caller*, and that assumption is false the moment OIDC is configured and the hostname is reachable, which is the normal case, not an edge case. ## What to discuss A few directions, not mutually exclusive: 1. **Fix the instructions, not the mechanism.** If "delete and recreate the TerdutServer" is wrong, say what actually works (reset/recreate the Postgres alongside it) — in both the event message and DESIGN.md §6. 2. **Give the operator a real way to self-heal this.** terdut-server exposes some admin-only "reset bootstrap state" endpoint the operator can call once it can prove ownership. A real new API surface with its own trust questions (similar in weight to the service-account work) — not small. 3. **Close the actual race (gap B).** Some way for the operator to claim the install before any human-facing path (OIDC, password login if ever enabled) can create the first user — e.g. terdut-server refusing OIDC sign-in entirely until `/api/bootstrap` has run once, or the operator's own networking/HTTPRoute work (still not implemented — DESIGN.md's own note on `NetworkingSpec`) deliberately not exposing the hostname until bootstrap succeeds. 4. **Document the real recovery as what it actually is.** At minimum: "delete the TerdutServer CR *and* the postgresql CR together" as the sanctioned path, even though that's more destructive than the current message implies. Probing reachability before calling `/api/bootstrap` (the direction the first version of this issue proposed) would **not** have caught this: the operator was perfectly able to reach the server before I fixed the NetworkPolicy at the Gateway/Envoy path, since that's what the human used to sign in. It only helps the connectivity-specific failure mode, not this one. Gap B is the one actually worth settling. The `NetworkPolicy` root cause from this specific incident is already fixed (Ryuvia/charts#276). This issue is about `BootstrapStateLost` staying unrecoverable without direct database access (gap A), and about nothing preventing a human from racing the operator's own bootstrap in the first place (gap B) — found this way, not by design review.
Sign in to join this conversation.
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: niklas/terdut-operator#1