BootstrapStateLost's documented recovery ("delete and recreate") doesn't actually work #1
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What happened
Standing up
terdut-demofor real (Ryuvia/charts#275) against this operator: aNetworkPolicygap in that chart (since fixed, Ryuvia/charts#276) silently refused everyPOST /api/bootstrapcall the operator's own controller made —connection refusedon every single reconcile, even once theTerdutServer's own Deployment wasReady. That policy only ever blocked the operator's traffic, though — Envoy's own route to the pod was allowed from the start, so the app was reachable from outside the whole time the operator couldn't reach it.Corrected from the first version of this issue: the orphaned user wasn't a bootstrap response that got cut off mid-flight — checked against source and that guess was wrong. The operator's own bootstrap call always creates a user named
terdut-operator-bootstrap/bootstrap@terdut-operator.local(terdutserver_controller.go's own constants). The orphaned row was a real human identity instead. terdut-server's OIDC callback (internal/api/oidc.go'screateSSOUser) auto-provisions a brand-new user on first sign-in with no check at all for whether/api/bootstraphas ever run — only that the identity is in an allowed group. Someone signed in via OIDC against the already-reachable (via Envoy) but not-yet-bootstrapped server, which auto-created the first user and consumed the single-shot slot/api/bootstrapgates on (SELECT COUNT(*) FROM users). Every subsequent operator bootstrap attempt was doomed to403from that point on, independent of and invisible to the NetworkPolicy fix.That's exactly DESIGN.md §6's own documented fail-closed case:
The real gap(s)
Two, now that the trigger is correctly understood:
A. The recovery instruction doesn't work. Deleting and recreating the
TerdutServerCR only re-runs the finalizer's own cleanup, which — per its own documented scope — tears down the generated credentials Secret and the owned Deployment/Service. It has no way to touch the underlying Postgres: that's a separatepostgresql.acid.zalan.doCR the operator only ever reads a connection string from, never owns. The occupying user row stays put, so a fresh/api/bootstrapcall403s immediately for the same reason the first one did.The only way we actually recovered this time was
DELETE FROM users WHERE id = 1run directly against the database — not something this operator, or terdut-server's own API, can do on its own, and not something anyone should need to do by hand against a real install.B. Nothing stops a human from winning the bootstrap race at all. This is the more fundamental gap, and the NetworkPolicy bug is almost incidental to it:
/api/bootstrap's single-shot gate and OIDC auto-provisioning both key off the sameuserstable with zero coordination between them. Even with perfect connectivity, anyone who can reach a freshly-createdTerdutServer's public hostname and is in an allowed OIDC group can sign in before the operator's very first reconcile completes and take the bootstrap slot — this operator's whole design (§1, §6) assumes it always wins that race because it's the only caller, and that assumption is false the moment OIDC is configured and the hostname is reachable, which is the normal case, not an edge case.What to discuss
A few directions, not mutually exclusive:
/api/bootstraphas run once, or the operator's own networking/HTTPRoute work (still not implemented — DESIGN.md's own note onNetworkingSpec) deliberately not exposing the hostname until bootstrap succeeds.Probing reachability before calling
/api/bootstrap(the direction the first version of this issue proposed) would not have caught this: the operator was perfectly able to reach the server before I fixed the NetworkPolicy at the Gateway/Envoy path, since that's what the human used to sign in. It only helps the connectivity-specific failure mode, not this one. Gap B is the one actually worth settling.The
NetworkPolicyroot cause from this specific incident is already fixed (Ryuvia/charts#276). This issue is aboutBootstrapStateLoststaying unrecoverable without direct database access (gap A), and about nothing preventing a human from racing the operator's own bootstrap in the first place (gap B) — found this way, not by design review.