TerdutTeam can wedge permanently in a 403 retry loop when its service-account mint succeeds but credential persistence fails #3
Reference in New Issue
Block a user
Delete Branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
Caught live while verifying
examples/demo/run-demo.shagainst a genuinelyfresh kind cluster (not a re-run, not a contrived scenario):
terdutteam-platformgot permanently stuck
Ready: Falsewhile its siblingterdutteam-payments(same manifest shape, same reconcile loop, same instant) converged fine.
Evidence, from the live cluster
service_accountstable — exactly one row for platform, proving theoriginal mint succeeded:
service_account_keys— exactly one key per account, including platform's:So
mintTeamCredential'sCreateTeamServiceAccountcall for platformreturned
201and minted key id 3 — same as payments. Yet the operator'sown log shows, at the very same timestamp:
409then403,with no way to ever reach a
201(create) or a successful adopt (mint).Permanent, unrecoverable
Ready: Falsefor thatTerdutTeamandeverything
teamRef-ing it (itsTerdutEscalationRule,TerdutDeadmanSwitch,TerdutAlertSourceall sat onWaitingForTeamfor the rest of the run).
Why this matters beyond the one unlucky reconcile
This is the same underlying gap as #23 on
terdut-server(authorizationlogic keyed on
userFromContext/callerIsAdmin, which a service-accountcaller can never satisfy) — but #23 only blocks admin-UI bootstrapping.
This one can permanently wedge a real production
TerdutTeamtheinstant its credential-persistence step has any hiccup at all (a pod
restart mid-reconcile, an apiserver conflict on the status update, anything
that separates "mint succeeded" from "Secret/status written" by even one
reconcile boundary) — with zero automatic recovery, since the fallback path
that's supposed to handle exactly this (
DESIGN.md §5's adopt-on-409 rule)is the thing that 403s.
Suggested fix directions (not prescriptive)
terdut-server(per #23): let aninstance-scoped service account satisfy
callerMayManageServiceAccountfor any team-scoped account, the same way it's already trusted to
create one.
terminal state distinct from "transient, keep requeuing" — a
CredentialMintStateLost-style condition (mirroringTerdutServer's ownbootstrapStateLostErrorhandling for the exact same class of problem)would at least make this visible/diagnosable instead of a silent infinite
error loop.
Repro
Not reliably deterministic on demand (didn't reproduce for
terdutteam-paymentsin the same run), but was hit on a genuinely first, fresh-cluster run of
examples/demo/run-demo.sh— no prior state, no re-run. Likely a timingrace around the Secret write / status update following a successful
service-account mint.
References
internal/controller/terdutteam_bootstrap.go—mintTeamCredential'screate-then-adopt-on-409 flow.
terdut-server/internal/api/service_accounts.go—callerMayManageServiceAccount,handleCreateServiceAccountKey.terdut-server/niklas#23— the sibling authorization gap on theadmin-settings/user-management routes.
DESIGN.md§5 (idempotent-create / adopt-on-conflict rule), §6 (bootstrapcredential lifecycle).