478ae6284adc943fb035ff93a8f09e57b1ec27e6
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a0ea13955e |
TerdutTeam: mint and surface a real invite link (spec.invite)
The actual fix for the human-onboarding gap niklas/terdut-server#23 found -- not a terdut-server change at all. A team-scoped credential is already owner-equivalent for POST/GET/DELETE /api/teams/{teamID}/invites (requireTeamOwner's synthetic-membership mechanism, ratified not accidental per that repo's SERVICE-ACCOUNTS.md), and invite redemption bypasses signup_mode entirely -- this TerdutTeam controller just never grew a feature to use either fact. New spec.invite{enabled, role (member|owner, default member), maxUses (1-100, default 1)} and status.inviteSecretRef. The Secret lives in the TerdutTeam's OWN namespace, not the operator's: unlike status.credentialsSecretRef (a durable, high-privilege credential, kept operator-side per DESIGN.md §6), an invite is bounded and limited-use, meant for this namespace's own human operators to read and hand out -- same precedent as TerdutAlertSource's status.webhookURLSecretRef, same- namespace and OwnerReference'd so deleting the TerdutTeam garbage-collects it automatically. internal/controller/terdutteam_invite.go: mints on first spec.invite.enabled, refreshes a day ahead of terdut-server's fixed 7-day TTL (reading the Secret's own stored expiresAt, no extra server round-trip per reconcile), revokes server-side and deletes the Secret when flipped back to false. A lost invite Secret is silently re-minted rather than treated as unrecoverable the way TerdutAlertSource's webhook key is -- nothing external holds a durable dependency on one specific invite link staying stable, it's read once by one human and handed out. New tdclient.Invite/CreateInvite/RevokeInvite. New envtest coverage: mint into the team's own namespace, refresh-before-expiry, revoke-on-disable (internal/controller/terdutteam_controller_test.go's new "spec.invite" Describe block), plus the fake server growing invite support (terdutserver_controller_test.go) -- its handleTeamSubPath dispatcher was split further (deadman switches into their own handleDeadmanSubPath, matching the existing handleIntegrationSubPath precedent) to stay under golangci-lint's gocyclo threshold with the new route added. examples/demo updated to prove this end to end: 02-team-platform.yaml turns on spec.invite; run-demo.sh's bootstrap_login/join_demo_teams (the psql signup_mode flip + a direct team_members INSERT) are replaced by redeem_platform_invite (reads status.inviteSecretRef, a real POST /api/signup with the invite token) and join_payments_team (POST /api/teams/{teamID}/members using Payments' own credential and alice's user id resolved via GET /api/users, deliberately not given its own spec.invite, so the demo shows both onboarding paths this feature unlocks) -- zero kubectl exec/psql calls remain anywhere in the script. README.md's "First login" section rewritten to match; it no longer documents the admin-token curl call that 403s against current terdut-server (niklas/terdut-server#23). Depends on niklas/terdut-server#24 (the callerMayManageServiceAccount fix for terdut-operator#3) being released before this is deployed for real -- not required to build or test this change itself, since the envtest fake never modeled that authorization gap to begin with. |
||
|
|
d9315322fc |
examples/demo: fix two real bugs this exact demo just hit live
1. Renamed every object this demo creates (TerdutServer, Postgres
Secret/Deployment/Service) from terdut-demo[-postgres] to
terdut-operator-demo[-postgres]. The user applied this kit into the
already-live "terdut-demo" namespace -- the real operator exercise
from earlier in this repo's own history -- and this demo's own
TerdutServer/Postgres objects shared that exact name. The TerdutServer
apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
while the live object already had postgresClusterRef violates "exactly
one of" and the API server refused it), and the real Postgres Service
was never touched (confirmed live: still Zalando's own spilo selector,
endpoint still the real StatefulSet pod) -- but the Postgres Secret and
Deployment, having no such protection, were created as brand new,
extra, crash-looping objects sitting right next to the real ones.
Prefixing every name this demo creates means a repeat of this exact
mistake no longer collides with anything, documented directly in
README.md now.
2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
(added responding to a PodSecurity "restricted" warning) took
CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
entrypoint needs to chown/chmod the data directory before it drops
privileges itself -- confirmed in a real crashed pod's logs: `chmod:
/var/run/postgresql: Operation not permitted`. kubectl apply
--dry-run=server, which is as far as this got verified before, only
checks admission policy; it was never actually booted. Removed the
capability drop and verified for real this time: applied just
00-postgres.yaml alone into a disposable namespace, waited for the pod
to go Ready, read its logs ("database system is ready to accept
connections"), then deleted that namespace.
|
||
|
|
ca2cd2c645 |
Add examples/demo: one of every CRD, plus a script to fire alerts at it
A self-contained demo kit: a TerdutServer against a throwaway, bare Postgres (bring-your-own DSN -- simplest path to stand up from nothing, ROADMAP.md Stage 1's own note), two TerdutTeams, and each team's own TerdutEscalationRule/TerdutDeadmanSwitch/TerdutAlertSource, so every CRD this operator manages is exercised together rather than in isolation the way config/samples' one-of-each already does. fire-alerts.sh sends terdut-server's own amPayload/amAlert shape (read from internal/api/alertmanager.go in that repo, not guessed from its docs) at whichever TerdutAlertSource's generated webhook Secret it reads the key out of -- high-cpu/disk-full/pod-crash scenarios to open and resolve incidents, and a heartbeat scenario matching each team's dead man's switch matcher, so stopping it demonstrates the switch noticing silence on its own. Verified server-side (kubectl apply --dry-run=server -k examples/demo) against this operator's own dev cluster, which already has these CRDs installed: every object validates. The one warning that cluster's "restricted" PodSecurity raises (postgres:17-alpine's entrypoint needs to start as root before it drops privileges itself) is noted inline in 00-postgres.yaml rather than worked around -- not a real production pattern, and this Postgres exists only to be thrown away with the rest of the demo namespace. README.md walks through: applying, watching status, why a few early CrashLoopBackOff restarts on terdut-demo itself are expected (this operator's Deployment template has no wait-for-postgres init container yet, unlike charts/terdut-server's chart as of v0.33.2), reaching the web UI (port-forward -- spec.networking.hostname is accepted but nothing creates an HTTPRoute for it yet), turning on open signup with the operator's own generated admin token since the bootstrap-created account has no password, firing alerts, and tearing down. |