e2c4475867b520c6995e5966e52f7ebc291d49f8
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e1103f2b7d |
Authenticate with a seeded operator key; fold escalation and switches into TerdutTeam
Credentials: the TerdutServer controller generates <name>-operator-key in the server's own namespace (owned by it) and hands it to the pods as TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it at every start. A replaced Secret rolls the pods. The bootstrap handshake, the checkpoint Secret, per-team service accounts and credentials Secrets, BootstrapStateLost and credentials.deletionPolicy are gone. CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[] on the team (matched by name, extras removed); team invites are removed. A team is created under the identity <namespace>/<name> (external_id), so a retry, a lost status or a deleted team heal by repeating the same call, and a display name owned by another team is TeamNameTaken instead of an adoption. The server resolves escalation usernames (UnknownUser condition). OIDC claim names and trustEmail are spec fields. Fixes: query values are URL-escaped; every delete treats 404 as success; deleting a team no longer depends on allowedTeams consent; a switch or integration deleted on the server is recreated; unnamed switches take the CR's name. Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo (run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays cluster-wide, now stated in DESIGN.md section 9. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN |
||
|
|
50ce5bcec0 |
examples/demo: run terdut-server v0.43.0, with alerts from two clusters
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the web UI since: the queue and incident layouts, the rota and escalation pages, the theme toggle, and the cluster chip, filter and page titles (v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the floor, which is what the replicas setting actually depends on. fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus external label plus `cluster` in Alertmanager's group_by (terdut-server's README, "Several clusters, one team"). It goes on the alert's labels and groupLabels, and into the group key and the fingerprint, so the same alert in two clusters is two incidents and not one. Unset, the payload is exactly what it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in both, so the queue has a chip and a filter to show. run-demo.sh also failed on its second run, though it says it is safe to re-run: it expected HTTP 409 when alice already exists, but a spent invite is answered with 403 "invite link is not usable" before the username is ever checked. It now tries to log alice in first and skips the signup if that works. Checked on the kind cluster: the server rolled to v0.43.0, every CR became Ready and Adopted (server, both teams, both escalation rules, both dead man's switches, both alert sources), and /api/incidents/clusters, /api/incidents?cluster=prod-us and the incident titles came back as expected. No operator code changed, so this needs no operator release. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
822c80dda6 |
examples/demo: split the Ready wait so escalation rules wait on alice too
wait_for_ready waited for every demo object at once, including terdutescalationrule-platform, which names alice as a level-1 target -- but alice does not exist yet at that point in main(): she is created by redeem_platform_invite, which ran after wait_for_ready. terdut-server resolves every named username at reconcile time, not just when an escalation actually fires, so that CR could never reach Ready before alice did, and main() had no step in between to create her. Split into wait_for_objects (the shared loop, now taking its object list as arguments) plus two callers: wait_for_teams_ready, covering just the server and the two teams redeem_platform_invite/join_payments_team need, run before alice exists; wait_for_remaining_ready, covering the escalation rules, dead man's switches and alert sources, run after. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
a0ea13955e |
TerdutTeam: mint and surface a real invite link (spec.invite)
The actual fix for the human-onboarding gap niklas/terdut-server#23 found -- not a terdut-server change at all. A team-scoped credential is already owner-equivalent for POST/GET/DELETE /api/teams/{teamID}/invites (requireTeamOwner's synthetic-membership mechanism, ratified not accidental per that repo's SERVICE-ACCOUNTS.md), and invite redemption bypasses signup_mode entirely -- this TerdutTeam controller just never grew a feature to use either fact. New spec.invite{enabled, role (member|owner, default member), maxUses (1-100, default 1)} and status.inviteSecretRef. The Secret lives in the TerdutTeam's OWN namespace, not the operator's: unlike status.credentialsSecretRef (a durable, high-privilege credential, kept operator-side per DESIGN.md §6), an invite is bounded and limited-use, meant for this namespace's own human operators to read and hand out -- same precedent as TerdutAlertSource's status.webhookURLSecretRef, same- namespace and OwnerReference'd so deleting the TerdutTeam garbage-collects it automatically. internal/controller/terdutteam_invite.go: mints on first spec.invite.enabled, refreshes a day ahead of terdut-server's fixed 7-day TTL (reading the Secret's own stored expiresAt, no extra server round-trip per reconcile), revokes server-side and deletes the Secret when flipped back to false. A lost invite Secret is silently re-minted rather than treated as unrecoverable the way TerdutAlertSource's webhook key is -- nothing external holds a durable dependency on one specific invite link staying stable, it's read once by one human and handed out. New tdclient.Invite/CreateInvite/RevokeInvite. New envtest coverage: mint into the team's own namespace, refresh-before-expiry, revoke-on-disable (internal/controller/terdutteam_controller_test.go's new "spec.invite" Describe block), plus the fake server growing invite support (terdutserver_controller_test.go) -- its handleTeamSubPath dispatcher was split further (deadman switches into their own handleDeadmanSubPath, matching the existing handleIntegrationSubPath precedent) to stay under golangci-lint's gocyclo threshold with the new route added. examples/demo updated to prove this end to end: 02-team-platform.yaml turns on spec.invite; run-demo.sh's bootstrap_login/join_demo_teams (the psql signup_mode flip + a direct team_members INSERT) are replaced by redeem_platform_invite (reads status.inviteSecretRef, a real POST /api/signup with the invite token) and join_payments_team (POST /api/teams/{teamID}/members using Payments' own credential and alice's user id resolved via GET /api/users, deliberately not given its own spec.invite, so the demo shows both onboarding paths this feature unlocks) -- zero kubectl exec/psql calls remain anywhere in the script. README.md's "First login" section rewritten to match; it no longer documents the admin-token curl call that 403s against current terdut-server (niklas/terdut-server#23). Depends on niklas/terdut-server#24 (the callerMayManageServiceAccount fix for terdut-operator#3) being released before this is deployed for real -- not required to build or test this change itself, since the envtest fake never modeled that authorization gap to begin with. |
||
|
|
2a08a8cd8e |
examples/demo: add run-demo.sh, an automated kind-cluster demo
One script, two modes (run-demo.sh / run-demo.sh --teardown), that takes a fresh empty kind cluster all the way to a working demo: creates the cluster if needed, helm-installs this chart, applies every CRD kind in this directory, waits for all nine objects to go Ready, then does what the README's own first-login section cannot (see niklas/terdut-server#23 and niklas/terdut-operator#3 -- no service-account credential this operator holds can ever call /api/admin/settings or POST /api/users) by reaching into the demo's own throwaway Postgres directly: flips signup_mode to open, signs alice up for real over the ordinary signup endpoint, and joins her to both Platform and Payments (open signup always creates its own new team, never joins an existing one by name, so without this she'd have a working login that can't see a single incident this demo fires -- /api/incidents and /api/alerts are both scoped to the caller's own team memberships). Finishes by port-forwarding the service and firing fire-alerts.sh at both teams, so a fresh run already has visible incidents waiting in the web UI. Verified end to end against a real kind cluster, including a second, genuinely-fresh run that hit niklas/terdut-operator#3 live (terdutteam- platform wedged in the 403 retry loop that issue describes) -- confirmed the script itself fails cleanly on that (clear FAILED message, correct exit code, no orphaned port-forward) rather than hanging or leaving a mess, which is the most this script can do about a bug in the operator it's driving. |