9b59e4bdbc27a62181abb62b27508bdca19036c2
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e1103f2b7d |
Authenticate with a seeded operator key; fold escalation and switches into TerdutTeam
Credentials: the TerdutServer controller generates <name>-operator-key in the server's own namespace (owned by it) and hands it to the pods as TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it at every start. A replaced Secret rolls the pods. The bootstrap handshake, the checkpoint Secret, per-team service accounts and credentials Secrets, BootstrapStateLost and credentials.deletionPolicy are gone. CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[] on the team (matched by name, extras removed); team invites are removed. A team is created under the identity <namespace>/<name> (external_id), so a retry, a lost status or a deleted team heal by repeating the same call, and a display name owned by another team is TeamNameTaken instead of an adoption. The server resolves escalation usernames (UnknownUser condition). OIDC claim names and trustEmail are spec fields. Fixes: query values are URL-escaped; every delete treats 404 as success; deleting a team no longer depends on allowedTeams consent; a switch or integration deleted on the server is recreated; unnamed switches take the CR's name. Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo (run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays cluster-wide, now stated in DESIGN.md section 9. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN |
||
|
|
50ce5bcec0 |
examples/demo: run terdut-server v0.43.0, with alerts from two clusters
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the web UI since: the queue and incident layouts, the rota and escalation pages, the theme toggle, and the cluster chip, filter and page titles (v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the floor, which is what the replicas setting actually depends on. fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus external label plus `cluster` in Alertmanager's group_by (terdut-server's README, "Several clusters, one team"). It goes on the alert's labels and groupLabels, and into the group key and the fingerprint, so the same alert in two clusters is two incidents and not one. Unset, the payload is exactly what it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in both, so the queue has a chip and a filter to show. run-demo.sh also failed on its second run, though it says it is safe to re-run: it expected HTTP 409 when alice already exists, but a spent invite is answered with 403 "invite link is not usable" before the username is ever checked. It now tries to log alice in first and skips the signup if that works. Checked on the kind cluster: the server rolled to v0.43.0, every CR became Ready and Adopted (server, both teams, both escalation rules, both dead man's switches, both alert sources), and /api/incidents/clusters, /api/incidents?cluster=prod-us and the incident titles came back as expected. No operator code changed, so this needs no operator release. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
d9315322fc |
examples/demo: fix two real bugs this exact demo just hit live
1. Renamed every object this demo creates (TerdutServer, Postgres
Secret/Deployment/Service) from terdut-demo[-postgres] to
terdut-operator-demo[-postgres]. The user applied this kit into the
already-live "terdut-demo" namespace -- the real operator exercise
from earlier in this repo's own history -- and this demo's own
TerdutServer/Postgres objects shared that exact name. The TerdutServer
apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
while the live object already had postgresClusterRef violates "exactly
one of" and the API server refused it), and the real Postgres Service
was never touched (confirmed live: still Zalando's own spilo selector,
endpoint still the real StatefulSet pod) -- but the Postgres Secret and
Deployment, having no such protection, were created as brand new,
extra, crash-looping objects sitting right next to the real ones.
Prefixing every name this demo creates means a repeat of this exact
mistake no longer collides with anything, documented directly in
README.md now.
2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
(added responding to a PodSecurity "restricted" warning) took
CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
entrypoint needs to chown/chmod the data directory before it drops
privileges itself -- confirmed in a real crashed pod's logs: `chmod:
/var/run/postgresql: Operation not permitted`. kubectl apply
--dry-run=server, which is as far as this got verified before, only
checks admission policy; it was never actually booted. Removed the
capability drop and verified for real this time: applied just
00-postgres.yaml alone into a disposable namespace, waited for the pod
to go Ready, read its logs ("database system is ready to accept
connections"), then deleted that namespace.
|
||
|
|
ca2cd2c645 |
Add examples/demo: one of every CRD, plus a script to fire alerts at it
A self-contained demo kit: a TerdutServer against a throwaway, bare Postgres (bring-your-own DSN -- simplest path to stand up from nothing, ROADMAP.md Stage 1's own note), two TerdutTeams, and each team's own TerdutEscalationRule/TerdutDeadmanSwitch/TerdutAlertSource, so every CRD this operator manages is exercised together rather than in isolation the way config/samples' one-of-each already does. fire-alerts.sh sends terdut-server's own amPayload/amAlert shape (read from internal/api/alertmanager.go in that repo, not guessed from its docs) at whichever TerdutAlertSource's generated webhook Secret it reads the key out of -- high-cpu/disk-full/pod-crash scenarios to open and resolve incidents, and a heartbeat scenario matching each team's dead man's switch matcher, so stopping it demonstrates the switch noticing silence on its own. Verified server-side (kubectl apply --dry-run=server -k examples/demo) against this operator's own dev cluster, which already has these CRDs installed: every object validates. The one warning that cluster's "restricted" PodSecurity raises (postgres:17-alpine's entrypoint needs to start as root before it drops privileges itself) is noted inline in 00-postgres.yaml rather than worked around -- not a real production pattern, and this Postgres exists only to be thrown away with the rest of the demo namespace. README.md walks through: applying, watching status, why a few early CrashLoopBackOff restarts on terdut-demo itself are expected (this operator's Deployment template has no wait-for-postgres init container yet, unlike charts/terdut-server's chart as of v0.33.2), reaching the web UI (port-forward -- spec.networking.hostname is accepted but nothing creates an HTTPRoute for it yet), turning on open signup with the operator's own generated admin token since the bootstrap-created account has no password, firing alerts, and tearing down. |