Commit Graph

3 Commits

Author SHA1 Message Date
Niklas Ye 50ce5bcec0 examples/demo: run terdut-server v0.43.0, with alerts from two clusters
CI / test (pull_request) Successful in 6m35s
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m6s
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the
web UI since: the queue and incident layouts, the rota and escalation
pages, the theme toggle, and the cluster chip, filter and page titles
(v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the
floor, which is what the replicas setting actually depends on.

fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus
external label plus `cluster` in Alertmanager's group_by (terdut-server's
README, "Several clusters, one team"). It goes on the alert's labels and
groupLabels, and into the group key and the fingerprint, so the same alert in
two clusters is two incidents and not one. Unset, the payload is exactly what
it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in
both, so the queue has a chip and a filter to show.

run-demo.sh also failed on its second run, though it says it is safe to
re-run: it expected HTTP 409 when alice already exists, but a spent invite
is answered with 403 "invite link is not usable" before the username is ever
checked. It now tries to log alice in first and skips the signup if that works.

Checked on the kind cluster: the server rolled to v0.43.0, every CR became
Ready and Adopted (server, both teams, both escalation rules, both dead man's
switches, both alert sources), and /api/incidents/clusters,
/api/incidents?cluster=prod-us and the incident titles came back as expected.
No operator code changed, so this needs no operator release.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 18:55:02 +02:00
Niklas Ye d9315322fc examples/demo: fix two real bugs this exact demo just hit live
CI / chart (push) Successful in 1s
CI / security (push) Successful in 3m24s
CI / test (push) Successful in 10m45s
1. Renamed every object this demo creates (TerdutServer, Postgres
   Secret/Deployment/Service) from terdut-demo[-postgres] to
   terdut-operator-demo[-postgres]. The user applied this kit into the
   already-live "terdut-demo" namespace -- the real operator exercise
   from earlier in this repo's own history -- and this demo's own
   TerdutServer/Postgres objects shared that exact name. The TerdutServer
   apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
   while the live object already had postgresClusterRef violates "exactly
   one of" and the API server refused it), and the real Postgres Service
   was never touched (confirmed live: still Zalando's own spilo selector,
   endpoint still the real StatefulSet pod) -- but the Postgres Secret and
   Deployment, having no such protection, were created as brand new,
   extra, crash-looping objects sitting right next to the real ones.
   Prefixing every name this demo creates means a repeat of this exact
   mistake no longer collides with anything, documented directly in
   README.md now.

2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
   (added responding to a PodSecurity "restricted" warning) took
   CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
   entrypoint needs to chown/chmod the data directory before it drops
   privileges itself -- confirmed in a real crashed pod's logs: `chmod:
   /var/run/postgresql: Operation not permitted`. kubectl apply
   --dry-run=server, which is as far as this got verified before, only
   checks admission policy; it was never actually booted. Removed the
   capability drop and verified for real this time: applied just
   00-postgres.yaml alone into a disposable namespace, waited for the pod
   to go Ready, read its logs ("database system is ready to accept
   connections"), then deleted that namespace.
2026-10-02 13:25:47 +02:00
Niklas Ye ca2cd2c645 Add examples/demo: one of every CRD, plus a script to fire alerts at it
CI / chart (push) Successful in 1s
CI / security (push) Failing after 1m7s
CI / test (push) Successful in 2m37s
A self-contained demo kit: a TerdutServer against a throwaway, bare
Postgres (bring-your-own DSN -- simplest path to stand up from nothing,
ROADMAP.md Stage 1's own note), two TerdutTeams, and each team's own
TerdutEscalationRule/TerdutDeadmanSwitch/TerdutAlertSource, so every CRD
this operator manages is exercised together rather than in isolation the
way config/samples' one-of-each already does.

fire-alerts.sh sends terdut-server's own amPayload/amAlert shape (read
from internal/api/alertmanager.go in that repo, not guessed from its
docs) at whichever TerdutAlertSource's generated webhook Secret it reads
the key out of -- high-cpu/disk-full/pod-crash scenarios to open and
resolve incidents, and a heartbeat scenario matching each team's dead
man's switch matcher, so stopping it demonstrates the switch noticing
silence on its own.

Verified server-side (kubectl apply --dry-run=server -k examples/demo)
against this operator's own dev cluster, which already has these CRDs
installed: every object validates. The one warning that cluster's
"restricted" PodSecurity raises (postgres:17-alpine's entrypoint needs to
start as root before it drops privileges itself) is noted inline in
00-postgres.yaml rather than worked around -- not a real production
pattern, and this Postgres exists only to be thrown away with the rest of
the demo namespace.

README.md walks through: applying, watching status, why a few early
CrashLoopBackOff restarts on terdut-demo itself are expected (this
operator's Deployment template has no wait-for-postgres init container
yet, unlike charts/terdut-server's chart as of v0.33.2), reaching the web
UI (port-forward -- spec.networking.hostname is accepted but nothing
creates an HTTPRoute for it yet), turning on open signup with the
operator's own generated admin token since the bootstrap-created account
has no password, firing alerts, and tearing down.
2026-10-02 10:03:49 +02:00