examples/demo: run terdut-server v0.43.0, with alerts from two clusters #6

Merged
niklas merged 3 commits from demo-matches-v0.4.0-crd into main 2026-10-08 17:10:07 +00:00
Owner

Brings the kind demo up to date with terdut-server v0.43.0. The branch has three commits, all under examples/demo; no operator code, CRD or chart changed, so this needs no operator release.

  • Server pin: the demo moves from terdut-server v0.36.0 to v0.43.0, so it shows the current web UI. v0.36.0 stays the documented floor for replicas: 2.
  • Two clusters: fire-alerts.sh takes an optional CLUSTER (a stand-in for a Prometheus external label plus cluster in Alertmanager's group_by), and run-demo.sh fires its alerts across prod-eu and prod-us. The queue then shows the cluster chip and the cluster filter. The same alert in two clusters is two incidents.
  • Re-run fix: run-demo.sh failed on its second run: it expected HTTP 409 for an existing user, but a spent invite is answered with 403 first. It now checks that alice can log in and skips the signup.
  • Earlier commits on the branch: the replicas default and server bump to v0.36.0 (0ee7ede), and the split Ready wait so escalation rules wait on alice (822c80d).

Checked

On the existing kind cluster, ./run-demo.sh rolled the server to v0.43.0 and every CR became Ready and Adopted (server, both teams, both escalation rules, both dead man's switches, both alert sources). /api/incidents/clusters returned ["prod-eu","prod-us"], ?cluster=prod-us returned only that cluster's incidents, and the same alert in the two clusters was two incidents. Only the API was checked, not the web UI in a browser.

Brings the kind demo up to date with terdut-server v0.43.0. The branch has three commits, all under `examples/demo`; no operator code, CRD or chart changed, so this needs no operator release. - **Server pin:** the demo moves from terdut-server v0.36.0 to v0.43.0, so it shows the current web UI. v0.36.0 stays the documented floor for `replicas: 2`. - **Two clusters:** `fire-alerts.sh` takes an optional `CLUSTER` (a stand-in for a Prometheus external label plus `cluster` in Alertmanager's `group_by`), and `run-demo.sh` fires its alerts across `prod-eu` and `prod-us`. The queue then shows the cluster chip and the cluster filter. The same alert in two clusters is two incidents. - **Re-run fix:** `run-demo.sh` failed on its second run: it expected HTTP 409 for an existing user, but a spent invite is answered with 403 first. It now checks that alice can log in and skips the signup. - **Earlier commits on the branch:** the `replicas` default and server bump to v0.36.0 (`0ee7ede`), and the split Ready wait so escalation rules wait on alice (`822c80d`). ## Checked On the existing kind cluster, `./run-demo.sh` rolled the server to v0.43.0 and every CR became Ready and Adopted (server, both teams, both escalation rules, both dead man's switches, both alert sources). `/api/incidents/clusters` returned `["prod-eu","prod-us"]`, `?cluster=prod-us` returned only that cluster's incidents, and the same alert in the two clusters was two incidents. Only the API was checked, not the web UI in a browser.
niklas added 3 commits 2026-10-08 16:59:49 +00:00
replicas: 1 and tag: v0.34.0 were both correct when written, but the CRD's
own default moved to 2 in v0.4.0 (same release this demo is meant to show
off), and v0.34.0 predates v0.36.0's advisory locks that make a second
replica safe instead of racing the first. Left as-is, the demo would have
been the one place in this repo demonstrating the exact unsafe combination
the CRD's own doc comment warns against: more than one replica against an
image that doesn't guard the sweeper/notifier/migration-runner singletons.

replicas is now stated explicitly as 2 rather than dropped to pick up the
default silently, matching every other field in this file's own habit of
spelling out what it depends on. tag moves to v0.36.0 specifically -- the
first version where the lock landed -- with the comment keeping v0.34.0's
original reasoning (the service-account race fix) alongside the new one,
since v0.36.0 still carries that fix forward.

Co-authored-by: Claude <noreply@anthropic.com>
wait_for_ready waited for every demo object at once, including
terdutescalationrule-platform, which names alice as a level-1 target --
but alice does not exist yet at that point in main(): she is created by
redeem_platform_invite, which ran after wait_for_ready. terdut-server
resolves every named username at reconcile time, not just when an
escalation actually fires, so that CR could never reach Ready before
alice did, and main() had no step in between to create her.

Split into wait_for_objects (the shared loop, now taking its object list
as arguments) plus two callers: wait_for_teams_ready, covering just the
server and the two teams redeem_platform_invite/join_payments_team
need, run before alice exists; wait_for_remaining_ready, covering the
escalation rules, dead man's switches and alert sources, run after.

Co-authored-by: Claude <noreply@anthropic.com>
examples/demo: run terdut-server v0.43.0, with alerts from two clusters
CI / test (pull_request) Successful in 6m35s
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m6s
50ce5bcec0
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the
web UI since: the queue and incident layouts, the rota and escalation
pages, the theme toggle, and the cluster chip, filter and page titles
(v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the
floor, which is what the replicas setting actually depends on.

fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus
external label plus `cluster` in Alertmanager's group_by (terdut-server's
README, "Several clusters, one team"). It goes on the alert's labels and
groupLabels, and into the group key and the fingerprint, so the same alert in
two clusters is two incidents and not one. Unset, the payload is exactly what
it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in
both, so the queue has a chip and a filter to show.

run-demo.sh also failed on its second run, though it says it is safe to
re-run: it expected HTTP 409 when alice already exists, but a spent invite
is answered with 403 "invite link is not usable" before the username is ever
checked. It now tries to log alice in first and skips the signup if that works.

Checked on the kind cluster: the server rolled to v0.43.0, every CR became
Ready and Adopted (server, both teams, both escalation rules, both dead man's
switches, both alert sources), and /api/incidents/clusters,
/api/incidents?cluster=prod-us and the incident titles came back as expected.
No operator code changed, so this needs no operator release.

Co-authored-by: Claude <noreply@anthropic.com>
niklas merged commit 6572f63157 into main 2026-10-08 17:10:07 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: niklas/terdut-operator#6