The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the
web UI since: the queue and incident layouts, the rota and escalation
pages, the theme toggle, and the cluster chip, filter and page titles
(v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the
floor, which is what the replicas setting actually depends on.
fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus
external label plus `cluster` in Alertmanager's group_by (terdut-server's
README, "Several clusters, one team"). It goes on the alert's labels and
groupLabels, and into the group key and the fingerprint, so the same alert in
two clusters is two incidents and not one. Unset, the payload is exactly what
it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in
both, so the queue has a chip and a filter to show.
run-demo.sh also failed on its second run, though it says it is safe to
re-run: it expected HTTP 409 when alice already exists, but a spent invite
is answered with 403 "invite link is not usable" before the username is ever
checked. It now tries to log alice in first and skips the signup if that works.
Checked on the kind cluster: the server rolled to v0.43.0, every CR became
Ready and Adopted (server, both teams, both escalation rules, both dead man's
switches, both alert sources), and /api/incidents/clusters,
/api/incidents?cluster=prod-us and the incident titles came back as expected.
No operator code changed, so this needs no operator release.
Co-authored-by: Claude <noreply@anthropic.com>
wait_for_ready waited for every demo object at once, including
terdutescalationrule-platform, which names alice as a level-1 target --
but alice does not exist yet at that point in main(): she is created by
redeem_platform_invite, which ran after wait_for_ready. terdut-server
resolves every named username at reconcile time, not just when an
escalation actually fires, so that CR could never reach Ready before
alice did, and main() had no step in between to create her.
Split into wait_for_objects (the shared loop, now taking its object list
as arguments) plus two callers: wait_for_teams_ready, covering just the
server and the two teams redeem_platform_invite/join_payments_team
need, run before alice exists; wait_for_remaining_ready, covering the
escalation rules, dead man's switches and alert sources, run after.
Co-authored-by: Claude <noreply@anthropic.com>
replicas: 1 and tag: v0.34.0 were both correct when written, but the CRD's
own default moved to 2 in v0.4.0 (same release this demo is meant to show
off), and v0.34.0 predates v0.36.0's advisory locks that make a second
replica safe instead of racing the first. Left as-is, the demo would have
been the one place in this repo demonstrating the exact unsafe combination
the CRD's own doc comment warns against: more than one replica against an
image that doesn't guard the sweeper/notifier/migration-runner singletons.
replicas is now stated explicitly as 2 rather than dropped to pick up the
default silently, matching every other field in this file's own habit of
spelling out what it depends on. tag moves to v0.36.0 specifically -- the
first version where the lock landed -- with the comment keeping v0.34.0's
original reasoning (the service-account race fix) alongside the new one,
since v0.36.0 still carries that fix forward.
Co-authored-by: Claude <noreply@anthropic.com>