The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the web UI since: the queue and incident layouts, the rota and escalation pages, the theme toggle, and the cluster chip, filter and page titles (v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the floor, which is what the replicas setting actually depends on. fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus external label plus `cluster` in Alertmanager's group_by (terdut-server's README, "Several clusters, one team"). It goes on the alert's labels and groupLabels, and into the group key and the fingerprint, so the same alert in two clusters is two incidents and not one. Unset, the payload is exactly what it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in both, so the queue has a chip and a filter to show. run-demo.sh also failed on its second run, though it says it is safe to re-run: it expected HTTP 409 when alice already exists, but a spent invite is answered with 403 "invite link is not usable" before the username is ever checked. It now tries to log alice in first and skips the signup if that works. Checked on the kind cluster: the server rolled to v0.43.0, every CR became Ready and Adopted (server, both teams, both escalation rules, both dead man's switches, both alert sources), and /api/incidents/clusters, /api/incidents?cluster=prod-us and the incident titles came back as expected. No operator code changed, so this needs no operator release. Co-authored-by: Claude <noreply@anthropic.com>
Terdut operator
Aims to expose most config as CRD's, so end users can self-service over gitops.
See DESIGN.md for the full design: CRD catalog and specs,
reconciliation semantics, bootstrap/auth, Postgres integration, RBAC, and the
relationship to charts/terdut-server. This README stays a short pitch; the
open questions it used to carry are now resolved decisions there (§2).
CRD's
terdutServers
Creates a server — Deployment, Service, database wiring, bootstrap, operator
credentials, and allowedTeams consent for cross-namespace teams. See
DESIGN.md §4.1, §4.6.
terdutTeams
- team name
- oidc groups
serverRef— explicit reference to itsTerdutServer, may be in a different namespace (one team owns the server, others self-service a team against it), gated by thatTerdutServer's ownallowedTeamsfield (DESIGN.md §2, §4.1, §4.2, §4.6)
terdutEscalationrules
- rule
teamRef— explicit reference to itsTerdutTeam(DESIGN.md §2, §4.3)
terdutDeadmansswitches
- rule
teamRef(DESIGN.md §4.4)
terdutAlertSources
teamRef(DESIGN.md §4.5)- URL/key are generated by the server at creation and surfaced only via a generated Secret, never set explicitly
Demo
examples/demo wires one of every CRD above together
— two teams, each with an escalation rule, a dead man's switch and an
alert source — plus a script that fires synthetic Alertmanager webhooks
at it, so you can watch real incidents open, escalate and resolve without
a real Alertmanager anywhere in the picture.