Files
terdut-operator/examples/demo/01-server.yaml
T
Niklas Ye 50ce5bcec0
CI / test (pull_request) Successful in 6m35s
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m6s
examples/demo: run terdut-server v0.43.0, with alerts from two clusters
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the
web UI since: the queue and incident layouts, the rota and escalation
pages, the theme toggle, and the cluster chip, filter and page titles
(v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the
floor, which is what the replicas setting actually depends on.

fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus
external label plus `cluster` in Alertmanager's group_by (terdut-server's
README, "Several clusters, one team"). It goes on the alert's labels and
groupLabels, and into the group key and the fingerprint, so the same alert in
two clusters is two incidents and not one. Unset, the payload is exactly what
it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in
both, so the queue has a chip and a filter to show.

run-demo.sh also failed on its second run, though it says it is safe to
re-run: it expected HTTP 409 when alice already exists, but a spent invite
is answered with 403 "invite link is not usable" before the username is ever
checked. It now tries to log alice in first and skips the signup if that works.

Checked on the kind cluster: the server rolled to v0.43.0, every CR became
Ready and Adopted (server, both teams, both escalation rules, both dead man's
switches, both alert sources), and /api/incidents/clusters,
/api/incidents?cluster=prod-us and the incident titles came back as expected.
No operator code changed, so this needs no operator release.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 18:55:02 +02:00

47 lines
2.1 KiB
YAML

# The one TerdutServer this whole demo runs against. Everything else in
# this directory (teams, escalation rules, dead man's switches, alert
# sources) references it by name.
#
# The operator never creates any ingress/HTTPRoute for this TerdutServer --
# that's a permanent non-goal (DESIGN.md §1, NetworkingSpec's own doc
# comment), not a missing feature. This demo reaches it only by
# port-forwarding its Service, same name as this object (see README.md);
# see ../networking for worked examples of exposing it yourself instead.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
name: terdut-operator-demo
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
# v0.36.0 is the floor now that replicas below is 2 (this demo pins
# the current release, v0.43.0, so it shows the current web UI too): that
# release put the sweeper, the notifier and the migration runner each
# behind a Postgres advisory lock, and gave incident creation its own
# conflict resolution, which is what makes a second replica safe
# instead of racing the first. (Still carries v0.34.0's fix too --
# callerMayManageServiceAccount, so an instance-scoped service account
# can adopt/rotate a key on a team-scoped account it didn't just create
# in the same call -- without which terdutteam-* can wedge permanently
# on the crash-window race this demo hit live, niklas/terdut-operator#3.)
tag: v0.43.0
# Matches this CRD's own spec.replicas default (v0.4.0) -- stated
# explicitly, like every other field in this file, rather than left to
# the default. RollingUpdate follows automatically; this operator does
# not expose Strategy as a spec field.
replicas: 2
networking:
hostname: terdut-operator-demo.example
servicePort: 8080
database:
dsn: "postgres://terdut@terdut-operator-demo-postgres:5432/terdut?sslmode=disable"
passwordSecretRef:
name: terdut-operator-demo-postgres
key: password
sweeper:
staleAfter: 6h
archiveAfter: 168h
# No oidc block: password login only, so there's nothing external to
# register a redirect URI with before this demo can sign in.
passwordLogin: true