Files
terdut-operator/examples/demo
Niklas Ye 50ce5bcec0
CI / test (pull_request) Successful in 6m35s
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m6s
examples/demo: run terdut-server v0.43.0, with alerts from two clusters
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the
web UI since: the queue and incident layouts, the rota and escalation
pages, the theme toggle, and the cluster chip, filter and page titles
(v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the
floor, which is what the replicas setting actually depends on.

fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus
external label plus `cluster` in Alertmanager's group_by (terdut-server's
README, "Several clusters, one team"). It goes on the alert's labels and
groupLabels, and into the group key and the fingerprint, so the same alert in
two clusters is two incidents and not one. Unset, the payload is exactly what
it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in
both, so the queue has a chip and a filter to show.

run-demo.sh also failed on its second run, though it says it is safe to
re-run: it expected HTTP 409 when alice already exists, but a spent invite
is answered with 403 "invite link is not usable" before the username is ever
checked. It now tries to log alice in first and skips the signup if that works.

Checked on the kind cluster: the server rolled to v0.43.0, every CR became
Ready and Adopted (server, both teams, both escalation rules, both dead man's
switches, both alert sources), and /api/incidents/clusters,
/api/incidents?cluster=prod-us and the incident titles came back as expected.
No operator code changed, so this needs no operator release.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 18:55:02 +02:00
..

Demo

Every CRD this operator reconciles, wired into one working install: one TerdutServer, two TerdutTeams (Platform and Payments) each with their own TerdutEscalationRule, TerdutDeadmanSwitch and TerdutAlertSource, plus a script that fires synthetic Alertmanager webhooks at it so you can watch real incidents appear, escalate and resolve.

This is a demo kit, not a reference deployment: 00-postgres.yaml runs Postgres with emptyDir storage and a password committed in this directory. Throw the whole namespace away when you're done.

Want this fully automated instead of walking through it by hand? ./run-demo.sh does everything below itself, against a fresh (or already-set-up) kind cluster — creates the cluster, installs the operator, applies every CR here, signs alice in for real, and fires a few alerts. ./run-demo.sh --help for the knobs, ./run-demo.sh --teardown to tear it back down. The rest of this file is the manual walkthrough it automates.

Prerequisites

  • The operator and its CRDs installed and running (make install deploy IMG=..., or charts/terdut-operator — see this repo's own README.md/DESIGN.md), pointed at a cluster you're fine creating throwaway resources in. A kind cluster is the easy choice.
  • kubectl, jq, curl on your path.

Apply this into a namespace of its own. Every object name in this directory is prefixed terdut-operator-demo specifically so applying it by mistake into some other namespace that already has unrelated objects doesn't collide with them -- but that only helps if this directory's own objects don't collide with each other across two applies. Applying it twice into two different namespaces is fine; applying it a second time into a namespace that already has something else named terdut-demo (a real install from following terdut-operator's own repo along, say) is exactly the mistake this prefix exists to avoid, and it only works if you don't override these names yourself.

Apply it

kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .

Listed and applied in dependency order (server → team → everything that teamRefs it), but you don't have to preserve that order yourself: every controller here re-queues and waits rather than failing when a ref isn't resolvable yet (kubectl describe shows Reason: TeamRefNotFound / WaitingForTeam while that settles).

Watch it converge:

kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
  -n terdut-operator-demo

The operator's own Deployment template now carries a wait-for-postgres init container (same fix as charts/terdut-server's chart as of v0.33.2), so terdut-operator-demo's pod should come up clean even against this brand-new Postgres doing its very first boot — no CrashLoopBackOff expected here.

Once terdut-operator-demo's own Ready condition is True, everything downstream of it should settle within a reconcile interval or two.

See the web UI

The operator never creates any external exposure for a TerdutServer -- that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this demo just reaches it the simplest way there is:

kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080

and open http://localhost:8080. See ../networking for worked examples of exposing it for real (Gateway API, Istio, or a plain Ingress) instead.

First login

The operator's own bootstrap (DESIGN.md §6) creates the first user through /api/bootstrap and immediately mints itself a service-account token from it, then discards the bootstrap user's own key — nobody ever signs in as that account, and signup_mode stays invite_only by default. Don't try to flip it via the operator's own token: that token is a service account, and /api/admin/settings is deliberately human-only on terdut-server (niklas/terdut-server#23 has the full reasoning — widening that gate was the wrong fix).

The real path in: 02-team-platform.yaml turns on spec.invite, so Platform's own TerdutTeam mints a real invite link with its own already-working team-scoped credential (the same reach that lets it manage its own escalation policy, dead man's switches and integrations — owner- equivalent, confirmed in terdut-server's SERVICE-ACCOUNTS.md). Invite redemption bypasses signup_mode entirely, so this needs no admin credential at all:

secretname=$(kubectl -n terdut-operator-demo get terdutteam terdutteam-platform \
  -o jsonpath='{.status.inviteSecretRef.name}')
url=$(kubectl -n terdut-operator-demo get secret "$secretname" -o jsonpath='{.data.url}' | base64 -d)
echo "$url"   # open this, or POST /api/signup with {"invite": "<the token after invite=>", ...}

04-escalation-platform.yaml names a user alice at its first escalation level — sign up as alice if you want that level to mean something rather than falling through to on-call after 5 minutes. run-demo.sh does exactly this automatically (and also joins alice to Payments, which deliberately has no spec.invite of its own — see that file's comment for the second onboarding path this demonstrates).

Fire some alerts

In another terminal, with the port-forward above still running:

export NAMESPACE=terdut-operator-demo

./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash

# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
# levels, and show up on the Platform/Payments team's own incident list.

./fire-alerts.sh platform high-cpu resolve

fire-alerts.sh -h (or any bad argument) prints the full scenario list. Each (team, scenario) pair is one stable fingerprint, so firing the same one twice updates the same alert (a real re-fire) and resolve closes exactly that one.

Set CLUSTER to send the alert as if it came from one of several clusters:

CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-us ./fire-alerts.sh platform high-cpu   # a second incident, not a join
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve

It stands in for a Prometheus external label plus cluster in Alertmanager's group_by (terdut-server's README, "Several clusters, one team"): the web UI then shows the cluster chip on each incident and a cluster filter in the queue. CLUSTER is part of the fingerprint, so resolve with the same value you fired with. ./run-demo.sh fires its alerts across prod-eu and prod-us.

Dead man's switches

06-deadman-platform.yaml / 07-deadman-payments.yaml expect a heartbeat alert on a 15-minute timeout:

./fire-alerts.sh platform heartbeat

Keep sending that (e.g. a watch -n 60) and nothing happens — that's the point. Stop sending it and, 15 minutes after the last one, terdut-server opens a critical incident on its own, with no webhook involved: proof the switch is watching for silence, not for a signal.

Tear down

kubectl delete namespace terdut-operator-demo

The operator's own finalizers clean up everything cross-namespace (credentials Secrets in the operator's namespace, server-side team/rule/ integration rows) before this namespace's objects actually disappear — give it a few seconds past the kubectl delete returning.