Files
terdut-operator/examples/demo
Niklas Ye 2a08a8cd8e
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 1m5s
CI / test (pull_request) Successful in 2m49s
examples/demo: add run-demo.sh, an automated kind-cluster demo
One script, two modes (run-demo.sh / run-demo.sh --teardown), that takes a
fresh empty kind cluster all the way to a working demo: creates the
cluster if needed, helm-installs this chart, applies every CRD kind in
this directory, waits for all nine objects to go Ready, then does what
the README's own first-login section cannot (see niklas/terdut-server#23
and niklas/terdut-operator#3 -- no service-account credential this
operator holds can ever call /api/admin/settings or POST /api/users) by
reaching into the demo's own throwaway Postgres directly: flips
signup_mode to open, signs alice up for real over the ordinary signup
endpoint, and joins her to both Platform and Payments (open signup always
creates its own new team, never joins an existing one by name, so
without this she'd have a working login that can't see a single incident
this demo fires -- /api/incidents and /api/alerts are both scoped to the
caller's own team memberships). Finishes by port-forwarding the service
and firing fire-alerts.sh at both teams, so a fresh run already has
visible incidents waiting in the web UI.

Verified end to end against a real kind cluster, including a second,
genuinely-fresh run that hit niklas/terdut-operator#3 live (terdutteam-
platform wedged in the 403 retry loop that issue describes) -- confirmed
the script itself fails cleanly on that (clear FAILED message, correct
exit code, no orphaned port-forward) rather than hanging or leaving a
mess, which is the most this script can do about a bug in the operator
it's driving.
2026-10-02 21:23:44 +02:00
..

Demo

Every CRD this operator reconciles, wired into one working install: one TerdutServer, two TerdutTeams (Platform and Payments) each with their own TerdutEscalationRule, TerdutDeadmanSwitch and TerdutAlertSource, plus a script that fires synthetic Alertmanager webhooks at it so you can watch real incidents appear, escalate and resolve.

This is a demo kit, not a reference deployment: 00-postgres.yaml runs Postgres with emptyDir storage and a password committed in this directory. Throw the whole namespace away when you're done.

Prerequisites

  • The operator and its CRDs installed and running (make install deploy IMG=..., or charts/terdut-operator — see this repo's own README.md/DESIGN.md), pointed at a cluster you're fine creating throwaway resources in. A kind cluster is the easy choice.
  • kubectl, jq, curl on your path.

Apply this into a namespace of its own. Every object name in this directory is prefixed terdut-operator-demo specifically so applying it by mistake into some other namespace that already has unrelated objects doesn't collide with them -- but that only helps if this directory's own objects don't collide with each other across two applies. Applying it twice into two different namespaces is fine; applying it a second time into a namespace that already has something else named terdut-demo (a real install from following terdut-operator's own repo along, say) is exactly the mistake this prefix exists to avoid, and it only works if you don't override these names yourself.

Apply it

kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .

Listed and applied in dependency order (server → team → everything that teamRefs it), but you don't have to preserve that order yourself: every controller here re-queues and waits rather than failing when a ref isn't resolvable yet (kubectl describe shows Reason: TeamRefNotFound / WaitingForTeam while that settles).

Watch it converge:

kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
  -n terdut-operator-demo

The operator's own Deployment template now carries a wait-for-postgres init container (same fix as charts/terdut-server's chart as of v0.33.2), so terdut-operator-demo's pod should come up clean even against this brand-new Postgres doing its very first boot — no CrashLoopBackOff expected here.

Once terdut-operator-demo's own Ready condition is True, everything downstream of it should settle within a reconcile interval or two.

See the web UI

The operator never creates any external exposure for a TerdutServer -- that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this demo just reaches it the simplest way there is:

kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080

and open http://localhost:8080. See ../networking for worked examples of exposing it for real (Gateway API, Istio, or a plain Ingress) instead.

First login

The operator's own bootstrap (DESIGN.md §6) creates the first user through /api/bootstrap and immediately mints itself a service-account token from it — that account has no password, so there's nothing to sign in with yet. signup_mode also defaults to invite_only, so open signup needs turning on first, using the admin token the operator generated for itself:

# Which namespace the operator itself runs in:
kubectl get deploy -A -l control-plane=controller-manager

# The Secret holding the operator's own admin token for this TerdutServer
# (cross-namespace from terdut-operator-demo, per DESIGN.md §7):
secretname=$(kubectl -n terdut-operator-demo get terdutserver terdut-operator-demo \
  -o jsonpath='{.status.credentialsSecretRef.name}')
token=$(kubectl -n <operator-namespace-from-above> get secret "$secretname" \
  -o jsonpath='{.data.token}' | base64 -d)

curl -X PUT http://localhost:8080/api/admin/settings \
  -H "Authorization: Bearer $token" -H 'Content-Type: application/json' \
  -d '{"signup_mode":"open"}'

Then sign up through the UI as a normal human account. 04-escalation-platform.yaml names a user alice at its first escalation level — sign up as alice if you want that level to mean something rather than falling through to on-call after 5 minutes.

Fire some alerts

In another terminal, with the port-forward above still running:

export NAMESPACE=terdut-operator-demo

./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash

# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
# levels, and show up on the Platform/Payments team's own incident list.

./fire-alerts.sh platform high-cpu resolve

fire-alerts.sh -h (or any bad argument) prints the full scenario list. Each (team, scenario) pair is one stable fingerprint, so firing the same one twice updates the same alert (a real re-fire) and resolve closes exactly that one.

Dead man's switches

06-deadman-platform.yaml / 07-deadman-payments.yaml expect a heartbeat alert on a 15-minute timeout:

./fire-alerts.sh platform heartbeat

Keep sending that (e.g. a watch -n 60) and nothing happens — that's the point. Stop sending it and, 15 minutes after the last one, terdut-server opens a critical incident on its own, with no webhook involved: proof the switch is watching for silence, not for a signal.

Tear down

kubectl delete namespace terdut-operator-demo

The operator's own finalizers clean up everything cross-namespace (credentials Secrets in the operator's namespace, server-side team/rule/ integration rows) before this namespace's objects actually disappear — give it a few seconds past the kubectl delete returning.