One script, two modes (run-demo.sh / run-demo.sh --teardown), that takes a fresh empty kind cluster all the way to a working demo: creates the cluster if needed, helm-installs this chart, applies every CRD kind in this directory, waits for all nine objects to go Ready, then does what the README's own first-login section cannot (see niklas/terdut-server#23 and niklas/terdut-operator#3 -- no service-account credential this operator holds can ever call /api/admin/settings or POST /api/users) by reaching into the demo's own throwaway Postgres directly: flips signup_mode to open, signs alice up for real over the ordinary signup endpoint, and joins her to both Platform and Payments (open signup always creates its own new team, never joins an existing one by name, so without this she'd have a working login that can't see a single incident this demo fires -- /api/incidents and /api/alerts are both scoped to the caller's own team memberships). Finishes by port-forwarding the service and firing fire-alerts.sh at both teams, so a fresh run already has visible incidents waiting in the web UI. Verified end to end against a real kind cluster, including a second, genuinely-fresh run that hit niklas/terdut-operator#3 live (terdutteam- platform wedged in the 403 retry loop that issue describes) -- confirmed the script itself fails cleanly on that (clear FAILED message, correct exit code, no orphaned port-forward) rather than hanging or leaving a mess, which is the most this script can do about a bug in the operator it's driving.
Demo
Every CRD this operator reconciles, wired into one working install: one
TerdutServer, two TerdutTeams (Platform and Payments) each with their
own TerdutEscalationRule, TerdutDeadmanSwitch and TerdutAlertSource,
plus a script that fires synthetic Alertmanager webhooks at it so you can
watch real incidents appear, escalate and resolve.
This is a demo kit, not a reference deployment: 00-postgres.yaml runs
Postgres with emptyDir storage and a password committed in this
directory. Throw the whole namespace away when you're done.
Prerequisites
- The operator and its CRDs installed and running (
make install deploy IMG=..., orcharts/terdut-operator— see this repo's own README.md/DESIGN.md), pointed at a cluster you're fine creating throwaway resources in. Akindcluster is the easy choice. kubectl,jq,curlon your path.
Apply this into a namespace of its own. Every object name in this
directory is prefixed terdut-operator-demo specifically so applying it
by mistake into some other namespace that already has unrelated objects
doesn't collide with them -- but that only helps if this directory's own
objects don't collide with each other across two applies. Applying it
twice into two different namespaces is fine; applying it a second time
into a namespace that already has something else named terdut-demo (a
real install from following terdut-operator's own repo along, say) is
exactly the mistake this prefix exists to avoid, and it only works if you
don't override these names yourself.
Apply it
kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .
Listed and applied in dependency order (server → team → everything that
teamRefs it), but you don't have to preserve that order yourself:
every controller here re-queues and waits rather than failing when a ref
isn't resolvable yet (kubectl describe shows Reason: TeamRefNotFound /
WaitingForTeam while that settles).
Watch it converge:
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
-n terdut-operator-demo
The operator's own Deployment template now carries a wait-for-postgres
init container (same fix as charts/terdut-server's chart as of v0.33.2),
so terdut-operator-demo's pod should come up clean even against this brand-new
Postgres doing its very first boot — no CrashLoopBackOff expected here.
Once terdut-operator-demo's own Ready condition is True, everything downstream
of it should settle within a reconcile interval or two.
See the web UI
The operator never creates any external exposure for a TerdutServer --
that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this
demo just reaches it the simplest way there is:
kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
and open http://localhost:8080. See ../networking for worked examples of
exposing it for real (Gateway API, Istio, or a plain Ingress) instead.
First login
The operator's own bootstrap (DESIGN.md §6) creates the first user through
/api/bootstrap and immediately mints itself a service-account token from
it — that account has no password, so there's nothing to sign in with yet.
signup_mode also defaults to invite_only, so open signup needs turning
on first, using the admin token the operator generated for itself:
# Which namespace the operator itself runs in:
kubectl get deploy -A -l control-plane=controller-manager
# The Secret holding the operator's own admin token for this TerdutServer
# (cross-namespace from terdut-operator-demo, per DESIGN.md §7):
secretname=$(kubectl -n terdut-operator-demo get terdutserver terdut-operator-demo \
-o jsonpath='{.status.credentialsSecretRef.name}')
token=$(kubectl -n <operator-namespace-from-above> get secret "$secretname" \
-o jsonpath='{.data.token}' | base64 -d)
curl -X PUT http://localhost:8080/api/admin/settings \
-H "Authorization: Bearer $token" -H 'Content-Type: application/json' \
-d '{"signup_mode":"open"}'
Then sign up through the UI as a normal human account. 04-escalation-platform.yaml
names a user alice at its first escalation level — sign up as alice if
you want that level to mean something rather than falling through to
on-call after 5 minutes.
Fire some alerts
In another terminal, with the port-forward above still running:
export NAMESPACE=terdut-operator-demo
./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
# levels, and show up on the Platform/Payments team's own incident list.
./fire-alerts.sh platform high-cpu resolve
fire-alerts.sh -h (or any bad argument) prints the full scenario list.
Each (team, scenario) pair is one stable fingerprint, so firing the same
one twice updates the same alert (a real re-fire) and resolve closes
exactly that one.
Dead man's switches
06-deadman-platform.yaml / 07-deadman-payments.yaml expect a heartbeat
alert on a 15-minute timeout:
./fire-alerts.sh platform heartbeat
Keep sending that (e.g. a watch -n 60) and nothing happens — that's the
point. Stop sending it and, 15 minutes after the last one, terdut-server
opens a critical incident on its own, with no webhook involved: proof
the switch is watching for silence, not for a signal.
Tear down
kubectl delete namespace terdut-operator-demo
The operator's own finalizers clean up everything cross-namespace
(credentials Secrets in the operator's namespace, server-side team/rule/
integration rows) before this namespace's objects actually disappear —
give it a few seconds past the kubectl delete returning.