Files
terdut-operator/examples/demo/README.md
T
Niklas Ye d9315322fc
CI / chart (push) Successful in 1s
CI / security (push) Successful in 3m24s
CI / test (push) Successful in 10m45s
examples/demo: fix two real bugs this exact demo just hit live
1. Renamed every object this demo creates (TerdutServer, Postgres
   Secret/Deployment/Service) from terdut-demo[-postgres] to
   terdut-operator-demo[-postgres]. The user applied this kit into the
   already-live "terdut-demo" namespace -- the real operator exercise
   from earlier in this repo's own history -- and this demo's own
   TerdutServer/Postgres objects shared that exact name. The TerdutServer
   apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
   while the live object already had postgresClusterRef violates "exactly
   one of" and the API server refused it), and the real Postgres Service
   was never touched (confirmed live: still Zalando's own spilo selector,
   endpoint still the real StatefulSet pod) -- but the Postgres Secret and
   Deployment, having no such protection, were created as brand new,
   extra, crash-looping objects sitting right next to the real ones.
   Prefixing every name this demo creates means a repeat of this exact
   mistake no longer collides with anything, documented directly in
   README.md now.

2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
   (added responding to a PodSecurity "restricted" warning) took
   CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
   entrypoint needs to chown/chmod the data directory before it drops
   privileges itself -- confirmed in a real crashed pod's logs: `chmod:
   /var/run/postgresql: Operation not permitted`. kubectl apply
   --dry-run=server, which is as far as this got verified before, only
   checks admission policy; it was never actually booted. Removed the
   capability drop and verified for real this time: applied just
   00-postgres.yaml alone into a disposable namespace, waited for the pod
   to go Ready, read its logs ("database system is ready to accept
   connections"), then deleted that namespace.
2026-10-02 13:25:47 +02:00

5.6 KiB

Demo

Every CRD this operator reconciles, wired into one working install: one TerdutServer, two TerdutTeams (Platform and Payments) each with their own TerdutEscalationRule, TerdutDeadmanSwitch and TerdutAlertSource, plus a script that fires synthetic Alertmanager webhooks at it so you can watch real incidents appear, escalate and resolve.

This is a demo kit, not a reference deployment: 00-postgres.yaml runs Postgres with emptyDir storage and a password committed in this directory. Throw the whole namespace away when you're done.

Prerequisites

  • The operator and its CRDs installed and running (make install deploy IMG=..., or charts/terdut-operator — see this repo's own README.md/DESIGN.md), pointed at a cluster you're fine creating throwaway resources in. A kind cluster is the easy choice.
  • kubectl, jq, curl on your path.

Apply this into a namespace of its own. Every object name in this directory is prefixed terdut-operator-demo specifically so applying it by mistake into some other namespace that already has unrelated objects doesn't collide with them -- but that only helps if this directory's own objects don't collide with each other across two applies. Applying it twice into two different namespaces is fine; applying it a second time into a namespace that already has something else named terdut-demo (a real install from following terdut-operator's own repo along, say) is exactly the mistake this prefix exists to avoid, and it only works if you don't override these names yourself.

Apply it

kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .

Listed and applied in dependency order (server → team → everything that teamRefs it), but you don't have to preserve that order yourself: every controller here re-queues and waits rather than failing when a ref isn't resolvable yet (kubectl describe shows Reason: TeamRefNotFound / WaitingForTeam while that settles).

Watch it converge:

kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
  -n terdut-operator-demo

The operator's own Deployment template now carries a wait-for-postgres init container (same fix as charts/terdut-server's chart as of v0.33.2), so terdut-operator-demo's pod should come up clean even against this brand-new Postgres doing its very first boot — no CrashLoopBackOff expected here.

Once terdut-operator-demo's own Ready condition is True, everything downstream of it should settle within a reconcile interval or two.

See the web UI

The operator doesn't create any external exposure yet (NetworkingSpec's own doc comment in api/v1alpha1/terdutserver_types.go — spec.networking.hostname is accepted but nothing acts on it), so:

kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080

and open http://localhost:8080.

First login

The operator's own bootstrap (DESIGN.md §6) creates the first user through /api/bootstrap and immediately mints itself a service-account token from it — that account has no password, so there's nothing to sign in with yet. signup_mode also defaults to invite_only, so open signup needs turning on first, using the admin token the operator generated for itself:

# Which namespace the operator itself runs in:
kubectl get deploy -A -l control-plane=controller-manager

# The Secret holding the operator's own admin token for this TerdutServer
# (cross-namespace from terdut-operator-demo, per DESIGN.md §7):
secretname=$(kubectl -n terdut-operator-demo get terdutserver terdut-operator-demo \
  -o jsonpath='{.status.credentialsSecretRef.name}')
token=$(kubectl -n <operator-namespace-from-above> get secret "$secretname" \
  -o jsonpath='{.data.token}' | base64 -d)

curl -X PUT http://localhost:8080/api/admin/settings \
  -H "Authorization: Bearer $token" -H 'Content-Type: application/json' \
  -d '{"signup_mode":"open"}'

Then sign up through the UI as a normal human account. 04-escalation-platform.yaml names a user alice at its first escalation level — sign up as alice if you want that level to mean something rather than falling through to on-call after 5 minutes.

Fire some alerts

In another terminal, with the port-forward above still running:

export NAMESPACE=terdut-operator-demo

./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash

# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
# levels, and show up on the Platform/Payments team's own incident list.

./fire-alerts.sh platform high-cpu resolve

fire-alerts.sh -h (or any bad argument) prints the full scenario list. Each (team, scenario) pair is one stable fingerprint, so firing the same one twice updates the same alert (a real re-fire) and resolve closes exactly that one.

Dead man's switches

06-deadman-platform.yaml / 07-deadman-payments.yaml expect a heartbeat alert on a 15-minute timeout:

./fire-alerts.sh platform heartbeat

Keep sending that (e.g. a watch -n 60) and nothing happens — that's the point. Stop sending it and, 15 minutes after the last one, terdut-server opens a critical incident on its own, with no webhook involved: proof the switch is watching for silence, not for a signal.

Tear down

kubectl delete namespace terdut-operator-demo

The operator's own finalizers clean up everything cross-namespace (credentials Secrets in the operator's namespace, server-side team/rule/ integration rows) before this namespace's objects actually disappear — give it a few seconds past the kubectl delete returning.