wait_for_ready waited for every demo object at once, including terdutescalationrule-platform, which names alice as a level-1 target -- but alice does not exist yet at that point in main(): she is created by redeem_platform_invite, which ran after wait_for_ready. terdut-server resolves every named username at reconcile time, not just when an escalation actually fires, so that CR could never reach Ready before alice did, and main() had no step in between to create her. Split into wait_for_objects (the shared loop, now taking its object list as arguments) plus two callers: wait_for_teams_ready, covering just the server and the two teams redeem_platform_invite/join_payments_team need, run before alice exists; wait_for_remaining_ready, covering the escalation rules, dead man's switches and alert sources, run after. Co-authored-by: Claude <noreply@anthropic.com>
Demo
Every CRD this operator reconciles, wired into one working install: one
TerdutServer, two TerdutTeams (Platform and Payments) each with their
own TerdutEscalationRule, TerdutDeadmanSwitch and TerdutAlertSource,
plus a script that fires synthetic Alertmanager webhooks at it so you can
watch real incidents appear, escalate and resolve.
This is a demo kit, not a reference deployment: 00-postgres.yaml runs
Postgres with emptyDir storage and a password committed in this
directory. Throw the whole namespace away when you're done.
Want this fully automated instead of walking through it by hand?
./run-demo.sh does everything below itself, against a fresh (or
already-set-up) kind cluster — creates the cluster, installs the
operator, applies every CR here, signs alice in for real, and fires a
few alerts. ./run-demo.sh --help for the knobs, ./run-demo.sh --teardown to tear it back down. The rest of this file is the manual
walkthrough it automates.
Prerequisites
- The operator and its CRDs installed and running (
make install deploy IMG=..., orcharts/terdut-operator— see this repo's own README.md/DESIGN.md), pointed at a cluster you're fine creating throwaway resources in. Akindcluster is the easy choice. kubectl,jq,curlon your path.
Apply this into a namespace of its own. Every object name in this
directory is prefixed terdut-operator-demo specifically so applying it
by mistake into some other namespace that already has unrelated objects
doesn't collide with them -- but that only helps if this directory's own
objects don't collide with each other across two applies. Applying it
twice into two different namespaces is fine; applying it a second time
into a namespace that already has something else named terdut-demo (a
real install from following terdut-operator's own repo along, say) is
exactly the mistake this prefix exists to avoid, and it only works if you
don't override these names yourself.
Apply it
kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .
Listed and applied in dependency order (server → team → everything that
teamRefs it), but you don't have to preserve that order yourself:
every controller here re-queues and waits rather than failing when a ref
isn't resolvable yet (kubectl describe shows Reason: TeamRefNotFound /
WaitingForTeam while that settles).
Watch it converge:
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
-n terdut-operator-demo
The operator's own Deployment template now carries a wait-for-postgres
init container (same fix as charts/terdut-server's chart as of v0.33.2),
so terdut-operator-demo's pod should come up clean even against this brand-new
Postgres doing its very first boot — no CrashLoopBackOff expected here.
Once terdut-operator-demo's own Ready condition is True, everything downstream
of it should settle within a reconcile interval or two.
See the web UI
The operator never creates any external exposure for a TerdutServer --
that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this
demo just reaches it the simplest way there is:
kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
and open http://localhost:8080. See ../networking for worked examples of
exposing it for real (Gateway API, Istio, or a plain Ingress) instead.
First login
The operator's own bootstrap (DESIGN.md §6) creates the first user through
/api/bootstrap and immediately mints itself a service-account token from
it, then discards the bootstrap user's own key — nobody ever signs in as
that account, and signup_mode stays invite_only by default. Don't try
to flip it via the operator's own token: that token is a service account,
and /api/admin/settings is deliberately human-only on terdut-server
(niklas/terdut-server#23 has the full reasoning — widening that gate was
the wrong fix).
The real path in: 02-team-platform.yaml turns on spec.invite, so
Platform's own TerdutTeam mints a real invite link with its own
already-working team-scoped credential (the same reach that lets it manage
its own escalation policy, dead man's switches and integrations — owner-
equivalent, confirmed in terdut-server's SERVICE-ACCOUNTS.md). Invite
redemption bypasses signup_mode entirely, so this needs no admin
credential at all:
secretname=$(kubectl -n terdut-operator-demo get terdutteam terdutteam-platform \
-o jsonpath='{.status.inviteSecretRef.name}')
url=$(kubectl -n terdut-operator-demo get secret "$secretname" -o jsonpath='{.data.url}' | base64 -d)
echo "$url" # open this, or POST /api/signup with {"invite": "<the token after invite=>", ...}
04-escalation-platform.yaml names a user alice at its first escalation
level — sign up as alice if you want that level to mean something rather
than falling through to on-call after 5 minutes. run-demo.sh does exactly
this automatically (and also joins alice to Payments, which deliberately
has no spec.invite of its own — see that file's comment for the second
onboarding path this demonstrates).
Fire some alerts
In another terminal, with the port-forward above still running:
export NAMESPACE=terdut-operator-demo
./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
# levels, and show up on the Platform/Payments team's own incident list.
./fire-alerts.sh platform high-cpu resolve
fire-alerts.sh -h (or any bad argument) prints the full scenario list.
Each (team, scenario) pair is one stable fingerprint, so firing the same
one twice updates the same alert (a real re-fire) and resolve closes
exactly that one.
Dead man's switches
06-deadman-platform.yaml / 07-deadman-payments.yaml expect a heartbeat
alert on a 15-minute timeout:
./fire-alerts.sh platform heartbeat
Keep sending that (e.g. a watch -n 60) and nothing happens — that's the
point. Stop sending it and, 15 minutes after the last one, terdut-server
opens a critical incident on its own, with no webhook involved: proof
the switch is watching for silence, not for a signal.
Tear down
kubectl delete namespace terdut-operator-demo
The operator's own finalizers clean up everything cross-namespace
(credentials Secrets in the operator's namespace, server-side team/rule/
integration rows) before this namespace's objects actually disappear —
give it a few seconds past the kubectl delete returning.