Files
terdut-operator/examples/demo/README.md
T
Niklas Ye e1103f2b7d Authenticate with a seeded operator key; fold escalation and switches into TerdutTeam
Credentials: the TerdutServer controller generates <name>-operator-key in
the server's own namespace (owned by it) and hands it to the pods as
TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it
at every start. A replaced Secret rolls the pods. The bootstrap handshake,
the checkpoint Secret, per-team service accounts and credentials Secrets,
BootstrapStateLost and credentials.deletionPolicy are gone.

CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule
and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[]
on the team (matched by name, extras removed); team invites are removed.
A team is created under the identity <namespace>/<name> (external_id), so a
retry, a lost status or a deleted team heal by repeating the same call, and
a display name owned by another team is TeamNameTaken instead of an
adoption. The server resolves escalation usernames (UnknownUser condition).
OIDC claim names and trustEmail are spec fields.

Fixes: query values are URL-escaped; every delete treats 404 as success;
deleting a team no longer depends on allowedTeams consent; a switch or
integration deleted on the server is recreated; unnamed switches take the
CR's name.

Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces
and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo
(run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays
cluster-wide, now stated in DESIGN.md section 9.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:22 +02:00

6.4 KiB

Demo

Every CRD this operator reconciles, wired into one working install: one TerdutServer, two TerdutTeams (Platform and Payments), each carrying its own escalation ladder and dead man's switches, and a TerdutAlertSource per team, plus a script that fires synthetic Alertmanager webhooks at it so you can watch real incidents appear, escalate and resolve.

This is a demo kit, not a reference deployment: 00-postgres.yaml runs Postgres with emptyDir storage and a password committed in this directory. Throw the whole namespace away when you're done.

Want this fully automated instead of walking through it by hand? ./run-demo.sh does everything below itself, against a fresh (or already-set-up) kind cluster — creates the cluster, installs the operator, applies every CR here, creates alice as the first user, and fires a few alerts. ./run-demo.sh --help for the knobs, ./run-demo.sh --teardown to tear it back down. The rest of this file is the manual walkthrough it automates.

Prerequisites

  • The operator and its CRDs installed and running (make install deploy IMG=..., or charts/terdut-operator — see this repo's own README.md/DESIGN.md), pointed at a cluster you're fine creating throwaway resources in. A kind cluster is the easy choice.
  • kubectl, jq, curl on your path.

Apply this into a namespace of its own. Every object name in this directory is prefixed terdut-operator-demo specifically so applying it by mistake into some other namespace that already has unrelated objects doesn't collide with them -- but that only helps if this directory's own objects don't collide with each other across two applies. Applying it twice into two different namespaces is fine; applying it a second time into a namespace that already has something else named terdut-demo (a real install from following terdut-operator's own repo along, say) is exactly the mistake this prefix exists to avoid, and it only works if you don't override these names yourself.

Apply it

kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .

Listed and applied in dependency order (server → team → alert source), but you don't have to preserve that order yourself: every controller here re-queues and waits rather than failing when a ref isn't resolvable yet (kubectl describe shows Reason: WaitingForServer / WaitingForTeam while that settles). Platform stays at Reason: UnknownUser until alice exists: its escalation ladder names her.

Watch it converge:

kubectl get terdutservers,terdutteams,terdutalertsources \
  -n terdut-operator-demo

The operator's own Deployment template now carries a wait-for-postgres init container (same fix as charts/terdut-server's chart as of v0.33.2), so terdut-operator-demo's pod should come up clean even against this brand-new Postgres doing its very first boot — no CrashLoopBackOff expected here.

Once terdut-operator-demo's own Ready condition is True, everything downstream of it should settle within a reconcile interval or two.

See the web UI

The operator never creates any external exposure for a TerdutServer -- that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this demo just reaches it the simplest way there is:

kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080

and open http://localhost:8080. See ../networking for worked examples of exposing it for real (Gateway API, Istio, or a plain Ingress) instead.

First login

The operator authenticates with a key of its own (the TerdutServer's <name>-operator-key Secret, handed to the server as TERDUT_OPERATOR_KEY) and never creates a user. A person gets in the way anyone does on a fresh terdut-server: /api/bootstrap creates the first user, an administrator, while no user exists yet:

curl -sS -X POST http://localhost:8080/api/bootstrap -H 'Content-Type: application/json' \
  -d '{"username":"alice","email":"alice@example.com","password":"a-long-demo-password"}'

The response carries an API key, shown once. An administrator can manage any team, so add alice to both teams with it (POST /api/teams/{teamID}/members; kubectl get terdutteam -o jsonpath='{.status.teamID}' gives the ids), or just use the UI's team pages. run-demo.sh does exactly this for you.

Fire some alerts

In another terminal, with the port-forward above still running:

export NAMESPACE=terdut-operator-demo

./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash

# Watch it open an incident, escalate per the escalation ladders in
# 02/03-team-*.yaml, and show up on the Platform/Payments team's own incident list.

./fire-alerts.sh platform high-cpu resolve

fire-alerts.sh -h (or any bad argument) prints the full scenario list. Each (team, scenario) pair is one stable fingerprint, so firing the same one twice updates the same alert (a real re-fire) and resolve closes exactly that one.

Set CLUSTER to send the alert as if it came from one of several clusters:

CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-us ./fire-alerts.sh platform high-cpu   # a second incident, not a join
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve

It stands in for a Prometheus external label plus cluster in Alertmanager's group_by (terdut-server's README, "Several clusters, one team"): the web UI then shows the cluster chip on each incident and a cluster filter in the queue. CLUSTER is part of the fingerprint, so resolve with the same value you fired with. ./run-demo.sh fires its alerts across prod-eu and prod-us.

Dead man's switches

The deadmanSwitches in 02-team-platform.yaml / 03-team-payments.yaml expect a heartbeat alert on a 15-minute timeout:

./fire-alerts.sh platform heartbeat

Keep sending that (e.g. a watch -n 60) and nothing happens — that's the point. Stop sending it and, 15 minutes after the last one, terdut-server opens a critical incident on its own, with no webhook involved: proof the switch is watching for silence, not for a signal.

Tear down

kubectl delete namespace terdut-operator-demo

The operator's finalizers delete each team on the server (with its escalation, switches and integrations) before this namespace's objects actually disappear — give it a few seconds past the kubectl delete returning. A team with open incidents is not deleted until they are resolved.