50ce5bcec0
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the web UI since: the queue and incident layouts, the rota and escalation pages, the theme toggle, and the cluster chip, filter and page titles (v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the floor, which is what the replicas setting actually depends on. fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus external label plus `cluster` in Alertmanager's group_by (terdut-server's README, "Several clusters, one team"). It goes on the alert's labels and groupLabels, and into the group key and the fingerprint, so the same alert in two clusters is two incidents and not one. Unset, the payload is exactly what it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in both, so the queue has a chip and a filter to show. run-demo.sh also failed on its second run, though it says it is safe to re-run: it expected HTTP 409 when alice already exists, but a spent invite is answered with 403 "invite link is not usable" before the username is ever checked. It now tries to log alice in first and skips the signup if that works. Checked on the kind cluster: the server rolled to v0.43.0, every CR became Ready and Adopted (server, both teams, both escalation rules, both dead man's switches, both alert sources), and /api/incidents/clusters, /api/incidents?cluster=prod-us and the incident titles came back as expected. No operator code changed, so this needs no operator release. Co-authored-by: Claude <noreply@anthropic.com>
174 lines
7.2 KiB
Markdown
174 lines
7.2 KiB
Markdown
# Demo
|
|
|
|
Every CRD this operator reconciles, wired into one working install: one
|
|
`TerdutServer`, two `TerdutTeam`s (Platform and Payments) each with their
|
|
own `TerdutEscalationRule`, `TerdutDeadmanSwitch` and `TerdutAlertSource`,
|
|
plus a script that fires synthetic Alertmanager webhooks at it so you can
|
|
watch real incidents appear, escalate and resolve.
|
|
|
|
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
|
|
Postgres with `emptyDir` storage and a password committed in this
|
|
directory. Throw the whole namespace away when you're done.
|
|
|
|
**Want this fully automated instead of walking through it by hand?**
|
|
`./run-demo.sh` does everything below itself, against a fresh (or
|
|
already-set-up) `kind` cluster — creates the cluster, installs the
|
|
operator, applies every CR here, signs `alice` in for real, and fires a
|
|
few alerts. `./run-demo.sh --help` for the knobs, `./run-demo.sh
|
|
--teardown` to tear it back down. The rest of this file is the manual
|
|
walkthrough it automates.
|
|
|
|
## Prerequisites
|
|
|
|
- The operator and its CRDs installed and running (`make install
|
|
deploy IMG=...`, or `charts/terdut-operator` — see this repo's own
|
|
README.md/DESIGN.md), pointed at a cluster you're fine creating
|
|
throwaway resources in. A `kind` cluster is the easy choice.
|
|
- `kubectl`, `jq`, `curl` on your path.
|
|
|
|
**Apply this into a namespace of its own.** Every object name in this
|
|
directory is prefixed `terdut-operator-demo` specifically so applying it
|
|
by mistake into some other namespace that already has unrelated objects
|
|
doesn't collide with them -- but that only helps if this directory's own
|
|
objects don't collide with *each other* across two applies. Applying it
|
|
twice into two different namespaces is fine; applying it a second time
|
|
into a namespace that already has something else named `terdut-demo` (a
|
|
real install from following `terdut-operator`'s own repo along, say) is
|
|
exactly the mistake this prefix exists to avoid, and it only works if you
|
|
don't override these names yourself.
|
|
|
|
## Apply it
|
|
|
|
```sh
|
|
kubectl create namespace terdut-operator-demo
|
|
kubectl apply -n terdut-operator-demo -k .
|
|
```
|
|
|
|
Listed and applied in dependency order (server → team → everything that
|
|
`teamRef`s it), but you don't have to preserve that order yourself:
|
|
every controller here re-queues and waits rather than failing when a ref
|
|
isn't resolvable yet (`kubectl describe` shows `Reason: TeamRefNotFound` /
|
|
`WaitingForTeam` while that settles).
|
|
|
|
Watch it converge:
|
|
|
|
```sh
|
|
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
|
|
-n terdut-operator-demo
|
|
```
|
|
|
|
The operator's own Deployment template now carries a `wait-for-postgres`
|
|
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
|
|
so `terdut-operator-demo`'s pod should come up clean even against this brand-new
|
|
Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
|
|
|
|
Once `terdut-operator-demo`'s own `Ready` condition is `True`, everything downstream
|
|
of it should settle within a reconcile interval or two.
|
|
|
|
## See the web UI
|
|
|
|
The operator never creates any external exposure for a `TerdutServer` --
|
|
that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this
|
|
demo just reaches it the simplest way there is:
|
|
|
|
```sh
|
|
kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
|
|
```
|
|
|
|
and open http://localhost:8080. See `../networking` for worked examples of
|
|
exposing it for real (Gateway API, Istio, or a plain `Ingress`) instead.
|
|
|
|
### First login
|
|
|
|
The operator's own bootstrap (DESIGN.md §6) creates the first user through
|
|
`/api/bootstrap` and immediately mints itself a service-account token from
|
|
it, then discards the bootstrap user's own key — nobody ever signs in as
|
|
that account, and `signup_mode` stays `invite_only` by default. **Don't try
|
|
to flip it via the operator's own token**: that token is a service account,
|
|
and `/api/admin/settings` is deliberately human-only on terdut-server
|
|
(`niklas/terdut-server#23` has the full reasoning — widening that gate was
|
|
the wrong fix).
|
|
|
|
The real path in: `02-team-platform.yaml` turns on `spec.invite`, so
|
|
Platform's own `TerdutTeam` mints a real invite link with its own
|
|
already-working team-scoped credential (the same reach that lets it manage
|
|
its own escalation policy, dead man's switches and integrations — owner-
|
|
equivalent, confirmed in terdut-server's `SERVICE-ACCOUNTS.md`). Invite
|
|
redemption bypasses `signup_mode` entirely, so this needs no admin
|
|
credential at all:
|
|
|
|
```sh
|
|
secretname=$(kubectl -n terdut-operator-demo get terdutteam terdutteam-platform \
|
|
-o jsonpath='{.status.inviteSecretRef.name}')
|
|
url=$(kubectl -n terdut-operator-demo get secret "$secretname" -o jsonpath='{.data.url}' | base64 -d)
|
|
echo "$url" # open this, or POST /api/signup with {"invite": "<the token after invite=>", ...}
|
|
```
|
|
|
|
`04-escalation-platform.yaml` names a user `alice` at its first escalation
|
|
level — sign up as `alice` if you want that level to mean something rather
|
|
than falling through to on-call after 5 minutes. `run-demo.sh` does exactly
|
|
this automatically (and also joins `alice` to Payments, which deliberately
|
|
has no `spec.invite` of its own — see that file's comment for the second
|
|
onboarding path this demonstrates).
|
|
|
|
## Fire some alerts
|
|
|
|
In another terminal, with the port-forward above still running:
|
|
|
|
```sh
|
|
export NAMESPACE=terdut-operator-demo
|
|
|
|
./fire-alerts.sh platform high-cpu
|
|
./fire-alerts.sh platform disk-full
|
|
./fire-alerts.sh payments pod-crash
|
|
|
|
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
|
|
# levels, and show up on the Platform/Payments team's own incident list.
|
|
|
|
./fire-alerts.sh platform high-cpu resolve
|
|
```
|
|
|
|
`fire-alerts.sh -h` (or any bad argument) prints the full scenario list.
|
|
Each `(team, scenario)` pair is one stable fingerprint, so firing the same
|
|
one twice updates the same alert (a real re-fire) and `resolve` closes
|
|
exactly that one.
|
|
|
|
Set `CLUSTER` to send the alert as if it came from one of several clusters:
|
|
|
|
```sh
|
|
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
|
|
CLUSTER=prod-us ./fire-alerts.sh platform high-cpu # a second incident, not a join
|
|
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
|
|
```
|
|
|
|
It stands in for a Prometheus external label plus `cluster` in Alertmanager's
|
|
`group_by` (terdut-server's README, "Several clusters, one team"): the web UI
|
|
then shows the cluster chip on each incident and a cluster filter in the
|
|
queue. `CLUSTER` is part of the fingerprint, so resolve with the same value you
|
|
fired with. `./run-demo.sh` fires its alerts across `prod-eu` and `prod-us`.
|
|
|
|
### Dead man's switches
|
|
|
|
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
|
|
alert on a 15-minute timeout:
|
|
|
|
```sh
|
|
./fire-alerts.sh platform heartbeat
|
|
```
|
|
|
|
Keep sending that (e.g. a `watch -n 60`) and nothing happens — that's the
|
|
point. Stop sending it and, 15 minutes after the last one, terdut-server
|
|
opens a `critical` incident on its own, with no webhook involved: proof
|
|
the switch is watching for silence, not for a signal.
|
|
|
|
## Tear down
|
|
|
|
```sh
|
|
kubectl delete namespace terdut-operator-demo
|
|
```
|
|
|
|
The operator's own finalizers clean up everything cross-namespace
|
|
(credentials Secrets in the operator's namespace, server-side team/rule/
|
|
integration rows) before this namespace's objects actually disappear —
|
|
give it a few seconds past the `kubectl delete` returning.
|