e1103f2b7d
Credentials: the TerdutServer controller generates <name>-operator-key in the server's own namespace (owned by it) and hands it to the pods as TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it at every start. A replaced Secret rolls the pods. The bootstrap handshake, the checkpoint Secret, per-team service accounts and credentials Secrets, BootstrapStateLost and credentials.deletionPolicy are gone. CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[] on the team (matched by name, extras removed); team invites are removed. A team is created under the identity <namespace>/<name> (external_id), so a retry, a lost status or a deleted team heal by repeating the same call, and a display name owned by another team is TeamNameTaken instead of an adoption. The server resolves escalation usernames (UnknownUser condition). OIDC claim names and trustEmail are spec fields. Fixes: query values are URL-escaped; every delete treats 404 as success; deleting a team no longer depends on allowedTeams consent; a switch or integration deleted on the server is recreated; unnamed switches take the CR's name. Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo (run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays cluster-wide, now stated in DESIGN.md section 9. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
160 lines
6.4 KiB
Markdown
160 lines
6.4 KiB
Markdown
# Demo
|
|
|
|
Every CRD this operator reconciles, wired into one working install: one
|
|
`TerdutServer`, two `TerdutTeam`s (Platform and Payments), each carrying its
|
|
own escalation ladder and dead man's switches, and a `TerdutAlertSource` per
|
|
team, plus a script that fires synthetic Alertmanager webhooks at it so you can
|
|
watch real incidents appear, escalate and resolve.
|
|
|
|
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
|
|
Postgres with `emptyDir` storage and a password committed in this
|
|
directory. Throw the whole namespace away when you're done.
|
|
|
|
**Want this fully automated instead of walking through it by hand?**
|
|
`./run-demo.sh` does everything below itself, against a fresh (or
|
|
already-set-up) `kind` cluster — creates the cluster, installs the
|
|
operator, applies every CR here, creates `alice` as the first user, and fires a
|
|
few alerts. `./run-demo.sh --help` for the knobs, `./run-demo.sh
|
|
--teardown` to tear it back down. The rest of this file is the manual
|
|
walkthrough it automates.
|
|
|
|
## Prerequisites
|
|
|
|
- The operator and its CRDs installed and running (`make install
|
|
deploy IMG=...`, or `charts/terdut-operator` — see this repo's own
|
|
README.md/DESIGN.md), pointed at a cluster you're fine creating
|
|
throwaway resources in. A `kind` cluster is the easy choice.
|
|
- `kubectl`, `jq`, `curl` on your path.
|
|
|
|
**Apply this into a namespace of its own.** Every object name in this
|
|
directory is prefixed `terdut-operator-demo` specifically so applying it
|
|
by mistake into some other namespace that already has unrelated objects
|
|
doesn't collide with them -- but that only helps if this directory's own
|
|
objects don't collide with *each other* across two applies. Applying it
|
|
twice into two different namespaces is fine; applying it a second time
|
|
into a namespace that already has something else named `terdut-demo` (a
|
|
real install from following `terdut-operator`'s own repo along, say) is
|
|
exactly the mistake this prefix exists to avoid, and it only works if you
|
|
don't override these names yourself.
|
|
|
|
## Apply it
|
|
|
|
```sh
|
|
kubectl create namespace terdut-operator-demo
|
|
kubectl apply -n terdut-operator-demo -k .
|
|
```
|
|
|
|
Listed and applied in dependency order (server → team → alert source), but
|
|
you don't have to preserve that order yourself: every controller here
|
|
re-queues and waits rather than failing when a ref isn't resolvable yet
|
|
(`kubectl describe` shows `Reason: WaitingForServer` / `WaitingForTeam` while
|
|
that settles). Platform stays at `Reason: UnknownUser` until `alice` exists:
|
|
its escalation ladder names her.
|
|
|
|
Watch it converge:
|
|
|
|
```sh
|
|
kubectl get terdutservers,terdutteams,terdutalertsources \
|
|
-n terdut-operator-demo
|
|
```
|
|
|
|
The operator's own Deployment template now carries a `wait-for-postgres`
|
|
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
|
|
so `terdut-operator-demo`'s pod should come up clean even against this brand-new
|
|
Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
|
|
|
|
Once `terdut-operator-demo`'s own `Ready` condition is `True`, everything downstream
|
|
of it should settle within a reconcile interval or two.
|
|
|
|
## See the web UI
|
|
|
|
The operator never creates any external exposure for a `TerdutServer` --
|
|
that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this
|
|
demo just reaches it the simplest way there is:
|
|
|
|
```sh
|
|
kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
|
|
```
|
|
|
|
and open http://localhost:8080. See `../networking` for worked examples of
|
|
exposing it for real (Gateway API, Istio, or a plain `Ingress`) instead.
|
|
|
|
### First login
|
|
|
|
The operator authenticates with a key of its own (the `TerdutServer`'s
|
|
`<name>-operator-key` Secret, handed to the server as `TERDUT_OPERATOR_KEY`) and
|
|
never creates a user. A person gets in the way anyone does on a fresh
|
|
terdut-server: `/api/bootstrap` creates the first user, an administrator,
|
|
while no user exists yet:
|
|
|
|
```sh
|
|
curl -sS -X POST http://localhost:8080/api/bootstrap -H 'Content-Type: application/json' \
|
|
-d '{"username":"alice","email":"alice@example.com","password":"a-long-demo-password"}'
|
|
```
|
|
|
|
The response carries an API key, shown once. An administrator can manage any
|
|
team, so add `alice` to both teams with it (`POST /api/teams/{teamID}/members`;
|
|
`kubectl get terdutteam -o jsonpath='{.status.teamID}'` gives the ids), or just
|
|
use the UI's team pages. `run-demo.sh` does exactly this for you.
|
|
|
|
## Fire some alerts
|
|
|
|
In another terminal, with the port-forward above still running:
|
|
|
|
```sh
|
|
export NAMESPACE=terdut-operator-demo
|
|
|
|
./fire-alerts.sh platform high-cpu
|
|
./fire-alerts.sh platform disk-full
|
|
./fire-alerts.sh payments pod-crash
|
|
|
|
# Watch it open an incident, escalate per the escalation ladders in
|
|
# 02/03-team-*.yaml, and show up on the Platform/Payments team's own incident list.
|
|
|
|
./fire-alerts.sh platform high-cpu resolve
|
|
```
|
|
|
|
`fire-alerts.sh -h` (or any bad argument) prints the full scenario list.
|
|
Each `(team, scenario)` pair is one stable fingerprint, so firing the same
|
|
one twice updates the same alert (a real re-fire) and `resolve` closes
|
|
exactly that one.
|
|
|
|
Set `CLUSTER` to send the alert as if it came from one of several clusters:
|
|
|
|
```sh
|
|
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
|
|
CLUSTER=prod-us ./fire-alerts.sh platform high-cpu # a second incident, not a join
|
|
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
|
|
```
|
|
|
|
It stands in for a Prometheus external label plus `cluster` in Alertmanager's
|
|
`group_by` (terdut-server's README, "Several clusters, one team"): the web UI
|
|
then shows the cluster chip on each incident and a cluster filter in the
|
|
queue. `CLUSTER` is part of the fingerprint, so resolve with the same value you
|
|
fired with. `./run-demo.sh` fires its alerts across `prod-eu` and `prod-us`.
|
|
|
|
### Dead man's switches
|
|
|
|
The `deadmanSwitches` in `02-team-platform.yaml` / `03-team-payments.yaml` expect a heartbeat
|
|
alert on a 15-minute timeout:
|
|
|
|
```sh
|
|
./fire-alerts.sh platform heartbeat
|
|
```
|
|
|
|
Keep sending that (e.g. a `watch -n 60`) and nothing happens — that's the
|
|
point. Stop sending it and, 15 minutes after the last one, terdut-server
|
|
opens a `critical` incident on its own, with no webhook involved: proof
|
|
the switch is watching for silence, not for a signal.
|
|
|
|
## Tear down
|
|
|
|
```sh
|
|
kubectl delete namespace terdut-operator-demo
|
|
```
|
|
|
|
The operator's finalizers delete each team on the server (with its
|
|
escalation, switches and integrations) before this namespace's objects
|
|
actually disappear — give it a few seconds past the `kubectl delete`
|
|
returning. A team with open incidents is not deleted until they are resolved.
|