Authenticate with a seeded operator key; fold escalation and switches into TerdutTeam
Credentials: the TerdutServer controller generates <name>-operator-key in the server's own namespace (owned by it) and hands it to the pods as TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it at every start. A replaced Secret rolls the pods. The bootstrap handshake, the checkpoint Secret, per-team service accounts and credentials Secrets, BootstrapStateLost and credentials.deletionPolicy are gone. CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[] on the team (matched by name, extras removed); team invites are removed. A team is created under the identity <namespace>/<name> (external_id), so a retry, a lost status or a deleted team heal by repeating the same call, and a display name owned by another team is TeamNameTaken instead of an adoption. The server resolves escalation usernames (UnknownUser condition). OIDC claim names and trustEmail are spec fields. Fixes: query values are URL-escaped; every delete treats 404 as success; deleting a team no longer depends on allowedTeams consent; a switch or integration deleted on the server is recreated; unnamed switches take the CR's name. Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo (run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays cluster-wide, now stated in DESIGN.md section 9. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
This commit is contained in:
+29
-43
@@ -1,9 +1,9 @@
|
||||
# Demo
|
||||
|
||||
Every CRD this operator reconciles, wired into one working install: one
|
||||
`TerdutServer`, two `TerdutTeam`s (Platform and Payments) each with their
|
||||
own `TerdutEscalationRule`, `TerdutDeadmanSwitch` and `TerdutAlertSource`,
|
||||
plus a script that fires synthetic Alertmanager webhooks at it so you can
|
||||
`TerdutServer`, two `TerdutTeam`s (Platform and Payments), each carrying its
|
||||
own escalation ladder and dead man's switches, and a `TerdutAlertSource` per
|
||||
team, plus a script that fires synthetic Alertmanager webhooks at it so you can
|
||||
watch real incidents appear, escalate and resolve.
|
||||
|
||||
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
|
||||
@@ -13,7 +13,7 @@ directory. Throw the whole namespace away when you're done.
|
||||
**Want this fully automated instead of walking through it by hand?**
|
||||
`./run-demo.sh` does everything below itself, against a fresh (or
|
||||
already-set-up) `kind` cluster — creates the cluster, installs the
|
||||
operator, applies every CR here, signs `alice` in for real, and fires a
|
||||
operator, applies every CR here, creates `alice` as the first user, and fires a
|
||||
few alerts. `./run-demo.sh --help` for the knobs, `./run-demo.sh
|
||||
--teardown` to tear it back down. The rest of this file is the manual
|
||||
walkthrough it automates.
|
||||
@@ -44,16 +44,17 @@ kubectl create namespace terdut-operator-demo
|
||||
kubectl apply -n terdut-operator-demo -k .
|
||||
```
|
||||
|
||||
Listed and applied in dependency order (server → team → everything that
|
||||
`teamRef`s it), but you don't have to preserve that order yourself:
|
||||
every controller here re-queues and waits rather than failing when a ref
|
||||
isn't resolvable yet (`kubectl describe` shows `Reason: TeamRefNotFound` /
|
||||
`WaitingForTeam` while that settles).
|
||||
Listed and applied in dependency order (server → team → alert source), but
|
||||
you don't have to preserve that order yourself: every controller here
|
||||
re-queues and waits rather than failing when a ref isn't resolvable yet
|
||||
(`kubectl describe` shows `Reason: WaitingForServer` / `WaitingForTeam` while
|
||||
that settles). Platform stays at `Reason: UnknownUser` until `alice` exists:
|
||||
its escalation ladder names her.
|
||||
|
||||
Watch it converge:
|
||||
|
||||
```sh
|
||||
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
|
||||
kubectl get terdutservers,terdutteams,terdutalertsources \
|
||||
-n terdut-operator-demo
|
||||
```
|
||||
|
||||
@@ -80,36 +81,21 @@ exposing it for real (Gateway API, Istio, or a plain `Ingress`) instead.
|
||||
|
||||
### First login
|
||||
|
||||
The operator's own bootstrap (DESIGN.md §6) creates the first user through
|
||||
`/api/bootstrap` and immediately mints itself a service-account token from
|
||||
it, then discards the bootstrap user's own key — nobody ever signs in as
|
||||
that account, and `signup_mode` stays `invite_only` by default. **Don't try
|
||||
to flip it via the operator's own token**: that token is a service account,
|
||||
and `/api/admin/settings` is deliberately human-only on terdut-server
|
||||
(`niklas/terdut-server#23` has the full reasoning — widening that gate was
|
||||
the wrong fix).
|
||||
|
||||
The real path in: `02-team-platform.yaml` turns on `spec.invite`, so
|
||||
Platform's own `TerdutTeam` mints a real invite link with its own
|
||||
already-working team-scoped credential (the same reach that lets it manage
|
||||
its own escalation policy, dead man's switches and integrations — owner-
|
||||
equivalent, confirmed in terdut-server's `SERVICE-ACCOUNTS.md`). Invite
|
||||
redemption bypasses `signup_mode` entirely, so this needs no admin
|
||||
credential at all:
|
||||
The operator authenticates with a key of its own (the `TerdutServer`'s
|
||||
`<name>-operator-key` Secret, handed to the server as `TERDUT_OPERATOR_KEY`) and
|
||||
never creates a user. A person gets in the way anyone does on a fresh
|
||||
terdut-server: `/api/bootstrap` creates the first user, an administrator,
|
||||
while no user exists yet:
|
||||
|
||||
```sh
|
||||
secretname=$(kubectl -n terdut-operator-demo get terdutteam terdutteam-platform \
|
||||
-o jsonpath='{.status.inviteSecretRef.name}')
|
||||
url=$(kubectl -n terdut-operator-demo get secret "$secretname" -o jsonpath='{.data.url}' | base64 -d)
|
||||
echo "$url" # open this, or POST /api/signup with {"invite": "<the token after invite=>", ...}
|
||||
curl -sS -X POST http://localhost:8080/api/bootstrap -H 'Content-Type: application/json' \
|
||||
-d '{"username":"alice","email":"alice@example.com","password":"a-long-demo-password"}'
|
||||
```
|
||||
|
||||
`04-escalation-platform.yaml` names a user `alice` at its first escalation
|
||||
level — sign up as `alice` if you want that level to mean something rather
|
||||
than falling through to on-call after 5 minutes. `run-demo.sh` does exactly
|
||||
this automatically (and also joins `alice` to Payments, which deliberately
|
||||
has no `spec.invite` of its own — see that file's comment for the second
|
||||
onboarding path this demonstrates).
|
||||
The response carries an API key, shown once. An administrator can manage any
|
||||
team, so add `alice` to both teams with it (`POST /api/teams/{teamID}/members`;
|
||||
`kubectl get terdutteam -o jsonpath='{.status.teamID}'` gives the ids), or just
|
||||
use the UI's team pages. `run-demo.sh` does exactly this for you.
|
||||
|
||||
## Fire some alerts
|
||||
|
||||
@@ -122,8 +108,8 @@ export NAMESPACE=terdut-operator-demo
|
||||
./fire-alerts.sh platform disk-full
|
||||
./fire-alerts.sh payments pod-crash
|
||||
|
||||
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
|
||||
# levels, and show up on the Platform/Payments team's own incident list.
|
||||
# Watch it open an incident, escalate per the escalation ladders in
|
||||
# 02/03-team-*.yaml, and show up on the Platform/Payments team's own incident list.
|
||||
|
||||
./fire-alerts.sh platform high-cpu resolve
|
||||
```
|
||||
@@ -149,7 +135,7 @@ fired with. `./run-demo.sh` fires its alerts across `prod-eu` and `prod-us`.
|
||||
|
||||
### Dead man's switches
|
||||
|
||||
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
|
||||
The `deadmanSwitches` in `02-team-platform.yaml` / `03-team-payments.yaml` expect a heartbeat
|
||||
alert on a 15-minute timeout:
|
||||
|
||||
```sh
|
||||
@@ -167,7 +153,7 @@ the switch is watching for silence, not for a signal.
|
||||
kubectl delete namespace terdut-operator-demo
|
||||
```
|
||||
|
||||
The operator's own finalizers clean up everything cross-namespace
|
||||
(credentials Secrets in the operator's namespace, server-side team/rule/
|
||||
integration rows) before this namespace's objects actually disappear —
|
||||
give it a few seconds past the `kubectl delete` returning.
|
||||
The operator's finalizers delete each team on the server (with its
|
||||
escalation, switches and integrations) before this namespace's objects
|
||||
actually disappear — give it a few seconds past the `kubectl delete`
|
||||
returning. A team with open incidents is not deleted until they are resolved.
|
||||
|
||||
Reference in New Issue
Block a user