Files
Niklas Ye a0ea13955e
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 52s
CI / test (pull_request) Successful in 2m43s
TerdutTeam: mint and surface a real invite link (spec.invite)
The actual fix for the human-onboarding gap niklas/terdut-server#23 found --
not a terdut-server change at all. A team-scoped credential is already
owner-equivalent for POST/GET/DELETE /api/teams/{teamID}/invites
(requireTeamOwner's synthetic-membership mechanism, ratified not
accidental per that repo's SERVICE-ACCOUNTS.md), and invite redemption
bypasses signup_mode entirely -- this TerdutTeam controller just never
grew a feature to use either fact.

New spec.invite{enabled, role (member|owner, default member), maxUses
(1-100, default 1)} and status.inviteSecretRef. The Secret lives in the
TerdutTeam's OWN namespace, not the operator's: unlike
status.credentialsSecretRef (a durable, high-privilege credential, kept
operator-side per DESIGN.md §6), an invite is bounded and limited-use,
meant for this namespace's own human operators to read and hand out --
same precedent as TerdutAlertSource's status.webhookURLSecretRef, same-
namespace and OwnerReference'd so deleting the TerdutTeam garbage-collects
it automatically.

internal/controller/terdutteam_invite.go: mints on first spec.invite.enabled,
refreshes a day ahead of terdut-server's fixed 7-day TTL (reading the
Secret's own stored expiresAt, no extra server round-trip per reconcile),
revokes server-side and deletes the Secret when flipped back to false. A
lost invite Secret is silently re-minted rather than treated as
unrecoverable the way TerdutAlertSource's webhook key is -- nothing
external holds a durable dependency on one specific invite link staying
stable, it's read once by one human and handed out.

New tdclient.Invite/CreateInvite/RevokeInvite. New envtest coverage: mint
into the team's own namespace, refresh-before-expiry, revoke-on-disable
(internal/controller/terdutteam_controller_test.go's new "spec.invite"
Describe block), plus the fake server growing invite support
(terdutserver_controller_test.go) -- its handleTeamSubPath dispatcher was
split further (deadman switches into their own handleDeadmanSubPath,
matching the existing handleIntegrationSubPath precedent) to stay under
golangci-lint's gocyclo threshold with the new route added.

examples/demo updated to prove this end to end: 02-team-platform.yaml
turns on spec.invite; run-demo.sh's bootstrap_login/join_demo_teams (the
psql signup_mode flip + a direct team_members INSERT) are replaced by
redeem_platform_invite (reads status.inviteSecretRef, a real POST
/api/signup with the invite token) and join_payments_team (POST
/api/teams/{teamID}/members using Payments' own credential and alice's
user id resolved via GET /api/users, deliberately not given its own
spec.invite, so the demo shows both onboarding paths this feature
unlocks) -- zero kubectl exec/psql calls remain anywhere in the script.
README.md's "First login" section rewritten to match; it no longer
documents the admin-token curl call that 403s against current
terdut-server (niklas/terdut-server#23).

Depends on niklas/terdut-server#24 (the callerMayManageServiceAccount fix
for terdut-operator#3) being released before this is deployed for real --
not required to build or test this change itself, since the envtest fake
never modeled that authorization gap to begin with.
2026-10-02 22:04:56 +02:00

160 lines
6.6 KiB
Markdown

# Demo
Every CRD this operator reconciles, wired into one working install: one
`TerdutServer`, two `TerdutTeam`s (Platform and Payments) each with their
own `TerdutEscalationRule`, `TerdutDeadmanSwitch` and `TerdutAlertSource`,
plus a script that fires synthetic Alertmanager webhooks at it so you can
watch real incidents appear, escalate and resolve.
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
Postgres with `emptyDir` storage and a password committed in this
directory. Throw the whole namespace away when you're done.
**Want this fully automated instead of walking through it by hand?**
`./run-demo.sh` does everything below itself, against a fresh (or
already-set-up) `kind` cluster — creates the cluster, installs the
operator, applies every CR here, signs `alice` in for real, and fires a
few alerts. `./run-demo.sh --help` for the knobs, `./run-demo.sh
--teardown` to tear it back down. The rest of this file is the manual
walkthrough it automates.
## Prerequisites
- The operator and its CRDs installed and running (`make install
deploy IMG=...`, or `charts/terdut-operator` — see this repo's own
README.md/DESIGN.md), pointed at a cluster you're fine creating
throwaway resources in. A `kind` cluster is the easy choice.
- `kubectl`, `jq`, `curl` on your path.
**Apply this into a namespace of its own.** Every object name in this
directory is prefixed `terdut-operator-demo` specifically so applying it
by mistake into some other namespace that already has unrelated objects
doesn't collide with them -- but that only helps if this directory's own
objects don't collide with *each other* across two applies. Applying it
twice into two different namespaces is fine; applying it a second time
into a namespace that already has something else named `terdut-demo` (a
real install from following `terdut-operator`'s own repo along, say) is
exactly the mistake this prefix exists to avoid, and it only works if you
don't override these names yourself.
## Apply it
```sh
kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .
```
Listed and applied in dependency order (server → team → everything that
`teamRef`s it), but you don't have to preserve that order yourself:
every controller here re-queues and waits rather than failing when a ref
isn't resolvable yet (`kubectl describe` shows `Reason: TeamRefNotFound` /
`WaitingForTeam` while that settles).
Watch it converge:
```sh
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
-n terdut-operator-demo
```
The operator's own Deployment template now carries a `wait-for-postgres`
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
so `terdut-operator-demo`'s pod should come up clean even against this brand-new
Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
Once `terdut-operator-demo`'s own `Ready` condition is `True`, everything downstream
of it should settle within a reconcile interval or two.
## See the web UI
The operator never creates any external exposure for a `TerdutServer` --
that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this
demo just reaches it the simplest way there is:
```sh
kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
```
and open http://localhost:8080. See `../networking` for worked examples of
exposing it for real (Gateway API, Istio, or a plain `Ingress`) instead.
### First login
The operator's own bootstrap (DESIGN.md §6) creates the first user through
`/api/bootstrap` and immediately mints itself a service-account token from
it, then discards the bootstrap user's own key — nobody ever signs in as
that account, and `signup_mode` stays `invite_only` by default. **Don't try
to flip it via the operator's own token**: that token is a service account,
and `/api/admin/settings` is deliberately human-only on terdut-server
(`niklas/terdut-server#23` has the full reasoning — widening that gate was
the wrong fix).
The real path in: `02-team-platform.yaml` turns on `spec.invite`, so
Platform's own `TerdutTeam` mints a real invite link with its own
already-working team-scoped credential (the same reach that lets it manage
its own escalation policy, dead man's switches and integrations — owner-
equivalent, confirmed in terdut-server's `SERVICE-ACCOUNTS.md`). Invite
redemption bypasses `signup_mode` entirely, so this needs no admin
credential at all:
```sh
secretname=$(kubectl -n terdut-operator-demo get terdutteam terdutteam-platform \
-o jsonpath='{.status.inviteSecretRef.name}')
url=$(kubectl -n terdut-operator-demo get secret "$secretname" -o jsonpath='{.data.url}' | base64 -d)
echo "$url" # open this, or POST /api/signup with {"invite": "<the token after invite=>", ...}
```
`04-escalation-platform.yaml` names a user `alice` at its first escalation
level — sign up as `alice` if you want that level to mean something rather
than falling through to on-call after 5 minutes. `run-demo.sh` does exactly
this automatically (and also joins `alice` to Payments, which deliberately
has no `spec.invite` of its own — see that file's comment for the second
onboarding path this demonstrates).
## Fire some alerts
In another terminal, with the port-forward above still running:
```sh
export NAMESPACE=terdut-operator-demo
./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
# levels, and show up on the Platform/Payments team's own incident list.
./fire-alerts.sh platform high-cpu resolve
```
`fire-alerts.sh -h` (or any bad argument) prints the full scenario list.
Each `(team, scenario)` pair is one stable fingerprint, so firing the same
one twice updates the same alert (a real re-fire) and `resolve` closes
exactly that one.
### Dead man's switches
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
alert on a 15-minute timeout:
```sh
./fire-alerts.sh platform heartbeat
```
Keep sending that (e.g. a `watch -n 60`) and nothing happens — that's the
point. Stop sending it and, 15 minutes after the last one, terdut-server
opens a `critical` incident on its own, with no webhook involved: proof
the switch is watching for silence, not for a signal.
## Tear down
```sh
kubectl delete namespace terdut-operator-demo
```
The operator's own finalizers clean up everything cross-namespace
(credentials Secrets in the operator's namespace, server-side team/rule/
integration rows) before this namespace's objects actually disappear —
give it a few seconds past the `kubectl delete` returning.