d9315322fc
1. Renamed every object this demo creates (TerdutServer, Postgres
Secret/Deployment/Service) from terdut-demo[-postgres] to
terdut-operator-demo[-postgres]. The user applied this kit into the
already-live "terdut-demo" namespace -- the real operator exercise
from earlier in this repo's own history -- and this demo's own
TerdutServer/Postgres objects shared that exact name. The TerdutServer
apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
while the live object already had postgresClusterRef violates "exactly
one of" and the API server refused it), and the real Postgres Service
was never touched (confirmed live: still Zalando's own spilo selector,
endpoint still the real StatefulSet pod) -- but the Postgres Secret and
Deployment, having no such protection, were created as brand new,
extra, crash-looping objects sitting right next to the real ones.
Prefixing every name this demo creates means a repeat of this exact
mistake no longer collides with anything, documented directly in
README.md now.
2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
(added responding to a PodSecurity "restricted" warning) took
CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
entrypoint needs to chown/chmod the data directory before it drops
privileges itself -- confirmed in a real crashed pod's logs: `chmod:
/var/run/postgresql: Operation not permitted`. kubectl apply
--dry-run=server, which is as far as this got verified before, only
checks admission policy; it was never actually booted. Removed the
capability drop and verified for real this time: applied just
00-postgres.yaml alone into a disposable namespace, waited for the pod
to go Ready, read its logs ("database system is ready to accept
connections"), then deleted that namespace.
147 lines
5.6 KiB
Markdown
147 lines
5.6 KiB
Markdown
# Demo
|
|
|
|
Every CRD this operator reconciles, wired into one working install: one
|
|
`TerdutServer`, two `TerdutTeam`s (Platform and Payments) each with their
|
|
own `TerdutEscalationRule`, `TerdutDeadmanSwitch` and `TerdutAlertSource`,
|
|
plus a script that fires synthetic Alertmanager webhooks at it so you can
|
|
watch real incidents appear, escalate and resolve.
|
|
|
|
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
|
|
Postgres with `emptyDir` storage and a password committed in this
|
|
directory. Throw the whole namespace away when you're done.
|
|
|
|
## Prerequisites
|
|
|
|
- The operator and its CRDs installed and running (`make install
|
|
deploy IMG=...`, or `charts/terdut-operator` — see this repo's own
|
|
README.md/DESIGN.md), pointed at a cluster you're fine creating
|
|
throwaway resources in. A `kind` cluster is the easy choice.
|
|
- `kubectl`, `jq`, `curl` on your path.
|
|
|
|
**Apply this into a namespace of its own.** Every object name in this
|
|
directory is prefixed `terdut-operator-demo` specifically so applying it
|
|
by mistake into some other namespace that already has unrelated objects
|
|
doesn't collide with them -- but that only helps if this directory's own
|
|
objects don't collide with *each other* across two applies. Applying it
|
|
twice into two different namespaces is fine; applying it a second time
|
|
into a namespace that already has something else named `terdut-demo` (a
|
|
real install from following `terdut-operator`'s own repo along, say) is
|
|
exactly the mistake this prefix exists to avoid, and it only works if you
|
|
don't override these names yourself.
|
|
|
|
## Apply it
|
|
|
|
```sh
|
|
kubectl create namespace terdut-operator-demo
|
|
kubectl apply -n terdut-operator-demo -k .
|
|
```
|
|
|
|
Listed and applied in dependency order (server → team → everything that
|
|
`teamRef`s it), but you don't have to preserve that order yourself:
|
|
every controller here re-queues and waits rather than failing when a ref
|
|
isn't resolvable yet (`kubectl describe` shows `Reason: TeamRefNotFound` /
|
|
`WaitingForTeam` while that settles).
|
|
|
|
Watch it converge:
|
|
|
|
```sh
|
|
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
|
|
-n terdut-operator-demo
|
|
```
|
|
|
|
The operator's own Deployment template now carries a `wait-for-postgres`
|
|
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
|
|
so `terdut-operator-demo`'s pod should come up clean even against this brand-new
|
|
Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
|
|
|
|
Once `terdut-operator-demo`'s own `Ready` condition is `True`, everything downstream
|
|
of it should settle within a reconcile interval or two.
|
|
|
|
## See the web UI
|
|
|
|
The operator doesn't create any external exposure yet
|
|
(`NetworkingSpec`'s own doc comment in `api/v1alpha1/terdutserver_types.go`
|
|
— `spec.networking.hostname` is accepted but nothing acts on it), so:
|
|
|
|
```sh
|
|
kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
|
|
```
|
|
|
|
and open http://localhost:8080.
|
|
|
|
### First login
|
|
|
|
The operator's own bootstrap (DESIGN.md §6) creates the first user through
|
|
`/api/bootstrap` and immediately mints itself a service-account token from
|
|
it — that account has no password, so there's nothing to sign in with yet.
|
|
`signup_mode` also defaults to `invite_only`, so open signup needs turning
|
|
on first, using the admin token the operator generated for itself:
|
|
|
|
```sh
|
|
# Which namespace the operator itself runs in:
|
|
kubectl get deploy -A -l control-plane=controller-manager
|
|
|
|
# The Secret holding the operator's own admin token for this TerdutServer
|
|
# (cross-namespace from terdut-operator-demo, per DESIGN.md §7):
|
|
secretname=$(kubectl -n terdut-operator-demo get terdutserver terdut-operator-demo \
|
|
-o jsonpath='{.status.credentialsSecretRef.name}')
|
|
token=$(kubectl -n <operator-namespace-from-above> get secret "$secretname" \
|
|
-o jsonpath='{.data.token}' | base64 -d)
|
|
|
|
curl -X PUT http://localhost:8080/api/admin/settings \
|
|
-H "Authorization: Bearer $token" -H 'Content-Type: application/json' \
|
|
-d '{"signup_mode":"open"}'
|
|
```
|
|
|
|
Then sign up through the UI as a normal human account. `04-escalation-platform.yaml`
|
|
names a user `alice` at its first escalation level — sign up as `alice` if
|
|
you want that level to mean something rather than falling through to
|
|
on-call after 5 minutes.
|
|
|
|
## Fire some alerts
|
|
|
|
In another terminal, with the port-forward above still running:
|
|
|
|
```sh
|
|
export NAMESPACE=terdut-operator-demo
|
|
|
|
./fire-alerts.sh platform high-cpu
|
|
./fire-alerts.sh platform disk-full
|
|
./fire-alerts.sh payments pod-crash
|
|
|
|
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
|
|
# levels, and show up on the Platform/Payments team's own incident list.
|
|
|
|
./fire-alerts.sh platform high-cpu resolve
|
|
```
|
|
|
|
`fire-alerts.sh -h` (or any bad argument) prints the full scenario list.
|
|
Each `(team, scenario)` pair is one stable fingerprint, so firing the same
|
|
one twice updates the same alert (a real re-fire) and `resolve` closes
|
|
exactly that one.
|
|
|
|
### Dead man's switches
|
|
|
|
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
|
|
alert on a 15-minute timeout:
|
|
|
|
```sh
|
|
./fire-alerts.sh platform heartbeat
|
|
```
|
|
|
|
Keep sending that (e.g. a `watch -n 60`) and nothing happens — that's the
|
|
point. Stop sending it and, 15 minutes after the last one, terdut-server
|
|
opens a `critical` incident on its own, with no webhook involved: proof
|
|
the switch is watching for silence, not for a signal.
|
|
|
|
## Tear down
|
|
|
|
```sh
|
|
kubectl delete namespace terdut-operator-demo
|
|
```
|
|
|
|
The operator's own finalizers clean up everything cross-namespace
|
|
(credentials Secrets in the operator's namespace, server-side team/rule/
|
|
integration rows) before this namespace's objects actually disappear —
|
|
give it a few seconds past the `kubectl delete` returning.
|