ae97d28444
This operator's own Deployment template crash-looped a few times against a from-scratch postgres-operator cluster still doing initdb and Patroni leader election -- exactly the gap examples/demo's own README just documented for it. terdut-server's ping-retry budget on startup (internal/db/db.go in that repo) is sized for a much shorter, different race (NetworkPolicy propagation, a few seconds), not genuine first-time cluster creation, so it exhausted and the process exited before ever binding its HTTP port -- a startupProbe cannot fix that, since the crash happens before there is anything to probe. Same root cause and same fix as charts/terdut-server's own deployment.yaml template as of that repo's v0.33.2. waitForPostgresContainer reuses dbEnv unchanged: both of resolveDatabaseEnv's two paths (DSN, postgresClusterRef) put TERDUT_DB_DSN first, so it's already exactly what pg_isready needs, and pg_isready needs no credentials, so dbEnv's optional PGPASSWORD riding along too is harmless rather than load-bearing. Covered by the existing envtest suite (asserts on Containers[0], the main container, unaffected by adding InitContainers) -- `make test` passes unchanged, 71.7% coverage on internal/controller. Updates examples/demo's own README, which no longer needs to warn about this.
136 lines
4.9 KiB
Markdown
136 lines
4.9 KiB
Markdown
# Demo
|
|
|
|
Every CRD this operator reconciles, wired into one working install: one
|
|
`TerdutServer`, two `TerdutTeam`s (Platform and Payments) each with their
|
|
own `TerdutEscalationRule`, `TerdutDeadmanSwitch` and `TerdutAlertSource`,
|
|
plus a script that fires synthetic Alertmanager webhooks at it so you can
|
|
watch real incidents appear, escalate and resolve.
|
|
|
|
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
|
|
Postgres with `emptyDir` storage and a password committed in this
|
|
directory. Throw the whole namespace away when you're done.
|
|
|
|
## Prerequisites
|
|
|
|
- The operator and its CRDs installed and running (`make install
|
|
deploy IMG=...`, or `charts/terdut-operator` — see this repo's own
|
|
README.md/DESIGN.md), pointed at a cluster you're fine creating
|
|
throwaway resources in. A `kind` cluster is the easy choice.
|
|
- `kubectl`, `jq`, `curl` on your path.
|
|
|
|
## Apply it
|
|
|
|
```sh
|
|
kubectl create namespace terdut-operator-demo
|
|
kubectl apply -n terdut-operator-demo -k .
|
|
```
|
|
|
|
Listed and applied in dependency order (server → team → everything that
|
|
`teamRef`s it), but you don't have to preserve that order yourself:
|
|
every controller here re-queues and waits rather than failing when a ref
|
|
isn't resolvable yet (`kubectl describe` shows `Reason: TeamRefNotFound` /
|
|
`WaitingForTeam` while that settles).
|
|
|
|
Watch it converge:
|
|
|
|
```sh
|
|
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
|
|
-n terdut-operator-demo
|
|
```
|
|
|
|
The operator's own Deployment template now carries a `wait-for-postgres`
|
|
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
|
|
so `terdut-demo`'s pod should come up clean even against this brand-new
|
|
Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
|
|
|
|
Once `terdut-demo`'s own `Ready` condition is `True`, everything downstream
|
|
of it should settle within a reconcile interval or two.
|
|
|
|
## See the web UI
|
|
|
|
The operator doesn't create any external exposure yet
|
|
(`NetworkingSpec`'s own doc comment in `api/v1alpha1/terdutserver_types.go`
|
|
— `spec.networking.hostname` is accepted but nothing acts on it), so:
|
|
|
|
```sh
|
|
kubectl -n terdut-operator-demo port-forward svc/terdut-demo 8080:8080
|
|
```
|
|
|
|
and open http://localhost:8080.
|
|
|
|
### First login
|
|
|
|
The operator's own bootstrap (DESIGN.md §6) creates the first user through
|
|
`/api/bootstrap` and immediately mints itself a service-account token from
|
|
it — that account has no password, so there's nothing to sign in with yet.
|
|
`signup_mode` also defaults to `invite_only`, so open signup needs turning
|
|
on first, using the admin token the operator generated for itself:
|
|
|
|
```sh
|
|
# Which namespace the operator itself runs in:
|
|
kubectl get deploy -A -l control-plane=controller-manager
|
|
|
|
# The Secret holding the operator's own admin token for this TerdutServer
|
|
# (cross-namespace from terdut-operator-demo, per DESIGN.md §7):
|
|
secretname=$(kubectl -n terdut-operator-demo get terdutserver terdut-demo \
|
|
-o jsonpath='{.status.credentialsSecretRef.name}')
|
|
token=$(kubectl -n <operator-namespace-from-above> get secret "$secretname" \
|
|
-o jsonpath='{.data.token}' | base64 -d)
|
|
|
|
curl -X PUT http://localhost:8080/api/admin/settings \
|
|
-H "Authorization: Bearer $token" -H 'Content-Type: application/json' \
|
|
-d '{"signup_mode":"open"}'
|
|
```
|
|
|
|
Then sign up through the UI as a normal human account. `04-escalation-platform.yaml`
|
|
names a user `alice` at its first escalation level — sign up as `alice` if
|
|
you want that level to mean something rather than falling through to
|
|
on-call after 5 minutes.
|
|
|
|
## Fire some alerts
|
|
|
|
In another terminal, with the port-forward above still running:
|
|
|
|
```sh
|
|
export NAMESPACE=terdut-operator-demo
|
|
|
|
./fire-alerts.sh platform high-cpu
|
|
./fire-alerts.sh platform disk-full
|
|
./fire-alerts.sh payments pod-crash
|
|
|
|
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
|
|
# levels, and show up on the Platform/Payments team's own incident list.
|
|
|
|
./fire-alerts.sh platform high-cpu resolve
|
|
```
|
|
|
|
`fire-alerts.sh -h` (or any bad argument) prints the full scenario list.
|
|
Each `(team, scenario)` pair is one stable fingerprint, so firing the same
|
|
one twice updates the same alert (a real re-fire) and `resolve` closes
|
|
exactly that one.
|
|
|
|
### Dead man's switches
|
|
|
|
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
|
|
alert on a 15-minute timeout:
|
|
|
|
```sh
|
|
./fire-alerts.sh platform heartbeat
|
|
```
|
|
|
|
Keep sending that (e.g. a `watch -n 60`) and nothing happens — that's the
|
|
point. Stop sending it and, 15 minutes after the last one, terdut-server
|
|
opens a `critical` incident on its own, with no webhook involved: proof
|
|
the switch is watching for silence, not for a signal.
|
|
|
|
## Tear down
|
|
|
|
```sh
|
|
kubectl delete namespace terdut-operator-demo
|
|
```
|
|
|
|
The operator's own finalizers clean up everything cross-namespace
|
|
(credentials Secrets in the operator's namespace, server-side team/rule/
|
|
integration rows) before this namespace's objects actually disappear —
|
|
give it a few seconds past the `kubectl delete` returning.
|