6a699d4341
spec.pod (api/v1alpha1/terdutserver_types.go): annotations, nodeSelector,
tolerations, affinity, topologySpreadConstraints, resources, pod and
container securityContext, serviceAccountName, extraEnv/extraEnvFrom,
extraVolumes/extraVolumeMounts, imagePullSecrets, and an optional
disruptionBudget. All direct corev1 passthrough -- no wrapper types buy
anything for any of these, matching how CloudNativePG and the Zalando
postgres-operator both expose the same knobs, and matching this repo's
own SweeperSpec precedent ("wrap only when a round-trip through a
different type buys something"). affinity is pure user-supplied
passthrough, not a toggle-plus-generated-default the way a multi-replica
cluster operator's pod anti-affinity usually is: this operator never
auto-generates one, since spec.replicas above 1 isn't a supported
topology (the sweeper/notifier singleton constraint). Considered and
declined for this round: priorityClassName, pod labels beyond
annotations, and a HorizontalPodAutoscaler -- the last of those would
directly contradict the singleton constraint above.
disruptionBudget is the one field here that isn't a plain PodTemplateSpec
knob: when set, the controller now reconciles a PodDisruptionBudget
selecting the TerdutServer's own pods (new terdutserver_pdb.go); clearing
it deletes any it previously created. New RBAC marker on
poddisruptionbudgets to match.
Driven by a public-release pass: looking past this project's own use case
at what a mature, general-purpose operator CRD exposes here (researched
against Zalando postgres-operator and CloudNativePG specifically), not
just the fields this install happened to need.
Separately, and found while answering a question about exposing
TerdutServer through Istio instead of Gateway API: spec.networking's own
doc comment quietly promised a Gateway API HTTPRoute this operator would
build eventually ("a near-term follow-up, not deferred"). That promise is
wrong for a public release -- an operator managing someone's ingress
mechanism for them is a worse default than not touching it at all, and a
surprise HTTPRoute appearing once that follow-up eventually landed would
have been exactly backwards for an Istio (or plain-Ingress, or
intentionally-unexposed) install. Made the non-goal explicit and
permanent instead (DESIGN.md §1), removed the dead `gatewayListener`
field it was the only consumer of (zero runtime call sites anywhere --
setting it already had no effect, so this is a schema cleanup, not a
behavior change), and corrected ROADMAP.md's framing. hostname/servicePort
stay: both are live (TERDUT_PUBLIC_URL, container/Service port), this
operator just never acts on hostname for exposure. Added
examples/networking (Gateway API HTTPRoute, Istio VirtualService) showing
how to expose the plain ClusterIP Service the operator already creates --
outside the operator itself, as illustrations, not as something
examples/demo applies automatically.
No new terdut-server version requirement: both changes are CRD/controller-
only, nothing about the API this operator's bootstrap flow depends on
changed.
148 lines
5.7 KiB
Markdown
148 lines
5.7 KiB
Markdown
# Demo
|
|
|
|
Every CRD this operator reconciles, wired into one working install: one
|
|
`TerdutServer`, two `TerdutTeam`s (Platform and Payments) each with their
|
|
own `TerdutEscalationRule`, `TerdutDeadmanSwitch` and `TerdutAlertSource`,
|
|
plus a script that fires synthetic Alertmanager webhooks at it so you can
|
|
watch real incidents appear, escalate and resolve.
|
|
|
|
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
|
|
Postgres with `emptyDir` storage and a password committed in this
|
|
directory. Throw the whole namespace away when you're done.
|
|
|
|
## Prerequisites
|
|
|
|
- The operator and its CRDs installed and running (`make install
|
|
deploy IMG=...`, or `charts/terdut-operator` — see this repo's own
|
|
README.md/DESIGN.md), pointed at a cluster you're fine creating
|
|
throwaway resources in. A `kind` cluster is the easy choice.
|
|
- `kubectl`, `jq`, `curl` on your path.
|
|
|
|
**Apply this into a namespace of its own.** Every object name in this
|
|
directory is prefixed `terdut-operator-demo` specifically so applying it
|
|
by mistake into some other namespace that already has unrelated objects
|
|
doesn't collide with them -- but that only helps if this directory's own
|
|
objects don't collide with *each other* across two applies. Applying it
|
|
twice into two different namespaces is fine; applying it a second time
|
|
into a namespace that already has something else named `terdut-demo` (a
|
|
real install from following `terdut-operator`'s own repo along, say) is
|
|
exactly the mistake this prefix exists to avoid, and it only works if you
|
|
don't override these names yourself.
|
|
|
|
## Apply it
|
|
|
|
```sh
|
|
kubectl create namespace terdut-operator-demo
|
|
kubectl apply -n terdut-operator-demo -k .
|
|
```
|
|
|
|
Listed and applied in dependency order (server → team → everything that
|
|
`teamRef`s it), but you don't have to preserve that order yourself:
|
|
every controller here re-queues and waits rather than failing when a ref
|
|
isn't resolvable yet (`kubectl describe` shows `Reason: TeamRefNotFound` /
|
|
`WaitingForTeam` while that settles).
|
|
|
|
Watch it converge:
|
|
|
|
```sh
|
|
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
|
|
-n terdut-operator-demo
|
|
```
|
|
|
|
The operator's own Deployment template now carries a `wait-for-postgres`
|
|
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
|
|
so `terdut-operator-demo`'s pod should come up clean even against this brand-new
|
|
Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
|
|
|
|
Once `terdut-operator-demo`'s own `Ready` condition is `True`, everything downstream
|
|
of it should settle within a reconcile interval or two.
|
|
|
|
## See the web UI
|
|
|
|
The operator never creates any external exposure for a `TerdutServer` --
|
|
that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this
|
|
demo just reaches it the simplest way there is:
|
|
|
|
```sh
|
|
kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
|
|
```
|
|
|
|
and open http://localhost:8080. See `../networking` for worked examples of
|
|
exposing it for real (Gateway API, Istio, or a plain `Ingress`) instead.
|
|
|
|
### First login
|
|
|
|
The operator's own bootstrap (DESIGN.md §6) creates the first user through
|
|
`/api/bootstrap` and immediately mints itself a service-account token from
|
|
it — that account has no password, so there's nothing to sign in with yet.
|
|
`signup_mode` also defaults to `invite_only`, so open signup needs turning
|
|
on first, using the admin token the operator generated for itself:
|
|
|
|
```sh
|
|
# Which namespace the operator itself runs in:
|
|
kubectl get deploy -A -l control-plane=controller-manager
|
|
|
|
# The Secret holding the operator's own admin token for this TerdutServer
|
|
# (cross-namespace from terdut-operator-demo, per DESIGN.md §7):
|
|
secretname=$(kubectl -n terdut-operator-demo get terdutserver terdut-operator-demo \
|
|
-o jsonpath='{.status.credentialsSecretRef.name}')
|
|
token=$(kubectl -n <operator-namespace-from-above> get secret "$secretname" \
|
|
-o jsonpath='{.data.token}' | base64 -d)
|
|
|
|
curl -X PUT http://localhost:8080/api/admin/settings \
|
|
-H "Authorization: Bearer $token" -H 'Content-Type: application/json' \
|
|
-d '{"signup_mode":"open"}'
|
|
```
|
|
|
|
Then sign up through the UI as a normal human account. `04-escalation-platform.yaml`
|
|
names a user `alice` at its first escalation level — sign up as `alice` if
|
|
you want that level to mean something rather than falling through to
|
|
on-call after 5 minutes.
|
|
|
|
## Fire some alerts
|
|
|
|
In another terminal, with the port-forward above still running:
|
|
|
|
```sh
|
|
export NAMESPACE=terdut-operator-demo
|
|
|
|
./fire-alerts.sh platform high-cpu
|
|
./fire-alerts.sh platform disk-full
|
|
./fire-alerts.sh payments pod-crash
|
|
|
|
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
|
|
# levels, and show up on the Platform/Payments team's own incident list.
|
|
|
|
./fire-alerts.sh platform high-cpu resolve
|
|
```
|
|
|
|
`fire-alerts.sh -h` (or any bad argument) prints the full scenario list.
|
|
Each `(team, scenario)` pair is one stable fingerprint, so firing the same
|
|
one twice updates the same alert (a real re-fire) and `resolve` closes
|
|
exactly that one.
|
|
|
|
### Dead man's switches
|
|
|
|
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
|
|
alert on a 15-minute timeout:
|
|
|
|
```sh
|
|
./fire-alerts.sh platform heartbeat
|
|
```
|
|
|
|
Keep sending that (e.g. a `watch -n 60`) and nothing happens — that's the
|
|
point. Stop sending it and, 15 minutes after the last one, terdut-server
|
|
opens a `critical` incident on its own, with no webhook involved: proof
|
|
the switch is watching for silence, not for a signal.
|
|
|
|
## Tear down
|
|
|
|
```sh
|
|
kubectl delete namespace terdut-operator-demo
|
|
```
|
|
|
|
The operator's own finalizers clean up everything cross-namespace
|
|
(credentials Secrets in the operator's namespace, server-side team/rule/
|
|
integration rows) before this namespace's objects actually disappear —
|
|
give it a few seconds past the `kubectl delete` returning.
|