7 Commits

Author SHA1 Message Date
Niklas Ye 5b45cf72e1 Bump go.opentelemetry.io/otel to v1.45.0: v1.44.0 carries GO-2026-6505
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m22s
CI / test (push) Successful in 5m0s
Release / test (push) Successful in 5m26s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 7m22s
Release / scan-image (push) Successful in 34s
Exporter config logging may leak endpoint URLs in info logs
(otlptrace/otlptracegrpc/sdk, transitively through grpc's own otel
instrumentation -- all indirect in go.mod, nothing imports these by
name). govulncheck flagged it reachable through real call chains
(tdclient.Client.DeleteIntegration, cmd/main.go's own init), caught by
ci.yaml's security job while cutting v0.1.2 (run 897) -- same pattern as
terdut-server's own da48814 for grpc's CVE-2026-84445.

go.opentelemetry.io/otel, /metric, /sdk, /sdk/metric, /trace,
/exporters/otlp/otlptrace, /exporters/otlp/otlptrace/otlptracegrpc all
moved 1.44.0 -> 1.45.0 together, plus go-logr/logr's own patch bump and
proto/otlp + genproto that go mod tidy pulled along with them. Verified:
go build, full test suite (71.7% coverage unchanged), golangci-lint,
helm-lint, and govulncheck itself now reporting zero reachable
vulnerabilities.
2026-10-02 12:59:40 +02:00
Niklas Ye 738c210505 Set the chart's placeholder version to 0.1.2
CI / chart (push) Successful in 1s
CI / security (push) Failing after 1m4s
CI / test (push) Successful in 2m11s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree heading
for v0.1.2 that still says 0.1.1 tells its reader something false. Same
as 0e3118d before it.
2026-10-02 12:56:16 +02:00
Niklas Ye 46ba0e8d5c CLAUDE.md: the wrapper-chart one-time step is done, not still pending
CI / chart (push) Successful in 1s
CI / test (push) Has been cancelled
CI / security (push) Has been cancelled
Stale since this repo's actual first release (v0.1.1, 2026-10-01) already
did it -- Ryuvia/charts/terdut-operator already exists and already pins
v0.1.1. Caught while about to repeat the same wrong assumption for this
release.
2026-10-02 12:55:23 +02:00
Niklas Ye ae97d28444 Add wait-for-postgres init container to the generated Deployment
CI / chart (push) Successful in 1s
CI / security (push) Failing after 57s
CI / test (push) Successful in 2m0s
This operator's own Deployment template crash-looped a few times against
a from-scratch postgres-operator cluster still doing initdb and Patroni
leader election -- exactly the gap examples/demo's own README just
documented for it. terdut-server's ping-retry budget on startup
(internal/db/db.go in that repo) is sized for a much shorter, different
race (NetworkPolicy propagation, a few seconds), not genuine first-time
cluster creation, so it exhausted and the process exited before ever
binding its HTTP port -- a startupProbe cannot fix that, since the crash
happens before there is anything to probe. Same root cause and same fix
as charts/terdut-server's own deployment.yaml template as of that repo's
v0.33.2.

waitForPostgresContainer reuses dbEnv unchanged: both of
resolveDatabaseEnv's two paths (DSN, postgresClusterRef) put
TERDUT_DB_DSN first, so it's already exactly what pg_isready needs, and
pg_isready needs no credentials, so dbEnv's optional PGPASSWORD riding
along too is harmless rather than load-bearing.

Covered by the existing envtest suite (asserts on Containers[0], the main
container, unaffected by adding InitContainers) -- `make test` passes
unchanged, 71.7% coverage on internal/controller. Updates examples/demo's
own README, which no longer needs to warn about this.
2026-10-02 10:11:55 +02:00
Niklas Ye ca2cd2c645 Add examples/demo: one of every CRD, plus a script to fire alerts at it
CI / chart (push) Successful in 1s
CI / security (push) Failing after 1m7s
CI / test (push) Successful in 2m37s
A self-contained demo kit: a TerdutServer against a throwaway, bare
Postgres (bring-your-own DSN -- simplest path to stand up from nothing,
ROADMAP.md Stage 1's own note), two TerdutTeams, and each team's own
TerdutEscalationRule/TerdutDeadmanSwitch/TerdutAlertSource, so every CRD
this operator manages is exercised together rather than in isolation the
way config/samples' one-of-each already does.

fire-alerts.sh sends terdut-server's own amPayload/amAlert shape (read
from internal/api/alertmanager.go in that repo, not guessed from its
docs) at whichever TerdutAlertSource's generated webhook Secret it reads
the key out of -- high-cpu/disk-full/pod-crash scenarios to open and
resolve incidents, and a heartbeat scenario matching each team's dead
man's switch matcher, so stopping it demonstrates the switch noticing
silence on its own.

Verified server-side (kubectl apply --dry-run=server -k examples/demo)
against this operator's own dev cluster, which already has these CRDs
installed: every object validates. The one warning that cluster's
"restricted" PodSecurity raises (postgres:17-alpine's entrypoint needs to
start as root before it drops privileges itself) is noted inline in
00-postgres.yaml rather than worked around -- not a real production
pattern, and this Postgres exists only to be thrown away with the rest of
the demo namespace.

README.md walks through: applying, watching status, why a few early
CrashLoopBackOff restarts on terdut-demo itself are expected (this
operator's Deployment template has no wait-for-postgres init container
yet, unlike charts/terdut-server's chart as of v0.33.2), reaching the web
UI (port-forward -- spec.networking.hostname is accepted but nothing
creates an HTTPRoute for it yet), turning on open signup with the
operator's own generated admin token since the bootstrap-created account
has no password, firing alerts, and tearing down.
2026-10-02 10:03:49 +02:00
Niklas Ye a75b23c4ad DESIGN.md: record the missing OIDC trustEmail field, found exercising a real second install
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 2m14s
Found while standing up terdut-demo (Ryuvia/charts#275), a second real
TerdutServer against the same Authentik provider as production: OIDCSpec
has no trustEmail override, so a demo install copying production's OIDC
config otherwise verbatim silently runs with the wrong default for it.
Not fixed here -- recorded in §13 as a real, found gap, not a decision,
same as the mid-life teamRef note already there.
2026-10-01 19:39:20 +02:00
Niklas Ye 4ab04d29a8 DESIGN.md: document operator mode and what it deliberately doesn't lock
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m3s
CI / test (push) Successful in 2m22s
No operator-mode section existed here before -- terdut-server's own
README.md documents the feature, but this repo's design doc never
mentioned it. Added as §6 point 7, confirmed against source
(internal/api/middleware.go's OperatorModeBlock, router.go's opMode
wrapper): it blocks human writes to exactly the resources this
operator's CRDs manage (team identity, OIDC-group binding, escalation,
dead man's switches, integrations), and nothing else -- team membership,
invites, and the on-call schedule/rota stay human-editable regardless,
confirmed from the router rather than assumed from the README's prose
alone.
2026-10-01 19:16:53 +02:00
20 changed files with 644 additions and 42 deletions
+6 -8
View File
@@ -36,11 +36,9 @@ charts --force` after `config/` changes, then re-review — `--force` does not t
importantly the optional `terdutServer` block, DESIGN.md §10). It installs the operator + importantly the optional `terdutServer` block, DESIGN.md §10). It installs the operator +
CRDs + RBAC, and optionally one `TerdutServer` CR (`terdutServer.enabled`, off by default). CRDs + RBAC, and optionally one `TerdutServer` CR (`terdutServer.enabled`, off by default).
**One manual step the release skill's own automation does not cover**: `release-preflight` That one-time manual step — hand-creating the initial `terdut-operator/` wrapper entry
expects an existing `terdut-operator/` entry under `Ryuvia/charts` to bump on release under `Ryuvia/charts`, since `chart-bump` only ever bumps an existing one — is done. It
(steps 8-10 of the skill). There is no such entry yet — this repo's first-ever release happened during this repo's actual first release (v0.1.1, 2026-10-01; v0.1.0 published but
can publish its own image and chart (the `test`/`image`/`chart`/`scan-image` jobs), but never deployed anywhere, after its own `scan-image` found a CVE in the grpc version it had
the wrapper-chart bump and PR will fail until someone creates that initial wrapper entry just bumped to). Every release since bumps that wrapper entry like any other onboarded
in `Ryuvia/charts` by hand, the same one-time step every other onboarded repo already had repo's.
done for it before its own first release. That's a deliberate decision to deploy this
operator for real, not something to do as a side effect of finishing this stage.
+36
View File
@@ -566,6 +566,27 @@ when nothing ever crosses into a tenant namespace in the first place.
Secret is updated in place. No DB-level workaround, no re-triggering a Secret is updated in place. No DB-level workaround, no re-triggering a
single-shot endpoint that can't fire twice (which is what made rotation single-shot endpoint that can't fire twice (which is what made rotation
unworkable under the old `/api/bootstrap`-only design). unworkable under the old `/api/bootstrap`-only design).
7. **Operator mode** (`TERDUT_OPERATOR_MODE`, terdut-server's own
deploy-time flag, off by default) is the complementary half of this
trust model: it makes terdut-server itself refuse a *human* write (a
session or a user's own API key) on a route, while a service account's
— this operator's — still goes through. Confirmed against source
(`internal/api/middleware.go`'s `OperatorModeBlock`,
`internal/api/router.go`'s `opMode` wrapper): the blocked set is exactly
team create/rename/delete, a team's OIDC-group binding, its escalation
policy, its dead man's switches, and its integrations — precisely the
resources `TerdutTeam`, `TerdutEscalationRule`, `TerdutDeadmanSwitch` and
`TerdutAlertSource` manage, and nothing more. **Deliberately not
blocked, confirmed against the same router**: team membership and
invites (the router's own comment: "membership is deliberately never
gitops-managed"), and the on-call schedule/rota
(`/api/teams/{teamID}/schedule`, `/api/schedule/current`) — neither
route carries the `opMode` wrapper at all. A human can still add or
remove a team member, or assign who's on call, on a server running in
operator mode; only the CRD-shaped resources above are locked to
GitOps. This isn't a gap to close — it's the same boundary §4.2 already
draws for membership, confirmed to hold on the server side too, not
just stated as an intent here.
## 7. Ownership, status, garbage collection ## 7. Ownership, status, garbage collection
@@ -746,6 +767,21 @@ what it was, a separate install, until someone deletes it.
integration/policy/switch to a different team in place regardless, so integration/policy/switch to a different team in place regardless, so
retargeting one onto a live child isn't a supported operation in v1 — retargeting one onto a live child isn't a supported operation in v1 —
delete and recreate the CR instead. delete and recreate the CR instead.
- `TerdutServerSpec.OIDC` has no `trustEmail` field (nor `usernameClaim`,
`emailClaim`, `groupsClaim` — the "rather than being added here
speculatively" fields its own doc comment already names), unlike
`charts/terdut-server`'s own chart, which sets `oidc.trustEmail: true`
for the production install specifically because Authentik reports
`email_verified: false` and without it a user's first SSO sign-in
creates a second, empty account instead of linking to their existing
one (`terdut-server/README.md`'s own account of this). Found by
actually trying to stand up a second real `TerdutServer` against the
same Authentik provider (`terdut-demo`, `Ryuvia/charts#275`), not by
inspection: that install's OIDC config is otherwise a straight copy of
production's and runs with terdut-server's own default (`trustEmail:
false`) regardless, since the CRD has nowhere to put the override.
Worth closing if a second real OIDC install becomes routine rather than
a one-off exercise.
- Gitops-managed team *membership* (see §4.2). - Gitops-managed team *membership* (see §4.2).
- Automatic Deployment restart on upstream Postgres credential rotation. - Automatic Deployment restart on upstream Postgres credential rotation.
- Admission webhooks / CEL-only validation limits (e.g. verifying a - Admission webhooks / CEL-only validation limits (e.g. verifying a
+8
View File
@@ -34,3 +34,11 @@ DESIGN.md §4.1, §4.6.
- `teamRef` (DESIGN.md §4.5) - `teamRef` (DESIGN.md §4.5)
- URL/key are generated by the server at creation and surfaced only via a - URL/key are generated by the server at creation and surfaced only via a
generated Secret, never set explicitly generated Secret, never set explicitly
## Demo
[`examples/demo`](./examples/demo) wires one of every CRD above together
— two teams, each with an escalation rule, a dead man's switch and an
alert source — plus a script that fires synthetic Alertmanager webhooks
at it, so you can watch real incidents open, escalate and resolve without
a real Alertmanager anywhere in the picture.
+2 -2
View File
@@ -6,8 +6,8 @@ type: application
# These fields decide nothing: `make helm-package` passes --version and # These fields decide nothing: `make helm-package` passes --version and
# --app-version from the release tag (same reasoning as terdut-server's own # --app-version from the release tag (same reasoning as terdut-server's own
# chart). They're for whoever reads the tree before a tag exists. # chart). They're for whoever reads the tree before a tag exists.
version: 0.1.1 version: 0.1.2
appVersion: "v0.1.1" appVersion: "v0.1.2"
keywords: keywords:
- kubernetes - kubernetes
+91
View File
@@ -0,0 +1,91 @@
# Demo-only Postgres: a bare Deployment+Service+Secret, not the Zalando
# postgres-operator path (DatabaseSpec.postgresClusterRef, DESIGN.md §8).
# Bring-your-own DSN is the simpler of the two paths to stand up from
# nothing (ROADMAP.md Stage 1's own note), which is all this needs to be.
#
# emptyDir, one replica, a password sitting in a plaintext Secret below --
# none of that is how you'd run Postgres for real. It exists only so
# 01-server.yaml has something to talk to. Throw the whole demo namespace
# away when you're done; nothing here is meant to survive that.
apiVersion: v1
kind: Secret
metadata:
name: terdut-demo-postgres
type: Opaque
stringData:
password: demo-not-a-real-password
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: terdut-demo-postgres
labels:
app: terdut-demo-postgres
spec:
replicas: 1
# Recreate, not RollingUpdate: emptyDir means a new pod starts with an
# empty database anyway, and two Postgres pods would never agree on one
# emptyDir each.
strategy:
type: Recreate
selector:
matchLabels:
app: terdut-demo-postgres
template:
metadata:
labels:
app: terdut-demo-postgres
spec:
containers:
- name: postgres
image: postgres:17-alpine
# Partial, deliberately: the official image's entrypoint needs to
# start as root to chown the data directory before it drops
# privileges itself (gosu, to the postgres user) -- forcing
# runAsNonRoot here would just refuse to start the container. A
# "restricted" PodSecurity namespace warns on that gap rather
# than blocking (confirmed server-side against this operator's
# own dev cluster), which is an acceptable tradeoff for Postgres
# that exists only to be thrown away with the rest of this demo.
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
seccompProfile:
type: RuntimeDefault
ports:
- name: postgres
containerPort: 5432
env:
- name: POSTGRES_USER
value: terdut
- name: POSTGRES_DB
value: terdut
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: terdut-demo-postgres
key: password
volumeMounts:
- name: data
mountPath: /var/lib/postgresql/data
subPath: pgdata
readinessProbe:
exec:
command: ["pg_isready", "-U", "terdut"]
initialDelaySeconds: 5
volumes:
- name: data
emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
name: terdut-demo-postgres
spec:
selector:
app: terdut-demo-postgres
ports:
- name: postgres
port: 5432
targetPort: postgres
+32
View File
@@ -0,0 +1,32 @@
# The one TerdutServer this whole demo runs against. Everything else in
# this directory (teams, escalation rules, dead man's switches, alert
# sources) references it by name.
#
# networking.hostname is accepted but not yet acted on: creating the
# HTTPRoute for it isn't implemented yet (api/v1alpha1/terdutserver_types.go,
# NetworkingSpec's own doc comment) -- this TerdutServer is reachable from
# outside the cluster only by port-forwarding its Service, same name as
# this object (see README.md).
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
name: terdut-demo
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: v0.33.2
replicas: 1
networking:
hostname: terdut-demo.example
servicePort: 8080
database:
dsn: "postgres://terdut@terdut-demo-postgres:5432/terdut?sslmode=disable"
passwordSecretRef:
name: terdut-demo-postgres
key: password
sweeper:
staleAfter: 6h
archiveAfter: 168h
# No oidc block: password login only, so there's nothing external to
# register a redirect URI with before this demo can sign in.
passwordLogin: true
+17
View File
@@ -0,0 +1,17 @@
# Two teams (this one and 03-team-payments.yaml) so the demo shows
# per-team isolation -- separate incident lists, separate escalation
# ladders, separate alert sources -- rather than one team standing in for
# everything.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
name: terdutteam-platform
spec:
# serverRef.namespace omitted: both this and terdut-demo (01-server.yaml)
# live in whatever namespace you apply this directory into, which is the
# common case and needs no allowedTeams consent on the TerdutServer side
# (DESIGN.md §4.1, §4.6).
serverRef:
name: terdut-demo
displayName: Platform
# No oidc block: this demo is password-login only (01-server.yaml).
+8
View File
@@ -0,0 +1,8 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
name: terdutteam-payments
spec:
serverRef:
name: terdut-demo
displayName: Payments
+24
View File
@@ -0,0 +1,24 @@
# One per team is the rule (DESIGN.md §4.3) -- a second TerdutEscalationRule
# naming the same teamRef would just clobber this one on the next reconcile.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutEscalationRule
metadata:
name: terdutescalationrule-platform
spec:
teamRef:
name: terdutteam-platform
repeatCount: 2
fallbackTopic: platform-fallback
levels:
# username is required iff kind is "user", rejected otherwise -- CEL
# validation at apply time (api/v1alpha1/terdutescalationrule_types.go).
# alice won't exist on a fresh demo install -- see README.md for
# creating a real user if you want this level to mean something, or
# just watch it fall through to oncall after 5m.
- timeout: 5m
targets:
- kind: user
username: alice
- timeout: 10m
targets:
- kind: oncall
+13
View File
@@ -0,0 +1,13 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutEscalationRule
metadata:
name: terdutescalationrule-payments
spec:
teamRef:
name: terdutteam-payments
repeatCount: 1
fallbackTopic: payments-fallback
levels:
- timeout: 5m
targets:
- kind: oncall
+17
View File
@@ -0,0 +1,17 @@
# A per-team dead man's switch (DESIGN.md §4.4) -- a different thing from
# 01-server.yaml's spec.deadman, which this demo leaves unset so this CRD
# is what you're actually seeing reconcile. fire-alerts.sh's "heartbeat"
# scenario sends a matching alert; stop sending it and terdut-server
# itself opens an incident once `timeout` passes with no heartbeat.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutDeadmanSwitch
metadata:
name: terdutdeadmanswitch-platform
spec:
teamRef:
name: terdutteam-platform
# name omitted -- terdut-server derives one from the matcher's own
# canonical form (DESIGN.md §4.4).
matcher: "alertname=PlatformWatchdog"
timeout: 15m
severity: critical
+10
View File
@@ -0,0 +1,10 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutDeadmanSwitch
metadata:
name: terdutdeadmanswitch-payments
spec:
teamRef:
name: terdutteam-payments
matcher: "alertname=PaymentsWatchdog"
timeout: 15m
severity: critical
@@ -0,0 +1,18 @@
# The webhook URL/key fire-alerts.sh sends to, for the Platform team.
# terdut-server shows the key exactly once, at creation, and never again
# (DESIGN.md §4.5) -- this object's status.webhookURLSecretRef names the
# generated Secret holding it (keys "url" and "key"), which is what
# fire-alerts.sh reads. See README.md before applying this: it's the one
# object in this directory whose Secret you can't just re-read if you
# miss it.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutAlertSource
metadata:
name: terdutalertsource-platform
spec:
teamRef:
name: terdutteam-platform
# kind defaults to "alertmanager" -- the only value terdut-server
# supports today.
kind: alertmanager
name: platform-demo-alertmanager
@@ -0,0 +1,9 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutAlertSource
metadata:
name: terdutalertsource-payments
spec:
teamRef:
name: terdutteam-payments
kind: alertmanager
name: payments-demo-alertmanager
+135
View File
@@ -0,0 +1,135 @@
# Demo
Every CRD this operator reconciles, wired into one working install: one
`TerdutServer`, two `TerdutTeam`s (Platform and Payments) each with their
own `TerdutEscalationRule`, `TerdutDeadmanSwitch` and `TerdutAlertSource`,
plus a script that fires synthetic Alertmanager webhooks at it so you can
watch real incidents appear, escalate and resolve.
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
Postgres with `emptyDir` storage and a password committed in this
directory. Throw the whole namespace away when you're done.
## Prerequisites
- The operator and its CRDs installed and running (`make install
deploy IMG=...`, or `charts/terdut-operator` — see this repo's own
README.md/DESIGN.md), pointed at a cluster you're fine creating
throwaway resources in. A `kind` cluster is the easy choice.
- `kubectl`, `jq`, `curl` on your path.
## Apply it
```sh
kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .
```
Listed and applied in dependency order (server → team → everything that
`teamRef`s it), but you don't have to preserve that order yourself:
every controller here re-queues and waits rather than failing when a ref
isn't resolvable yet (`kubectl describe` shows `Reason: TeamRefNotFound` /
`WaitingForTeam` while that settles).
Watch it converge:
```sh
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
-n terdut-operator-demo
```
The operator's own Deployment template now carries a `wait-for-postgres`
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
so `terdut-demo`'s pod should come up clean even against this brand-new
Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
Once `terdut-demo`'s own `Ready` condition is `True`, everything downstream
of it should settle within a reconcile interval or two.
## See the web UI
The operator doesn't create any external exposure yet
(`NetworkingSpec`'s own doc comment in `api/v1alpha1/terdutserver_types.go`
— `spec.networking.hostname` is accepted but nothing acts on it), so:
```sh
kubectl -n terdut-operator-demo port-forward svc/terdut-demo 8080:8080
```
and open http://localhost:8080.
### First login
The operator's own bootstrap (DESIGN.md §6) creates the first user through
`/api/bootstrap` and immediately mints itself a service-account token from
it — that account has no password, so there's nothing to sign in with yet.
`signup_mode` also defaults to `invite_only`, so open signup needs turning
on first, using the admin token the operator generated for itself:
```sh
# Which namespace the operator itself runs in:
kubectl get deploy -A -l control-plane=controller-manager
# The Secret holding the operator's own admin token for this TerdutServer
# (cross-namespace from terdut-operator-demo, per DESIGN.md §7):
secretname=$(kubectl -n terdut-operator-demo get terdutserver terdut-demo \
-o jsonpath='{.status.credentialsSecretRef.name}')
token=$(kubectl -n <operator-namespace-from-above> get secret "$secretname" \
-o jsonpath='{.data.token}' | base64 -d)
curl -X PUT http://localhost:8080/api/admin/settings \
-H "Authorization: Bearer $token" -H 'Content-Type: application/json' \
-d '{"signup_mode":"open"}'
```
Then sign up through the UI as a normal human account. `04-escalation-platform.yaml`
names a user `alice` at its first escalation level — sign up as `alice` if
you want that level to mean something rather than falling through to
on-call after 5 minutes.
## Fire some alerts
In another terminal, with the port-forward above still running:
```sh
export NAMESPACE=terdut-operator-demo
./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
# levels, and show up on the Platform/Payments team's own incident list.
./fire-alerts.sh platform high-cpu resolve
```
`fire-alerts.sh -h` (or any bad argument) prints the full scenario list.
Each `(team, scenario)` pair is one stable fingerprint, so firing the same
one twice updates the same alert (a real re-fire) and `resolve` closes
exactly that one.
### Dead man's switches
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
alert on a 15-minute timeout:
```sh
./fire-alerts.sh platform heartbeat
```
Keep sending that (e.g. a `watch -n 60`) and nothing happens — that's the
point. Stop sending it and, 15 minutes after the last one, terdut-server
opens a `critical` incident on its own, with no webhook involved: proof
the switch is watching for silence, not for a signal.
## Tear down
```sh
kubectl delete namespace terdut-operator-demo
```
The operator's own finalizers clean up everything cross-namespace
(credentials Secrets in the operator's namespace, server-side team/rule/
integration rows) before this namespace's objects actually disappear —
give it a few seconds past the `kubectl delete` returning.
+138
View File
@@ -0,0 +1,138 @@
#!/usr/bin/env bash
# Sends a synthetic Alertmanager v4 webhook payload at the "platform" or
# "payments" demo team's TerdutAlertSource, so terdut-server opens (or
# resolves) an incident exactly the way it would for a real Alertmanager.
#
# The payload shape here is amPayload/amAlert, read straight out of
# terdut-server's own internal/api/alertmanager.go rather than guessed from
# its docs -- version/status/groupKey/groupLabels, and alerts[] carrying
# status/labels/annotations/startsAt/endsAt/generatorURL/fingerprint.
#
# Why this reads the webhook key out of a kubectl Secret instead of using
# the "url" key already in it: that URL is built from spec.networking.hostname
# (TERDUT_PUBLIC_URL), and nothing in this demo stands up real ingress for
# it (01-server.yaml's own comment) -- so it resolves nowhere. The key
# alone, against whatever you've actually port-forwarded BASE_URL to below,
# is the one part of that URL still usable here.
#
# Usage:
# ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
#
# Prerequisites: kubectl context pointed at the demo namespace, jq, curl,
# and (in another terminal) a running:
# kubectl port-forward svc/terdut-demo 8080:8080
set -euo pipefail
NAMESPACE="${NAMESPACE:-}"
BASE_URL="${BASE_URL:-http://localhost:8080}"
usage() {
cat >&2 <<'EOF'
usage: fire-alerts.sh <platform|payments> <scenario> [resolve]
scenarios:
high-cpu warning -- CPU usage above 90% for 10 minutes
disk-full critical -- disk usage above 95%
pod-crash error -- a pod crash-looping
heartbeat critical -- the team's dead man's switch heartbeat
(matches the matcher in 06/07-deadman-*.yaml -- send this
repeatedly to keep the switch alive, or stop sending it and
watch terdut-server open an incident on its own once
`timeout` passes with no heartbeat. "resolve" is not a valid
third argument for this scenario: a heartbeat is only ever
firing.)
env vars:
NAMESPACE kubectl -n for reading the webhook Secret (required)
BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080)
EOF
exit 1
}
[ $# -ge 2 ] || usage
team="$1" scenario="$2" verb="${3:-fire}"
[ -n "$NAMESPACE" ] || { echo "fire-alerts.sh: set NAMESPACE" >&2; exit 1; }
case "$team" in
platform|payments) ;;
*) usage ;;
esac
case "$scenario" in
high-cpu) alertname=TerdutDemoHighCPU severity=warning summary="CPU usage above 90% for 10 minutes" ;;
disk-full) alertname=TerdutDemoDiskFull severity=critical summary="Disk usage above 95% on /data" ;;
pod-crash) alertname=TerdutDemoPodCrashLooping severity=error summary="Pod web-7f8b9 is crash-looping (5 restarts in 10m)" ;;
heartbeat)
# Must match 06-deadman-platform.yaml / 07-deadman-payments.yaml's own
# matcher exactly -- that's what makes this a heartbeat rather than a
# third ordinary alert.
case "$team" in
platform) alertname=PlatformWatchdog ;;
payments) alertname=PaymentsWatchdog ;;
esac
severity=critical summary="demo heartbeat"
[ "$verb" = fire ] || { echo "fire-alerts.sh: heartbeat is only ever fired, never resolved -- just stop sending it" >&2; exit 1; }
;;
*) usage ;;
esac
case "$verb" in
fire) status=firing ;;
resolve) status=resolved ;;
*) usage ;;
esac
secret_name="terdutalertsource-${team}-terdut-webhook"
key="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.key}' | base64 -d)"
[ -n "$key" ] || { echo "fire-alerts.sh: empty key read from Secret $secret_name -- has 08/09-alertsource-*.yaml reconciled yet?" >&2; exit 1; }
# Stable per (team, scenario) so a resolve targets the same alert a fire
# created: terdut-server correlates on (team_id, fingerprint), not on
# anything else in the payload. Real Alertmanager computes this from the
# alert's label set; a fixed string plays the same role here.
fingerprint="demo-${team}-${scenario}"
now="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
if [ "$status" = firing ]; then
ends_at="0001-01-01T00:00:00Z" # Alertmanager's own "not resolved" zero value
else
ends_at="$now"
fi
payload="$(jq -n \
--arg status "$status" \
--arg groupKey "demo:${team}:${scenario}" \
--arg alertname "$alertname" \
--arg team "$team" \
--arg severity "$severity" \
--arg summary "$summary" \
--arg startsAt "$now" \
--arg endsAt "$ends_at" \
--arg fingerprint "$fingerprint" \
'{
version: "4",
status: $status,
groupKey: $groupKey,
groupLabels: { alertname: $alertname, team: $team },
alerts: [{
status: $status,
labels: { alertname: $alertname, severity: $severity, team: $team, instance: "demo" },
annotations: { summary: $summary },
startsAt: $startsAt,
endsAt: $endsAt,
generatorURL: "https://example.com/demo",
fingerprint: $fingerprint
}]
}')"
url="${BASE_URL}/api/integrations/${key}/alertmanager"
echo "POST $url (team=$team scenario=$scenario status=$status)" >&2
code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \
-X POST "$url" -H 'Content-Type: application/json' -d "$payload")"
echo "-> HTTP $code" >&2
cat /tmp/fire-alerts-response.json >&2
echo >&2
if [ "$code" != "200" ]; then
exit 1
fi
+17
View File
@@ -0,0 +1,17 @@
## kubectl apply -n <your-demo-namespace> -k examples/demo
##
## Listed in apply order even though kustomize itself doesn't need that --
## a human reading this file top-to-bottom should see the same dependency
## order the controllers themselves require (server before team, team
## before everything that teamRefs it).
resources:
- 00-postgres.yaml
- 01-server.yaml
- 02-team-platform.yaml
- 03-team-payments.yaml
- 04-escalation-platform.yaml
- 05-escalation-payments.yaml
- 06-deadman-platform.yaml
- 07-deadman-payments.yaml
- 08-alertsource-platform.yaml
- 09-alertsource-payments.yaml
+10 -10
View File
@@ -25,7 +25,7 @@ require (
github.com/felixge/httpsnoop v1.0.4 // indirect github.com/felixge/httpsnoop v1.0.4 // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/fxamacker/cbor/v2 v2.9.1 // indirect github.com/fxamacker/cbor/v2 v2.9.1 // indirect
github.com/go-logr/logr v1.4.3 // indirect github.com/go-logr/logr v1.4.4 // indirect
github.com/go-logr/stdr v1.2.2 // indirect github.com/go-logr/stdr v1.2.2 // indirect
github.com/go-logr/zapr v1.3.0 // indirect github.com/go-logr/zapr v1.3.0 // indirect
github.com/go-openapi/jsonpointer v1.0.0 // indirect github.com/go-openapi/jsonpointer v1.0.0 // indirect
@@ -64,13 +64,13 @@ require (
github.com/x448/float16 v0.8.4 // indirect github.com/x448/float16 v0.8.4 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect
go.opentelemetry.io/otel v1.44.0 // indirect go.opentelemetry.io/otel v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 // indirect go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 // indirect go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 // indirect
go.opentelemetry.io/otel/metric v1.44.0 // indirect go.opentelemetry.io/otel/metric v1.45.0 // indirect
go.opentelemetry.io/otel/sdk v1.44.0 // indirect go.opentelemetry.io/otel/sdk v1.45.0 // indirect
go.opentelemetry.io/otel/trace v1.44.0 // indirect go.opentelemetry.io/otel/trace v1.45.0 // indirect
go.opentelemetry.io/proto/otlp v1.10.0 // indirect go.opentelemetry.io/proto/otlp v1.11.0 // indirect
go.uber.org/multierr v1.11.0 // indirect go.uber.org/multierr v1.11.0 // indirect
go.uber.org/zap v1.27.1 // indirect go.uber.org/zap v1.27.1 // indirect
go.yaml.in/yaml/v2 v2.4.4 // indirect go.yaml.in/yaml/v2 v2.4.4 // indirect
@@ -86,8 +86,8 @@ require (
golang.org/x/time v0.15.0 // indirect golang.org/x/time v0.15.0 // indirect
golang.org/x/tools v0.48.0 // indirect golang.org/x/tools v0.48.0 // indirect
gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa // indirect google.golang.org/genproto/googleapis/api v0.0.0-20260803160001-6ac0973c030d // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa // indirect google.golang.org/genproto/googleapis/rpc v0.0.0-20260803160001-6ac0973c030d // indirect
google.golang.org/grpc v1.83.2 // indirect google.golang.org/grpc v1.83.2 // indirect
google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af // indirect google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af // indirect
gopkg.in/evanphx/json-patch.v4 v4.13.0 // indirect gopkg.in/evanphx/json-patch.v4 v4.13.0 // indirect
+22 -22
View File
@@ -36,8 +36,8 @@ github.com/gkampitakis/go-diff v1.3.2/go.mod h1:LLgOrpqleQe26cte8s36HTWcTmMEur6O
github.com/gkampitakis/go-snaps v0.5.15 h1:amyJrvM1D33cPHwVrjo9jQxX8g/7E2wYdZ+01KS3zGE= github.com/gkampitakis/go-snaps v0.5.15 h1:amyJrvM1D33cPHwVrjo9jQxX8g/7E2wYdZ+01KS3zGE=
github.com/gkampitakis/go-snaps v0.5.15/go.mod h1:HNpx/9GoKisdhw9AFOBT1N7DBs9DiHo/hGheFGBZ+mc= github.com/gkampitakis/go-snaps v0.5.15/go.mod h1:HNpx/9GoKisdhw9AFOBT1N7DBs9DiHo/hGheFGBZ+mc=
github.com/go-logr/logr v1.2.2/go.mod h1:jdQByPbusPIv2/zmleS9BjJVeZ6kBagPoEUsqbVz/1A= github.com/go-logr/logr v1.2.2/go.mod h1:jdQByPbusPIv2/zmleS9BjJVeZ6kBagPoEUsqbVz/1A=
github.com/go-logr/logr v1.4.3 h1:CjnDlHq8ikf6E492q6eKboGOC0T8CDaOvkHCIg8idEI= github.com/go-logr/logr v1.4.4 h1:tG4xh9yMsRCAiodLVTxyrkzSZ9+o0L1Kg/+cPVcbP/8=
github.com/go-logr/logr v1.4.3/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY= github.com/go-logr/logr v1.4.4/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY=
github.com/go-logr/stdr v1.2.2 h1:hSWxHoqTgW2S2qGc0LTAI563KZ5YKYRhT3MFKZMbjag= github.com/go-logr/stdr v1.2.2 h1:hSWxHoqTgW2S2qGc0LTAI563KZ5YKYRhT3MFKZMbjag=
github.com/go-logr/stdr v1.2.2/go.mod h1:mMo/vtBO5dYbehREoey6XUKy/eSumjCCveDpRre4VKE= github.com/go-logr/stdr v1.2.2/go.mod h1:mMo/vtBO5dYbehREoey6XUKy/eSumjCCveDpRre4VKE=
github.com/go-logr/zapr v1.3.0 h1:XGdV8XW8zdwFiwOA2Dryh1gj2KRQyOOoNmBy4EplIcQ= github.com/go-logr/zapr v1.3.0 h1:XGdV8XW8zdwFiwOA2Dryh1gj2KRQyOOoNmBy4EplIcQ=
@@ -168,22 +168,22 @@ go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ
go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y= go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo= go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI= go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI=
go.opentelemetry.io/otel v1.44.0 h1:JjwHmHpA4iZ3wBxluu2fbbE7j4kqlE8jXyAyPXH7HqU= go.opentelemetry.io/otel v1.45.0 h1:pdrWmLHofpubmArBv1LgFSv1Z0Ie/ppdZzu+kUN5EeU=
go.opentelemetry.io/otel v1.44.0/go.mod h1:BMgjTHL9WPRlRjL2oZCBTL4whCGtXch2H4BhOPIAyYc= go.opentelemetry.io/otel v1.45.0/go.mod h1:XZxIqPapzEYnhNSScF5DIqXhm/rYi0FzCe2XddAwZfQ=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 h1:4YsVu3B8+3qtWYYrsUYgn0OG78pN0rnNPRGX4SbokQI= go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 h1:QRefszxJmfPdjXUUm3j6iDzY03mTPXMjqErFqQ67vUg=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0/go.mod h1:+wnlSn0mD1ADVMe3v9Z/WIaiz6q6gL2J/ejaAmdmv80= go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0/go.mod h1:Tiz03lTBVBrm7eWZBOidzEaYaJa8tjwGUGv6d8mlTyk=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 h1:qazEJlUOQzhCpzQpFETGby7EdqjI1wsd0W+6Gg1SCTU= go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 h1:fG5MCxGz8+2VtrN/WgqSpJFctVz24gpxj8CxkKmc8Ww=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0/go.mod h1:fOD2Yefuxixkx3ahVNf0O/PERb6r4OlbxfATVnYvzCo= go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0/go.mod h1:BmAYTn+3ysbRe+IU2msxmf5Rx3g6DHvex+tWI3LdhYI=
go.opentelemetry.io/otel/metric v1.44.0 h1:1w0gILTcHdr3YI+ixLyjemwrVnsMURbTZFrSYCdDdmc= go.opentelemetry.io/otel/metric v1.45.0 h1:7Eg1uH7CJ5cXv9is6tnBe1FI6rj1nwUdbFypRm3br/M=
go.opentelemetry.io/otel/metric v1.44.0/go.mod h1:8O7hanEPBNgEMmybD3s2VBKcgWOCsA6tzHBPODAiquo= go.opentelemetry.io/otel/metric v1.45.0/go.mod h1:HAPbm1nd3p1PmFH7v2dR+6BjXxw+Lq4a2+pndMAm08s=
go.opentelemetry.io/otel/sdk v1.44.0 h1:nHYwb9lK+fJPU/dnT6s7W7Z8itMWyqrnVfbheVYrZ58= go.opentelemetry.io/otel/sdk v1.45.0 h1:4VVSMgQ83dUgW2aoX5f6JgLvHwIvzcuLnF9lUdCSpCw=
go.opentelemetry.io/otel/sdk v1.44.0/go.mod h1:Osuydd3Se74nqjAKxid74N5eC+jfEqfTegHRnq58oK0= go.opentelemetry.io/otel/sdk v1.45.0/go.mod h1:Sr40LgXV7DsKMMJMKOhUWOgMWTfAaqvm2kF0g7ilwuA=
go.opentelemetry.io/otel/sdk/metric v1.44.0 h1:3LlKgI+VjbVsjNRFZJZAJ30WjXC5VkNRks6si09iEfI= go.opentelemetry.io/otel/sdk/metric v1.45.0 h1:oVFszMfyj1Am6s24Vtc7wBb8BKLcwepJjNEYILuiE3o=
go.opentelemetry.io/otel/sdk/metric v1.44.0/go.mod h1:5B5pMARnXxKhltooO4xUuCBorl65a4EpnTalObqOigA= go.opentelemetry.io/otel/sdk/metric v1.45.0/go.mod h1:vUWUxDZvu1WVRj8JA8S0AdhsPrZoDpA2DdZauIh4mDA=
go.opentelemetry.io/otel/trace v1.44.0 h1:jxF5CsGYCe74MCRx2X4g7WsY/VBKRqqpNvXlX/6gtIk= go.opentelemetry.io/otel/trace v1.45.0 h1:l/mP6Uv7oNO7/TblbhpbgMidxhq1uO/rPsikOyVhxag=
go.opentelemetry.io/otel/trace v1.44.0/go.mod h1:oLl1jrMQAVo6v3GAggN+1VH9VIz9iUSvW53sW1Q8PIE= go.opentelemetry.io/otel/trace v1.45.0/go.mod h1:qoJJA2xNMnxRrdISU/kLtfUH2wNeQbiv+jhs/CxI8bc=
go.opentelemetry.io/proto/otlp v1.10.0 h1:IQRWgT5srOCYfiWnpqUYz9CVmbO8bFmKcwYxpuCSL2g= go.opentelemetry.io/proto/otlp v1.11.0 h1:5rrYs0Ykyj50sdU/JU0x8etU+LubXWb+gED6TbEdMIk=
go.opentelemetry.io/proto/otlp v1.10.0/go.mod h1:/CV4QoCR/S9yaPj8utp3lvQPoqMtxXdzn7ozvvozVqk= go.opentelemetry.io/proto/otlp v1.11.0/go.mod h1:SmVizdCOAm3XBtG1g1NnOdhW6jtddT72hLMhv8VwA8E=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto= go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE= go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
go.uber.org/multierr v1.11.0 h1:blXXJkSxSSfBVBlC76pxqeO+LN3aDfLQo+309xJstO0= go.uber.org/multierr v1.11.0 h1:blXXJkSxSSfBVBlC76pxqeO+LN3aDfLQo+309xJstO0=
@@ -218,10 +218,10 @@ gomodules.xyz/jsonpatch/v2 v2.4.0 h1:Ci3iUJyx9UeRx7CeFN8ARgGbkESwJK+KB9lLcWxY/Zw
gomodules.xyz/jsonpatch/v2 v2.4.0/go.mod h1:AH3dM2RI6uoBZxn3LVrfvJ3E0/9dG4cSrbuBJT4moAY= gomodules.xyz/jsonpatch/v2 v2.4.0/go.mod h1:AH3dM2RI6uoBZxn3LVrfvJ3E0/9dG4cSrbuBJT4moAY=
gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4= gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4=
gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E= gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E=
google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa h1:Kjn0N0tCrDgiAFW+lGO4JZ3ck44CehvJQMAwj9QF0G8= google.golang.org/genproto/googleapis/api v0.0.0-20260803160001-6ac0973c030d h1:FarXi840EJWSHYTN3ERkADbPWjl307+FGrA22KAVjjc=
google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa/go.mod h1:q4lMZS6kskjT5HvCPrnnypcDPVJqT/f4nfxmkE7gryY= google.golang.org/genproto/googleapis/api v0.0.0-20260803160001-6ac0973c030d/go.mod h1:K/+WGbmBY7aNW1HDw1fJnKYo10i0DkAX6pows00dLig=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa h1:mZHHdPZl0dbGHCflZgAq/Q468DWVFcU2whhB2KAo8fk= google.golang.org/genproto/googleapis/rpc v0.0.0-20260803160001-6ac0973c030d h1:IL4hdHzcUv2l/gcg98/Rj3FbtE6axwqslOW8SW0C+S0=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8= google.golang.org/genproto/googleapis/rpc v0.0.0-20260803160001-6ac0973c030d/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8=
google.golang.org/grpc v1.83.2 h1:EManeRomTObA0BU7I8vXgg/78uE5MJ9M8B39EX2WscU= google.golang.org/grpc v1.83.2 h1:EManeRomTObA0BU7I8vXgg/78uE5MJ9M8B39EX2WscU=
google.golang.org/grpc v1.83.2/go.mod h1:YPI1hK3kDked6iHvgX3tR0y+nX/qpMFKhPgFsokw1S8= google.golang.org/grpc v1.83.2/go.mod h1:YPI1hK3kDked6iHvgX3tR0y+nX/qpMFKhPgFsokw1S8=
google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af h1:+5/Sw3GsDNlEmu7TfklWKPdQ0Ykja5VEmq2i817+jbI= google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af h1:+5/Sw3GsDNlEmu7TfklWKPdQ0Ykja5VEmq2i817+jbI=
@@ -52,6 +52,7 @@ func (r *TerdutServerReconciler) reconcileDeployment(
ObjectMeta: metav1.ObjectMeta{Labels: labels}, ObjectMeta: metav1.ObjectMeta{Labels: labels},
Spec: corev1.PodSpec{ Spec: corev1.PodSpec{
EnableServiceLinks: new(false), EnableServiceLinks: new(false),
InitContainers: []corev1.Container{waitForPostgresContainer(dbEnv)},
Containers: []corev1.Container{{ Containers: []corev1.Container{{
Name: "terdut-server", Name: "terdut-server",
Image: fmt.Sprintf("%s:%s", srv.Spec.Image.Repository, srv.Spec.Image.Tag), Image: fmt.Sprintf("%s:%s", srv.Spec.Image.Repository, srv.Spec.Image.Tag),
@@ -104,6 +105,36 @@ func servicePort(srv *terdutv1alpha1.TerdutServer) int32 {
return srv.Spec.Networking.ServicePort return srv.Spec.Networking.ServicePort
} }
// waitForPostgresContainer blocks the main container from starting until
// Postgres accepts connections, matching charts/terdut-server's own
// deployment.yaml template as of v0.33.2 (that repo's CLAUDE.md/release
// notes) -- that chart grew this the moment this exact Deployment, created
// by this controller, crash-looped a few times against a from-scratch
// postgres-operator cluster still doing initdb and Patroni leader election:
// terdut-server's own ping-retry budget on startup (internal/db/db.go) is
// sized for a much shorter, different race (NetworkPolicy propagation, a
// few seconds), not for genuine first-time cluster creation, so it
// exhausted and the process exited before ever binding its HTTP port -- a
// startupProbe cannot help there, since the crash happens before there is
// anything to probe.
//
// Reuses dbEnv unchanged: both of resolveDatabaseEnv's paths put
// TERDUT_DB_DSN first (terdutserver_database.go), so it's already exactly
// what pg_isready needs, and pg_isready needs no credentials -- it reports
// PQPING_OK on anything that amounts to a Postgres backend answering,
// including an auth challenge -- so including dbEnv's optional PGPASSWORD
// here too is harmless rather than load-bearing.
func waitForPostgresContainer(dbEnv []corev1.EnvVar) corev1.Container {
return corev1.Container{
Name: "wait-for-postgres",
Image: "postgres:17-alpine",
Env: dbEnv,
Command: []string{"sh", "-c",
`until pg_isready -d "$TERDUT_DB_DSN"; do echo "wait-for-postgres: not ready yet, retrying in 2s"; sleep 2; done`,
},
}
}
func healthzProbe() *corev1.Probe { func healthzProbe() *corev1.Probe {
return &corev1.Probe{ return &corev1.Probe{
ProbeHandler: corev1.ProbeHandler{ ProbeHandler: corev1.ProbeHandler{