Niklas Ye d9315322fc
CI / chart (push) Successful in 1s
CI / security (push) Successful in 3m24s
CI / test (push) Successful in 10m45s
examples/demo: fix two real bugs this exact demo just hit live
1. Renamed every object this demo creates (TerdutServer, Postgres
   Secret/Deployment/Service) from terdut-demo[-postgres] to
   terdut-operator-demo[-postgres]. The user applied this kit into the
   already-live "terdut-demo" namespace -- the real operator exercise
   from earlier in this repo's own history -- and this demo's own
   TerdutServer/Postgres objects shared that exact name. The TerdutServer
   apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
   while the live object already had postgresClusterRef violates "exactly
   one of" and the API server refused it), and the real Postgres Service
   was never touched (confirmed live: still Zalando's own spilo selector,
   endpoint still the real StatefulSet pod) -- but the Postgres Secret and
   Deployment, having no such protection, were created as brand new,
   extra, crash-looping objects sitting right next to the real ones.
   Prefixing every name this demo creates means a repeat of this exact
   mistake no longer collides with anything, documented directly in
   README.md now.

2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
   (added responding to a PodSecurity "restricted" warning) took
   CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
   entrypoint needs to chown/chmod the data directory before it drops
   privileges itself -- confirmed in a real crashed pod's logs: `chmod:
   /var/run/postgresql: Operation not permitted`. kubectl apply
   --dry-run=server, which is as far as this got verified before, only
   checks admission policy; it was never actually booted. Removed the
   capability drop and verified for real this time: applied just
   00-postgres.yaml alone into a disposable namespace, waited for the pod
   to go Ready, read its logs ("database system is ready to accept
   connections"), then deleted that namespace.
2026-10-02 13:25:47 +02:00
2026-10-01 14:13:19 +02:00
2026-10-01 14:13:19 +02:00
2026-09-30 19:22:09 +02:00
2026-09-30 19:22:09 +02:00

Terdut operator

Aims to expose most config as CRD's, so end users can self-service over gitops.

See DESIGN.md for the full design: CRD catalog and specs, reconciliation semantics, bootstrap/auth, Postgres integration, RBAC, and the relationship to charts/terdut-server. This README stays a short pitch; the open questions it used to carry are now resolved decisions there (§2).

CRD's

terdutServers

Creates a server — Deployment, Service, database wiring, bootstrap, operator credentials, and allowedTeams consent for cross-namespace teams. See DESIGN.md §4.1, §4.6.

terdutTeams

  • team name
  • oidc groups
  • serverRef — explicit reference to its TerdutServer, may be in a different namespace (one team owns the server, others self-service a team against it), gated by that TerdutServer's own allowedTeams field (DESIGN.md §2, §4.1, §4.2, §4.6)

terdutEscalationrules

  • rule
  • teamRef — explicit reference to its TerdutTeam (DESIGN.md §2, §4.3)

terdutDeadmansswitches

  • rule
  • teamRef (DESIGN.md §4.4)

terdutAlertSources

  • teamRef (DESIGN.md §4.5)
  • URL/key are generated by the server at creation and surfaced only via a generated Secret, never set explicitly

Demo

examples/demo wires one of every CRD above together — two teams, each with an escalation rule, a dead man's switch and an alert source — plus a script that fires synthetic Alertmanager webhooks at it, so you can watch real incidents open, escalate and resolve without a real Alertmanager anywhere in the picture.

S
Description
No description provided
Readme 1.5 MiB
Languages
Go 91.1%
Makefile 6.3%
Shell 1.4%
Go Template 0.7%
Dockerfile 0.5%