# terdut-operator design This is the design reference for implementing terdut-operator. It exists so implementation can start from settled decisions instead of re-litigating them mid-PR. It is deliberately more detailed than the README; the README stays as the short pitch and now points here instead of carrying open questions. Written against terdut-server as of the Postgres-only, per-team-resources version (teams, escalation policies, dead man's switches and integrations are all rows scoped to a team, managed over `internal/api/*` — not env/config-file driven; see that repo's `charts/terdut-server` for the current deploy story this operator supersedes). ## 1. Goals & non-goals **Goal:** let a terdut-server install — the server itself, its teams, their escalation policies, dead man's switches and alert-source integrations — be fully described as Kubernetes objects and managed through gitops, following controller-runtime / Kubebuilder conventions. **Non-goals (v1):** - Not a general-purpose Postgres operator. It *consumes* a database that either the Zalando `postgres-operator` or something else already provides. - Not managing Alertmanager itself, or the routing rules that decide which alerts reach which integration webhook — only the terdut-server side (creating the integration and handing back its URL/key). - Not OLM packaging. Plain Kubebuilder manifests + Helm chart for install, matching how terdut-server itself ships. - Cross-namespace references are limited to exactly one edge: `TerdutTeam.spec.serverRef` may name a `TerdutServer` in a different namespace, gated by an explicit consent object in the target namespace (§4.2, §4.6) — this is the multi-tenant shape the operator exists for (one platform team owns a `TerdutServer`; other teams self-service a `TerdutTeam` against it without needing write access to the server's namespace). Every other reference (`teamRef` on the escalation rule/dead-man-switch/alert-source CRDs) stays same-namespace-as-its-`TerdutTeam` only, in v1 — those manage a specific team's own resources and are expected to live alongside it. - No validating/mutating admission webhooks in v1. CEL validation rules on the CRDs (OpenAPI `x-kubernetes-validations`) cover what they can; anything that needs a live look at another object (e.g. "does this teamRef exist") is a status condition, not an admission rejection — keeps v1 to a controller-only deployment with no cert-manager/webhook dependency. ## 2. The README's open questions, resolved > How do team-crd connect with server-crd? Explicit `spec.serverRef: {name, namespace}` on `TerdutTeam` — same reasoning as before (explicit, greppable, trivially validated), but **`namespace` is deliberately part of the reference**: one team can run and own a `TerdutServer`, and other teams — in their own namespaces, without any write access to the server-owning team's namespace — self-service a `TerdutTeam` against it. `namespace` defaults to the `TerdutTeam`'s own namespace when omitted, so the common single-tenant case (`serverRef: {name: terdut}`) is unchanged. This cross-namespace edge needs the target namespace's explicit consent — otherwise any namespace in the cluster could point a `TerdutTeam` at someone else's `TerdutServer` and have the operator provision a team + credentials against it, which is a namespace-boundary violation, not a gitops convenience. §4.6 covers the consent object (`TerdutServerReferenceGrant`), modeled directly on Gateway API's `ReferenceGrant` — the established Kubernetes pattern for exactly this "object A in namespace X wants to reference object B in namespace Y" problem. No other reference in this design (`teamRef` on the child CRDs) crosses a namespace boundary, so this is the only place `ReferenceGrant`-style consent is needed (§1, §4.2, §5, §6, §9). > How does escalationrules, switches and alertsources connect to a team? Explicit `spec.teamRef: {name}` on each of `TerdutEscalationRule`, `TerdutDeadmanSwitch`, `TerdutAlertSource` — same reasoning, and it mirrors terdut-server's own data model, where every one of these rows carries a `team_id` foreign key already. A matcher/selector on `TerdutTeam` would be inventing a second source of truth for an association the server already models as a plain reference. > Support for both postgres-operator (Zalando) and bring-your-own, how do we > design that to be user friendly? `TerdutServer.spec.database` is a oneOf, mirroring the chart's existing `database.dsn` / `database.passwordSecret` contract (see §8): - `dsn` + `passwordSecretRef` — bring-your-own, exactly today's chart inputs. - `postgresClusterRef` — points at a Zalando `postgresql.acid.zalan.do` CR; the operator derives the DSN and resolves the generated credentials Secret itself (see §8). CloudNativePG support is a natural follow-up using the same shape and is called out as deferred (§13), not designed in detail now. ## 3. API group, versions, CRD catalog - Group: `terdut.ryuvia.com`, version: `v1alpha1` (matches the `ryuvia.com` domain terdut-server already uses; bump to `v1beta1`/`v1` per the normal Kubernetes API graduation criteria once the shapes below have proven stable against a real install). - Module: `git.ryuvia.com/niklas/terdut-operator`, scaffolded with Kubebuilder (controller-runtime), matching terdut-server's Go toolchain and house style. | Kind | Scope | Purpose | |---|---|---| | `TerdutServer` | Namespaced | One terdut-server install: Deployment, Service, database wiring, bootstrap, operator credentials. | | `TerdutTeam` | Namespaced | One team on a `TerdutServer`, possibly in another namespace: name, OIDC group mapping. | | `TerdutServerReferenceGrant` | Namespaced (lives with the `TerdutServer`) | Consent for a `TerdutTeam` in another namespace to reference this namespace's `TerdutServer`(s). | | `TerdutEscalationRule` | Namespaced | A team's escalation policy (levels, targets, repeat). | | `TerdutDeadmanSwitch` | Namespaced | One dead man's switch on a team. | | `TerdutAlertSource` | Namespaced | One alert-ingest integration on a team (currently: Alertmanager webhook). | ## 4. Per-CRD spec ### 4.1 `TerdutServer` ```yaml apiVersion: terdut.ryuvia.com/v1alpha1 kind: TerdutServer metadata: name: terdut namespace: oncall spec: image: repository: git.ryuvia.com/niklas/terdut-server tag: v0.9.3 replicas: 1 # terdut-server is not horizontally-scale-tested; keep the field, default 1 networking: hostname: terdut.example.com servicePort: 8080 gatewayListener: "" # same semantics as chart's networking.listener database: dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password} # --- OR --- postgresClusterRef: {name: terdut-postgres} # Zalando `postgresql` CR in the same namespace sweeper: staleAfter: 6h archiveAfter: 168h deadman: matchers: "alertname=Watchdog" timeout: 15m severity: critical notify: ntfyURL: "http://ntfy.ntfy.svc.cluster.local" fallbackTopic: "" repeatEvery: 15m tokenSecretRef: {name: "", key: token} oidc: enabled: false issuer: "" clientID: "" clientSecretRef: {name: "", key: client-secret} name: SSO scopes: "openid profile email" allowedGroups: [] adminGroup: "" sessionMaxAge: 12h passwordLogin: true status: conditions: [...] # Ready, DatabaseReady, Bootstrapped observedGeneration: 3 serviceName: terdut operatorCredentialsSecretRef: {name: terdut-operator-credentials} # see §6 ``` Field-for-field this is the chart's `values.yaml` reshaped as a spec — the operator absorbs the chart's Deployment/Service/bootstrap-job templates, so existing installs have a direct mapping when migrating (see §10). `spec.database` fields are `+kubebuilder:validation:XValidation` guarded to be mutually exclusive (`dsn` xor `postgresClusterRef`); mirrors the chart's "chart provisions no database" stance — this operator provisions no database either, only wires up one that exists. ### 4.2 `TerdutTeam` ```yaml spec: serverRef: name: terdut namespace: platform-oncall # optional; defaults to this TerdutTeam's own namespace. # Cross-namespace requires a matching TerdutServerReferenceGrant # in that namespace (§4.6) — otherwise Ready: False, reason: RefNotPermitted. displayName: "Platform" # -> POST /api/teams {"name": ...}; server assigns the ID oidc: memberGroup: "terdut-platform-members" ownerGroup: "terdut-platform-owners" status: conditions: [...] teamID: 42 # the server-side ID; needed by every child object's controller observedGeneration: 1 ``` Team *membership* (which users belong, `team_members`) is explicitly **not** modeled as a CRD field in v1: terdut-server already manages membership via OIDC group sync at login for SSO installs, and manual membership for password-login installs is a people-management action, not infrastructure — forcing it through gitops would mean a human's team change goes through a PR review. Flagged in §13 as revisitable if a real gitops-membership need shows up. ### 4.3 `TerdutEscalationRule` ```yaml spec: teamRef: {name: platform-team} repeatCount: 2 fallbackTopic: "platform-oncall" levels: - timeout: 5m targets: - kind: oncall # "oncall" or "user" - kind: user username: alice # resolved to a user ID by the controller at apply time - timeout: 15m targets: - kind: oncall status: conditions: [...] observedGeneration: 1 ``` One `TerdutEscalationRule` per team — the server itself models a policy as one row (`escalation_policies`) with an owned list of levels, so a one-CRD-to-one-policy mapping (not one-CRD-per-level) matches the server's own aggregate and lets the whole thing be reconciled with the single `PUT /api/teams/{teamID}/escalation` the API actually exposes (see §5). ### 4.4 `TerdutDeadmanSwitch` ```yaml spec: teamRef: {name: platform-team} name: "prod-watchdog" matcher: "alertname=Watchdog,cluster=prod" timeout: 15m severity: critical status: conditions: [...] switchID: 7 ``` ### 4.5 `TerdutAlertSource` ```yaml spec: teamRef: {name: platform-team} kind: alertmanager name: "prod-alertmanager" status: conditions: [...] integrationID: 3 webhookURLSecretRef: {name: prod-alertmanager-terdut-webhook} # url + key, never in status/spec ``` The integration key is shown by the API exactly once, at creation (`Integration.Key`/`URL` in terdut-server's own model) — never re-readable, same shape as the bootstrap admin key. The controller writes it straight into a generated, owner-referenced Secret on the create it caused and never logs or stores it anywhere else; the CR's `status` carries only the Secret reference, matching how e.g. cert-manager's `Certificate` exposes `spec.secretName` rather than the key material itself. ### 4.6 `TerdutServerReferenceGrant` Consent object living in the *`TerdutServer`'s* namespace, modeled directly on Gateway API's `ReferenceGrant` (`gateway.networking.k8s.io/v1beta1`): without one, no `TerdutTeam` outside this namespace may resolve a `serverRef` into it, no matter what it names. ```yaml apiVersion: terdut.ryuvia.com/v1alpha1 kind: TerdutServerReferenceGrant metadata: name: allow-app-teams namespace: platform-oncall # the TerdutServer's namespace spec: from: - group: terdut.ryuvia.com kind: TerdutTeam namespace: team-checkout # one entry per namespace permitted to reference in - group: terdut.ryuvia.com kind: TerdutTeam namespace: team-payments to: - group: terdut.ryuvia.com kind: TerdutServer name: terdut # optional: omit to permit any TerdutServer in this namespace ``` - **No selector/wildcard-across-namespaces field** (e.g. no `namespaceSelector`) in v1, matching upstream `ReferenceGrant`'s own choice: the namespace owner enumerates exactly which namespaces may reference in, which is the whole point of requiring affirmative, auditable consent rather than an implicit or pattern-matched grant. - A `TerdutTeam`'s controller checks for a matching grant (namespace + optionally name) on every reconcile before it will resolve `serverRef` cross-namespace, exactly mirroring how `HTTPRoute`/`Certificate` controllers check `ReferenceGrant` before honouring a cross-namespace `backendRef`/`secretRef`. Same-namespace `serverRef` never needs a grant. - Deleting the grant is a live revocation: the next reconcile of any `TerdutTeam` it used to authorize finds no matching grant, flips `Ready: False, reason: RefNotPermitted`, and — deliberately — does **not** delete the team server-side on revocation alone; it stops reconciling further changes until access is restored or the `TerdutTeam` CR itself is deleted (whose finalizer still needs the credentials Secret described in §6 to clean up, so a revoked grant blocking *new* changes rather than forcing an immediate, possibly credential-less deletion is the safer failure mode). ## 5. Reconciliation semantics terdut-server's REST surface (`internal/api/router.go`) does not give every resource a full update verb, so reconciliation strategy is per-resource: | Resource | Verbs available | Strategy | |---|---|---| | Team | POST create, PUT rename, DELETE, PUT oidc-groups | Real update-in-place: diff spec vs. last-applied, PUT the changed pieces. | | Escalation policy | GET/PUT whole-policy | Update-in-place: PUT the full desired policy every reconcile that finds drift; cheap because whole-policy is small and already loaded whole server-side. | | Dead man's switch | POST create, DELETE — **no PUT** | Delete-and-recreate on any spec diff other than `name`. The controller diffs against `status` (which mirrors what was last successfully applied) rather than re-reading the server every reconcile, to avoid a spurious recreate from field reordering. | | Integration (alert source) | POST create, PATCH rename, DELETE | Rename via PATCH; any other spec change (kind) is delete-and-recreate, which **rotates the webhook key** — called out loudly in the CRD's field docs and in a `Warning` event, since it breaks whatever sends to the old URL/key until the new Secret is picked up. | General rules for every controller: - **Idempotent create**: before POSTing, check `status.` is unset; if the server already has a same-named object from a previous partial reconcile (e.g. after a crash between POST and status-write), treat a 409/name-conflict as "adopt" — GET-by-name and populate status, rather than erroring forever. terdut-server's list endpoints in each of these areas return objects by name, so this is a straightforward correlation. - **Periodic resync** in addition to watch-triggered reconciles (Kubebuilder default `RequeueAfter` on success, e.g. every 5–10 minutes) to catch drift from **someone changing state directly against the server's API/UI**, since gitops correctness means the CR wins, not "first write wins". - **Finalizers** on every CRD that has a server-side counterpart, so deletion calls the corresponding DELETE before the Kubernetes object disappears. Failure to delete server-side (e.g. server unreachable) blocks finalizer removal and surfaces as a `Degraded` condition + event, rather than silently orphaning a row. - **Owner chain for status resolution, not API calls**: `TerdutTeam`'s controller does not call any other controller; every child CRD's controller independently resolves its own `teamRef` → `TerdutTeam.status.teamID` and `serverRef` chain down to `TerdutServer.status.operatorCredentialsSecretRef`, the way any two independent controller-runtime reconcilers would. If the referenced parent isn't `Ready` yet, the child requeues with backoff and reports `Ready: False, reason: WaitingForTeam` — no cross-controller RPC. - **Cross-namespace `serverRef` is re-checked every reconcile, not just at creation**: `TerdutTeam`'s controller lists `TerdutServerReferenceGrant` objects in the target namespace on every pass (cheap: a namespaced List with a field/label index, not a full-cluster scan) before touching the cross-namespace `TerdutServer` — revocation (§4.6) takes effect on the team's very next reconcile, not just when the CR is first applied. ## 6. Bootstrap & authentication to terdut-server's API terdut-server has no first-class "service account" token type — API keys belong to a real user row (`api_keys.user_id`). The chart's current answer is a one-shot Job that calls `POST /api/bootstrap`, gets a one-time admin key back, and stores it in a Secret (`bootstrap-job.yaml`). The operator absorbs this rather than shelling out to curl: 1. On a `TerdutServer`'s first reconcile after its Deployment reports Ready (`/healthz` reachable through the Service), the controller calls `POST /api/bootstrap` itself with a fixed, recognizable identity — e.g. username `terdut-operator`, email `terdut-operator@` — distinct from `spec.bootstrap.username/email` used for the *human* first admin the chart bootstraps today. Two separate bootstrap identities are needed only if terdut-server's `/api/bootstrap` is single-shot per install; if it already returns 403 after the first caller regardless of identity, the operator instead **creates its bot user via `POST /api/users`** once a human admin exists — this needs confirming against the endpoint's actual behavior before implementation and is flagged as a first task, not assumed here. 2. The returned API key is written to a generated Secret (`-operator-credentials`), owner-referenced to the `TerdutServer`, referenced back from `status.operatorCredentialsSecretRef`. 3. Every other **same-namespace** controller (EscalationRule, DeadmanSwitch, AlertSource, and any same-namespace `TerdutTeam`) reads that Secret directly to call the API — never its own credentials. 4. **Cross-namespace `TerdutTeam` never reads the source Secret directly.** Granting arbitrary consuming namespaces `get`/`list`/`watch` on a Secret in the server's namespace would mean *any* future workload in that namespace with Secret-read RBAC could be pointed at it too — the `TerdutServerReferenceGrant` (§4.6) only authorizes the `TerdutTeam` *kind*, not "read this Secret". Instead, once a `TerdutServer` observes at least one valid grant for a namespace, its own controller (which already holds the real credentials, and whose ServiceAccount is the only thing with legitimate cross-namespace write access — see §9) mirrors a copy of the Secret into that consenting namespace, named `.-terdut-credentials`, owned not by an `OwnerReference` (those can't cross namespaces) but tracked in the `TerdutServer`'s status and cleaned up explicitly when the grant authorizing that namespace is removed. The remote `TerdutTeam`'s controller reads only this local mirror, never the original. - This is one shared mirrored Secret per (server, consuming namespace) pair, not one per `TerdutTeam` — terdut-server's own API key isn't team-scoped (§13), so there is nothing finer to hand out; "scoped" here means scoped by *namespace boundary*, not by team permission. 5. **Rotation**: the key is a bearer credential with no expiry modeled server-side today. Rotation is manual (delete the Secret + the `api_keys` row via `DELETE /api/users/{id}/api-keys/{keyID}`, let the controller re-bootstrap) until/unless terdut-server grows key expiry; a rotation invalidates every mirror too, which the `TerdutServer` controller re-copies on its own next reconcile. Documented as an operational runbook note, not automated in v1. 6. This is explicitly a stand-in for a real scoped service-account token type; see §13. ## 7. Ownership, status, garbage collection - Every generated object (Deployment, Service, credentials Secret, webhook Secret) carries a `metav1.OwnerReference` to the CR that caused it, in the same namespace — standard GC, no finalizer needed for these (only for the server-side REST resources, per §5). - Status conditions follow the standard `metav1.Condition` shape with at least `Ready` on every kind, plus kind-specific ones (`TerdutServer`: `DatabaseReady`, `Bootstrapped`; children: `Synced`). - `status.observedGeneration` on every kind, bumped only after a successful reconcile against that generation's spec — the standard way a client (or `kubectl wait`) tells "applied" from "seen". - No cluster-scoped aggregation object (e.g. no cluster-wide "all servers" status) in v1 — `kubectl get terdutservers -A` is the aggregate view. ## 8. Postgres integration Mirrors the chart's existing two-path contract (`values.yaml` `database.dsn`/`passwordSecret`), because that contract is already documented and tested operationally: - **Bring-your-own**: `spec.database.dsn` (no password) + `spec.database.passwordSecretRef` — the operator sets `PGPASSWORD` on the Deployment's container env from that Secret, exactly like the chart does today, and does nothing else. No connectivity check beyond what the Deployment's own readiness probe already gives. - **Zalando `postgres-operator`**: `spec.database.postgresClusterRef` names a `postgresql.acid.zalan.do` CR in the same namespace. The controller: - Reads that CR's status for the primary Service name/port to build the DSN host — `..svc:5432` — and database name convention. - Resolves the generated credentials Secret (`..credentials.postgresql.acid.zalan.do`) the same way the chart's comment already documents, and wires it in as `PGPASSWORD` the same way. - Watches that Secret (not just the `postgresql` CR) so a credential rotation triggers a requeue — the chart today requires a manual pod restart for this; the operator can at least detect and report it via a condition even if restarting on rotation is left as a §13 follow-up rather than done automatically (a rolling restart on credential change is a behavior change worth its own design pass, not folded in here). - Requires read RBAC on `postgresql.acid.zalan.do` (optional CRD — the operator's ClusterRole/Role should not hard-fail if the CRD isn't installed and a given `TerdutServer` uses BYO DSN instead). ## 9. RBAC - The operator's own ServiceAccount needs, per namespace it's granted: `get/list/watch/create/update/patch/delete` on `Deployments`, `Services`, `Secrets` it owns, and `get/list/watch` on `postgresql.acid.zalan.do` (optional, degrade gracefully if absent per §8), plus `get/list/watch` on `TerdutServerReferenceGrant` (§4.6). - No cluster-scoped resources are created by this operator (namespaced CRDs only, per §1) — a `Role` + `RoleBinding` per watched namespace is sufficient; a `ClusterRole` is only needed for watching CRDs across all namespaces, which is the normal Kubebuilder multi-tenant-operator default and doesn't imply cluster-scoped *managed* resources. - **The operator's ServiceAccount is the only thing that ever holds cross-namespace `Secret` write.** It is a single Deployment/binary already watching every namespace it's granted (the normal Kubebuilder shape), so mirroring a credentials Secret into a consenting namespace (§6) is not a new privilege *boundary* — it's the same ServiceAccount that already reconciles objects there — but it is new *scope* (`create`/`update` on `Secrets` cluster-wide rather than only within each object's own namespace), and should be called out explicitly in the operator's ClusterRole/RBAC review, not left implicit. No human or team's own RBAC is ever granted cross-namespace Secret access by this design — `TerdutServerReferenceGrant` only ever authorizes the operator to act, never a person. - terdut-server's own RBAC is unaffected — the operator talks to it purely over HTTP with the bot user's API key, never via the Kubernetes API for app-level state. ## 10. Relationship to `charts/terdut-server` **Recommendation** (flagged explicitly as a decision to confirm before implementation starts, not settled by this document alone): the chart's Deployment/Service/bootstrap-job templates become redundant once `TerdutServer` exists — running both would mean two controllers (Helm and this operator) reconciling the same Deployment, which is exactly the conflict Kubernetes operators exist to avoid. Proposed path: - The chart is repurposed into an **installer chart**: it installs the operator + CRDs (and optionally one `TerdutServer` CR from `values.yaml`, for users who want "helm install and get a server" without hand-writing a CR) rather than templating the Deployment directly. - Existing installs migrate by: `helm template` the current release's `values.yaml` into an equivalent `TerdutServer` CR (mechanical, since §4.1 is deliberately shaped to make that mapping 1:1), install the operator, apply the CR, then let Helm's release be uninstalled or reduced to just the CRD/operator subchart. - This is a breaking change to the chart's contract and needs its own migration guide and probably a major chart version bump — out of scope for this design doc beyond flagging it; do not start that migration work without separately confirming this recommendation. ## 11. Testing strategy - `envtest` (controller-runtime's fake API server) for every controller's reconcile logic against the Kubernetes side. - terdut-server's REST API is faked with a small `httptest.Server` per controller test driven by fixtures matching the real handlers' request/response shapes (already well-documented in `internal/api/*_test.go` on the server side) — no real Postgres or real terdut-server binary needed for controller unit tests. - A smaller number of true end-to-end tests (`kind` cluster + real terdut-server image + real Postgres) covering the golden path per CRD: create `TerdutServer` → `TerdutTeam` → one of each child kind → verify via terdut-server's own API that the objects exist with the right shape → delete the CR → verify the server-side object is gone. ## 12. Observability - Standard controller-runtime metrics (reconcile duration/error counts) are enough for v1 — no custom metrics. - Every externally-visible action (bootstrap, key rotation-needed, delete- and-recreate on the no-PUT resources, adopt-on-conflict) emits a Kubernetes `Event` on the CR, since that's what shows up in `kubectl describe` and gitops tooling (Argo CD/Flux) surfaces without extra wiring. ## 13. Deferred / explicitly out of scope for this design - A real scoped service-account/token type in terdut-server (least-privilege operator identity instead of a bot admin user) — worth raising as a terdut-server feature request, not designed here. - CloudNativePG support — same `spec.database` shape as Zalando should extend to it, but the concrete field/Secret-naming conventions need their own look. - Cross-namespace `teamRef` on the child CRDs (`TerdutEscalationRule`, `TerdutDeadmanSwitch`, `TerdutAlertSource`) — only `TerdutTeam.serverRef` crosses namespaces in v1 (§2, §4.2, §4.6); these stay same-namespace as their `TerdutTeam` until a real need for splitting them out shows up. - A `TerdutServerReferenceGrant`-style namespace selector (rather than an enumerated list of namespaces) — deliberately omitted to match upstream `ReferenceGrant`'s own choice of explicit, auditable consent over pattern-matching (§4.6); revisit only if enumerating namespaces becomes genuinely unworkable at scale. - Gitops-managed team *membership* (see §4.2). - Automatic Deployment restart on upstream Postgres credential rotation. - Admission webhooks / CEL-only validation limits (e.g. verifying a `teamRef` exists at admission time rather than surfacing it as a status condition after the fact). - OLM packaging, Helm chart migration execution (§10 is a recommendation, not a plan to execute).