Files
terdut-operator/DESIGN.md
T
2026-09-29 15:16:15 +02:00

28 KiB
Raw Blame History

terdut-operator design

This is the design reference for implementing terdut-operator. It exists so implementation can start from settled decisions instead of re-litigating them mid-PR. It is deliberately more detailed than the README; the README stays as the short pitch and now points here instead of carrying open questions.

Written against terdut-server as of the Postgres-only, per-team-resources version (teams, escalation policies, dead man's switches and integrations are all rows scoped to a team, managed over internal/api/* — not env/config-file driven; see that repo's charts/terdut-server for the current deploy story this operator supersedes).

1. Goals & non-goals

Goal: let a terdut-server install — the server itself, its teams, their escalation policies, dead man's switches and alert-source integrations — be fully described as Kubernetes objects and managed through gitops, following controller-runtime / Kubebuilder conventions.

Non-goals (v1):

  • Not a general-purpose Postgres operator. It consumes a database that either the Zalando postgres-operator or something else already provides.
  • Not managing Alertmanager itself, or the routing rules that decide which alerts reach which integration webhook — only the terdut-server side (creating the integration and handing back its URL/key).
  • Not OLM packaging. Plain Kubebuilder manifests + Helm chart for install, matching how terdut-server itself ships.
  • Cross-namespace references are limited to exactly one edge: TerdutTeam.spec.serverRef may name a TerdutServer in a different namespace, gated by an explicit consent object in the target namespace (§4.2, §4.6) — this is the multi-tenant shape the operator exists for (one platform team owns a TerdutServer; other teams self-service a TerdutTeam against it without needing write access to the server's namespace). Every other reference (teamRef on the escalation rule/dead-man-switch/alert-source CRDs) stays same-namespace-as-its-TerdutTeam only, in v1 — those manage a specific team's own resources and are expected to live alongside it.
  • No validating/mutating admission webhooks in v1. CEL validation rules on the CRDs (OpenAPI x-kubernetes-validations) cover what they can; anything that needs a live look at another object (e.g. "does this teamRef exist") is a status condition, not an admission rejection — keeps v1 to a controller-only deployment with no cert-manager/webhook dependency.

2. The README's open questions, resolved

How do team-crd connect with server-crd?

Explicit spec.serverRef: {name, namespace} on TerdutTeam — same reasoning as before (explicit, greppable, trivially validated), but namespace is deliberately part of the reference: one team can run and own a TerdutServer, and other teams — in their own namespaces, without any write access to the server-owning team's namespace — self-service a TerdutTeam against it. namespace defaults to the TerdutTeam's own namespace when omitted, so the common single-tenant case (serverRef: {name: terdut}) is unchanged.

This cross-namespace edge needs the target namespace's explicit consent — otherwise any namespace in the cluster could point a TerdutTeam at someone else's TerdutServer and have the operator provision a team + credentials against it, which is a namespace-boundary violation, not a gitops convenience. §4.6 covers the consent object (TerdutServerReferenceGrant), modeled directly on Gateway API's ReferenceGrant — the established Kubernetes pattern for exactly this "object A in namespace X wants to reference object B in namespace Y" problem. No other reference in this design (teamRef on the child CRDs) crosses a namespace boundary, so this is the only place ReferenceGrant-style consent is needed (§1, §4.2, §5, §6, §9).

How does escalationrules, switches and alertsources connect to a team?

Explicit spec.teamRef: {name} on each of TerdutEscalationRule, TerdutDeadmanSwitch, TerdutAlertSource — same reasoning, and it mirrors terdut-server's own data model, where every one of these rows carries a team_id foreign key already. A matcher/selector on TerdutTeam would be inventing a second source of truth for an association the server already models as a plain reference.

Support for both postgres-operator (Zalando) and bring-your-own, how do we design that to be user friendly?

TerdutServer.spec.database is a oneOf, mirroring the chart's existing database.dsn / database.passwordSecret contract (see §8):

  • dsn + passwordSecretRef — bring-your-own, exactly today's chart inputs.
  • postgresClusterRef — points at a Zalando postgresql.acid.zalan.do CR; the operator derives the DSN and resolves the generated credentials Secret itself (see §8). CloudNativePG support is a natural follow-up using the same shape and is called out as deferred (§13), not designed in detail now.

3. API group, versions, CRD catalog

  • Group: terdut.ryuvia.com, version: v1alpha1 (matches the ryuvia.com domain terdut-server already uses; bump to v1beta1/v1 per the normal Kubernetes API graduation criteria once the shapes below have proven stable against a real install).
  • Module: git.ryuvia.com/niklas/terdut-operator, scaffolded with Kubebuilder (controller-runtime), matching terdut-server's Go toolchain and house style.
Kind Scope Purpose
TerdutServer Namespaced One terdut-server install: Deployment, Service, database wiring, bootstrap, operator credentials.
TerdutTeam Namespaced One team on a TerdutServer, possibly in another namespace: name, OIDC group mapping.
TerdutServerReferenceGrant Namespaced (lives with the TerdutServer) Consent for a TerdutTeam in another namespace to reference this namespace's TerdutServer(s).
TerdutEscalationRule Namespaced A team's escalation policy (levels, targets, repeat).
TerdutDeadmanSwitch Namespaced One dead man's switch on a team.
TerdutAlertSource Namespaced One alert-ingest integration on a team (currently: Alertmanager webhook).

4. Per-CRD spec

4.1 TerdutServer

apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
  name: terdut
  namespace: oncall
spec:
  image:
    repository: git.ryuvia.com/niklas/terdut-server
    tag: v0.9.3
  replicas: 1                      # terdut-server is not horizontally-scale-tested; keep the field, default 1
  networking:
    hostname: terdut.example.com
    servicePort: 8080
    gatewayListener: ""            # same semantics as chart's networking.listener
  database:
    dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require"   # mutually exclusive with postgresClusterRef
    passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password}
    # --- OR ---
    postgresClusterRef: {name: terdut-postgres}   # Zalando `postgresql` CR in the same namespace
  sweeper:
    staleAfter: 6h
    archiveAfter: 168h
  deadman:
    matchers: "alertname=Watchdog"
    timeout: 15m
    severity: critical
  notify:
    ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
    fallbackTopic: ""
    repeatEvery: 15m
    tokenSecretRef: {name: "", key: token}
  oidc:
    enabled: false
    issuer: ""
    clientID: ""
    clientSecretRef: {name: "", key: client-secret}
    name: SSO
    scopes: "openid profile email"
    allowedGroups: []
    adminGroup: ""
    sessionMaxAge: 12h
  passwordLogin: true
status:
  conditions: [...]                 # Ready, DatabaseReady, Bootstrapped
  observedGeneration: 3
  serviceName: terdut
  operatorCredentialsSecretRef: {name: terdut-operator-credentials}  # see §6

Field-for-field this is the chart's values.yaml reshaped as a spec — the operator absorbs the chart's Deployment/Service/bootstrap-job templates, so existing installs have a direct mapping when migrating (see §10).

spec.database fields are +kubebuilder:validation:XValidation guarded to be mutually exclusive (dsn xor postgresClusterRef); mirrors the chart's "chart provisions no database" stance — this operator provisions no database either, only wires up one that exists.

4.2 TerdutTeam

spec:
  serverRef:
    name: terdut
    namespace: platform-oncall      # optional; defaults to this TerdutTeam's own namespace.
                                     # Cross-namespace requires a matching TerdutServerReferenceGrant
                                     # in that namespace (§4.6) — otherwise Ready: False, reason: RefNotPermitted.
  displayName: "Platform"          # -> POST /api/teams {"name": ...}; server assigns the ID
  oidc:
    memberGroup: "terdut-platform-members"
    ownerGroup: "terdut-platform-owners"
status:
  conditions: [...]
  teamID: 42                       # the server-side ID; needed by every child object's controller
  observedGeneration: 1

Team membership (which users belong, team_members) is explicitly not modeled as a CRD field in v1: terdut-server already manages membership via OIDC group sync at login for SSO installs, and manual membership for password-login installs is a people-management action, not infrastructure — forcing it through gitops would mean a human's team change goes through a PR review. Flagged in §13 as revisitable if a real gitops-membership need shows up.

4.3 TerdutEscalationRule

spec:
  teamRef: {name: platform-team}
  repeatCount: 2
  fallbackTopic: "platform-oncall"
  levels:
    - timeout: 5m
      targets:
        - kind: oncall            # "oncall" or "user"
        - kind: user
          username: alice         # resolved to a user ID by the controller at apply time
    - timeout: 15m
      targets:
        - kind: oncall
status:
  conditions: [...]
  observedGeneration: 1

One TerdutEscalationRule per team — the server itself models a policy as one row (escalation_policies) with an owned list of levels, so a one-CRD-to-one-policy mapping (not one-CRD-per-level) matches the server's own aggregate and lets the whole thing be reconciled with the single PUT /api/teams/{teamID}/escalation the API actually exposes (see §5).

4.4 TerdutDeadmanSwitch

spec:
  teamRef: {name: platform-team}
  name: "prod-watchdog"
  matcher: "alertname=Watchdog,cluster=prod"
  timeout: 15m
  severity: critical
status:
  conditions: [...]
  switchID: 7

4.5 TerdutAlertSource

spec:
  teamRef: {name: platform-team}
  kind: alertmanager
  name: "prod-alertmanager"
status:
  conditions: [...]
  integrationID: 3
  webhookURLSecretRef: {name: prod-alertmanager-terdut-webhook}   # url + key, never in status/spec

The integration key is shown by the API exactly once, at creation (Integration.Key/URL in terdut-server's own model) — never re-readable, same shape as the bootstrap admin key. The controller writes it straight into a generated, owner-referenced Secret on the create it caused and never logs or stores it anywhere else; the CR's status carries only the Secret reference, matching how e.g. cert-manager's Certificate exposes spec.secretName rather than the key material itself.

4.6 TerdutServerReferenceGrant

Consent object living in the TerdutServer's namespace, modeled directly on Gateway API's ReferenceGrant (gateway.networking.k8s.io/v1beta1): without one, no TerdutTeam outside this namespace may resolve a serverRef into it, no matter what it names.

apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServerReferenceGrant
metadata:
  name: allow-app-teams
  namespace: platform-oncall        # the TerdutServer's namespace
spec:
  from:
    - group: terdut.ryuvia.com
      kind: TerdutTeam
      namespace: team-checkout      # one entry per namespace permitted to reference in
    - group: terdut.ryuvia.com
      kind: TerdutTeam
      namespace: team-payments
  to:
    - group: terdut.ryuvia.com
      kind: TerdutServer
      name: terdut                  # optional: omit to permit any TerdutServer in this namespace
  • No selector/wildcard-across-namespaces field (e.g. no namespaceSelector) in v1, matching upstream ReferenceGrant's own choice: the namespace owner enumerates exactly which namespaces may reference in, which is the whole point of requiring affirmative, auditable consent rather than an implicit or pattern-matched grant.
  • A TerdutTeam's controller checks for a matching grant (namespace + optionally name) on every reconcile before it will resolve serverRef cross-namespace, exactly mirroring how HTTPRoute/Certificate controllers check ReferenceGrant before honouring a cross-namespace backendRef/secretRef. Same-namespace serverRef never needs a grant.
  • Deleting the grant is a live revocation: the next reconcile of any TerdutTeam it used to authorize finds no matching grant, flips Ready: False, reason: RefNotPermitted, and — deliberately — does not delete the team server-side on revocation alone; it stops reconciling further changes until access is restored or the TerdutTeam CR itself is deleted (whose finalizer still needs the credentials Secret described in §6 to clean up, so a revoked grant blocking new changes rather than forcing an immediate, possibly credential-less deletion is the safer failure mode).

5. Reconciliation semantics

terdut-server's REST surface (internal/api/router.go) does not give every resource a full update verb, so reconciliation strategy is per-resource:

Resource Verbs available Strategy
Team POST create, PUT rename, DELETE, PUT oidc-groups Real update-in-place: diff spec vs. last-applied, PUT the changed pieces.
Escalation policy GET/PUT whole-policy Update-in-place: PUT the full desired policy every reconcile that finds drift; cheap because whole-policy is small and already loaded whole server-side.
Dead man's switch POST create, DELETE — no PUT Delete-and-recreate on any spec diff other than name. The controller diffs against status (which mirrors what was last successfully applied) rather than re-reading the server every reconcile, to avoid a spurious recreate from field reordering.
Integration (alert source) POST create, PATCH rename, DELETE Rename via PATCH; any other spec change (kind) is delete-and-recreate, which rotates the webhook key — called out loudly in the CRD's field docs and in a Warning event, since it breaks whatever sends to the old URL/key until the new Secret is picked up.

General rules for every controller:

  • Idempotent create: before POSTing, check status.<serverSideID> is unset; if the server already has a same-named object from a previous partial reconcile (e.g. after a crash between POST and status-write), treat a 409/name-conflict as "adopt" — GET-by-name and populate status, rather than erroring forever. terdut-server's list endpoints in each of these areas return objects by name, so this is a straightforward correlation.
  • Periodic resync in addition to watch-triggered reconciles (Kubebuilder default RequeueAfter on success, e.g. every 5–10 minutes) to catch drift from someone changing state directly against the server's API/UI, since gitops correctness means the CR wins, not "first write wins".
  • Finalizers on every CRD that has a server-side counterpart, so deletion calls the corresponding DELETE before the Kubernetes object disappears. Failure to delete server-side (e.g. server unreachable) blocks finalizer removal and surfaces as a Degraded condition + event, rather than silently orphaning a row.
  • Owner chain for status resolution, not API calls: TerdutTeam's controller does not call any other controller; every child CRD's controller independently resolves its own teamRef → TerdutTeam.status.teamID and serverRef chain down to TerdutServer.status.operatorCredentialsSecretRef, the way any two independent controller-runtime reconcilers would. If the referenced parent isn't Ready yet, the child requeues with backoff and reports Ready: False, reason: WaitingForTeam — no cross-controller RPC.
  • Cross-namespace serverRef is re-checked every reconcile, not just at creation: TerdutTeam's controller lists TerdutServerReferenceGrant objects in the target namespace on every pass (cheap: a namespaced List with a field/label index, not a full-cluster scan) before touching the cross-namespace TerdutServer — revocation (§4.6) takes effect on the team's very next reconcile, not just when the CR is first applied.

6. Bootstrap & authentication to terdut-server's API

terdut-server has no first-class "service account" token type — API keys belong to a real user row (api_keys.user_id). The chart's current answer is a one-shot Job that calls POST /api/bootstrap, gets a one-time admin key back, and stores it in a Secret (bootstrap-job.yaml). The operator absorbs this rather than shelling out to curl:

  1. On a TerdutServer's first reconcile after its Deployment reports Ready (/healthz reachable through the Service), the controller calls POST /api/bootstrap itself with a fixed, recognizable identity — e.g. username terdut-operator, email terdut-operator@<serverRef> — distinct from spec.bootstrap.username/email used for the human first admin the chart bootstraps today. Two separate bootstrap identities are needed only if terdut-server's /api/bootstrap is single-shot per install; if it already returns 403 after the first caller regardless of identity, the operator instead creates its bot user via POST /api/users once a human admin exists — this needs confirming against the endpoint's actual behavior before implementation and is flagged as a first task, not assumed here.
  2. The returned API key is written to a generated Secret (<name>-operator-credentials), owner-referenced to the TerdutServer, referenced back from status.operatorCredentialsSecretRef.
  3. Every other same-namespace controller (EscalationRule, DeadmanSwitch, AlertSource, and any same-namespace TerdutTeam) reads that Secret directly to call the API — never its own credentials.
  4. Cross-namespace TerdutTeam never reads the source Secret directly. Granting arbitrary consuming namespaces get/list/watch on a Secret in the server's namespace would mean any future workload in that namespace with Secret-read RBAC could be pointed at it too — the TerdutServerReferenceGrant (§4.6) only authorizes the TerdutTeam kind, not "read this Secret". Instead, once a TerdutServer observes at least one valid grant for a namespace, its own controller (which already holds the real credentials, and whose ServiceAccount is the only thing with legitimate cross-namespace write access — see §9) mirrors a copy of the Secret into that consenting namespace, named <serverRef.name>.<serverRef.namespace>-terdut-credentials, owned not by an OwnerReference (those can't cross namespaces) but tracked in the TerdutServer's status and cleaned up explicitly when the grant authorizing that namespace is removed. The remote TerdutTeam's controller reads only this local mirror, never the original.
    • This is one shared mirrored Secret per (server, consuming namespace) pair, not one per TerdutTeam — terdut-server's own API key isn't team-scoped (§13), so there is nothing finer to hand out; "scoped" here means scoped by namespace boundary, not by team permission.
  5. Rotation: the key is a bearer credential with no expiry modeled server-side today. Rotation is manual (delete the Secret + the api_keys row via DELETE /api/users/{id}/api-keys/{keyID}, let the controller re-bootstrap) until/unless terdut-server grows key expiry; a rotation invalidates every mirror too, which the TerdutServer controller re-copies on its own next reconcile. Documented as an operational runbook note, not automated in v1.
  6. This is explicitly a stand-in for a real scoped service-account token type; see §13.

7. Ownership, status, garbage collection

  • Every generated object (Deployment, Service, credentials Secret, webhook Secret) carries a metav1.OwnerReference to the CR that caused it, in the same namespace — standard GC, no finalizer needed for these (only for the server-side REST resources, per §5).
  • Status conditions follow the standard metav1.Condition shape with at least Ready on every kind, plus kind-specific ones (TerdutServer: DatabaseReady, Bootstrapped; children: Synced).
  • status.observedGeneration on every kind, bumped only after a successful reconcile against that generation's spec — the standard way a client (or kubectl wait) tells "applied" from "seen".
  • No cluster-scoped aggregation object (e.g. no cluster-wide "all servers" status) in v1 — kubectl get terdutservers -A is the aggregate view.

8. Postgres integration

Mirrors the chart's existing two-path contract (values.yaml database.dsn/passwordSecret), because that contract is already documented and tested operationally:

  • Bring-your-own: spec.database.dsn (no password) + spec.database.passwordSecretRef — the operator sets PGPASSWORD on the Deployment's container env from that Secret, exactly like the chart does today, and does nothing else. No connectivity check beyond what the Deployment's own readiness probe already gives.
  • Zalando postgres-operator: spec.database.postgresClusterRef names a postgresql.acid.zalan.do CR in the same namespace. The controller:
    • Reads that CR's status for the primary Service name/port to build the DSN host — <cluster>.<namespace>.svc:5432 — and database name convention.
    • Resolves the generated credentials Secret (<user>.<cluster>.credentials.postgresql.acid.zalan.do) the same way the chart's comment already documents, and wires it in as PGPASSWORD the same way.
    • Watches that Secret (not just the postgresql CR) so a credential rotation triggers a requeue — the chart today requires a manual pod restart for this; the operator can at least detect and report it via a condition even if restarting on rotation is left as a §13 follow-up rather than done automatically (a rolling restart on credential change is a behavior change worth its own design pass, not folded in here).
    • Requires read RBAC on postgresql.acid.zalan.do (optional CRD — the operator's ClusterRole/Role should not hard-fail if the CRD isn't installed and a given TerdutServer uses BYO DSN instead).

9. RBAC

  • The operator's own ServiceAccount needs, per namespace it's granted: get/list/watch/create/update/patch/delete on Deployments, Services, Secrets it owns, and get/list/watch on postgresql.acid.zalan.do (optional, degrade gracefully if absent per §8), plus get/list/watch on TerdutServerReferenceGrant (§4.6).
  • No cluster-scoped resources are created by this operator (namespaced CRDs only, per §1) — a Role + RoleBinding per watched namespace is sufficient; a ClusterRole is only needed for watching CRDs across all namespaces, which is the normal Kubebuilder multi-tenant-operator default and doesn't imply cluster-scoped managed resources.
  • The operator's ServiceAccount is the only thing that ever holds cross-namespace Secret write. It is a single Deployment/binary already watching every namespace it's granted (the normal Kubebuilder shape), so mirroring a credentials Secret into a consenting namespace (§6) is not a new privilege boundary — it's the same ServiceAccount that already reconciles objects there — but it is new scope (create/update on Secrets cluster-wide rather than only within each object's own namespace), and should be called out explicitly in the operator's ClusterRole/RBAC review, not left implicit. No human or team's own RBAC is ever granted cross-namespace Secret access by this design — TerdutServerReferenceGrant only ever authorizes the operator to act, never a person.
  • terdut-server's own RBAC is unaffected — the operator talks to it purely over HTTP with the bot user's API key, never via the Kubernetes API for app-level state.

10. Relationship to charts/terdut-server

Recommendation (flagged explicitly as a decision to confirm before implementation starts, not settled by this document alone): the chart's Deployment/Service/bootstrap-job templates become redundant once TerdutServer exists — running both would mean two controllers (Helm and this operator) reconciling the same Deployment, which is exactly the conflict Kubernetes operators exist to avoid. Proposed path:

  • The chart is repurposed into an installer chart: it installs the operator + CRDs (and optionally one TerdutServer CR from values.yaml, for users who want "helm install and get a server" without hand-writing a CR) rather than templating the Deployment directly.
  • Existing installs migrate by: helm template the current release's values.yaml into an equivalent TerdutServer CR (mechanical, since §4.1 is deliberately shaped to make that mapping 1:1), install the operator, apply the CR, then let Helm's release be uninstalled or reduced to just the CRD/operator subchart.
  • This is a breaking change to the chart's contract and needs its own migration guide and probably a major chart version bump — out of scope for this design doc beyond flagging it; do not start that migration work without separately confirming this recommendation.

11. Testing strategy

  • envtest (controller-runtime's fake API server) for every controller's reconcile logic against the Kubernetes side.
  • terdut-server's REST API is faked with a small httptest.Server per controller test driven by fixtures matching the real handlers' request/response shapes (already well-documented in internal/api/*_test.go on the server side) — no real Postgres or real terdut-server binary needed for controller unit tests.
  • A smaller number of true end-to-end tests (kind cluster + real terdut-server image + real Postgres) covering the golden path per CRD: create TerdutServer → TerdutTeam → one of each child kind → verify via terdut-server's own API that the objects exist with the right shape → delete the CR → verify the server-side object is gone.

12. Observability

  • Standard controller-runtime metrics (reconcile duration/error counts) are enough for v1 — no custom metrics.
  • Every externally-visible action (bootstrap, key rotation-needed, delete- and-recreate on the no-PUT resources, adopt-on-conflict) emits a Kubernetes Event on the CR, since that's what shows up in kubectl describe and gitops tooling (Argo CD/Flux) surfaces without extra wiring.

13. Deferred / explicitly out of scope for this design

  • A real scoped service-account/token type in terdut-server (least-privilege operator identity instead of a bot admin user) — worth raising as a terdut-server feature request, not designed here.
  • CloudNativePG support — same spec.database shape as Zalando should extend to it, but the concrete field/Secret-naming conventions need their own look.
  • Cross-namespace teamRef on the child CRDs (TerdutEscalationRule, TerdutDeadmanSwitch, TerdutAlertSource) — only TerdutTeam.serverRef crosses namespaces in v1 (§2, §4.2, §4.6); these stay same-namespace as their TerdutTeam until a real need for splitting them out shows up.
  • A TerdutServerReferenceGrant-style namespace selector (rather than an enumerated list of namespaces) — deliberately omitted to match upstream ReferenceGrant's own choice of explicit, auditable consent over pattern-matching (§4.6); revisit only if enumerating namespaces becomes genuinely unworkable at scale.
  • Gitops-managed team membership (see §4.2).
  • Automatic Deployment restart on upstream Postgres credential rotation.
  • Admission webhooks / CEL-only validation limits (e.g. verifying a teamRef exists at admission time rather than surfacing it as a status condition after the fact).
  • OLM packaging, Helm chart migration execution (§10 is a recommendation, not a plan to execute).