Files
terdut-operator/DESIGN.md
T
Niklas Ye e1103f2b7d Authenticate with a seeded operator key; fold escalation and switches into TerdutTeam
Credentials: the TerdutServer controller generates <name>-operator-key in
the server's own namespace (owned by it) and hands it to the pods as
TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it
at every start. A replaced Secret rolls the pods. The bootstrap handshake,
the checkpoint Secret, per-team service accounts and credentials Secrets,
BootstrapStateLost and credentials.deletionPolicy are gone.

CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule
and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[]
on the team (matched by name, extras removed); team invites are removed.
A team is created under the identity <namespace>/<name> (external_id), so a
retry, a lost status or a deleted team heal by repeating the same call, and
a display name owned by another team is TeamNameTaken instead of an
adoption. The server resolves escalation usernames (UnknownUser condition).
OIDC claim names and trustEmail are spec fields.

Fixes: query values are URL-escaped; every delete treats 404 as success;
deleting a team no longer depends on allowedTeams consent; a switch or
integration deleted on the server is recreated; unnamed switches take the
CR's name.

Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces
and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo
(run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays
cluster-wide, now stated in DESIGN.md section 9.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:22 +02:00

23 KiB

terdut-operator design

The design of terdut-operator as built. terdut-server is the system of record; this operator makes one install of it — the server, its teams, their escalation ladders, dead man's switches and alert-source integrations — describable as Kubernetes objects and manageable through gitops.

1. Goals & non-goals

Goal: a terdut-server install fully described as Kubernetes objects, following controller-runtime / Kubebuilder conventions.

The operator creates and owns every TerdutServer it manages. It never adopts a pre-existing, independently-deployed terdut-server, whether deployed by hand or by charts/terdut-server. There is no migration path from a chart-based install (§10): starting with the operator means applying a fresh TerdutServer.

Non-goals (v1):

  • Not a Postgres operator. It consumes a database that the Zalando postgres-operator or something else provides (§8).
  • Not managing Alertmanager or its routing, only the terdut-server side (creating the integration and handing back its URL/key).
  • Not OLM packaging: plain Kubebuilder manifests and a Helm chart, like terdut-server.
  • Cross-namespace references are limited to one edge: TerdutTeam.spec.serverRef may name a TerdutServer in another namespace, gated by that server's spec.allowedTeams (§4.6). A TerdutAlertSource lives beside its TerdutTeam.
  • No admission webhooks. CEL validation covers what it can; anything needing a live look at another object is a status condition, not an admission rejection.
  • Never manages external exposure for a TerdutServer (Ingress, HTTPRoute, VirtualService) — a permanent non-goal. The operator creates a plain ClusterIP Service; see examples/networking.

2. Credentials and CRD shape (the 2026-10 redesign)

What the design is, and why — each point replaced something heavier:

  • No bootstrap handshake. The operator generates a key into a Secret named <TerdutServer>-operator-key, in the TerdutServer's own namespace and owned by it (a pod can only mount Secrets of its own namespace, and an owner reference replaces the old finalizer and credentials.deletionPolicy). The Deployment hands it to terdut-server as TERDUT_OPERATOR_KEY; the server creates or re-keys its instance-scoped service account terdut-operator from it at every start. A replaced Secret rolls the pods (key-hash annotation). /api/bootstrap stays free for the first human administrator. Gone with it: the checkpoint Secret, BootstrapStateLost, DatabaseReady/Bootstrapped conditions.
  • One credential per server. An instance-scoped account acts as owner of every team's configuration (not a member, so it reads no incidents). The per-team service accounts, Secrets and status.credentialsSecretRef/serverEndpoint on TerdutTeam are gone.
  • Team identity is external_id. The operator creates a team with external_id: <namespace>/<name> of its CR; the server returns the existing team for a known id (200) instead of creating one, so crash recovery, a lost status and a deleted team all heal by repeating the same call, and a display name that belongs to another team is a 409 (TeamNameTaken) instead of an adoption.
  • Three CRDs. TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule and TerdutDeadmanSwitch are spec.escalation and spec.deadmanSwitches[] (matched by name, unique per team server-side; ones not listed are removed) on the team: they were one-to-one children with the team's lifecycle, and folding them removes teamRef, the two-rules-clobber footgun and three controllers. TerdutAlertSource stays separate because it owns a webhook Secret in its own namespace.
  • Usernames are resolved by the server (PUT .../escalation accepts username), so an unknown user is a UnknownUser condition, not a list-and-match.
  • Invites are removed. Membership is not modelled; people get in through the server's own signup/OIDC.
  • Env mirrors the chart's where it matters: OIDC claim names and trustEmail are spec fields; TERDUT_DEADMAN_* no longer exist on the server.

3. CRD catalog

Group terdut.ryuvia.com, version v1alpha1; module git.ryuvia.com/niklas/terdut-operator.

Kind Purpose
TerdutServer One terdut-server install: Deployment, Service, database wiring, operator key, cross-namespace team consent.
TerdutTeam One team on a server, possibly in another namespace: name, OIDC groups, escalation ladder, dead man's switches.
TerdutAlertSource One alert-ingest integration on a team; owns a webhook Secret.

4. Per-CRD spec

(The escalation and dead man's switch specs that used to be §4.3 and §4.4 are part of §4.2; the section numbers are kept because code comments cite them.)

4.1 TerdutServer

apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
  name: terdut
  namespace: oncall
spec:
  image:
    repository: git.ryuvia.com/niklas/terdut-server
    tag: v0.9.3
  replicas: 2                      # default since terdut-server v0.36.0's advisory locks; see TerdutServerSpec.Replicas
  networking:
    hostname: terdut.example.com
    servicePort: 8080
  database:
    dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require"   # mutually exclusive with postgresClusterRef
    passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password}
    # --- OR ---
    postgresClusterRef: {name: terdut-postgres}   # Zalando `postgresql` CR in the same namespace
  sweeper:
    staleAfter: 6h
    archiveAfter: 168h
  notify:
    ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
    fallbackTopic: ""
    repeatEvery: 15m
    tokenSecretRef: {name: "", key: token}
  oidc:
    enabled: false
    issuer: ""
    clientID: ""
    clientSecretRef: {name: "", key: client-secret}
    name: SSO
    scopes: "openid profile email"
    allowedGroups: []
    adminGroup: ""
    usernameClaim: preferred_username   # also emailClaim, groupsClaim
    trustEmail: false
    sessionMaxAge: 12h
  passwordLogin: true
  # Consent for TerdutTeams in OTHER namespaces to set serverRef at this
  # TerdutServer. Same-namespace TerdutTeams never need this. Modeled on
  # Gateway API's Gateway.spec.allowedListeners.namespaces (the ListenerSet
  # attachment pattern, not ReferenceGrant — see §2 for why).
  allowedTeams:
    namespaces:
      from: None        # None (default) | Same | All | Selector
      # selector:       # required, and only meaningful, when from: Selector
      #   matchLabels:
      #     terdut.ryuvia.com/allowed: "true"
  # Pod-level customization of the Deployment -- all optional, direct corev1
  # passthrough throughout (see PodSpec's own doc comment). A representative
  # subset:
  pod:
    resources:
      requests: {cpu: 100m, memory: 128Mi}
      limits: {memory: 256Mi}
    tolerations:
      - key: dedicated
        operator: Equal
        value: terdut
        effect: NoSchedule
    disruptionBudget:
      minAvailable: 1       # mutually exclusive with maxUnavailable
status:
  conditions: [...]                 # Ready
  observedGeneration: 3
  credentialsSecretRef: {name: terdut-operator-key, key: token}   # see §6; pure output, in this TerdutServer's own namespace

Field-for-field this is the chart's values.yaml reshaped as a spec — not for migrating an existing chart-based install (§1: there is no such path), just because the shape is already familiar from the chart, and the operator absorbs what the chart's Deployment/Service/bootstrap-job templates used to do.

spec.database fields are +kubebuilder:validation:XValidation guarded to be mutually exclusive (dsn xor postgresClusterRef); mirrors the chart's "chart provisions no database" stance — this operator provisions no database either, only wires up one that exists.

spec.allowedTeams.namespaces.from defaults to None, matching allowedListeners's own default — a fresh TerdutServer accepts no cross-namespace TerdutTeam until its owner opts in, same-namespace TerdutTeams are unaffected either way. Selector deliberately has no per-name allowlist (no "and only these teams") — namespace-level consent is the right granularity here, same reasoning as ListenerSet: the namespace is the tenancy boundary, not the object.

spec.pod is pod-level customization of the Deployment, all optional and directly reusing corev1 types wherever corev1 already models the knob exactly (tolerations, affinity, topologySpreadConstraints, resources, securityContext/containerSecurityContext, extraEnv/ extraEnvFrom, extraVolumes/extraVolumeMounts, imagePullSecrets) — no custom wrapper buys anything for any of these, matching how CloudNativePG and the Zalando postgres-operator both expose the same knobs. affinity is pure user-supplied passthrough, not a toggle-plus-generated-default the way a multi-replica-aware operator's pod anti-affinity typically is: even though replicas now defaults to 2 (terdut-server v0.36.0's advisory locks made that safe, §4.1's own illustrative YAML comment), this operator still never auto-generates affinity of its own. spec.pod.disruptionBudget is the one field here that isn't a straight PodTemplateSpec knob — when set, the controller reconciles a PodDisruptionBudget selecting this TerdutServer's pods; clearing it deletes any it previously created (§7). minAvailable/maxUnavailable are mutually exclusive, +kubebuilder:validation:XValidation-guarded the same way as spec.database's own dsn/postgresClusterRef rule.

4.2 TerdutTeam

spec:
  serverRef: {name: terdut, namespace: platform-oncall}   # namespace optional; cross-namespace needs allowedTeams
  displayName: Platform
  oidc: {memberGroup: terdut-platform-members, ownerGroup: terdut-platform-owners}
  escalation:
    repeatCount: 2
    fallbackTopic: platform-oncall
    levels:
      - timeout: 5m
        targets: [{kind: oncall}, {kind: user, username: alice}]
      - timeout: 15m
        targets: [{kind: oncall}]
  deadmanSwitches:
    - {name: watchdog, matcher: "alertname=Watchdog,cluster=prod", timeout: 15m, severity: critical}
status:
  conditions: [...]      # Ready; reasons: ServerRefNotFound, RefNotPermitted, WaitingForServer,
                         # TeamNameTaken, UnknownUser, InvalidSpec, Adopted
  teamID: 42
  observedGeneration: 1

The team is created under the identity <namespace>/<name> of its CR (external_id), so a retry, a lost status or a team deleted behind the operator's back all heal by repeating the same call; a displayName that belongs to a different team is TeamNameTaken, never an adoption. Every reconcile applies, in order: create-or-find, rename, OIDC groups, the escalation ladder (one PUT of the whole ladder; absent spec.escalation clears it), and the switches. Switches are matched by name (unique per team on the server): missing ones are created, changed ones updated in place, and any not listed are deleted, since in operator mode nobody else can add one. Durations are validated by a CRD pattern and re-checked (InvalidSpec). A username the server does not know is UnknownUser until that person exists.

Team membership is not modelled: the server manages it through OIDC group sync and its own UI.

4.5 TerdutAlertSource

spec:
  teamRef: {name: platform-team}
  kind: alertmanager
  name: "prod-alertmanager"
status:
  conditions: [...]
  integrationID: 3
  webhookURLSecretRef: {name: prod-alertmanager-terdut-webhook}   # url + key, never in status/spec

The integration key is shown by the API exactly once, at creation (Integration.Key/URL in terdut-server's own model) — never re-readable, a one-shot value. The controller writes it straight into a generated, owner-referenced Secret on the create it caused and never logs or stores it anywhere else; the CR's status carries only the Secret reference, matching how e.g. cert-manager's Certificate exposes spec.secretName rather than the key material itself.

No separate CRD — the consent lives on TerdutServer itself (§4.1), following Gateway API's Gateway.spec.allowedListeners (ListenerSet attachment) rather than its ReferenceGrant, since a TerdutTeam attaching to a shared TerdutServer is the same shape of relationship: a bounded set of expected children attaching to a parent explicitly designed to be shared, not an arbitrary cross-namespace backend reference (see §2 for the full comparison of both patterns).

  • from: None (default) — no cross-namespace TerdutTeam may resolve a serverRef into this TerdutServer. Same-namespace TerdutTeams are always allowed regardless of this field.
  • from: Same — equivalent to None in effect (same-namespace is already unrestricted); kept for parity with the upstream enum.
  • from: All — any namespace in the cluster may reference in. Appropriate for a genuinely shared, cluster-wide TerdutServer; the audit trail is "check allowedTeams plus who has RBAC to create a TerdutTeam anywhere," which is materially weaker than Selector.
  • from: Selector — only namespaces matching spec.allowedTeams.namespaces.selector (a standard metav1.LabelSelector over Namespace objects, exactly like allowedListeners's own selector) may reference in. This is the recommended mode for the platform-team-owns-a-shared-server scenario this design targets: label the consuming namespaces once (e.g. terdut.ryuvia.com/allowed-server: platform-oncall/terdut) and new namespaces opt in by carrying the label, without editing the TerdutServer again.
  • A TerdutTeam's controller re-evaluates allowedTeams on every reconcile before it will resolve a cross-namespace serverRef — for Selector, this means a Get on its own Namespace object plus reading the target TerdutServer's spec, not a List across the cluster. Same-namespace serverRef never consults this field at all.
  • Narrowing or clearing allowedTeams (or unlabeling a namespace, under Selector) is a live revocation: the next reconcile of any TerdutTeam it used to authorize finds itself no longer permitted, flips Ready: False, reason: RefNotPermitted, and — deliberately — does not delete the team server-side on revocation alone; it stops reconciling further changes until access is restored. Deleting the TerdutTeam CR still deletes the team: consent gates what the operator starts acting on, not whether it may clean up after itself.

5. Reconciliation semantics

Resource Server verbs Strategy
Team POST (idempotent on external_id), PUT rename, PUT oidc-groups, DELETE Repeat the idempotent POST, then PUT the rest, every reconcile. Delete runs in a finalizer; the server refuses (409) while the team has open incidents, which is retried.
Escalation ladder PUT whole ladder PUT the full desired ladder every reconcile; the server resolves usernames.
Dead man's switch list, POST, PUT, DELETE; names unique per team Diff by name against the list.
Integration (alert source) POST, PATCH rename, DELETE Rename via PATCH; a kind change is delete-and-recreate, which rotates the webhook key (Warning event). An integration deleted on the server is recreated with a new key (Warning event).

General rules:

  • Periodic resync (5 minutes) besides watch-triggered reconciles, to catch someone changing state directly against the server: the CR wins.
  • Finalizers on TerdutTeam and TerdutAlertSource call the server's DELETE first. A 404 is success; any other failure blocks removal and surfaces as an event rather than orphaning a row. TerdutServer needs none: its Secret is owned and its database is never touched.
  • A webhook Secret that is lost fails closed. The key is never re-readable from the server, so the controller does not mint a replacement for a live URL with no spec change to explain it: Ready: False, reason: WebhookSecretLost; delete and recreate the TerdutAlertSource.
  • Children resolve through the team. A TerdutAlertSource finds its TerdutTeam (same namespace), requires it Ready, then reads the server named by the team's serverRef and that server's operator key. A TerdutTeam re-reconciles when its TerdutServer changes.
  • Cross-namespace consent is re-checked every reconcile (§4.6), so revocation takes effect on the team's next pass.

6. Authentication to terdut-server's API

The operator talks to the server over HTTP with one bearer key per TerdutServer:

  1. The TerdutServer controller creates the Secret <name>-operator-key (data key token, tdsa_ + 48 hex characters) in the server's own namespace, owned by the TerdutServer. It is created once and never overwritten while it exists: the running server was seeded with it.
  2. The Deployment hands it to the pods as TERDUT_OPERATOR_KEY through a secretKeyRef (the key never appears in the pod spec), plus an annotation holding a hash of it so a replaced Secret rolls the pods.
  3. At every start the server creates or re-keys its instance-scoped service account terdut-operator from that value. An instance-scoped account acts as owner of every team's configuration but is not a member of any team, so it reads no incidents and is never an administrator.
  4. status.credentialsSecretRef points at the Secret; TerdutTeam and TerdutAlertSource controllers read it from the server's namespace.

There is no bootstrap handshake: /api/bootstrap stays free for the first human administrator. Deleting the TerdutServer deletes the key with it; a recreated one gets a new key and the server re-seeds on its next start. Rotating by hand means deleting the Secret: the next reconcile makes a new one and rolls the pods.

7. Ownership, status, garbage collection

  • Everything the controllers generate in a CR's namespace carries an OwnerReference (Deployment, Service, PodDisruptionBudget, the operator key Secret, the webhook Secret), so GC cleans up and no finalizer is needed. The PDB exists only while spec.pod.disruptionBudget is set; the controller deletes it itself when the field is cleared.
  • Every kind has a Ready condition and status.observedGeneration.
  • No cluster-scoped aggregate object: kubectl get terdutservers -A is the overview.

8. Postgres integration

Mirrors the chart's existing two-path contract (values.yaml database.dsn/passwordSecret), because that contract is already documented and tested operationally:

  • Bring-your-own: spec.database.dsn (no password) + spec.database.passwordSecretRef — the operator sets PGPASSWORD on the Deployment's container env from that Secret, exactly like the chart does today, and does nothing else. No connectivity check beyond what the Deployment's own readiness probe already gives.
  • Zalando postgres-operator: spec.database.postgresClusterRef names a postgresql.acid.zalan.do CR in the same namespace. The controller:
    • Reads that CR's status for the primary Service name/port to build the DSN host — <cluster>.<namespace>.svc:5432 — and database name convention.
    • Resolves the generated credentials Secret (<user>.<cluster>.credentials.postgresql.acid.zalan.do) the same way the chart's comment already documents, and wires it in as PGPASSWORD the same way.
    • Does not watch that Secret: a rotated credential is noticed at the next 5-minute resync, and restarting pods on rotation is a §13 follow-up.
    • Requires read RBAC on postgresql.acid.zalan.do (optional CRD — the operator's ClusterRole/Role should not hard-fail if the CRD isn't installed and a given TerdutServer uses BYO DSN instead).

9. RBAC

  • The operator needs get/list/watch/create/update/patch/delete on Deployments, Services, PodDisruptionBudgets and Secrets it owns, get/list/watch on postgresql.acid.zalan.do (optional; degraded gracefully when the CRD is absent), and get/list/watch on Namespaces (labels only, for allowedTeams: {from: Selector}).
  • Secrets live in the TerdutServer's namespace (the operator key) and in each TerdutAlertSource's namespace (the webhook Secret), so the operator needs Secret access in every tenant namespace. Kubernetes RBAC has no owner-scoped grant finer than the namespace.
  • As built (2026-10): broader than the above. The shipped default is a ClusterRole with full verbs on Secrets in every namespace (the chart's rbac.namespaced: true gives a Role in the release namespace only, which cannot serve tenant namespaces), and the manager's cache is not restricted to watched namespaces. A per-namespace Role split (a Role and RoleBinding per watched namespace, with the cache restricted to them) is the intended end state, not implemented, and the decision (2026-10) is to stay cluster-wide for now: treat this operator as able to read every Secret in the cluster.
  • terdut-server's own RBAC is unaffected: the operator uses only its HTTP API (§6), never the Kubernetes API for app-level state.

10. Relationship to charts/terdut-server

The server chart's Deployment, Service and bootstrap Job are redundant once TerdutServer exists: running both would have two controllers reconciling the same Deployment. The operator's chart (charts/terdut-operator) installs the operator, CRDs and RBAC, and optionally one TerdutServer from values.yaml (terdutServer.enabled). There is no migration from a chart-based install and none is planned (§1); whatever the server chart deployed stays a separate install until someone deletes it.

The operator's Deployment builder mirrors the chart's env block for the knobs both expose (the chart's deployment.yaml and buildEnv in internal/controller/terdutserver_deployment.go); a new server setting is added in config.go, the chart, and the operator, in that order.

11. Testing strategy

  • envtest (controller-runtime's fake API server) for every controller's reconcile logic against the Kubernetes side.
  • terdut-server's REST API is faked with a small httptest.Server per controller test driven by fixtures matching the real handlers' request/response shapes (already well-documented in internal/api/*_test.go on the server side) — no real Postgres or real terdut-server binary needed for controller unit tests.
  • The golden path (kind cluster + real terdut-server image + real Postgres: create TerdutServer → TerdutTeam → TerdutAlertSource, verify through terdut-server's own API, delete, verify it is gone) is a manual pass via examples/demo/run-demo.sh, not a CI job.

12. Observability

  • Standard controller-runtime metrics (reconcile duration/error counts) are enough for v1 — no custom metrics.
  • Every externally-visible action (a team created or renamed, an integration recreated, a delete that failed) emits a Kubernetes Event on the CR, since that's what shows up in kubectl describe and gitops tooling (Argo CD/Flux) surfaces without extra wiring.

13. Deferred / out of scope

  • CloudNativePG support, alongside the Zalando postgresClusterRef (same spec.database shape).
  • Gitops-managed team membership (§4.2).
  • Per-namespace RBAC with a restricted cache (§9).
  • Automatic Deployment restart on upstream Postgres credential rotation.
  • spec.pod.priorityClassName, pod labels beyond annotations, and an HPA for TerdutServer.
  • Admission webhooks beyond CEL (for example, checking a teamRef exists at admission time).
  • A shared API types module or generated client for terdut-server, terdut-operator and terdut-tui.
  • OLM packaging.