Credentials: the TerdutServer controller generates <name>-operator-key in the server's own namespace (owned by it) and hands it to the pods as TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it at every start. A replaced Secret rolls the pods. The bootstrap handshake, the checkpoint Secret, per-team service accounts and credentials Secrets, BootstrapStateLost and credentials.deletionPolicy are gone. CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[] on the team (matched by name, extras removed); team invites are removed. A team is created under the identity <namespace>/<name> (external_id), so a retry, a lost status or a deleted team heal by repeating the same call, and a display name owned by another team is TeamNameTaken instead of an adoption. The server resolves escalation usernames (UnknownUser condition). OIDC claim names and trustEmail are spec fields. Fixes: query values are URL-escaped; every delete treats 404 as success; deleting a team no longer depends on allowedTeams consent; a switch or integration deleted on the server is recreated; unnamed switches take the CR's name. Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo (run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays cluster-wide, now stated in DESIGN.md section 9. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
23 KiB
terdut-operator design
The design of terdut-operator as built. terdut-server is the system of record; this operator makes one install of it — the server, its teams, their escalation ladders, dead man's switches and alert-source integrations — describable as Kubernetes objects and manageable through gitops.
1. Goals & non-goals
Goal: a terdut-server install fully described as Kubernetes objects, following controller-runtime / Kubebuilder conventions.
The operator creates and owns every TerdutServer it manages. It never adopts a pre-existing,
independently-deployed terdut-server, whether deployed by hand or by charts/terdut-server.
There is no migration path from a chart-based install (§10): starting with the operator means
applying a fresh TerdutServer.
Non-goals (v1):
- Not a Postgres operator. It consumes a database that the Zalando
postgres-operatoror something else provides (§8). - Not managing Alertmanager or its routing, only the terdut-server side (creating the integration and handing back its URL/key).
- Not OLM packaging: plain Kubebuilder manifests and a Helm chart, like terdut-server.
- Cross-namespace references are limited to one edge:
TerdutTeam.spec.serverRefmay name aTerdutServerin another namespace, gated by that server'sspec.allowedTeams(§4.6). ATerdutAlertSourcelives beside itsTerdutTeam. - No admission webhooks. CEL validation covers what it can; anything needing a live look at another object is a status condition, not an admission rejection.
- Never manages external exposure for a
TerdutServer(Ingress, HTTPRoute, VirtualService) — a permanent non-goal. The operator creates a plainClusterIPService; seeexamples/networking.
2. Credentials and CRD shape (the 2026-10 redesign)
What the design is, and why — each point replaced something heavier:
- No bootstrap handshake. The operator generates a key into a Secret named
<TerdutServer>-operator-key, in the TerdutServer's own namespace and owned by it (a pod can only mount Secrets of its own namespace, and an owner reference replaces the old finalizer andcredentials.deletionPolicy). The Deployment hands it to terdut-server asTERDUT_OPERATOR_KEY; the server creates or re-keys its instance-scoped service accountterdut-operatorfrom it at every start. A replaced Secret rolls the pods (key-hash annotation)./api/bootstrapstays free for the first human administrator. Gone with it: the checkpoint Secret,BootstrapStateLost,DatabaseReady/Bootstrappedconditions. - One credential per server. An instance-scoped account acts as owner of every
team's configuration (not a member, so it reads no incidents). The per-team
service accounts, Secrets and
status.credentialsSecretRef/serverEndpointonTerdutTeamare gone. - Team identity is
external_id. The operator creates a team withexternal_id: <namespace>/<name>of its CR; the server returns the existing team for a known id (200) instead of creating one, so crash recovery, a lost status and a deleted team all heal by repeating the same call, and a display name that belongs to another team is a 409 (TeamNameTaken) instead of an adoption. - Three CRDs.
TerdutServer,TerdutTeamandTerdutAlertSource.TerdutEscalationRuleandTerdutDeadmanSwitcharespec.escalationandspec.deadmanSwitches[](matched by name, unique per team server-side; ones not listed are removed) on the team: they were one-to-one children with the team's lifecycle, and folding them removesteamRef, the two-rules-clobber footgun and three controllers.TerdutAlertSourcestays separate because it owns a webhook Secret in its own namespace. - Usernames are resolved by the server (
PUT .../escalationacceptsusername), so an unknown user is aUnknownUsercondition, not a list-and-match. - Invites are removed. Membership is not modelled; people get in through the server's own signup/OIDC.
- Env mirrors the chart's where it matters: OIDC claim names and
trustEmailare spec fields;TERDUT_DEADMAN_*no longer exist on the server.
3. CRD catalog
Group terdut.ryuvia.com, version v1alpha1; module git.ryuvia.com/niklas/terdut-operator.
| Kind | Purpose |
|---|---|
TerdutServer |
One terdut-server install: Deployment, Service, database wiring, operator key, cross-namespace team consent. |
TerdutTeam |
One team on a server, possibly in another namespace: name, OIDC groups, escalation ladder, dead man's switches. |
TerdutAlertSource |
One alert-ingest integration on a team; owns a webhook Secret. |
4. Per-CRD spec
(The escalation and dead man's switch specs that used to be §4.3 and §4.4 are part of §4.2; the section numbers are kept because code comments cite them.)
4.1 TerdutServer
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
name: terdut
namespace: oncall
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: v0.9.3
replicas: 2 # default since terdut-server v0.36.0's advisory locks; see TerdutServerSpec.Replicas
networking:
hostname: terdut.example.com
servicePort: 8080
database:
dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef
passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password}
# --- OR ---
postgresClusterRef: {name: terdut-postgres} # Zalando `postgresql` CR in the same namespace
sweeper:
staleAfter: 6h
archiveAfter: 168h
notify:
ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
fallbackTopic: ""
repeatEvery: 15m
tokenSecretRef: {name: "", key: token}
oidc:
enabled: false
issuer: ""
clientID: ""
clientSecretRef: {name: "", key: client-secret}
name: SSO
scopes: "openid profile email"
allowedGroups: []
adminGroup: ""
usernameClaim: preferred_username # also emailClaim, groupsClaim
trustEmail: false
sessionMaxAge: 12h
passwordLogin: true
# Consent for TerdutTeams in OTHER namespaces to set serverRef at this
# TerdutServer. Same-namespace TerdutTeams never need this. Modeled on
# Gateway API's Gateway.spec.allowedListeners.namespaces (the ListenerSet
# attachment pattern, not ReferenceGrant — see §2 for why).
allowedTeams:
namespaces:
from: None # None (default) | Same | All | Selector
# selector: # required, and only meaningful, when from: Selector
# matchLabels:
# terdut.ryuvia.com/allowed: "true"
# Pod-level customization of the Deployment -- all optional, direct corev1
# passthrough throughout (see PodSpec's own doc comment). A representative
# subset:
pod:
resources:
requests: {cpu: 100m, memory: 128Mi}
limits: {memory: 256Mi}
tolerations:
- key: dedicated
operator: Equal
value: terdut
effect: NoSchedule
disruptionBudget:
minAvailable: 1 # mutually exclusive with maxUnavailable
status:
conditions: [...] # Ready
observedGeneration: 3
credentialsSecretRef: {name: terdut-operator-key, key: token} # see §6; pure output, in this TerdutServer's own namespace
Field-for-field this is the chart's values.yaml reshaped as a spec — not
for migrating an existing chart-based install (§1: there is no such path),
just because the shape is already familiar from the chart, and the operator
absorbs what the chart's Deployment/Service/bootstrap-job templates used to
do.
spec.database fields are +kubebuilder:validation:XValidation guarded to
be mutually exclusive (dsn xor postgresClusterRef); mirrors the chart's
"chart provisions no database" stance — this operator provisions no database
either, only wires up one that exists.
spec.allowedTeams.namespaces.from defaults to None, matching
allowedListeners's own default — a fresh TerdutServer accepts no
cross-namespace TerdutTeam until its owner opts in, same-namespace
TerdutTeams are unaffected either way. Selector deliberately has no
per-name allowlist (no "and only these teams") — namespace-level consent is
the right granularity here, same reasoning as ListenerSet: the namespace
is the tenancy boundary, not the object.
spec.pod is pod-level customization of the Deployment, all optional and
directly reusing corev1 types wherever corev1 already models the knob
exactly (tolerations, affinity, topologySpreadConstraints,
resources, securityContext/containerSecurityContext, extraEnv/
extraEnvFrom, extraVolumes/extraVolumeMounts, imagePullSecrets) —
no custom wrapper buys anything for any of these, matching how
CloudNativePG and the Zalando postgres-operator both expose the same
knobs. affinity is pure user-supplied passthrough, not a
toggle-plus-generated-default the way a multi-replica-aware operator's
pod anti-affinity typically is: even though replicas now defaults to 2
(terdut-server v0.36.0's advisory locks made that safe, §4.1's own
illustrative YAML comment), this operator still never auto-generates
affinity of its own. spec.pod.disruptionBudget is the one field here that isn't a
straight PodTemplateSpec knob — when set, the controller reconciles a
PodDisruptionBudget selecting this TerdutServer's pods; clearing it
deletes any it previously created (§7). minAvailable/maxUnavailable
are mutually exclusive, +kubebuilder:validation:XValidation-guarded the
same way as spec.database's own dsn/postgresClusterRef rule.
4.2 TerdutTeam
spec:
serverRef: {name: terdut, namespace: platform-oncall} # namespace optional; cross-namespace needs allowedTeams
displayName: Platform
oidc: {memberGroup: terdut-platform-members, ownerGroup: terdut-platform-owners}
escalation:
repeatCount: 2
fallbackTopic: platform-oncall
levels:
- timeout: 5m
targets: [{kind: oncall}, {kind: user, username: alice}]
- timeout: 15m
targets: [{kind: oncall}]
deadmanSwitches:
- {name: watchdog, matcher: "alertname=Watchdog,cluster=prod", timeout: 15m, severity: critical}
status:
conditions: [...] # Ready; reasons: ServerRefNotFound, RefNotPermitted, WaitingForServer,
# TeamNameTaken, UnknownUser, InvalidSpec, Adopted
teamID: 42
observedGeneration: 1
The team is created under the identity <namespace>/<name> of its CR (external_id), so a retry,
a lost status or a team deleted behind the operator's back all heal by repeating the same call; a
displayName that belongs to a different team is TeamNameTaken, never an adoption. Every
reconcile applies, in order: create-or-find, rename, OIDC groups, the escalation ladder (one PUT
of the whole ladder; absent spec.escalation clears it), and the switches. Switches are matched
by name (unique per team on the server): missing ones are created, changed ones updated in
place, and any not listed are deleted, since in operator mode nobody else can add one. Durations
are validated by a CRD pattern and re-checked (InvalidSpec). A username the server does not
know is UnknownUser until that person exists.
Team membership is not modelled: the server manages it through OIDC group sync and its own UI.
4.5 TerdutAlertSource
spec:
teamRef: {name: platform-team}
kind: alertmanager
name: "prod-alertmanager"
status:
conditions: [...]
integrationID: 3
webhookURLSecretRef: {name: prod-alertmanager-terdut-webhook} # url + key, never in status/spec
The integration key is shown by the API exactly once, at creation
(Integration.Key/URL in terdut-server's own model) — never re-readable,
a one-shot value. The controller writes it straight into
a generated, owner-referenced Secret on the create it caused and never logs
or stores it anywhere else; the CR's status carries only the Secret
reference, matching how e.g. cert-manager's Certificate exposes
spec.secretName rather than the key material itself.
4.6 Cross-namespace consent: TerdutServer.spec.allowedTeams
No separate CRD — the consent lives on TerdutServer itself (§4.1), following
Gateway API's Gateway.spec.allowedListeners (ListenerSet attachment)
rather than its ReferenceGrant, since a TerdutTeam attaching to a shared
TerdutServer is the same shape of relationship: a bounded set of expected
children attaching to a parent explicitly designed to be shared, not an
arbitrary cross-namespace backend reference (see §2 for the full comparison
of both patterns).
from: None(default) — no cross-namespaceTerdutTeammay resolve aserverRefinto thisTerdutServer. Same-namespaceTerdutTeams are always allowed regardless of this field.from: Same— equivalent toNonein effect (same-namespace is already unrestricted); kept for parity with the upstream enum.from: All— any namespace in the cluster may reference in. Appropriate for a genuinely shared, cluster-wideTerdutServer; the audit trail is "checkallowedTeamsplus who has RBAC to create aTerdutTeamanywhere," which is materially weaker thanSelector.from: Selector— only namespaces matchingspec.allowedTeams.namespaces.selector(a standardmetav1.LabelSelectoroverNamespaceobjects, exactly likeallowedListeners's ownselector) may reference in. This is the recommended mode for the platform-team-owns-a-shared-server scenario this design targets: label the consuming namespaces once (e.g.terdut.ryuvia.com/allowed-server: platform-oncall/terdut) and new namespaces opt in by carrying the label, without editing theTerdutServeragain.- A
TerdutTeam's controller re-evaluatesallowedTeamson every reconcile before it will resolve a cross-namespaceserverRef— forSelector, this means aGeton its ownNamespaceobject plus reading the targetTerdutServer's spec, not a List across the cluster. Same-namespaceserverRefnever consults this field at all. - Narrowing or clearing
allowedTeams(or unlabeling a namespace, underSelector) is a live revocation: the next reconcile of anyTerdutTeamit used to authorize finds itself no longer permitted, flipsReady: False, reason: RefNotPermitted, and — deliberately — does not delete the team server-side on revocation alone; it stops reconciling further changes until access is restored. Deleting theTerdutTeamCR still deletes the team: consent gates what the operator starts acting on, not whether it may clean up after itself.
5. Reconciliation semantics
| Resource | Server verbs | Strategy |
|---|---|---|
| Team | POST (idempotent on external_id), PUT rename, PUT oidc-groups, DELETE |
Repeat the idempotent POST, then PUT the rest, every reconcile. Delete runs in a finalizer; the server refuses (409) while the team has open incidents, which is retried. |
| Escalation ladder | PUT whole ladder | PUT the full desired ladder every reconcile; the server resolves usernames. |
| Dead man's switch | list, POST, PUT, DELETE; names unique per team | Diff by name against the list. |
| Integration (alert source) | POST, PATCH rename, DELETE | Rename via PATCH; a kind change is delete-and-recreate, which rotates the webhook key (Warning event). An integration deleted on the server is recreated with a new key (Warning event). |
General rules:
- Periodic resync (5 minutes) besides watch-triggered reconciles, to catch someone changing state directly against the server: the CR wins.
- Finalizers on
TerdutTeamandTerdutAlertSourcecall the server's DELETE first. A 404 is success; any other failure blocks removal and surfaces as an event rather than orphaning a row.TerdutServerneeds none: its Secret is owned and its database is never touched. - A webhook Secret that is lost fails closed. The key is never re-readable from the server, so
the controller does not mint a replacement for a live URL with no spec change to explain it:
Ready: False, reason: WebhookSecretLost; delete and recreate theTerdutAlertSource. - Children resolve through the team. A
TerdutAlertSourcefinds itsTerdutTeam(same namespace), requires it Ready, then reads the server named by the team'sserverRefand that server's operator key. ATerdutTeamre-reconciles when itsTerdutServerchanges. - Cross-namespace consent is re-checked every reconcile (§4.6), so revocation takes effect on the team's next pass.
6. Authentication to terdut-server's API
The operator talks to the server over HTTP with one bearer key per TerdutServer:
- The
TerdutServercontroller creates the Secret<name>-operator-key(data keytoken,tdsa_+ 48 hex characters) in the server's own namespace, owned by theTerdutServer. It is created once and never overwritten while it exists: the running server was seeded with it. - The Deployment hands it to the pods as
TERDUT_OPERATOR_KEYthrough asecretKeyRef(the key never appears in the pod spec), plus an annotation holding a hash of it so a replaced Secret rolls the pods. - At every start the server creates or re-keys its instance-scoped service account
terdut-operatorfrom that value. An instance-scoped account acts as owner of every team's configuration but is not a member of any team, so it reads no incidents and is never an administrator. status.credentialsSecretRefpoints at the Secret;TerdutTeamandTerdutAlertSourcecontrollers read it from the server's namespace.
There is no bootstrap handshake: /api/bootstrap stays free for the first human administrator.
Deleting the TerdutServer deletes the key with it; a recreated one gets a new key and the server
re-seeds on its next start. Rotating by hand means deleting the Secret: the next reconcile makes
a new one and rolls the pods.
7. Ownership, status, garbage collection
- Everything the controllers generate in a CR's namespace carries an
OwnerReference(Deployment, Service, PodDisruptionBudget, the operator key Secret, the webhook Secret), so GC cleans up and no finalizer is needed. The PDB exists only whilespec.pod.disruptionBudgetis set; the controller deletes it itself when the field is cleared. - Every kind has a
Readycondition andstatus.observedGeneration. - No cluster-scoped aggregate object:
kubectl get terdutservers -Ais the overview.
8. Postgres integration
Mirrors the chart's existing two-path contract (values.yaml
database.dsn/passwordSecret), because that contract is already
documented and tested operationally:
- Bring-your-own:
spec.database.dsn(no password) +spec.database.passwordSecretRef— the operator setsPGPASSWORDon the Deployment's container env from that Secret, exactly like the chart does today, and does nothing else. No connectivity check beyond what the Deployment's own readiness probe already gives. - Zalando
postgres-operator:spec.database.postgresClusterRefnames apostgresql.acid.zalan.doCR in the same namespace. The controller:- Reads that CR's status for the primary Service name/port to build the DSN
host —
<cluster>.<namespace>.svc:5432— and database name convention. - Resolves the generated credentials Secret
(
<user>.<cluster>.credentials.postgresql.acid.zalan.do) the same way the chart's comment already documents, and wires it in asPGPASSWORDthe same way. - Does not watch that Secret: a rotated credential is noticed at the next 5-minute resync, and restarting pods on rotation is a §13 follow-up.
- Requires read RBAC on
postgresql.acid.zalan.do(optional CRD — the operator's ClusterRole/Role should not hard-fail if the CRD isn't installed and a givenTerdutServeruses BYO DSN instead).
- Reads that CR's status for the primary Service name/port to build the DSN
host —
9. RBAC
- The operator needs
get/list/watch/create/update/patch/deleteonDeployments,Services,PodDisruptionBudgetsandSecretsit owns,get/list/watchonpostgresql.acid.zalan.do(optional; degraded gracefully when the CRD is absent), andget/list/watchonNamespaces(labels only, forallowedTeams: {from: Selector}). - Secrets live in the
TerdutServer's namespace (the operator key) and in eachTerdutAlertSource's namespace (the webhook Secret), so the operator needs Secret access in every tenant namespace. Kubernetes RBAC has no owner-scoped grant finer than the namespace. - As built (2026-10): broader than the above. The shipped default is a
ClusterRolewith full verbs onSecretsin every namespace (the chart'srbac.namespaced: truegives aRolein the release namespace only, which cannot serve tenant namespaces), and the manager's cache is not restricted to watched namespaces. A per-namespaceRolesplit (a Role and RoleBinding per watched namespace, with the cache restricted to them) is the intended end state, not implemented, and the decision (2026-10) is to stay cluster-wide for now: treat this operator as able to read every Secret in the cluster. - terdut-server's own RBAC is unaffected: the operator uses only its HTTP API (§6), never the Kubernetes API for app-level state.
10. Relationship to charts/terdut-server
The server chart's Deployment, Service and bootstrap Job are redundant once TerdutServer
exists: running both would have two controllers reconciling the same Deployment. The operator's
chart (charts/terdut-operator) installs the operator, CRDs and RBAC, and optionally one
TerdutServer from values.yaml (terdutServer.enabled). There is no migration from a
chart-based install and none is planned (§1); whatever the server chart deployed stays a separate
install until someone deletes it.
The operator's Deployment builder mirrors the chart's env block for the knobs both expose (the
chart's deployment.yaml and buildEnv in internal/controller/terdutserver_deployment.go);
a new server setting is added in config.go, the chart, and the operator, in that order.
11. Testing strategy
envtest(controller-runtime's fake API server) for every controller's reconcile logic against the Kubernetes side.- terdut-server's REST API is faked with a small
httptest.Serverper controller test driven by fixtures matching the real handlers' request/response shapes (already well-documented ininternal/api/*_test.goon the server side) — no real Postgres or real terdut-server binary needed for controller unit tests. - The golden path (
kindcluster + real terdut-server image + real Postgres: createTerdutServer→TerdutTeam→TerdutAlertSource, verify through terdut-server's own API, delete, verify it is gone) is a manual pass viaexamples/demo/run-demo.sh, not a CI job.
12. Observability
- Standard controller-runtime metrics (reconcile duration/error counts) are enough for v1 — no custom metrics.
- Every externally-visible action (a team created or renamed, an integration recreated, a delete
that failed) emits a
Kubernetes
Eventon the CR, since that's what shows up inkubectl describeand gitops tooling (Argo CD/Flux) surfaces without extra wiring.
13. Deferred / out of scope
- CloudNativePG support, alongside the Zalando
postgresClusterRef(samespec.databaseshape). - Gitops-managed team membership (§4.2).
- Per-namespace RBAC with a restricted cache (§9).
- Automatic Deployment restart on upstream Postgres credential rotation.
spec.pod.priorityClassName, pod labels beyond annotations, and an HPA forTerdutServer.- Admission webhooks beyond CEL (for example, checking a
teamRefexists at admission time). - A shared API types module or generated client for terdut-server, terdut-operator and terdut-tui.
- OLM packaging.