It's always the operator's own namespace by construction now (§6), never anything else, so there was nothing for the field to vary -- key varies instead (fixed 'token' when self-generated, whatever a human chose when adopted from spec.credentialsSecretRef).
41 KiB
terdut-operator design
This is the design reference for implementing terdut-operator. It exists so implementation can start from settled decisions instead of re-litigating them mid-PR. It is deliberately more detailed than the README; the README stays as the short pitch and now points here instead of carrying open questions.
Written against terdut-server as of the Postgres-only, per-team-resources
version (teams, escalation policies, dead man's switches and integrations are
all rows scoped to a team, managed over internal/api/* — not env/config-file
driven; see that repo's charts/terdut-server for the current deploy story
this operator supersedes).
1. Goals & non-goals
Goal: let a terdut-server install — the server itself, its teams, their escalation policies, dead man's switches and alert-source integrations — be fully described as Kubernetes objects and managed through gitops, following controller-runtime / Kubebuilder conventions.
Non-goals (v1):
- Not a general-purpose Postgres operator. It consumes a database that
either the Zalando
postgres-operatoror something else already provides. - Not managing Alertmanager itself, or the routing rules that decide which alerts reach which integration webhook — only the terdut-server side (creating the integration and handing back its URL/key).
- Not OLM packaging. Plain Kubebuilder manifests + Helm chart for install, matching how terdut-server itself ships.
- Cross-namespace references are limited to exactly one edge:
TerdutTeam.spec.serverRefmay name aTerdutServerin a different namespace, gated by thatTerdutServer's ownspec.allowedTeamsconsent field (§4.1, §4.2) — this is the multi-tenant shape the operator exists for (one platform team owns aTerdutServer; other teams self-service aTerdutTeamagainst it without needing write access to the server's namespace). Every other reference (teamRefon the escalation rule/dead-man-switch/alert-source CRDs) stays same-namespace-as-its-TerdutTeamonly, in v1 — those manage a specific team's own resources and are expected to live alongside it. - No validating/mutating admission webhooks in v1. CEL validation rules on
the CRDs (OpenAPI
x-kubernetes-validations) cover what they can; anything that needs a live look at another object (e.g. "does this teamRef exist") is a status condition, not an admission rejection — keeps v1 to a controller-only deployment with no cert-manager/webhook dependency.
2. The README's open questions, resolved
How do team-crd connect with server-crd?
Explicit spec.serverRef: {name, namespace} on TerdutTeam — same reasoning
as before (explicit, greppable, trivially validated), but namespace is
deliberately part of the reference: one team can run and own a
TerdutServer, and other teams — in their own namespaces, without any write
access to the server-owning team's namespace — self-service a TerdutTeam
against it. namespace defaults to the TerdutTeam's own namespace when
omitted, so the common single-tenant case (serverRef: {name: terdut}) is
unchanged.
This cross-namespace edge needs the target namespace's explicit consent —
otherwise any namespace in the cluster could point a TerdutTeam at
someone else's TerdutServer and have the operator provision a team on
its behalf, which is a namespace-boundary violation, not a gitops
convenience. (This consent gate is about which TerdutTeams the operator
will act on, not about credential exposure — no terdut-server credential
is ever placed in a TerdutTeam's own namespace regardless of this
setting; see §6.) Kubernetes has two established patterns for this kind
of consent, and Gateway API itself uses both, for two different
relationships:
ReferenceGrant(used for a Route reaching into an arbitrary Service/Secret): a separate object, living in the target namespace, enumerating exact{fromNamespace, fromKind} → {toKind, toName}pairs. No wildcard, no selector — every permitted namespace is spelled out.- An inline field on the parent (used for a
ListenerSetattaching to a sharedGateway, GA since Gateway API v1.5): the parent carriesspec.allowedListeners.namespaces: {from: None|Same|All|Selector, selector}directly, no separate CRD.
TerdutTeam attaching to a shared TerdutServer is structurally the
second case, not the first — a bounded set of expected children attaching
to a parent they were deliberately made shareable, not an arbitrary
backend reference — so this design follows the ListenerSet precedent:
TerdutServer.spec.allowedTeams (§4.1), no extra CRD. §4.2 covers how
TerdutTeam resolves against it. No other reference in this design
(teamRef on the child CRDs) crosses a namespace boundary, so this is the
only place cross-namespace consent is needed at all (§1, §5, §6, §9).
How does escalationrules, switches and alertsources connect to a team?
Explicit spec.teamRef: {name} on each of TerdutEscalationRule,
TerdutDeadmanSwitch, TerdutAlertSource — same reasoning, and it mirrors
terdut-server's own data model, where every one of these rows carries a
team_id foreign key already. A matcher/selector on TerdutTeam would be
inventing a second source of truth for an association the server already
models as a plain reference.
Support for both postgres-operator (Zalando) and bring-your-own, how do we design that to be user friendly?
TerdutServer.spec.database is a oneOf, mirroring the chart's existing
database.dsn / database.passwordSecret contract (see §8):
dsn+passwordSecretRef— bring-your-own, exactly today's chart inputs.postgresClusterRef— points at a Zalandopostgresql.acid.zalan.doCR; the operator derives the DSN and resolves the generated credentials Secret itself (see §8). CloudNativePG support is a natural follow-up using the same shape and is called out as deferred (§13), not designed in detail now.
3. API group, versions, CRD catalog
- Group:
terdut.ryuvia.com, version:v1alpha1(matches theryuvia.comdomain terdut-server already uses; bump tov1beta1/v1per the normal Kubernetes API graduation criteria once the shapes below have proven stable against a real install). - Module:
git.ryuvia.com/niklas/terdut-operator, scaffolded with Kubebuilder (controller-runtime), matching terdut-server's Go toolchain and house style.
| Kind | Scope | Purpose |
|---|---|---|
TerdutServer |
Namespaced | One terdut-server install: Deployment, Service, database wiring, bootstrap, operator credentials, cross-namespace team consent. |
TerdutTeam |
Namespaced | One team on a TerdutServer, possibly in another namespace: name, OIDC group mapping. |
TerdutEscalationRule |
Namespaced | A team's escalation policy (levels, targets, repeat). |
TerdutDeadmanSwitch |
Namespaced | One dead man's switch on a team. |
TerdutAlertSource |
Namespaced | One alert-ingest integration on a team (currently: Alertmanager webhook). |
4. Per-CRD spec
4.1 TerdutServer
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
name: terdut
namespace: oncall
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: v0.9.3
replicas: 1 # terdut-server is not horizontally-scale-tested; keep the field, default 1
networking:
hostname: terdut.example.com
servicePort: 8080
gatewayListener: "" # same semantics as chart's networking.listener
database:
dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef
passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password}
# --- OR ---
postgresClusterRef: {name: terdut-postgres} # Zalando `postgresql` CR in the same namespace
sweeper:
staleAfter: 6h
archiveAfter: 168h
deadman:
matchers: "alertname=Watchdog"
timeout: 15m
severity: critical
notify:
ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
fallbackTopic: ""
repeatEvery: 15m
tokenSecretRef: {name: "", key: token}
oidc:
enabled: false
issuer: ""
clientID: ""
clientSecretRef: {name: "", key: client-secret}
name: SSO
scopes: "openid profile email"
allowedGroups: []
adminGroup: ""
sessionMaxAge: 12h
passwordLogin: true
# Bring-your-own instance credential: a human mints this once, manually, with
# their own admin session (`POST /api/service-accounts {name: ..., scope:
# instance}`) and creates this Secret themselves, in the OPERATOR's own
# namespace (same namespace status.credentialsSecretRef below would otherwise
# point into). When set and the Secret exists, the controller adopts it
# directly and skips bootstrap entirely -- see §6 for why this is required,
# not optional polish, whenever terdut-server was already bootstrapped by its
# chart or a human before this CR existed (the normal case, not an edge one).
credentialsSecretRef: {name: "", key: token}
# Consent for TerdutTeams in OTHER namespaces to set serverRef at this
# TerdutServer. Same-namespace TerdutTeams never need this. Modeled on
# Gateway API's Gateway.spec.allowedListeners.namespaces (the ListenerSet
# attachment pattern, not ReferenceGrant — see §2 for why).
allowedTeams:
namespaces:
from: None # None (default) | Same | All | Selector
# selector: # required, and only meaningful, when from: Selector
# matchLabels:
# terdut.ryuvia.com/allowed: "true"
status:
conditions: [...] # Ready, DatabaseReady, Bootstrapped
observedGeneration: 3
serviceName: terdut
credentialsSecretRef: {name: terdut.platform-oncall-instance-credentials, key: token} # see §6; lives in the OPERATOR's namespace (always, implicitly -- not stored here), not this TerdutServer's. No `namespace` field: unlike an earlier draft, it's never anything other than the operator's own, so there's nothing to record. `key` replaces it, since that *does* vary -- fixed ("token") when the controller generated this Secret itself, whatever the human chose when it was adopted from spec.credentialsSecretRef instead.
Field-for-field this is the chart's values.yaml reshaped as a spec — the
operator absorbs the chart's Deployment/Service/bootstrap-job templates, so
existing installs have a direct mapping when migrating (see §10).
spec.credentialsSecretRef is an input (bring-your-own), distinct from
status.credentialsSecretRef's output (generated-by-the-controller) —
when the input is set, the controller treats it as the credential outright
and echoes its name back into status.credentialsSecretRef rather than
generating a second Secret alongside it. Left unset, the controller falls
back to the self-registration flow (§6) — which only actually completes if
this TerdutServer's own first reconcile is the one that wins the
/api/bootstrap race against a genuinely empty install; see §6 for why
that fallback is the exception, not the common case, and why this field
exists at all rather than being deferred as nice-to-have.
spec.database fields are +kubebuilder:validation:XValidation guarded to
be mutually exclusive (dsn xor postgresClusterRef); mirrors the chart's
"chart provisions no database" stance — this operator provisions no database
either, only wires up one that exists.
spec.allowedTeams.namespaces.from defaults to None, matching
allowedListeners's own default — a fresh TerdutServer accepts no
cross-namespace TerdutTeam until its owner opts in, same-namespace
TerdutTeams are unaffected either way. Selector deliberately has no
per-name allowlist (no "and only these teams") — namespace-level consent is
the right granularity here, same reasoning as ListenerSet: the namespace
is the tenancy boundary, not the object.
4.2 TerdutTeam
spec:
serverRef:
name: terdut
namespace: platform-oncall # optional; defaults to this TerdutTeam's own namespace.
# Cross-namespace requires that TerdutServer's spec.allowedTeams
# (§4.1) to admit this namespace — otherwise Ready: False, reason: RefNotPermitted.
displayName: "Platform" # -> POST /api/teams {"name": ...}; server assigns the ID
oidc:
memberGroup: "terdut-platform-members"
ownerGroup: "terdut-platform-owners"
status:
conditions: [...]
teamID: 42 # the server-side ID; needed by every child object's controller
credentialsSecretRef: {name: platform-oncall.platform-team-credentials, namespace: terdut-operator-system} # see §6; this team's own scoped key, in the OPERATOR's namespace
observedGeneration: 1
Team membership (which users belong, team_members) is explicitly not
modeled as a CRD field in v1: terdut-server already manages membership via
OIDC group sync at login for SSO installs, and manual membership for
password-login installs is a people-management action, not infrastructure —
forcing it through gitops would mean a human's team change goes through a PR
review. Flagged in §13 as revisitable if a real gitops-membership need shows up.
4.3 TerdutEscalationRule
spec:
teamRef: {name: platform-team}
repeatCount: 2
fallbackTopic: "platform-oncall"
levels:
- timeout: 5m
targets:
- kind: oncall # "oncall" or "user"
- kind: user
username: alice # resolved to a user ID by the controller at apply time
- timeout: 15m
targets:
- kind: oncall
status:
conditions: [...]
observedGeneration: 1
One TerdutEscalationRule per team — the server itself models a policy as
one row (escalation_policies) with an owned list of levels, so a
one-CRD-to-one-policy mapping (not one-CRD-per-level) matches the server's
own aggregate and lets the whole thing be reconciled with the single
PUT /api/teams/{teamID}/escalation the API actually exposes (see §5).
4.4 TerdutDeadmanSwitch
spec:
teamRef: {name: platform-team}
name: "prod-watchdog"
matcher: "alertname=Watchdog,cluster=prod"
timeout: 15m
severity: critical
status:
conditions: [...]
switchID: 7
4.5 TerdutAlertSource
spec:
teamRef: {name: platform-team}
kind: alertmanager
name: "prod-alertmanager"
status:
conditions: [...]
integrationID: 3
webhookURLSecretRef: {name: prod-alertmanager-terdut-webhook} # url + key, never in status/spec
The integration key is shown by the API exactly once, at creation
(Integration.Key/URL in terdut-server's own model) — never re-readable,
same shape as the bootstrap admin key. The controller writes it straight into
a generated, owner-referenced Secret on the create it caused and never logs
or stores it anywhere else; the CR's status carries only the Secret
reference, matching how e.g. cert-manager's Certificate exposes
spec.secretName rather than the key material itself.
4.6 Cross-namespace consent: TerdutServer.spec.allowedTeams
No separate CRD — the consent lives on TerdutServer itself (§4.1), following
Gateway API's Gateway.spec.allowedListeners (ListenerSet attachment)
rather than its ReferenceGrant, since a TerdutTeam attaching to a shared
TerdutServer is the same shape of relationship: a bounded set of expected
children attaching to a parent explicitly designed to be shared, not an
arbitrary cross-namespace backend reference (see §2 for the full comparison
of both patterns).
from: None(default) — no cross-namespaceTerdutTeammay resolve aserverRefinto thisTerdutServer. Same-namespaceTerdutTeams are always allowed regardless of this field.from: Same— equivalent toNonein effect (same-namespace is already unrestricted) but kept for parity with the upstream enum and to make the policy self-documenting in a diff.from: All— any namespace in the cluster may reference in. Appropriate for a genuinely shared, cluster-wideTerdutServer; the audit trail is "checkallowedTeamsplus who has RBAC to create aTerdutTeamanywhere," which is materially weaker thanSelector.from: Selector— only namespaces matchingspec.allowedTeams.namespaces.selector(a standardmetav1.LabelSelectoroverNamespaceobjects, exactly likeallowedListeners's ownselector) may reference in. This is the recommended mode for the platform-team-owns-a-shared-server scenario this design targets: label the consuming namespaces once (e.g.terdut.ryuvia.com/allowed-server: platform-oncall/terdut) and new namespaces opt in by carrying the label, without editing theTerdutServeragain.- A
TerdutTeam's controller re-evaluatesallowedTeamson every reconcile before it will resolve a cross-namespaceserverRef— forSelector, this means aGeton its ownNamespaceobject plus reading the targetTerdutServer's spec, not a List across the cluster. Same-namespaceserverRefnever consults this field at all. - Narrowing or clearing
allowedTeams(or unlabeling a namespace, underSelector) is a live revocation: the next reconcile of anyTerdutTeamit used to authorize finds itself no longer permitted, flipsReady: False, reason: RefNotPermitted, and — deliberately — does not delete the team server-side on revocation alone; it stops reconciling further changes until access is restored or theTerdutTeamCR itself is deleted (whose finalizer still needs the credentials Secret described in §6 to clean up, so blocking new changes rather than forcing an immediate, possibly credential-less deletion is the safer failure mode).
5. Reconciliation semantics
terdut-server's REST surface (internal/api/router.go) does not give every
resource a full update verb, so reconciliation strategy is per-resource:
| Resource | Verbs available | Strategy |
|---|---|---|
| Team | POST create, PUT rename, DELETE, PUT oidc-groups | Real update-in-place: diff spec vs. last-applied, PUT the changed pieces. |
| Escalation policy | GET/PUT whole-policy | Update-in-place: PUT the full desired policy every reconcile that finds drift; cheap because whole-policy is small and already loaded whole server-side. |
| Dead man's switch | POST create, DELETE — no PUT | Delete-and-recreate on any spec diff other than name. The controller diffs against status (which mirrors what was last successfully applied) rather than re-reading the server every reconcile, to avoid a spurious recreate from field reordering. |
| Integration (alert source) | POST create, PATCH rename, DELETE | Rename via PATCH; any other spec change (kind) is delete-and-recreate, which rotates the webhook key — called out loudly in the CRD's field docs and in a Warning event, since it breaks whatever sends to the old URL/key until the new Secret is picked up. (See the general rule below for what happens if the Secret is lost with no spec change.) |
General rules for every controller:
- Idempotent create: before POSTing, check
status.<serverSideID>is unset; if the server already has a same-named object from a previous partial reconcile (e.g. after a crash between POST and status-write), treat a 409/name-conflict as "adopt" — GET-by-name and populate status, rather than erroring forever. terdut-server's list endpoints in each of these areas return objects by name, so this is a straightforward correlation. - Periodic resync in addition to watch-triggered reconciles (Kubebuilder
default
RequeueAfteron success, e.g. every 5–10 minutes) to catch drift from someone changing state directly against the server's API/UI, since gitops correctness means the CR wins, not "first write wins". - Finalizers on every CRD that has a server-side counterpart, so deletion
calls the corresponding DELETE before the Kubernetes object disappears.
Failure to delete server-side (e.g. server unreachable) blocks finalizer
removal and surfaces as a
Degradedcondition + event, rather than silently orphaning a row. - Generated Secrets holding unrecoverable server-issued material are
watched, and their loss is fail-closed, not self-healed. Currently this
is just
TerdutAlertSource's webhook Secret (§4.5): the controller adds it to itsOwns()watches, not just the CR. If it disappears whilestatus.integrationIDis still set, the controller does not attempt to recreate it — the key is genuinely gone (§4.5: never stored anywhere but that one Secret), so silently minting a replacement would rotate a live production webhook URL with no corresponding spec change to explain why. Instead it flipsReady: False, reason: WebhookSecretLostand fires aWarningevent telling the operator to delete and recreate theTerdutAlertSource. No new mechanism is needed for recovery: deleting the CR runs the existing finalizer (DELETE the still-live integration server-side, above), and recreating it runs the existing idempotent-create path (this same section) — a fresh POST, a new key, a new Secret. This is deliberately the same recovery motion as the kind-change rotation above, just human-triggered instead of spec-triggered. - Owner chain for status resolution, not API calls:
TerdutTeam's controller does not call any other controller; every child CRD's controller independently resolves its ownteamRef→TerdutTeam.statusfor both theteamIDand thecredentialsSecretRefit needs to call the API (§6) — it never needs to chain further up toTerdutServerat all, since the team's own scoped credential is everything a child resource's controller requires. If the referencedTerdutTeamisn'tReadyyet (which includes not having acredentialsSecretRefset), the child requeues with backoff and reportsReady: False, reason: WaitingForTeam— no cross-controller RPC. - Cross-namespace
serverRefis re-checked every reconcile, not just at creation:TerdutTeam's controller reads the targetTerdutServer'sspec.allowedTeams(and, underSelector, aGeton its ownNamespaceobject for labels) on every pass before touching a cross-namespaceTerdutServer— revocation (§4.6) takes effect on the team's very next reconcile, not just when the CR is first applied.
6. Bootstrap & authentication to terdut-server's API
terdut-server's scoped service-account credential type
(terdut-server's SERVICE-ACCOUNTS.md) has shipped — confirmed against
source: internal/api/service_accounts.go, migration
014_service_accounts.sql, and internal/api/router.go wiring it in under
AuthMiddleware. This section is no longer blocked on it; the "v1-blocking"
framing here was accurate when this section was first written and is stale
now.
GET /api/service-accounts?name= is not an unauthenticated lookup, unlike
an earlier draft of this section assumed — confirmed against
internal/api/middleware.go's AuthMiddleware, which hard-rejects any
request carrying neither a Bearer token nor a session cookie with 401
before any handler (including this one's own internal, more permissive
name-filter check) ever runs. This matters beyond a technicality: it means
the self-registration flow below (point 1) only ever completes for the
TerdutServer whose own controller happens to win the /api/bootstrap race
on a genuinely empty install. Every other case — including Stage 1's own
setup (ROADMAP.md): terdut-server deployed by its existing chart, which
runs its own bootstrap Job, before the TerdutServer CR or its controller
ever exist — leaves the controller with no credential and no authenticated
way to get one. spec.credentialsSecretRef (§4.1) exists to make that the
normal path, not an unhandled edge case: a human mints an instance-scoped
service account once, manually, with their own admin session, and hands the
controller that Secret directly.
Every credential the operator holds — the one instance-scoped key per
TerdutServer, and one team-scoped key per TerdutTeam — lives in a Secret
in the operator's own namespace, never in the namespace of the CR it
authenticates for. Reconciliation happens entirely inside the operator's
controller loop, which is a single Deployment/ServiceAccount already
watching every namespace it's granted (§9); nothing about calling
terdut-server's API on a CR's behalf requires the credential to be
physically located near that CR, and no CR owner (human or otherwise) ever
needs to see, hold, or have RBAC to read a terdut-server credential. This is
a straight simplification of an earlier draft of this section, which mirrored
a shared credential into each consenting namespace instead — that version
conflated "the CR's owner never needs to see this" (true, and preserved
here) with "so the credential must live in the CR's namespace" (a
non-sequitur once you don't need to grant anyone else namespace-local
read access). Dropping that assumption also removes an entire class of
complexity: no on-demand mirroring, no garbage-collecting an orphaned copy
when allowedTeams narrows, no "OwnerReferences can't cross namespaces so
track it in status instead" workaround — none of that machinery is needed
when nothing ever crosses into a tenant namespace in the first place.
- First reconcile checks
spec.credentialsSecretRefbefore anything else. Set and the Secret exists: adopt it as-is, setstatus.credentialsSecretRefto the same reference,Bootstrapped: True, done — no API call made at all. This is the path every Stage 1 install actually takes (ROADMAP.md): terdut-server deployed by its existing chart, bootstrapped by that chart's own Job, before this CR exists. Unset (or the named Secret doesn't exist yet): fall through to self-registration, confirmed against source (internal/api/users.go'shandleBootstrap):/api/bootstrapis single-shot per install, gated onSELECT COUNT(*) FROM users— once non-zero, every call 403s regardless of identity, exactly ascharts/terdut-server's ownbootstrap-job.yamlalready assumes (403 → "already bootstrapped, nothing to do", exit 0). Call/api/bootstrap; on201, its response ({"user": ..., "api_key": {"key": "<raw>", ...}}) hands back a real, usable admin key directly — use it for exactly one further call,POST /api/service-accounts {name: "terdut-operator", scope: "instance"}, and keep that key, not the raw admin one, as the lasting credential. On403, there is no further fallback: per this section's opening note,GET /api/service-accounts?name=needs a credential this controller does not have, so self-lookup cannot run unauthenticated. SetReady: False, reason: WaitingForCredentialwith an event telling the human to mint an instance-scoped service account with their own admin session and setspec.credentialsSecretRef, and requeue with backoff — this is the expected, steady-state outcome whenever the operator loses (or never entered) the bootstrap race, not a transient error to retry past. - On the self-registration path only (point 1's bring-your-own path
generates nothing — it adopts the human-provided Secret directly): the
resulting instance-scoped key is written to a generated Secret in the
operator's own namespace (e.g.
<serverRef.namespace>.<serverRef.name>-instance-credentials, under a fixed data key,token), referenced back fromTerdutServer.status.credentialsSecretRef: {name, key}(§4.1). NoOwnerReference(those can't cross namespaces, and this Secret doesn't share a namespace with theTerdutServerthat caused it); theTerdutServer's finalizer deletes this Secret directly as part of its own teardown, the same way it already has to clean up the server-side resources it created (§5's general finalizer rule extends naturally to this Secret). - When a
TerdutTeamfirst becomesReady(itsserverRefresolved,allowedTeamssatisfied if cross-namespace), its controller uses theTerdutServer's instance-scoped credential (read from the operator's own namespace, resolved via the owner chain in §5) to mint a team-scoped service account for itself:POST /api/service-accountswithscope: team, teamID: <status.teamID>. The resulting key is written to its own generated Secret, again in the operator's own namespace (e.g.<teamNamespace>.<teamName>-team-credentials), referenced fromTerdutTeam.status.credentialsSecretRef(§4.2). Same finalizer pattern as point 2: theTerdutTeam's finalizer deletes this Secret as part of its own teardown. - Every child controller (
TerdutEscalationRule,TerdutDeadmanSwitch,TerdutAlertSource) reads its team'scredentialsSecretRef— resolved through itsteamRef→TerdutTeam.status(§5) — and never touches the instance-scoped credential at all. Since child CRDs stay same-namespace- as-their-TerdutTeamin v1 (§1), and the credential itself lives in the operator's namespace regardless of where theTerdutTeamor its children are, this works identically whether theTerdutTeamis same-namespace or cross-namespace relative to itsTerdutServer— there is no separate cross-namespace case to handle here at all, unlike the mirroring design this replaced. - Blast radius: a team-scoped key can only touch its own
team_id's escalation policy, dead-man switches, integrations, schedule and OIDC group bindings server-side (enforced by terdut-server itself, perSERVICE-ACCOUNTS.md) — compromising one such Secret (e.g. a bug that leaks operator-namespace Secrets, or an overly broad RBAC grant on that one namespace) exposes exactly one team, never the whole server. This is the real fix for what an earlier draft of this section called out as its weak point (every mirrored copy being server-admin-equivalent); it falls out of team-scoped credentials existing at all, independent of where they're stored — the operator-private storage described above closes the RBAC-footprint half of the problem, team scoping closes the credential- privilege half. - Rotation:
POST /api/service-accounts/{id}/keysmints a new key on the existing account without recreating it; the old key is revoked viaDELETE /api/service-accounts/{id}/keys/{keyID}; the operator's local Secret is updated in place. No DB-level workaround, no re-triggering a single-shot endpoint that can't fire twice (which is what made rotation unworkable under the old/api/bootstrap-only design).
7. Ownership, status, garbage collection
- Every generated object that lives in the same namespace as the CR that
caused it (Deployment, Service, webhook Secret) carries a standard
metav1.OwnerReference— GC handles these, no finalizer needed. The two credential Secrets from §6 are the one exception: they live in the operator's own namespace regardless of where their owning CR lives, soOwnerReferencedoesn't apply (cross-namespace) and cleanup instead runs through that CR's finalizer directly, alongside the server-side DELETE it already has to issue (§5). - Status conditions follow the standard
metav1.Conditionshape with at leastReadyon every kind, plus kind-specific ones (TerdutServer:DatabaseReady,Bootstrapped; children:Synced). status.observedGenerationon every kind, bumped only after a successful reconcile against that generation's spec — the standard way a client (orkubectl wait) tells "applied" from "seen".- No cluster-scoped aggregation object (e.g. no cluster-wide "all servers"
status) in v1 —
kubectl get terdutservers -Ais the aggregate view.
8. Postgres integration
Mirrors the chart's existing two-path contract (values.yaml
database.dsn/passwordSecret), because that contract is already
documented and tested operationally:
- Bring-your-own:
spec.database.dsn(no password) +spec.database.passwordSecretRef— the operator setsPGPASSWORDon the Deployment's container env from that Secret, exactly like the chart does today, and does nothing else. No connectivity check beyond what the Deployment's own readiness probe already gives. - Zalando
postgres-operator:spec.database.postgresClusterRefnames apostgresql.acid.zalan.doCR in the same namespace. The controller:- Reads that CR's status for the primary Service name/port to build the DSN
host —
<cluster>.<namespace>.svc:5432— and database name convention. - Resolves the generated credentials Secret
(
<user>.<cluster>.credentials.postgresql.acid.zalan.do) the same way the chart's comment already documents, and wires it in asPGPASSWORDthe same way. - Watches that Secret (not just the
postgresqlCR) so a credential rotation triggers a requeue — the chart today requires a manual pod restart for this; the operator can at least detect and report it via a condition even if restarting on rotation is left as a §13 follow-up rather than done automatically (a rolling restart on credential change is a behavior change worth its own design pass, not folded in here). - Requires read RBAC on
postgresql.acid.zalan.do(optional CRD — the operator's ClusterRole/Role should not hard-fail if the CRD isn't installed and a givenTerdutServeruses BYO DSN instead).
- Reads that CR's status for the primary Service name/port to build the DSN
host —
9. RBAC
- The operator's own ServiceAccount needs, per namespace it's granted:
get/list/watch/create/update/patch/deleteonDeployments,Servicesit owns, andget/list/watchonpostgresql.acid.zalan.do(optional, degrade gracefully if absent per §8), plus cluster-wideget/listonNamespace(labels only, forallowedTeams: {from: Selector}evaluation — §4.1, §4.6). - Two different
Secretscopes, not one — corrected from an earlier draft of this section. That earlier draft saidSecretaccess was "scoped to the operator's own namespace only... nowhere else," reasoning that with every credential held privately in the operator's own namespace (§6) there was no legitimate reason to touch aSecretanywhere else. That was wrong once §4.5 existed:- The §6 credential Secrets (one instance-scoped key per
TerdutServer, one team-scoped key perTerdutTeam) do live in, and are only ever touched from, the operator's own namespace —get/list/watch/create/update/patch/deletethere, nowhere else. The rest of the original reasoning stands for these Secrets specifically: no human or team's own RBAC is ever granted access to a terdut-server credential by this design, and the operator itself never needs cross-namespace access to reach them. - The §4.5 webhook Secret is different: it's owned by and lives beside
its
TerdutAlertSource, in that CR's own tenant namespace, not the operator's. The per-namespaceRolealready granted forDeployments/Servicesin each watched namespace (below) must carry the sameSecretverbs there too, or the controller cannot create, watch, or even detect the loss of (§5) that Secret at all. - This necessarily widens the operator's footprint in each watched
tenant namespace to "any
Secretin that namespace," not just the ones it created — Kubernetes RBAC has no owner-scoped grant finer than the namespace itself, and the design already accepts this same granularity for Deployments/Services there. Flagged as an accepted trade-off, not a silent gap (§13).
- The §6 credential Secrets (one instance-scoped key per
- No cluster-scoped resources are created by this operator (namespaced CRDs
only, per §1) — a
Role+RoleBindingper watched (tenant) namespace is sufficient for Deployments/Services/the webhookSecret/the optional Zalando CRD, plus a separateRole+RoleBindingin the operator's own namespace for the §6 credential Secrets; aClusterRoleis only needed for watching CRDs across all namespaces (the normal Kubebuilder multi-tenant-operator default) — none of theSecretaccess above needs to be cluster-scoped. - terdut-server's own RBAC is unaffected — the operator talks to it purely over HTTP with service-account API keys (§6), never via the Kubernetes API for app-level state.
10. Relationship to charts/terdut-server
Recommendation (flagged explicitly as a decision to confirm before
implementation starts, not settled by this document alone): the chart's
Deployment/Service/bootstrap-job templates become redundant once
TerdutServer exists — running both would mean two controllers (Helm and
this operator) reconciling the same Deployment, which is exactly the
conflict Kubernetes operators exist to avoid. Proposed path:
- The chart is repurposed into an installer chart: it installs the
operator + CRDs (and optionally one
TerdutServerCR fromvalues.yaml, for users who want "helm install and get a server" without hand-writing a CR) rather than templating the Deployment directly. - Existing installs migrate by:
helm templatethe current release'svalues.yamlinto an equivalentTerdutServerCR (mechanical, since §4.1 is deliberately shaped to make that mapping 1:1), install the operator, apply the CR, then let Helm's release be uninstalled or reduced to just the CRD/operator subchart. - This is a breaking change to the chart's contract and needs its own migration guide and probably a major chart version bump — out of scope for this design doc beyond flagging it; do not start that migration work without separately confirming this recommendation.
- This decision isn't only about the migration — it also decides who
owns bootstrap. §6 assumed the operator could always get its own
bootstrap identity separately from the chart's; §6 point 1 shows that's
false, so whichever of {chart's Job, operator controller} is expected to
call
/api/bootstrapfirst has to be settled explicitly (a one-paragraph call, not the full migration plan) before writing any operator bootstrap/credential code, not deferred alongside the rest of this section.
11. Testing strategy
envtest(controller-runtime's fake API server) for every controller's reconcile logic against the Kubernetes side.- terdut-server's REST API is faked with a small
httptest.Serverper controller test driven by fixtures matching the real handlers' request/response shapes (already well-documented ininternal/api/*_test.goon the server side) — no real Postgres or real terdut-server binary needed for controller unit tests. - A smaller number of true end-to-end tests (
kindcluster + real terdut-server image + real Postgres) covering the golden path per CRD: createTerdutServer→TerdutTeam→ one of each child kind → verify via terdut-server's own API that the objects exist with the right shape → delete the CR → verify the server-side object is gone.
12. Observability
- Standard controller-runtime metrics (reconcile duration/error counts) are enough for v1 — no custom metrics.
- Every externally-visible action (bootstrap, key rotation-needed, delete-
and-recreate on the no-PUT resources, adopt-on-conflict) emits a
Kubernetes
Eventon the CR, since that's what shows up inkubectl describeand gitops tooling (Argo CD/Flux) surfaces without extra wiring.
13. Deferred / explicitly out of scope for this design
- A real scoped service-account/token type in terdut-server — not merely
deferred, this is v1-blocking for §6 as written (verified: without it,
§6's bootstrap flow has no working credential-rotation path and no clean
answer to the chart-vs-operator bootstrap race; see §6 points 1, 5, 6 and
§10). Sequence this server-side change before implementing the
TerdutServercontroller's bootstrap logic, not after. - A version-discovery endpoint on terdut-server (e.g.
GET /api/version). Neither this operator nor terdut-tui has one today — both independently detect capability by probing specific routes (terdut-tui viaGET /api/teams404-checking; this operator would otherwise need to invent its own equivalent probe). An unattended reconciler is more exposed to a silent breaking API change than an interactive TUI a human is watching; raising this alongside the service-account request rather than inventing another route-probe here. - CloudNativePG support — same
spec.databaseshape as Zalando should extend to it, but the concrete field/Secret-naming conventions need their own look. - Cross-namespace
teamRefon the child CRDs (TerdutEscalationRule,TerdutDeadmanSwitch,TerdutAlertSource) — onlyTerdutTeam.serverRefcrosses namespaces in v1 (§2, §4.2, §4.6); these stay same-namespace as theirTerdutTeamuntil a real need for splitting them out shows up. - Narrower-than-namespace RBAC for the §4.5 webhook Secret (Kubernetes RBAC
has no owner-scoped grant below the namespace itself, per §9) — revisit
if the widened per-tenant-namespace
Secretaccess proves too broad in practice. - Gitops-managed team membership (see §4.2).
- Automatic Deployment restart on upstream Postgres credential rotation.
- Admission webhooks / CEL-only validation limits (e.g. verifying a
teamRefexists at admission time rather than surfacing it as a status condition after the fact). - OLM packaging, Helm chart migration execution (§10 is a recommendation, not a plan to execute).