# terdut-operator design This is the design reference for implementing terdut-operator. It exists so implementation can start from settled decisions instead of re-litigating them mid-PR. It is deliberately more detailed than the README; the README stays as the short pitch and now points here instead of carrying open questions. Written against terdut-server as of the Postgres-only, per-team-resources version (teams, escalation policies, dead man's switches and integrations are all rows scoped to a team, managed over `internal/api/*` — not env/config-file driven; see that repo's `charts/terdut-server` for the current deploy story this operator supersedes). ## 1. Goals & non-goals **Goal:** let a terdut-server install — the server itself, its teams, their escalation policies, dead man's switches and alert-source integrations — be fully described as Kubernetes objects and managed through gitops, following controller-runtime / Kubebuilder conventions. **Non-goals (v1):** - Not a general-purpose Postgres operator. It *consumes* a database that either the Zalando `postgres-operator` or something else already provides. - Not managing Alertmanager itself, or the routing rules that decide which alerts reach which integration webhook — only the terdut-server side (creating the integration and handing back its URL/key). - Not OLM packaging. Plain Kubebuilder manifests + Helm chart for install, matching how terdut-server itself ships. - Cross-namespace references are limited to exactly one edge: `TerdutTeam.spec.serverRef` may name a `TerdutServer` in a different namespace, gated by that `TerdutServer`'s own `spec.allowedTeams` consent field (§4.1, §4.2) — this is the multi-tenant shape the operator exists for (one platform team owns a `TerdutServer`; other teams self-service a `TerdutTeam` against it without needing write access to the server's namespace). Every other reference (`teamRef` on the escalation rule/dead-man-switch/alert-source CRDs) stays same-namespace-as-its-`TerdutTeam` only, in v1 — those manage a specific team's own resources and are expected to live alongside it. - No validating/mutating admission webhooks in v1. CEL validation rules on the CRDs (OpenAPI `x-kubernetes-validations`) cover what they can; anything that needs a live look at another object (e.g. "does this teamRef exist") is a status condition, not an admission rejection — keeps v1 to a controller-only deployment with no cert-manager/webhook dependency. ## 2. The README's open questions, resolved > How do team-crd connect with server-crd? Explicit `spec.serverRef: {name, namespace}` on `TerdutTeam` — same reasoning as before (explicit, greppable, trivially validated), but **`namespace` is deliberately part of the reference**: one team can run and own a `TerdutServer`, and other teams — in their own namespaces, without any write access to the server-owning team's namespace — self-service a `TerdutTeam` against it. `namespace` defaults to the `TerdutTeam`'s own namespace when omitted, so the common single-tenant case (`serverRef: {name: terdut}`) is unchanged. This cross-namespace edge needs the target namespace's explicit consent — otherwise any namespace in the cluster could point a `TerdutTeam` at someone else's `TerdutServer` and have the operator provision a team + credentials against it, which is a namespace-boundary violation, not a gitops convenience. Kubernetes has two established patterns for this kind of consent, and Gateway API itself uses both, for two different relationships: - **`ReferenceGrant`** (used for a Route reaching into an arbitrary Service/Secret): a separate object, living in the *target* namespace, enumerating exact `{fromNamespace, fromKind} → {toKind, toName}` pairs. No wildcard, no selector — every permitted namespace is spelled out. - **An inline field on the parent** (used for a `ListenerSet` attaching to a shared `Gateway`, GA since Gateway API v1.5): the parent carries `spec.allowedListeners.namespaces: {from: None|Same|All|Selector, selector}` directly, no separate CRD. `TerdutTeam` attaching to a shared `TerdutServer` is structurally the second case, not the first — a bounded set of expected children attaching to a parent they were deliberately made shareable, not an arbitrary backend reference — so this design follows the `ListenerSet` precedent: `TerdutServer.spec.allowedTeams` (§4.1), no extra CRD. §4.2 covers how `TerdutTeam` resolves against it. No other reference in this design (`teamRef` on the child CRDs) crosses a namespace boundary, so this is the only place cross-namespace consent is needed at all (§1, §5, §6, §9). > How does escalationrules, switches and alertsources connect to a team? Explicit `spec.teamRef: {name}` on each of `TerdutEscalationRule`, `TerdutDeadmanSwitch`, `TerdutAlertSource` — same reasoning, and it mirrors terdut-server's own data model, where every one of these rows carries a `team_id` foreign key already. A matcher/selector on `TerdutTeam` would be inventing a second source of truth for an association the server already models as a plain reference. > Support for both postgres-operator (Zalando) and bring-your-own, how do we > design that to be user friendly? `TerdutServer.spec.database` is a oneOf, mirroring the chart's existing `database.dsn` / `database.passwordSecret` contract (see §8): - `dsn` + `passwordSecretRef` — bring-your-own, exactly today's chart inputs. - `postgresClusterRef` — points at a Zalando `postgresql.acid.zalan.do` CR; the operator derives the DSN and resolves the generated credentials Secret itself (see §8). CloudNativePG support is a natural follow-up using the same shape and is called out as deferred (§13), not designed in detail now. ## 3. API group, versions, CRD catalog - Group: `terdut.ryuvia.com`, version: `v1alpha1` (matches the `ryuvia.com` domain terdut-server already uses; bump to `v1beta1`/`v1` per the normal Kubernetes API graduation criteria once the shapes below have proven stable against a real install). - Module: `git.ryuvia.com/niklas/terdut-operator`, scaffolded with Kubebuilder (controller-runtime), matching terdut-server's Go toolchain and house style. | Kind | Scope | Purpose | |---|---|---| | `TerdutServer` | Namespaced | One terdut-server install: Deployment, Service, database wiring, bootstrap, operator credentials, cross-namespace team consent. | | `TerdutTeam` | Namespaced | One team on a `TerdutServer`, possibly in another namespace: name, OIDC group mapping. | | `TerdutEscalationRule` | Namespaced | A team's escalation policy (levels, targets, repeat). | | `TerdutDeadmanSwitch` | Namespaced | One dead man's switch on a team. | | `TerdutAlertSource` | Namespaced | One alert-ingest integration on a team (currently: Alertmanager webhook). | ## 4. Per-CRD spec ### 4.1 `TerdutServer` ```yaml apiVersion: terdut.ryuvia.com/v1alpha1 kind: TerdutServer metadata: name: terdut namespace: oncall spec: image: repository: git.ryuvia.com/niklas/terdut-server tag: v0.9.3 replicas: 1 # terdut-server is not horizontally-scale-tested; keep the field, default 1 networking: hostname: terdut.example.com servicePort: 8080 gatewayListener: "" # same semantics as chart's networking.listener database: dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password} # --- OR --- postgresClusterRef: {name: terdut-postgres} # Zalando `postgresql` CR in the same namespace sweeper: staleAfter: 6h archiveAfter: 168h deadman: matchers: "alertname=Watchdog" timeout: 15m severity: critical notify: ntfyURL: "http://ntfy.ntfy.svc.cluster.local" fallbackTopic: "" repeatEvery: 15m tokenSecretRef: {name: "", key: token} oidc: enabled: false issuer: "" clientID: "" clientSecretRef: {name: "", key: client-secret} name: SSO scopes: "openid profile email" allowedGroups: [] adminGroup: "" sessionMaxAge: 12h passwordLogin: true # Consent for TerdutTeams in OTHER namespaces to set serverRef at this # TerdutServer. Same-namespace TerdutTeams never need this. Modeled on # Gateway API's Gateway.spec.allowedListeners.namespaces (the ListenerSet # attachment pattern, not ReferenceGrant — see §2 for why). allowedTeams: namespaces: from: None # None (default) | Same | All | Selector # selector: # required, and only meaningful, when from: Selector # matchLabels: # terdut.ryuvia.com/allowed: "true" status: conditions: [...] # Ready, DatabaseReady, Bootstrapped observedGeneration: 3 serviceName: terdut operatorCredentialsSecretRef: {name: terdut-operator-credentials} # see §6 ``` Field-for-field this is the chart's `values.yaml` reshaped as a spec — the operator absorbs the chart's Deployment/Service/bootstrap-job templates, so existing installs have a direct mapping when migrating (see §10). `spec.database` fields are `+kubebuilder:validation:XValidation` guarded to be mutually exclusive (`dsn` xor `postgresClusterRef`); mirrors the chart's "chart provisions no database" stance — this operator provisions no database either, only wires up one that exists. `spec.allowedTeams.namespaces.from` defaults to `None`, matching `allowedListeners`'s own default — a fresh `TerdutServer` accepts no cross-namespace `TerdutTeam` until its owner opts in, same-namespace `TerdutTeam`s are unaffected either way. `Selector` deliberately has no per-name allowlist (no "and only these teams") — namespace-level consent is the right granularity here, same reasoning as `ListenerSet`: the namespace is the tenancy boundary, not the object. ### 4.2 `TerdutTeam` ```yaml spec: serverRef: name: terdut namespace: platform-oncall # optional; defaults to this TerdutTeam's own namespace. # Cross-namespace requires that TerdutServer's spec.allowedTeams # (§4.1) to admit this namespace — otherwise Ready: False, reason: RefNotPermitted. displayName: "Platform" # -> POST /api/teams {"name": ...}; server assigns the ID oidc: memberGroup: "terdut-platform-members" ownerGroup: "terdut-platform-owners" status: conditions: [...] teamID: 42 # the server-side ID; needed by every child object's controller observedGeneration: 1 ``` Team *membership* (which users belong, `team_members`) is explicitly **not** modeled as a CRD field in v1: terdut-server already manages membership via OIDC group sync at login for SSO installs, and manual membership for password-login installs is a people-management action, not infrastructure — forcing it through gitops would mean a human's team change goes through a PR review. Flagged in §13 as revisitable if a real gitops-membership need shows up. ### 4.3 `TerdutEscalationRule` ```yaml spec: teamRef: {name: platform-team} repeatCount: 2 fallbackTopic: "platform-oncall" levels: - timeout: 5m targets: - kind: oncall # "oncall" or "user" - kind: user username: alice # resolved to a user ID by the controller at apply time - timeout: 15m targets: - kind: oncall status: conditions: [...] observedGeneration: 1 ``` One `TerdutEscalationRule` per team — the server itself models a policy as one row (`escalation_policies`) with an owned list of levels, so a one-CRD-to-one-policy mapping (not one-CRD-per-level) matches the server's own aggregate and lets the whole thing be reconciled with the single `PUT /api/teams/{teamID}/escalation` the API actually exposes (see §5). ### 4.4 `TerdutDeadmanSwitch` ```yaml spec: teamRef: {name: platform-team} name: "prod-watchdog" matcher: "alertname=Watchdog,cluster=prod" timeout: 15m severity: critical status: conditions: [...] switchID: 7 ``` ### 4.5 `TerdutAlertSource` ```yaml spec: teamRef: {name: platform-team} kind: alertmanager name: "prod-alertmanager" status: conditions: [...] integrationID: 3 webhookURLSecretRef: {name: prod-alertmanager-terdut-webhook} # url + key, never in status/spec ``` The integration key is shown by the API exactly once, at creation (`Integration.Key`/`URL` in terdut-server's own model) — never re-readable, same shape as the bootstrap admin key. The controller writes it straight into a generated, owner-referenced Secret on the create it caused and never logs or stores it anywhere else; the CR's `status` carries only the Secret reference, matching how e.g. cert-manager's `Certificate` exposes `spec.secretName` rather than the key material itself. ### 4.6 Cross-namespace consent: `TerdutServer.spec.allowedTeams` No separate CRD — the consent lives on `TerdutServer` itself (§4.1), following Gateway API's `Gateway.spec.allowedListeners` (`ListenerSet` attachment) rather than its `ReferenceGrant`, since a `TerdutTeam` attaching to a shared `TerdutServer` is the same shape of relationship: a bounded set of expected children attaching to a parent explicitly designed to be shared, not an arbitrary cross-namespace backend reference (see §2 for the full comparison of both patterns). - `from: None` (default) — no cross-namespace `TerdutTeam` may resolve a `serverRef` into this `TerdutServer`. Same-namespace `TerdutTeam`s are always allowed regardless of this field. - `from: Same` — equivalent to `None` in effect (same-namespace is already unrestricted) but kept for parity with the upstream enum and to make the policy self-documenting in a diff. - `from: All` — any namespace in the cluster may reference in. Appropriate for a genuinely shared, cluster-wide `TerdutServer`; the audit trail is "check `allowedTeams` plus who has RBAC to create a `TerdutTeam` anywhere," which is materially weaker than `Selector`. - `from: Selector` — only namespaces matching `spec.allowedTeams.namespaces.selector` (a standard `metav1.LabelSelector` over `Namespace` objects, exactly like `allowedListeners`'s own `selector`) may reference in. This is the recommended mode for the platform-team-owns-a-shared-server scenario this design targets: label the consuming namespaces once (e.g. `terdut.ryuvia.com/allowed-server: platform-oncall/terdut`) and new namespaces opt in by carrying the label, without editing the `TerdutServer` again. - A `TerdutTeam`'s controller re-evaluates `allowedTeams` on every reconcile before it will resolve a cross-namespace `serverRef` — for `Selector`, this means a `Get` on its own `Namespace` object plus reading the target `TerdutServer`'s spec, not a List across the cluster. Same-namespace `serverRef` never consults this field at all. - Narrowing or clearing `allowedTeams` (or unlabeling a namespace, under `Selector`) is a live revocation: the next reconcile of any `TerdutTeam` it used to authorize finds itself no longer permitted, flips `Ready: False, reason: RefNotPermitted`, and — deliberately — does **not** delete the team server-side on revocation alone; it stops reconciling further changes until access is restored or the `TerdutTeam` CR itself is deleted (whose finalizer still needs the credentials Secret described in §6 to clean up, so blocking *new* changes rather than forcing an immediate, possibly credential-less deletion is the safer failure mode). ## 5. Reconciliation semantics terdut-server's REST surface (`internal/api/router.go`) does not give every resource a full update verb, so reconciliation strategy is per-resource: | Resource | Verbs available | Strategy | |---|---|---| | Team | POST create, PUT rename, DELETE, PUT oidc-groups | Real update-in-place: diff spec vs. last-applied, PUT the changed pieces. | | Escalation policy | GET/PUT whole-policy | Update-in-place: PUT the full desired policy every reconcile that finds drift; cheap because whole-policy is small and already loaded whole server-side. | | Dead man's switch | POST create, DELETE — **no PUT** | Delete-and-recreate on any spec diff other than `name`. The controller diffs against `status` (which mirrors what was last successfully applied) rather than re-reading the server every reconcile, to avoid a spurious recreate from field reordering. | | Integration (alert source) | POST create, PATCH rename, DELETE | Rename via PATCH; any other spec change (kind) is delete-and-recreate, which **rotates the webhook key** — called out loudly in the CRD's field docs and in a `Warning` event, since it breaks whatever sends to the old URL/key until the new Secret is picked up. | General rules for every controller: - **Idempotent create**: before POSTing, check `status.` is unset; if the server already has a same-named object from a previous partial reconcile (e.g. after a crash between POST and status-write), treat a 409/name-conflict as "adopt" — GET-by-name and populate status, rather than erroring forever. terdut-server's list endpoints in each of these areas return objects by name, so this is a straightforward correlation. - **Periodic resync** in addition to watch-triggered reconciles (Kubebuilder default `RequeueAfter` on success, e.g. every 5–10 minutes) to catch drift from **someone changing state directly against the server's API/UI**, since gitops correctness means the CR wins, not "first write wins". - **Finalizers** on every CRD that has a server-side counterpart, so deletion calls the corresponding DELETE before the Kubernetes object disappears. Failure to delete server-side (e.g. server unreachable) blocks finalizer removal and surfaces as a `Degraded` condition + event, rather than silently orphaning a row. - **Owner chain for status resolution, not API calls**: `TerdutTeam`'s controller does not call any other controller; every child CRD's controller independently resolves its own `teamRef` → `TerdutTeam.status.teamID` and `serverRef` chain down to `TerdutServer.status.operatorCredentialsSecretRef`, the way any two independent controller-runtime reconcilers would. If the referenced parent isn't `Ready` yet, the child requeues with backoff and reports `Ready: False, reason: WaitingForTeam` — no cross-controller RPC. - **Cross-namespace `serverRef` is re-checked every reconcile, not just at creation**: `TerdutTeam`'s controller reads the target `TerdutServer`'s `spec.allowedTeams` (and, under `Selector`, a `Get` on its own `Namespace` object for labels) on every pass before touching a cross-namespace `TerdutServer` — revocation (§4.6) takes effect on the team's very next reconcile, not just when the CR is first applied. ## 6. Bootstrap & authentication to terdut-server's API terdut-server has no first-class "service account" token type — API keys belong to a real user row (`api_keys.user_id`). The chart's current answer is a one-shot Job that calls `POST /api/bootstrap`, gets a one-time admin key back, and stores it in a Secret (`bootstrap-job.yaml`). The operator absorbs this rather than shelling out to curl: 1. **Confirmed against source** (`internal/api/users.go`'s `handleBootstrap`): `/api/bootstrap` is single-shot *per install*, not per identity — it gates on `SELECT COUNT(*) FROM users`, so any call once one user exists 403s regardless of who's asking, exactly as `charts/terdut-server`'s own `bootstrap-job.yaml` already assumes (403 → "already bootstrapped, nothing to do", exit 0). This rules out the two-identity plan this section originally described: there is no way for the operator to get its *own* bootstrap identity once the chart (or a human) has already bootstrapped the server. The operator's real first-reconcile flow has to be: call `/api/bootstrap` **only if nothing has bootstrapped yet** (an empty-DB fresh install with no chart bootstrap job enabled), and otherwise obtain its credential through a dedicated, repeatable service-account endpoint — see the note at the end of this section. This also surfaces an unhandled race worth designing around explicitly once §10 is settled: if both the chart's bootstrap Job and this controller call `/api/bootstrap` against the same fresh install, exactly one gets the 201 and the other must treat 403 as "someone else already bootstrapped, go get my own credential the other way" rather than as an error. 2. The returned API key is written to a generated Secret (`-operator-credentials`), owner-referenced to the `TerdutServer`, referenced back from `status.operatorCredentialsSecretRef`. 3. Every other **same-namespace** controller (EscalationRule, DeadmanSwitch, AlertSource, and any same-namespace `TerdutTeam`) reads that Secret directly to call the API — never its own credentials. 4. **Cross-namespace `TerdutTeam` never reads the source Secret directly.** Granting arbitrary consuming namespaces `get`/`list`/`watch` on a Secret in the server's namespace would mean *any* future workload in that namespace with Secret-read RBAC could be pointed at it too — `allowedTeams` (§4.1, §4.6) only authorizes the `TerdutTeam` *kind* to resolve a reference, not "read this Secret". Instead, the `TerdutServer` controller (which already holds the real credentials, and whose ServiceAccount is the only thing with legitimate cross-namespace write access — see §9) watches `TerdutTeam` objects across the cluster for ones that both name it and pass its `allowedTeams` check, and mirrors a copy of the credentials Secret into each such namespace on demand — not proactively into every namespace a `Selector`/`All` policy *could* admit, only into ones an actual permitted `TerdutTeam` currently references. The mirror is named `.-terdut-credentials`, owned not by an `OwnerReference` (those can't cross namespaces) but tracked in the `TerdutServer`'s status and cleaned up once no permitted `TerdutTeam` in that namespace references it any more (revoked `allowedTeams`, or the last referencing `TerdutTeam` deleted). The remote `TerdutTeam`'s controller reads only this local mirror, never the original. - This is one shared mirrored Secret per (server, consuming namespace) pair, not one per `TerdutTeam` — terdut-server's own API key isn't team-scoped (§13), so there is nothing finer to hand out; "scoped" here means scoped by *namespace boundary*, not by team permission. **This is the design's real weak point, not the mirroring mechanism itself:** the RBAC argument above (mirror rather than grant broad cross-namespace Secret-read) is sound on its own terms, but every mirrored copy is still server-admin-equivalent regardless of which team's namespace it lands in — `allowedTeams` gates whether a namespace may attach a `TerdutTeam` at all, it does nothing to bound what that namespace's copy of the credential can then do to every *other* team on the same server. This goes away, mirroring included, once team-scoped service- account tokens exist (see the note below): mint one key per `TerdutTeam`, directly into its own namespace, owner-referenced to the CR. No mirror, no shared-per-namespace blast radius — a leaked Secret compromises exactly one team. 5. **Rotation**: the key is a bearer credential with no expiry modeled server-side today. This section's original plan — delete the Secret + the `api_keys` row, let the controller re-bootstrap — **does not work**: deleting an `api_keys` row doesn't reduce `users` to zero, so the next `/api/bootstrap` call still 403s (see point 1 above). Until terdut-server grows a real credential-issuance endpoint, there is no working rotation story here at all; do not implement this as written. 6. **This entire section is a stand-in for a real scoped service-account token type, and more than a nice-to-have**: it's the dependency that makes points 1 and 5 above actually resolvable. Recommended shape (raised as a terdut-server feature request, tracked in §13): a `POST /api/service-accounts` (instance-scoped, admin-only, safely callable repeatedly — unlike `/api/bootstrap`) to create the operator's own identity and mint its first key, `POST /api/service-accounts/{id}/keys` to rotate without recreating the account, and `GET /api/service-accounts?name=` so a 403 from a stale lookup resolves to "fetch my existing account" instead of an unhandled error. Team-scoped accounts (rather than the one instance-scoped operator identity) are what let point 4 above mint a key per `TerdutTeam` instead of mirroring. Until this lands server-side, treat this section's bootstrap flow as v1-blocking, not v1-shippable — see §10's note on sequencing. ## 7. Ownership, status, garbage collection - Every generated object (Deployment, Service, credentials Secret, webhook Secret) carries a `metav1.OwnerReference` to the CR that caused it, in the same namespace — standard GC, no finalizer needed for these (only for the server-side REST resources, per §5). - Status conditions follow the standard `metav1.Condition` shape with at least `Ready` on every kind, plus kind-specific ones (`TerdutServer`: `DatabaseReady`, `Bootstrapped`; children: `Synced`). - `status.observedGeneration` on every kind, bumped only after a successful reconcile against that generation's spec — the standard way a client (or `kubectl wait`) tells "applied" from "seen". - No cluster-scoped aggregation object (e.g. no cluster-wide "all servers" status) in v1 — `kubectl get terdutservers -A` is the aggregate view. ## 8. Postgres integration Mirrors the chart's existing two-path contract (`values.yaml` `database.dsn`/`passwordSecret`), because that contract is already documented and tested operationally: - **Bring-your-own**: `spec.database.dsn` (no password) + `spec.database.passwordSecretRef` — the operator sets `PGPASSWORD` on the Deployment's container env from that Secret, exactly like the chart does today, and does nothing else. No connectivity check beyond what the Deployment's own readiness probe already gives. - **Zalando `postgres-operator`**: `spec.database.postgresClusterRef` names a `postgresql.acid.zalan.do` CR in the same namespace. The controller: - Reads that CR's status for the primary Service name/port to build the DSN host — `..svc:5432` — and database name convention. - Resolves the generated credentials Secret (`..credentials.postgresql.acid.zalan.do`) the same way the chart's comment already documents, and wires it in as `PGPASSWORD` the same way. - Watches that Secret (not just the `postgresql` CR) so a credential rotation triggers a requeue — the chart today requires a manual pod restart for this; the operator can at least detect and report it via a condition even if restarting on rotation is left as a §13 follow-up rather than done automatically (a rolling restart on credential change is a behavior change worth its own design pass, not folded in here). - Requires read RBAC on `postgresql.acid.zalan.do` (optional CRD — the operator's ClusterRole/Role should not hard-fail if the CRD isn't installed and a given `TerdutServer` uses BYO DSN instead). ## 9. RBAC - The operator's own ServiceAccount needs, per namespace it's granted: `get/list/watch/create/update/patch/delete` on `Deployments`, `Services`, `Secrets` it owns, and `get/list/watch` on `postgresql.acid.zalan.do` (optional, degrade gracefully if absent per §8), plus cluster-wide `get/list` on `Namespace` (labels only, for `allowedTeams: {from: Selector}` evaluation — §4.1, §4.6). - No cluster-scoped resources are created by this operator (namespaced CRDs only, per §1) — a `Role` + `RoleBinding` per watched namespace is sufficient; a `ClusterRole` is only needed for watching CRDs across all namespaces, which is the normal Kubebuilder multi-tenant-operator default and doesn't imply cluster-scoped *managed* resources. - **The operator's ServiceAccount is the only thing that ever holds cross-namespace `Secret` write.** It is a single Deployment/binary already watching every namespace it's granted (the normal Kubebuilder shape), so mirroring a credentials Secret into a consenting namespace (§6) is not a new Kubernetes RBAC *boundary* — it's the same ServiceAccount that already reconciles objects there — but it is new *scope* (`create`/`update` on `Secrets` cluster-wide rather than only within each object's own namespace), and should be called out explicitly in the operator's ClusterRole/RBAC review, not left implicit. No human or team's own RBAC is ever granted cross-namespace Secret access by this design — `TerdutServer.spec.allowedTeams` only ever authorizes the operator to act on a `TerdutTeam`'s behalf, never a person or a workload directly. **This is a statement about Kubernetes RBAC only, though — it says nothing about what the mirrored terdut-server credential itself can do once it's there.** Today that credential is the server-admin-equivalent bot key (§6), so a compromised or over-read namespace can reach every team on the server, not just its own; `allowedTeams` bounds who may *attach*, not what an attached namespace's copy of the credential can then *do*. That's a real privilege-boundary gap, and it closes once §6's team-scoped service-account keys exist and mirroring is dropped in favor of a key minted directly per `TerdutTeam`. - terdut-server's own RBAC is unaffected — the operator talks to it purely over HTTP with the bot user's API key, never via the Kubernetes API for app-level state. ## 10. Relationship to `charts/terdut-server` **Recommendation** (flagged explicitly as a decision to confirm before implementation starts, not settled by this document alone): the chart's Deployment/Service/bootstrap-job templates become redundant once `TerdutServer` exists — running both would mean two controllers (Helm and this operator) reconciling the same Deployment, which is exactly the conflict Kubernetes operators exist to avoid. Proposed path: - The chart is repurposed into an **installer chart**: it installs the operator + CRDs (and optionally one `TerdutServer` CR from `values.yaml`, for users who want "helm install and get a server" without hand-writing a CR) rather than templating the Deployment directly. - Existing installs migrate by: `helm template` the current release's `values.yaml` into an equivalent `TerdutServer` CR (mechanical, since §4.1 is deliberately shaped to make that mapping 1:1), install the operator, apply the CR, then let Helm's release be uninstalled or reduced to just the CRD/operator subchart. - This is a breaking change to the chart's contract and needs its own migration guide and probably a major chart version bump — out of scope for this design doc beyond flagging it; do not start that migration work without separately confirming this recommendation. - **This decision isn't only about the migration — it also decides who owns bootstrap.** §6 assumed the operator could always get its own bootstrap identity separately from the chart's; §6 point 1 shows that's false, so whichever of {chart's Job, operator controller} is expected to call `/api/bootstrap` first has to be settled explicitly (a one-paragraph call, not the full migration plan) *before* writing any operator bootstrap/credential code, not deferred alongside the rest of this section. ## 11. Testing strategy - `envtest` (controller-runtime's fake API server) for every controller's reconcile logic against the Kubernetes side. - terdut-server's REST API is faked with a small `httptest.Server` per controller test driven by fixtures matching the real handlers' request/response shapes (already well-documented in `internal/api/*_test.go` on the server side) — no real Postgres or real terdut-server binary needed for controller unit tests. - A smaller number of true end-to-end tests (`kind` cluster + real terdut-server image + real Postgres) covering the golden path per CRD: create `TerdutServer` → `TerdutTeam` → one of each child kind → verify via terdut-server's own API that the objects exist with the right shape → delete the CR → verify the server-side object is gone. ## 12. Observability - Standard controller-runtime metrics (reconcile duration/error counts) are enough for v1 — no custom metrics. - Every externally-visible action (bootstrap, key rotation-needed, delete- and-recreate on the no-PUT resources, adopt-on-conflict) emits a Kubernetes `Event` on the CR, since that's what shows up in `kubectl describe` and gitops tooling (Argo CD/Flux) surfaces without extra wiring. ## 13. Deferred / explicitly out of scope for this design - **A real scoped service-account/token type in terdut-server — not merely deferred, this is v1-blocking for §6 as written** (verified: without it, §6's bootstrap flow has no working credential-rotation path and no clean answer to the chart-vs-operator bootstrap race; see §6 points 1, 5, 6 and §10). Sequence this server-side change *before* implementing the `TerdutServer` controller's bootstrap logic, not after. - **A version-discovery endpoint on terdut-server** (e.g. `GET /api/version`). Neither this operator nor terdut-tui has one today — both independently detect capability by probing specific routes (terdut-tui via `GET /api/teams` 404-checking; this operator would otherwise need to invent its own equivalent probe). An unattended reconciler is more exposed to a silent breaking API change than an interactive TUI a human is watching; raising this alongside the service-account request rather than inventing another route-probe here. - CloudNativePG support — same `spec.database` shape as Zalando should extend to it, but the concrete field/Secret-naming conventions need their own look. - Cross-namespace `teamRef` on the child CRDs (`TerdutEscalationRule`, `TerdutDeadmanSwitch`, `TerdutAlertSource`) — only `TerdutTeam.serverRef` crosses namespaces in v1 (§2, §4.2, §4.6); these stay same-namespace as their `TerdutTeam` until a real need for splitting them out shows up. - Gitops-managed team *membership* (see §4.2). - Automatic Deployment restart on upstream Postgres credential rotation. - Admission webhooks / CEL-only validation limits (e.g. verifying a `teamRef` exists at admission time rather than surfacing it as a status condition after the fact). - OLM packaging, Helm chart migration execution (§10 is a recommendation, not a plan to execute).