Files
terdut-operator/DESIGN.md
T
2026-09-29 15:16:15 +02:00

546 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# terdut-operator design
This is the design reference for implementing terdut-operator. It exists so
implementation can start from settled decisions instead of re-litigating them
mid-PR. It is deliberately more detailed than the README; the README stays as
the short pitch and now points here instead of carrying open questions.
Written against terdut-server as of the Postgres-only, per-team-resources
version (teams, escalation policies, dead man's switches and integrations are
all rows scoped to a team, managed over `internal/api/*` — not env/config-file
driven; see that repo's `charts/terdut-server` for the current deploy story
this operator supersedes).
## 1. Goals & non-goals
**Goal:** let a terdut-server install — the server itself, its teams, their
escalation policies, dead man's switches and alert-source integrations — be
fully described as Kubernetes objects and managed through gitops, following
controller-runtime / Kubebuilder conventions.
**Non-goals (v1):**
- Not a general-purpose Postgres operator. It *consumes* a database that
either the Zalando `postgres-operator` or something else already provides.
- Not managing Alertmanager itself, or the routing rules that decide which
alerts reach which integration webhook — only the terdut-server side
(creating the integration and handing back its URL/key).
- Not OLM packaging. Plain Kubebuilder manifests + Helm chart for
install, matching how terdut-server itself ships.
- Cross-namespace references are limited to exactly one edge:
`TerdutTeam.spec.serverRef` may name a `TerdutServer` in a different
namespace, gated by an explicit consent object in the target namespace
(§4.2, §4.6) — this is the multi-tenant shape the operator exists for
(one platform team owns a `TerdutServer`; other teams self-service a
`TerdutTeam` against it without needing write access to the server's
namespace). Every other reference (`teamRef` on the escalation
rule/dead-man-switch/alert-source CRDs) stays same-namespace-as-its-`TerdutTeam`
only, in v1 — those manage a specific team's own resources and are
expected to live alongside it.
- No validating/mutating admission webhooks in v1. CEL validation rules on
the CRDs (OpenAPI `x-kubernetes-validations`) cover what they can; anything
that needs a live look at another object (e.g. "does this teamRef exist")
is a status condition, not an admission rejection — keeps v1 to a
controller-only deployment with no cert-manager/webhook dependency.
## 2. The README's open questions, resolved
> How do team-crd connect with server-crd?
Explicit `spec.serverRef: {name, namespace}` on `TerdutTeam` — same reasoning
as before (explicit, greppable, trivially validated), but **`namespace` is
deliberately part of the reference**: one team can run and own a
`TerdutServer`, and other teams — in their own namespaces, without any write
access to the server-owning team's namespace — self-service a `TerdutTeam`
against it. `namespace` defaults to the `TerdutTeam`'s own namespace when
omitted, so the common single-tenant case (`serverRef: {name: terdut}`) is
unchanged.
This cross-namespace edge needs the target namespace's explicit consent —
otherwise any namespace in the cluster could point a `TerdutTeam` at
someone else's `TerdutServer` and have the operator provision a team +
credentials against it, which is a namespace-boundary violation, not a
gitops convenience. §4.6 covers the consent object
(`TerdutServerReferenceGrant`), modeled directly on Gateway API's
`ReferenceGrant` — the established Kubernetes pattern for exactly this
"object A in namespace X wants to reference object B in namespace Y"
problem. No other reference in this design (`teamRef` on the child CRDs)
crosses a namespace boundary, so this is the only place `ReferenceGrant`-style
consent is needed (§1, §4.2, §5, §6, §9).
> How does escalationrules, switches and alertsources connect to a team?
Explicit `spec.teamRef: {name}` on each of `TerdutEscalationRule`,
`TerdutDeadmanSwitch`, `TerdutAlertSource` — same reasoning, and it mirrors
terdut-server's own data model, where every one of these rows carries a
`team_id` foreign key already. A matcher/selector on `TerdutTeam` would be
inventing a second source of truth for an association the server already
models as a plain reference.
> Support for both postgres-operator (Zalando) and bring-your-own, how do we
> design that to be user friendly?
`TerdutServer.spec.database` is a oneOf, mirroring the chart's existing
`database.dsn` / `database.passwordSecret` contract (see §8):
- `dsn` + `passwordSecretRef` — bring-your-own, exactly today's chart inputs.
- `postgresClusterRef` — points at a Zalando `postgresql.acid.zalan.do` CR;
the operator derives the DSN and resolves the generated credentials Secret
itself (see §8). CloudNativePG support is a natural follow-up using the
same shape and is called out as deferred (§13), not designed in detail now.
## 3. API group, versions, CRD catalog
- Group: `terdut.ryuvia.com`, version: `v1alpha1` (matches the `ryuvia.com`
domain terdut-server already uses; bump to `v1beta1`/`v1` per the normal
Kubernetes API graduation criteria once the shapes below have proven
stable against a real install).
- Module: `git.ryuvia.com/niklas/terdut-operator`, scaffolded with
Kubebuilder (controller-runtime), matching terdut-server's Go toolchain
and house style.
| Kind | Scope | Purpose |
|---|---|---|
| `TerdutServer` | Namespaced | One terdut-server install: Deployment, Service, database wiring, bootstrap, operator credentials. |
| `TerdutTeam` | Namespaced | One team on a `TerdutServer`, possibly in another namespace: name, OIDC group mapping. |
| `TerdutServerReferenceGrant` | Namespaced (lives with the `TerdutServer`) | Consent for a `TerdutTeam` in another namespace to reference this namespace's `TerdutServer`(s). |
| `TerdutEscalationRule` | Namespaced | A team's escalation policy (levels, targets, repeat). |
| `TerdutDeadmanSwitch` | Namespaced | One dead man's switch on a team. |
| `TerdutAlertSource` | Namespaced | One alert-ingest integration on a team (currently: Alertmanager webhook). |
## 4. Per-CRD spec
### 4.1 `TerdutServer`
```yaml
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
name: terdut
namespace: oncall
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: v0.9.3
replicas: 1 # terdut-server is not horizontally-scale-tested; keep the field, default 1
networking:
hostname: terdut.example.com
servicePort: 8080
gatewayListener: "" # same semantics as chart's networking.listener
database:
dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef
passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password}
# --- OR ---
postgresClusterRef: {name: terdut-postgres} # Zalando `postgresql` CR in the same namespace
sweeper:
staleAfter: 6h
archiveAfter: 168h
deadman:
matchers: "alertname=Watchdog"
timeout: 15m
severity: critical
notify:
ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
fallbackTopic: ""
repeatEvery: 15m
tokenSecretRef: {name: "", key: token}
oidc:
enabled: false
issuer: ""
clientID: ""
clientSecretRef: {name: "", key: client-secret}
name: SSO
scopes: "openid profile email"
allowedGroups: []
adminGroup: ""
sessionMaxAge: 12h
passwordLogin: true
status:
conditions: [...] # Ready, DatabaseReady, Bootstrapped
observedGeneration: 3
serviceName: terdut
operatorCredentialsSecretRef: {name: terdut-operator-credentials} # see §6
```
Field-for-field this is the chart's `values.yaml` reshaped as a spec — the
operator absorbs the chart's Deployment/Service/bootstrap-job templates, so
existing installs have a direct mapping when migrating (see §10).
`spec.database` fields are `+kubebuilder:validation:XValidation` guarded to
be mutually exclusive (`dsn` xor `postgresClusterRef`); mirrors the chart's
"chart provisions no database" stance — this operator provisions no database
either, only wires up one that exists.
### 4.2 `TerdutTeam`
```yaml
spec:
serverRef:
name: terdut
namespace: platform-oncall # optional; defaults to this TerdutTeam's own namespace.
# Cross-namespace requires a matching TerdutServerReferenceGrant
# in that namespace (§4.6) — otherwise Ready: False, reason: RefNotPermitted.
displayName: "Platform" # -> POST /api/teams {"name": ...}; server assigns the ID
oidc:
memberGroup: "terdut-platform-members"
ownerGroup: "terdut-platform-owners"
status:
conditions: [...]
teamID: 42 # the server-side ID; needed by every child object's controller
observedGeneration: 1
```
Team *membership* (which users belong, `team_members`) is explicitly **not**
modeled as a CRD field in v1: terdut-server already manages membership via
OIDC group sync at login for SSO installs, and manual membership for
password-login installs is a people-management action, not infrastructure —
forcing it through gitops would mean a human's team change goes through a PR
review. Flagged in §13 as revisitable if a real gitops-membership need shows up.
### 4.3 `TerdutEscalationRule`
```yaml
spec:
teamRef: {name: platform-team}
repeatCount: 2
fallbackTopic: "platform-oncall"
levels:
- timeout: 5m
targets:
- kind: oncall # "oncall" or "user"
- kind: user
username: alice # resolved to a user ID by the controller at apply time
- timeout: 15m
targets:
- kind: oncall
status:
conditions: [...]
observedGeneration: 1
```
One `TerdutEscalationRule` per team — the server itself models a policy as
one row (`escalation_policies`) with an owned list of levels, so a
one-CRD-to-one-policy mapping (not one-CRD-per-level) matches the server's
own aggregate and lets the whole thing be reconciled with the single
`PUT /api/teams/{teamID}/escalation` the API actually exposes (see §5).
### 4.4 `TerdutDeadmanSwitch`
```yaml
spec:
teamRef: {name: platform-team}
name: "prod-watchdog"
matcher: "alertname=Watchdog,cluster=prod"
timeout: 15m
severity: critical
status:
conditions: [...]
switchID: 7
```
### 4.5 `TerdutAlertSource`
```yaml
spec:
teamRef: {name: platform-team}
kind: alertmanager
name: "prod-alertmanager"
status:
conditions: [...]
integrationID: 3
webhookURLSecretRef: {name: prod-alertmanager-terdut-webhook} # url + key, never in status/spec
```
The integration key is shown by the API exactly once, at creation
(`Integration.Key`/`URL` in terdut-server's own model) — never re-readable,
same shape as the bootstrap admin key. The controller writes it straight into
a generated, owner-referenced Secret on the create it caused and never logs
or stores it anywhere else; the CR's `status` carries only the Secret
reference, matching how e.g. cert-manager's `Certificate` exposes
`spec.secretName` rather than the key material itself.
### 4.6 `TerdutServerReferenceGrant`
Consent object living in the *`TerdutServer`'s* namespace, modeled directly
on Gateway API's `ReferenceGrant` (`gateway.networking.k8s.io/v1beta1`):
without one, no `TerdutTeam` outside this namespace may resolve a
`serverRef` into it, no matter what it names.
```yaml
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServerReferenceGrant
metadata:
name: allow-app-teams
namespace: platform-oncall # the TerdutServer's namespace
spec:
from:
- group: terdut.ryuvia.com
kind: TerdutTeam
namespace: team-checkout # one entry per namespace permitted to reference in
- group: terdut.ryuvia.com
kind: TerdutTeam
namespace: team-payments
to:
- group: terdut.ryuvia.com
kind: TerdutServer
name: terdut # optional: omit to permit any TerdutServer in this namespace
```
- **No selector/wildcard-across-namespaces field** (e.g. no `namespaceSelector`)
in v1, matching upstream `ReferenceGrant`'s own choice: the namespace owner
enumerates exactly which namespaces may reference in, which is the whole
point of requiring affirmative, auditable consent rather than an implicit
or pattern-matched grant.
- A `TerdutTeam`'s controller checks for a matching grant (namespace +
optionally name) on every reconcile before it will resolve `serverRef`
cross-namespace, exactly mirroring how `HTTPRoute`/`Certificate`
controllers check `ReferenceGrant` before honouring a cross-namespace
`backendRef`/`secretRef`. Same-namespace `serverRef` never needs a grant.
- Deleting the grant is a live revocation: the next reconcile of any
`TerdutTeam` it used to authorize finds no matching grant, flips
`Ready: False, reason: RefNotPermitted`, and — deliberately — does **not**
delete the team server-side on revocation alone; it stops reconciling
further changes until access is restored or the `TerdutTeam` CR itself is
deleted (whose finalizer still needs the credentials Secret described in
§6 to clean up, so a revoked grant blocking *new* changes rather than
forcing an immediate, possibly credential-less deletion is the safer
failure mode).
## 5. Reconciliation semantics
terdut-server's REST surface (`internal/api/router.go`) does not give every
resource a full update verb, so reconciliation strategy is per-resource:
| Resource | Verbs available | Strategy |
|---|---|---|
| Team | POST create, PUT rename, DELETE, PUT oidc-groups | Real update-in-place: diff spec vs. last-applied, PUT the changed pieces. |
| Escalation policy | GET/PUT whole-policy | Update-in-place: PUT the full desired policy every reconcile that finds drift; cheap because whole-policy is small and already loaded whole server-side. |
| Dead man's switch | POST create, DELETE — **no PUT** | Delete-and-recreate on any spec diff other than `name`. The controller diffs against `status` (which mirrors what was last successfully applied) rather than re-reading the server every reconcile, to avoid a spurious recreate from field reordering. |
| Integration (alert source) | POST create, PATCH rename, DELETE | Rename via PATCH; any other spec change (kind) is delete-and-recreate, which **rotates the webhook key** — called out loudly in the CRD's field docs and in a `Warning` event, since it breaks whatever sends to the old URL/key until the new Secret is picked up. |
General rules for every controller:
- **Idempotent create**: before POSTing, check `status.<serverSideID>` is
unset; if the server already has a same-named object from a previous
partial reconcile (e.g. after a crash between POST and status-write), treat
a 409/name-conflict as "adopt" — GET-by-name and populate status, rather
than erroring forever. terdut-server's list endpoints in each of these
areas return objects by name, so this is a straightforward correlation.
- **Periodic resync** in addition to watch-triggered reconciles (Kubebuilder
default `RequeueAfter` on success, e.g. every 5–10 minutes) to catch drift
from **someone changing state directly against the server's API/UI**,
since gitops correctness means the CR wins, not "first write wins".
- **Finalizers** on every CRD that has a server-side counterpart, so deletion
calls the corresponding DELETE before the Kubernetes object disappears.
Failure to delete server-side (e.g. server unreachable) blocks finalizer
removal and surfaces as a `Degraded` condition + event, rather than
silently orphaning a row.
- **Owner chain for status resolution, not API calls**: `TerdutTeam`'s
controller does not call any other controller; every child CRD's
controller independently resolves its own `teamRef` → `TerdutTeam.status.teamID`
and `serverRef` chain down to `TerdutServer.status.operatorCredentialsSecretRef`,
the way any two independent controller-runtime reconcilers would. If the
referenced parent isn't `Ready` yet, the child requeues with backoff and
reports `Ready: False, reason: WaitingForTeam` — no cross-controller RPC.
- **Cross-namespace `serverRef` is re-checked every reconcile, not just at
creation**: `TerdutTeam`'s controller lists `TerdutServerReferenceGrant`
objects in the target namespace on every pass (cheap: a namespaced List
with a field/label index, not a full-cluster scan) before touching the
cross-namespace `TerdutServer` — revocation (§4.6) takes effect on the
team's very next reconcile, not just when the CR is first applied.
## 6. Bootstrap & authentication to terdut-server's API
terdut-server has no first-class "service account" token type — API keys
belong to a real user row (`api_keys.user_id`). The chart's current answer is
a one-shot Job that calls `POST /api/bootstrap`, gets a one-time admin key
back, and stores it in a Secret (`bootstrap-job.yaml`). The operator absorbs
this rather than shelling out to curl:
1. On a `TerdutServer`'s first reconcile after its Deployment reports Ready
(`/healthz` reachable through the Service), the controller calls
`POST /api/bootstrap` itself with a fixed, recognizable identity —
e.g. username `terdut-operator`, email `terdut-operator@<serverRef>` —
distinct from `spec.bootstrap.username/email` used for the *human* first
admin the chart bootstraps today. Two separate bootstrap identities are
needed only if terdut-server's `/api/bootstrap` is single-shot per
install; if it already returns 403 after the first caller regardless of
identity, the operator instead **creates its bot user via
`POST /api/users`** once a human admin exists — this needs confirming
against the endpoint's actual behavior before implementation and is
flagged as a first task, not assumed here.
2. The returned API key is written to a generated Secret
(`<name>-operator-credentials`), owner-referenced to the `TerdutServer`,
referenced back from `status.operatorCredentialsSecretRef`.
3. Every other **same-namespace** controller (EscalationRule, DeadmanSwitch,
AlertSource, and any same-namespace `TerdutTeam`) reads that Secret
directly to call the API — never its own credentials.
4. **Cross-namespace `TerdutTeam` never reads the source Secret directly.**
Granting arbitrary consuming namespaces `get`/`list`/`watch` on a Secret
in the server's namespace would mean *any* future workload in that
namespace with Secret-read RBAC could be pointed at it too — the
`TerdutServerReferenceGrant` (§4.6) only authorizes the `TerdutTeam`
*kind*, not "read this Secret". Instead, once a `TerdutServer` observes at
least one valid grant for a namespace, its own controller (which already
holds the real credentials, and whose ServiceAccount is the only thing
with legitimate cross-namespace write access — see §9) mirrors a copy of
the Secret into that consenting namespace, named
`<serverRef.name>.<serverRef.namespace>-terdut-credentials`, owned not by
an `OwnerReference` (those can't cross namespaces) but tracked in the
`TerdutServer`'s status and cleaned up explicitly when the grant
authorizing that namespace is removed. The remote `TerdutTeam`'s
controller reads only this local mirror, never the original.
- This is one shared mirrored Secret per (server, consuming namespace)
pair, not one per `TerdutTeam` — terdut-server's own API key isn't
team-scoped (§13), so there is nothing finer to hand out; "scoped" here
means scoped by *namespace boundary*, not by team permission.
5. **Rotation**: the key is a bearer credential with no expiry modeled
server-side today. Rotation is manual (delete the Secret + the
`api_keys` row via `DELETE /api/users/{id}/api-keys/{keyID}`, let the
controller re-bootstrap) until/unless terdut-server grows key expiry; a
rotation invalidates every mirror too, which the `TerdutServer` controller
re-copies on its own next reconcile.
Documented as an operational runbook note, not automated in v1.
6. This is explicitly a stand-in for a real scoped service-account token
type; see §13.
## 7. Ownership, status, garbage collection
- Every generated object (Deployment, Service, credentials Secret, webhook
Secret) carries a `metav1.OwnerReference` to the CR that caused it, in the
same namespace — standard GC, no finalizer needed for these (only for the
server-side REST resources, per §5).
- Status conditions follow the standard `metav1.Condition` shape with at
least `Ready` on every kind, plus kind-specific ones (`TerdutServer`:
`DatabaseReady`, `Bootstrapped`; children: `Synced`).
- `status.observedGeneration` on every kind, bumped only after a successful
reconcile against that generation's spec — the standard way a client
(or `kubectl wait`) tells "applied" from "seen".
- No cluster-scoped aggregation object (e.g. no cluster-wide "all servers"
status) in v1 — `kubectl get terdutservers -A` is the aggregate view.
## 8. Postgres integration
Mirrors the chart's existing two-path contract (`values.yaml`
`database.dsn`/`passwordSecret`), because that contract is already
documented and tested operationally:
- **Bring-your-own**: `spec.database.dsn` (no password) +
`spec.database.passwordSecretRef` — the operator sets `PGPASSWORD` on the
Deployment's container env from that Secret, exactly like the chart does
today, and does nothing else. No connectivity check beyond what the
Deployment's own readiness probe already gives.
- **Zalando `postgres-operator`**: `spec.database.postgresClusterRef` names a
`postgresql.acid.zalan.do` CR in the same namespace. The controller:
- Reads that CR's status for the primary Service name/port to build the DSN
host — `<cluster>.<namespace>.svc:5432` — and database name convention.
- Resolves the generated credentials Secret
(`<user>.<cluster>.credentials.postgresql.acid.zalan.do`) the same way
the chart's comment already documents, and wires it in as
`PGPASSWORD` the same way.
- Watches that Secret (not just the `postgresql` CR) so a credential
rotation triggers a requeue — the chart today requires a manual pod
restart for this; the operator can at least detect and report it via a
condition even if restarting on rotation is left as a §13 follow-up
rather than done automatically (a rolling restart on credential change
is a behavior change worth its own design pass, not folded in here).
- Requires read RBAC on `postgresql.acid.zalan.do` (optional CRD — the
operator's ClusterRole/Role should not hard-fail if the CRD isn't
installed and a given `TerdutServer` uses BYO DSN instead).
## 9. RBAC
- The operator's own ServiceAccount needs, per namespace it's granted:
`get/list/watch/create/update/patch/delete` on `Deployments`, `Services`,
`Secrets` it owns, and `get/list/watch` on `postgresql.acid.zalan.do`
(optional, degrade gracefully if absent per §8), plus `get/list/watch` on
`TerdutServerReferenceGrant` (§4.6).
- No cluster-scoped resources are created by this operator (namespaced CRDs
only, per §1) — a `Role` + `RoleBinding` per watched namespace is
sufficient; a `ClusterRole` is only needed for watching CRDs across all
namespaces, which is the normal Kubebuilder multi-tenant-operator default
and doesn't imply cluster-scoped *managed* resources.
- **The operator's ServiceAccount is the only thing that ever holds
cross-namespace `Secret` write.** It is a single Deployment/binary already
watching every namespace it's granted (the normal Kubebuilder shape), so
mirroring a credentials Secret into a consenting namespace (§6) is not a
new privilege *boundary* — it's the same ServiceAccount that already
reconciles objects there — but it is new *scope* (`create`/`update` on
`Secrets` cluster-wide rather than only within each object's own
namespace), and should be called out explicitly in the operator's
ClusterRole/RBAC review, not left implicit. No human or team's own
RBAC is ever granted cross-namespace Secret access by this design —
`TerdutServerReferenceGrant` only ever authorizes the operator to act,
never a person.
- terdut-server's own RBAC is unaffected — the operator talks to it purely
over HTTP with the bot user's API key, never via the Kubernetes API for
app-level state.
## 10. Relationship to `charts/terdut-server`
**Recommendation** (flagged explicitly as a decision to confirm before
implementation starts, not settled by this document alone): the chart's
Deployment/Service/bootstrap-job templates become redundant once
`TerdutServer` exists — running both would mean two controllers (Helm and
this operator) reconciling the same Deployment, which is exactly the
conflict Kubernetes operators exist to avoid. Proposed path:
- The chart is repurposed into an **installer chart**: it installs the
operator + CRDs (and optionally one `TerdutServer` CR from `values.yaml`,
for users who want "helm install and get a server" without hand-writing a
CR) rather than templating the Deployment directly.
- Existing installs migrate by: `helm template` the current release's
`values.yaml` into an equivalent `TerdutServer` CR (mechanical, since §4.1
is deliberately shaped to make that mapping 1:1), install the operator,
apply the CR, then let Helm's release be uninstalled or reduced to just
the CRD/operator subchart.
- This is a breaking change to the chart's contract and needs its own
migration guide and probably a major chart version bump — out of scope
for this design doc beyond flagging it; do not start that migration work
without separately confirming this recommendation.
## 11. Testing strategy
- `envtest` (controller-runtime's fake API server) for every controller's
reconcile logic against the Kubernetes side.
- terdut-server's REST API is faked with a small `httptest.Server` per
controller test driven by fixtures matching the real handlers'
request/response shapes (already well-documented in
`internal/api/*_test.go` on the server side) — no real Postgres or real
terdut-server binary needed for controller unit tests.
- A smaller number of true end-to-end tests (`kind` cluster + real
terdut-server image + real Postgres) covering the golden path per CRD:
create `TerdutServer` → `TerdutTeam` → one of each child kind → verify via
terdut-server's own API that the objects exist with the right shape →
delete the CR → verify the server-side object is gone.
## 12. Observability
- Standard controller-runtime metrics (reconcile duration/error counts) are
enough for v1 — no custom metrics.
- Every externally-visible action (bootstrap, key rotation-needed, delete-
and-recreate on the no-PUT resources, adopt-on-conflict) emits a
Kubernetes `Event` on the CR, since that's what shows up in `kubectl
describe` and gitops tooling (Argo CD/Flux) surfaces without extra wiring.
## 13. Deferred / explicitly out of scope for this design
- A real scoped service-account/token type in terdut-server (least-privilege
operator identity instead of a bot admin user) — worth raising as a
terdut-server feature request, not designed here.
- CloudNativePG support — same `spec.database` shape as Zalando should
extend to it, but the concrete field/Secret-naming conventions need their
own look.
- Cross-namespace `teamRef` on the child CRDs (`TerdutEscalationRule`,
`TerdutDeadmanSwitch`, `TerdutAlertSource`) — only `TerdutTeam.serverRef`
crosses namespaces in v1 (§2, §4.2, §4.6); these stay same-namespace as
their `TerdutTeam` until a real need for splitting them out shows up.
- A `TerdutServerReferenceGrant`-style namespace selector (rather than an
enumerated list of namespaces) — deliberately omitted to match upstream
`ReferenceGrant`'s own choice of explicit, auditable consent over
pattern-matching (§4.6); revisit only if enumerating namespaces becomes
genuinely unworkable at scale.
- Gitops-managed team *membership* (see §4.2).
- Automatic Deployment restart on upstream Postgres credential rotation.
- Admission webhooks / CEL-only validation limits (e.g. verifying a
`teamRef` exists at admission time rather than surfacing it as a status
condition after the fact).
- OLM packaging, Helm chart migration execution (§10 is a recommendation,
not a plan to execute).