Files
terdut-operator/ROADMAP.md
T
Niklas Ye 6a699d4341 Let TerdutServer customize its pod, and never manage its own ingress
spec.pod (api/v1alpha1/terdutserver_types.go): annotations, nodeSelector,
tolerations, affinity, topologySpreadConstraints, resources, pod and
container securityContext, serviceAccountName, extraEnv/extraEnvFrom,
extraVolumes/extraVolumeMounts, imagePullSecrets, and an optional
disruptionBudget. All direct corev1 passthrough -- no wrapper types buy
anything for any of these, matching how CloudNativePG and the Zalando
postgres-operator both expose the same knobs, and matching this repo's
own SweeperSpec precedent ("wrap only when a round-trip through a
different type buys something"). affinity is pure user-supplied
passthrough, not a toggle-plus-generated-default the way a multi-replica
cluster operator's pod anti-affinity usually is: this operator never
auto-generates one, since spec.replicas above 1 isn't a supported
topology (the sweeper/notifier singleton constraint). Considered and
declined for this round: priorityClassName, pod labels beyond
annotations, and a HorizontalPodAutoscaler -- the last of those would
directly contradict the singleton constraint above.

disruptionBudget is the one field here that isn't a plain PodTemplateSpec
knob: when set, the controller now reconciles a PodDisruptionBudget
selecting the TerdutServer's own pods (new terdutserver_pdb.go); clearing
it deletes any it previously created. New RBAC marker on
poddisruptionbudgets to match.

Driven by a public-release pass: looking past this project's own use case
at what a mature, general-purpose operator CRD exposes here (researched
against Zalando postgres-operator and CloudNativePG specifically), not
just the fields this install happened to need.

Separately, and found while answering a question about exposing
TerdutServer through Istio instead of Gateway API: spec.networking's own
doc comment quietly promised a Gateway API HTTPRoute this operator would
build eventually ("a near-term follow-up, not deferred"). That promise is
wrong for a public release -- an operator managing someone's ingress
mechanism for them is a worse default than not touching it at all, and a
surprise HTTPRoute appearing once that follow-up eventually landed would
have been exactly backwards for an Istio (or plain-Ingress, or
intentionally-unexposed) install. Made the non-goal explicit and
permanent instead (DESIGN.md §1), removed the dead `gatewayListener`
field it was the only consumer of (zero runtime call sites anywhere --
setting it already had no effect, so this is a schema cleanup, not a
behavior change), and corrected ROADMAP.md's framing. hostname/servicePort
stay: both are live (TERDUT_PUBLIC_URL, container/Service port), this
operator just never acts on hostname for exposure. Added
examples/networking (Gateway API HTTPRoute, Istio VirtualService) showing
how to expose the plain ClusterIP Service the operator already creates --
outside the operator itself, as illustrations, not as something
examples/demo applies automatically.

No new terdut-server version requirement: both changes are CRD/controller-
only, nothing about the API this operator's bootstrap flow depends on
changed.
2026-10-02 18:50:38 +02:00

14 KiB

terdut-operator build roadmap

This is the staging plan for implementing the operator against DESIGN.md's settled decisions. It exists for the same reason DESIGN.md and SERVICE-ACCOUNTS.md do: so each stage starts from an agreed sequencing instead of re-litigating "what do we build first" mid-PR.

Sequencing call this roadmap makes

The operator creates and owns every TerdutServer it manages — it never adopts one deployed independently, by hand or by charts/terdut-server (DESIGN.md §1). An earlier version of this roadmap staged TerdutServer's Deployment/Service/bootstrap takeover separately (old Stage 5), behind a hand-deployed server the simpler CRDs could be proven against first. That staging existed only because a credential-less operator couldn't /api/bootstrap its way into a server something else had already bootstrapped (DESIGN.md §6's original gap). With no server to adopt at all, that split has nothing left to justify it: TerdutServer now builds its full lifecycle — Deployment, Service, database wiring, bootstrap, credentials — in one stage, Stage 1, since bootstrap only has something to bootstrap once the Deployment exists.

Stage 0 — Scaffolding & CI

  • go.mod (git.ryuvia.com/niklas/terdut-operator) + Kubebuilder v4 scaffold (cmd/main.go, config/, Makefile, PROJECT), matching terdut-server's Go toolchain and house style (§3). Kubebuilder's own scaffolded Makefile already wires manifests/generate (controller-gen) and setup-envtest into test, and golangci-lint into lint, all fetched on demand into bin/ — no separate install step needed beyond what make test/make lint already do.
  • .release.conf deliberately not added yet: it names a HELM_CHART this repo doesn't have until Stage 5. Adding it now would either be a stub that lies about what's releasable or dead config nobody can run — it lands in Stage 5, alongside the chart it describes.
  • Gitea Actions CI calling fmt lint test, mirroring terdut-server's ci.yaml convention (its CLAUDE.md: "a green gate here and a green pipeline are the same code, not two descriptions of it") minus the chart/security jobs, which need a chart (Stage 5) and real controller code (Stage 1+) respectively to have anything to check.
  • Drop kubebuilder's default .github/workflows/* scaffold — this org runs on Gitea, not GitHub; .gitea/workflows/ci.yaml is the only CI this repo has.
  • Housekeeping: drop the stray .DESIGN.md.swp (leftover vim swapfile, shouldn't be committed); correct DESIGN.md §6/§13's "v1-blocking, not v1-shippable" language — the service-account feature it was blocking on has since shipped in terdut-server.

Done when: CI is green on an otherwise-empty scaffold.

Stage 1 — TerdutServer, full lifecycle

Supersedes the Stage 1 shipped before this redesign (commit 1be7cf2) outright — that TerdutServerSpec/Status/controller/tests implemented the now-removed bring-your-own path and get replaced wholesale, not extended. New commits build forward over the old ones; no git history rewrite.

  • Full §4.1 spec: image, replicas, networking, database, sweeper, deadman, notify, oidc, passwordLogin, allowedTeams, all together — no narrowing, since bootstrap needs the Deployment it's narrowed away from in the version this replaces.
  • Controller manages the Deployment + Service, both Postgres paths from §8 at once (bring-your-own DSN and the Zalando postgres-operator postgresClusterRef integration — not sequenced, per the user's call), and bootstrap/credentials per §6's self-registration flow: /api/bootstrap once the Deployment has a ready replica, checkpoint the admin key, mint the instance-scoped service account, generated credentials Secret in the operator's own namespace, status.credentialsSecretRef. Finalizer cleans up that Secret (and the checkpoint, if one's still there) on delete — there's no server-side row to clean up alongside it: terdut-server's API has no way to delete a user or a service account, only to revoke individual keys, so there's nothing to undo there regardless.
  • RBAC: read-only watch on postgresql.acid.zalan.do, degrading gracefully if that CRD isn't installed (§8, §9).
  • Shipped, scoped down from §8's full ambition in one way, called out in code rather than silently dropped: no live watch on the Zalando- generated credentials Secret for rotation (relies on the periodic resync to notice eventually, higher latency than a watch). A near-term follow-up, not deferred to a later stage.
  • External exposure (a Gateway API HTTPRoute from spec.networking.hostname) was originally sketched here too, as a second near-term follow-up alongside the one above. It's since become an explicit, permanent non-goal instead (DESIGN.md §1): the operator will never manage ingress/exposure for TerdutServer in any form. See examples/networking for how to do that yourself.
  • envtest covering Deployment/Service reconciliation and both database paths — the Zalando path needs that CRD's schema vendored into the test environment (there's no real postgres-operator controller in envtest, only the CRD shape to create fixture objects against) — plus the self-registration flow against an httptest.Server fake of /api/bootstrap and /api/service-accounts (§11), including the adopt-on-409 recovery path and the one fail-closed case (BootstrapStateLost), not just the happy path.
  • Done, 2026-10-01: a real kind end-to-end pass (bring-your-own DSN, real terdut-server v0.33.0 image, operator built into a real image and deployed as a real Pod, not go run against the cluster). TerdutServer went Ready; the generated credential authenticated and exercised its real capability against the actual server (GET/POST /api/teams → 200/201, confirmed from terdut-server's own access log). Caught two real bugs no envtest suite could have (its client bypasses RBAC): .dockerignore's !**/*.go not working under podman, and missing RBAC for events.k8s.io (the new events API GetEventRecorder uses) — both fixed. The Zalando path can additionally be validated for real against the org's own cluster later, where postgres-operator already runs, rather than only in a disposable kind stand-in — not done in this pass.

Stage 2 — TerdutTeam

  • §4.2: serverRef resolution, real update-in-place (POST create / PUT rename / PUT oidc-groups), team-scoped service-account minting once Ready (§6 point 3), finalizer that DELETEs the team server-side and its credential Secret.
  • First place the "every child resolves its own teamRef → TerdutTeam.status, never chains up to TerdutServer" pattern (§5) gets proven end to end.

Stage 3 — TerdutEscalationRule + TerdutDeadmanSwitch

  • Built together: both stay same-namespace-as-their-TerdutTeam (§1), so neither exercises cross-namespace complexity, but together they cover the two different reconciliation shapes §5's table calls out — whole-policy PUT-upsert for the escalation policy (no separate create step at all), real create/update-in-place/delete for the dead man's switch (PUT added in terdut-server v0.33.0 specifically for this operator) — against the same shared create/finalizer/resync scaffolding Stage 2 already built.
  • New shared resolveTeamAndClient helper (childref.go) implements §5's "every child resolves its own teamRef → TerdutTeam.status, never chains up to TerdutServer" rule once, for both controllers — TerdutTeam.status.serverEndpoint, added in this stage, is what makes that literally true rather than just a stated intent.
  • TerdutEscalationRule resolves each "user" target's username to a user_id via GET /api/users (confirmed open to any authenticated caller) and reports Ready: False, reason: UnknownUser if it doesn't resolve. No DELETE exists for this resource, so its delete path PUTs an empty policy as the closest available undo.
  • TerdutDeadmanSwitch has no unique-name constraint server-side, so its idempotent-create is GET-list-and-match-by-name rather than adopt-on-409 (unlike every other resource in this operator).
  • Done, 2026-10-01: envtest coverage for both controllers' happy path, TeamRefNotFound/WaitingForTeam, UnknownUser, list-and-match adoption, update-in-place on spec drift, and deletion. make fmt lint test build all clean; internal/controller envtest coverage 50.5% → 71.7%. No kind e2e pass for this stage — Stage 1's already proved the real-cluster mechanics (RBAC, image, bootstrap) these two controllers reuse unchanged, and neither introduces a new mechanism that pass would exercise differently (same reasoning Stage 2 used to skip one).

Stage 4 — TerdutAlertSource

  • Last of the children on purpose: it has the subtlest failure mode of the four. Covers webhook Secret generation/ownership (§4.5, §7), the WebhookSecretLost fail-closed condition + Warning event (§5, added 2026-09-30), and the kind-change delete-and-recreate rotation path — all easier to get right with the other three controllers' patterns already in place to build on.
  • Idempotent-create here is neither adopt-on-409 (Team/service-account) nor list-and-match-by-name (TerdutDeadmanSwitch): terdut-server shows the webhook key exactly once, at creation, and never again, so no server-side lookup could ever recover it after a crash. The generated webhook Secret itself — written immediately after the POST, before status is ever touched — is this CR's only durable record that a create already succeeded; found on a later reconcile with status.integrationID still unset, it's read back directly rather than POSTing again. Found missing with status.integrationID set, that's the already-designed WebhookSecretLost fail-closed case instead.
  • Done, 2026-10-01: envtest coverage for the happy path, rename (PATCH, no key rotation), a kind change (delete-and-recreate, new id and key), crash recovery between POST and the Secret write, WebhookSecretLost, TeamRefNotFound/WaitingForTeam, and deletion. make fmt lint test build all clean; internal/controller envtest coverage holds at 71.6%. No kind e2e pass for this stage, same reasoning as Stage 3 (reuses Stage 1's already-proven real-cluster mechanics unchanged).

Stage 5 — Installer chart + real release

  • Package CRDs + the operator's own Deployment/RBAC into the installer chart §10 describes; wire .release.conf/release-vars the same way terdut-server does; run it through the release skill for a real first cut.
  • Full kind end-to-end test per §11: create TerdutServer → TerdutTeam → one of each child kind → verify via terdut-server's own API that each object exists with the right shape → delete the CR → verify the server-side object is gone.
  • Chart built via kubebuilder's own helm/v2-alpha plugin from config/'s kustomize output (charts/terdut-operator), not hand-rolled -- CRDs + manager Deployment/RBAC come from the same markers/manifests every other stage already generates, so there's exactly one source of truth for them. Hand-added on top: the optional terdutServer values block (§10's "helm install and get a server" path), .release.conf, and the release-vars/helm-lint/push/helm-package/helm-push/release Makefile targets .gitea/workflows/release.yaml calls, mirroring terdut-server's own shape end to end (same registry/namespace convention, same multi-arch buildx push, same trivy/govulncheck/gitleaks scans). Also fixed while wiring this: the Dockerfile's builder stage didn't pin --platform=$BUILDPLATFORM, which would have made a multi-arch release build fail outright on this org's runners (no binfmt registration) -- caught before it ever shipped, not discovered mid-release; and govulncheck surfaced one real, reachable finding (google.golang.org/grpc v1.82.1, transitive via controller-runtime's otel exporter), fixed by bumping to v1.83.1.
  • Done, 2026-10-01: the full golden-path pass above, run for real against a kind cluster, installed via helm install (not raw kustomize/kubectl apply -- the first time the chart itself, not just config/, was exercised): TerdutServer (real terdut-server v0.33.0 image, bring-your-own DSN against a throwaway in-cluster Postgres) → TerdutTeam → one TerdutEscalationRule + TerdutDeadmanSwitch + TerdutAlertSource, each confirmed Ready and then confirmed a second way, independent of the operator's own status: a curl pod inside the cluster, authenticated with the generated team credential, hit terdut-server's real API directly (GET /api/teams/{id}/escalation, .../deadman/switches, .../integrations) and got back exactly the policy/switch/integration each spec declared. Deleting every CR in reverse order was verified the same way: the escalation policy came back empty (its only available "undo"), the switch and the integration were both gone from their list endpoints, the team no longer resolved by name, and the Deployment/Service/every generated Secret were gone from the cluster. No new bugs found this pass -- Stage 1's own kind e2e pass already caught the two issues (events.k8s.io RBAC, the podman .dockerignore fix) a real cluster catches and envtest can't, and nothing since has touched that surface.
  • Not done in this pass, deliberately: an actual tagged release. make release-vars/helm-lint/push/helm-package/helm-push all work locally and .gitea/workflows/release.yaml is wired, but release-preflight found there is no terdut-operator/ entry under Ryuvia/charts yet to bump -- every other onboarded repo had that one-time wrapper-chart bootstrap done for it before its own first release, and this one doesn't, since deploying this operator for real is a decision for whoever runs the cluster, not a side effect of finishing this stage. Cutting the first real release (and creating that wrapper entry) is therefore the next action, not yet taken.

Deferred (§13, unchanged by this roadmap)

Cross-namespace allowedTeams exercised against a real second namespace, CloudNativePG support, narrower-than-namespace Secret RBAC for the webhook Secret, gitops-managed team membership, automatic Deployment restart on upstream Postgres credential rotation, admission webhooks/CEL-only validation limits, OLM packaging.