a0ea13955e8ce109f4b9487d8f5471c64ddcc57d
10 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6a699d4341 |
Let TerdutServer customize its pod, and never manage its own ingress
spec.pod (api/v1alpha1/terdutserver_types.go): annotations, nodeSelector,
tolerations, affinity, topologySpreadConstraints, resources, pod and
container securityContext, serviceAccountName, extraEnv/extraEnvFrom,
extraVolumes/extraVolumeMounts, imagePullSecrets, and an optional
disruptionBudget. All direct corev1 passthrough -- no wrapper types buy
anything for any of these, matching how CloudNativePG and the Zalando
postgres-operator both expose the same knobs, and matching this repo's
own SweeperSpec precedent ("wrap only when a round-trip through a
different type buys something"). affinity is pure user-supplied
passthrough, not a toggle-plus-generated-default the way a multi-replica
cluster operator's pod anti-affinity usually is: this operator never
auto-generates one, since spec.replicas above 1 isn't a supported
topology (the sweeper/notifier singleton constraint). Considered and
declined for this round: priorityClassName, pod labels beyond
annotations, and a HorizontalPodAutoscaler -- the last of those would
directly contradict the singleton constraint above.
disruptionBudget is the one field here that isn't a plain PodTemplateSpec
knob: when set, the controller now reconciles a PodDisruptionBudget
selecting the TerdutServer's own pods (new terdutserver_pdb.go); clearing
it deletes any it previously created. New RBAC marker on
poddisruptionbudgets to match.
Driven by a public-release pass: looking past this project's own use case
at what a mature, general-purpose operator CRD exposes here (researched
against Zalando postgres-operator and CloudNativePG specifically), not
just the fields this install happened to need.
Separately, and found while answering a question about exposing
TerdutServer through Istio instead of Gateway API: spec.networking's own
doc comment quietly promised a Gateway API HTTPRoute this operator would
build eventually ("a near-term follow-up, not deferred"). That promise is
wrong for a public release -- an operator managing someone's ingress
mechanism for them is a worse default than not touching it at all, and a
surprise HTTPRoute appearing once that follow-up eventually landed would
have been exactly backwards for an Istio (or plain-Ingress, or
intentionally-unexposed) install. Made the non-goal explicit and
permanent instead (DESIGN.md §1), removed the dead `gatewayListener`
field it was the only consumer of (zero runtime call sites anywhere --
setting it already had no effect, so this is a schema cleanup, not a
behavior change), and corrected ROADMAP.md's framing. hostname/servicePort
stay: both are live (TERDUT_PUBLIC_URL, container/Service port), this
operator just never acts on hostname for exposure. Added
examples/networking (Gateway API HTTPRoute, Istio VirtualService) showing
how to expose the plain ClusterIP Service the operator already creates --
outside the operator itself, as illustrations, not as something
examples/demo applies automatically.
No new terdut-server version requirement: both changes are CRD/controller-
only, nothing about the API this operator's bootstrap flow depends on
changed.
|
||
|
|
b4ccdb09d5 |
Stage 5: installer chart + release infra, kind e2e pass through the chart
Release / test (push) Successful in 2m48s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m3s
CI / test (push) Successful in 2m5s
Release / chart (push) Successful in 4s
Release / image (push) Successful in 7m6s
Release / scan-image (push) Failing after 33s
Chart (charts/terdut-operator) generated via kubebuilder's own helm/v2-alpha plugin from config/'s kustomize output -- CRDs + manager Deployment/RBAC come from the same markers every other stage already generates, one source of truth. Hand-added on top: the optional terdutServer values block (DESIGN.md §10's "helm install and get a server" path, off by default) and the release-skill plumbing -- .release.conf, release-vars/helm-lint/push/ helm-package/helm-push/release Makefile targets, .gitea/workflows/release.yaml (test -> image/chart -> scan-image) -- mirroring terdut-server's own shape (registry/namespace convention, multi-arch buildx push, trivy/govulncheck/ gitleaks scans). ci.yaml gains security and chart jobs to match. Two real issues caught while wiring this, fixed before either shipped: - Dockerfile's builder stage didn't pin --platform=$BUILDPLATFORM, which would have made a multi-arch release build fail outright on this org's runners (no binfmt registration) -- same fix terdut-server's own Dockerfile already needed for the same reason. - govulncheck found one real, reachable finding: google.golang.org/grpc v1.82.1 (transitive via controller-runtime's otel exporter), fixed by bumping to v1.83.1. Full golden-path kind e2e pass, this time through `helm install` rather than raw kustomize: TerdutServer (real terdut-server v0.33.0 image) -> TerdutTeam -> one of each child kind, each confirmed Ready and then independently confirmed against terdut-server's own API from inside the cluster (not just the operator's own status). Deleted every CR in reverse order and confirmed server-side cleanup the same independent way for all three child kinds, the team, and the server. No new bugs found -- Stage 1's own kind pass already caught what a real cluster catches that envtest can't. Also dropped the kubebuilder helm plugin's default .github/workflows/ scaffold, same as Stage 0 already did for the main scaffold: this org runs on Gitea, not GitHub. Not done here, deliberately: an actual tagged release. release-preflight found no terdut-operator/ entry under Ryuvia/charts yet to bump -- that one-time wrapper bootstrap is a decision about deploying this operator for real, not a side effect of finishing this stage. make fmt lint test helm-lint build all clean. |
||
|
|
048f4448c4 |
Stage 4: TerdutAlertSource
CI / test (push) Successful in 1m34s
Covers webhook Secret generation/ownership (DESIGN.md §4.5, §7), the
WebhookSecretLost fail-closed condition, and the kind-change
delete-and-recreate rotation path.
Idempotent-create here is deliberately neither adopt-on-409
(Team/service-account) nor list-and-match-by-name (TerdutDeadmanSwitch):
terdut-server shows the webhook key exactly once, at creation, and never
again, so no server-side lookup could ever recover it after a crash.
Instead the generated webhook Secret itself -- written immediately after
the POST, before status is ever touched -- is this CR's only durable
record that a create already succeeded; found with status.integrationID
still unset on a later reconcile, it's read back directly rather than
POSTing a second, orphaned integration. Found missing with
status.integrationID *set* instead, that's the already-designed
WebhookSecretLost case: fail closed, not self-healed, since the key is
genuinely gone and recreating it would rotate a live webhook URL with no
spec change to explain why.
Renaming (PATCH) never touches the key, so it's applied unconditionally
every reconcile, same as the escalation policy's whole-policy PUT. A
spec.kind change is the one case with no in-place update verb at all:
DELETE the old integration, delete the stale webhook Secret, then run the
same create path fresh -- fires a Warning event since this breaks whatever
still sends to the old URL.
Also: fakeTerdutServer grows POST/PATCH/DELETE .../integrations routes
behind a new handleIntegrationSubPath, split out of handleTeamSubPath to
stay under gocyclo's threshold; three goconst-flagged test literals
("does-not-exist", "unready") and one unparam-flagged test helper
parameter (bootstrapReadyTerdutServer's always-"default" namespace) get
shared/removed now that a fourth same-shaped caller made the repetition
concrete enough for the linter to flag.
DESIGN.md §13 gains one honest gap found while grounding this stage, not
introduced by it: no child CRD specially detects a mid-life teamRef
change; all three always resolve spec.teamRef fresh and trust the
already-stored server-side id remains valid there.
make fmt lint test build all clean; internal/controller envtest coverage
holds at 71.6%.
|
||
|
|
b0d50e305a |
ROADMAP.md: mark Stage 3 done, fix stale dead man's switch reconciliation note
CI / test (push) Successful in 1m35s
|
||
|
|
b263b48510 |
ROADMAP.md: mark Stage 1's kind e2e pass done, with what it found
CI / test (push) Successful in 1m39s
|
||
|
|
8064876cb1 |
Stage 1: TerdutServer full lifecycle (Deployment, Service, both database
CI / test (push) Successful in 1m46s
paths, self-registration bootstrap)
Replaces the bring-your-own-only Stage 1 (commit
|
||
|
|
fc68ee7256 |
ROADMAP.md: merge Stage 1 + old Stage 5, renumber
CI / test (push) Successful in 1m33s
Follows DESIGN.md's redesign (previous commit): with no hand-deployed
server to prove the simpler CRDs against, there's no reason left to defer
TerdutServer's Deployment/Service/database management behind a separate
later stage. Stage 1 now covers TerdutServer's full lifecycle --
Deployment, Service, both Postgres paths from §8 at once (bring-your-own
DSN and Zalando, per the user's call, not sequenced), bootstrap,
credentials -- built together, since bootstrap only has something to
bootstrap once the Deployment exists.
Old Stage 5 (TerdutServer absorbs Deployment/Service/bootstrap) is gone,
folded into Stage 1. Old Stage 6 (installer chart + release) renumbers to
Stage 5. Stages 2-4 (TerdutTeam, EscalationRule+DeadmanSwitch,
AlertSource) are unchanged in content, renumbering only where old Stage 5
disappears from ahead of them.
Explicitly supersedes the Stage 1 shipped before this redesign (commit
|
||
|
|
ffc2e6441e |
ROADMAP.md Stage 1: scope down to bring-your-own only, defer self-registration to Stage 5
CI / test (push) Successful in 1m23s
Stage 1's own setup (chart bootstraps before the CR exists) never exercises the self-registration fallback, and the narrowed spec has nowhere to put the username/email /api/bootstrap needs anyway. Matches DESIGN.md's §6 rewrite (bring-your-own is now the primary path, not an equal alternative). |
||
|
|
ba253b7bf7 |
Add CI, repo CLAUDE.md, and finish Stage 0
- .gitea/workflows/ci.yaml: fmt/lint/test, same no-actions/checkout-and-manual-clone
shape as terdut-server's ci.yaml, and the same reasoning for why (Node/ES2022
incompatibility on the runner image). No chart/security jobs yet -- nothing for
either to check until Stage 6 / real controller code exists.
- CLAUDE.md: Checks + Release sections, matching the sibling repos' convention from
the workspace-level CLAUDE.md ("each repo has its own CLAUDE.md... read it before
working in that repo"). Release is explicitly marked not-wired-yet rather than
copying terdut-server's, since there's no chart to release against until Stage 6.
- ROADMAP.md: moved the .release.conf bullet out of Stage 0 (it names a HELM_CHART
this repo doesn't have yet) -- it was already duplicated into Stage 6, which is
where it actually belongs.
Stage 0 done: `make fmt lint test` verified green locally. Real open question the CI
workflow's comments flag rather than assume past: whether storage.googleapis.com
(envtest's binary source) is reachable from this Gitea runner's container network the
way proxy.golang.org is -- terdut-server's own ci.yaml notes get.helm.sh/github.com are
not. Only running the workflow for real will confirm; the comment names the fallback
(move the job out of `container:`, like terdut-server's chart job) if it isn't.
|
||
|
|
ee39b8e668 |
Add build roadmap
Stages the operator's implementation: TerdutServer stays bootstrap/credentials-only (no Deployment/Service takeover) until Stage 5, so every earlier stage targets a hand-deployed terdut-server in a disposable dev namespace instead of forcing the chart-migration decision (§10) up front. |