4007f54279914c920ecc4838968514fb30b32328
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4007f54279 |
Default TerdutServer.spec.replicas to 2 and switch to RollingUpdate
Mirrors charts/terdut-server's own deployment.yaml change: v0.36.0 put the sweeper, the notifier and the migration runner each behind a Postgres advisory lock, and gave incident creation its own conflict resolution, so the Recreate strategy and replicas-stays-at-1 guidance this controller carried (explicitly tracking that chart's comment) are no longer load-bearing. spec.replicas' +kubebuilder:default moves 1 -> 2 (config/crd/bases and the chart's CRD template regenerated via controller-gen and kubebuilder's helm plugin respectively, then hand-verified identical to the generator's own output rather than trusting a bulk regen -- the plugin's --output-dir charts writes a fresh charts/chart scaffold rather than updating charts/terdut-operator in place, so only the diff was taken, not the whole tree). terdutserver_deployment.go's same-value fallback (reachable only for a TerdutServer stored before this default existed) moves with it, and its Strategy changes from Recreate to RollingUpdate with no explicit maxUnavailable/maxSurge -- the 25%/25% default rounds to 0/1 at replicas: 2, already zero-downtime. DESIGN.md's three places asserting multi-replica isn't a supported topology (the illustrative spec.replicas YAML, spec.pod.affinity's rationale, and the HPA deferred-feature note) are corrected to match; the HPA note now gives its own standing reason (no scaling metric or bounds decided yet) rather than a contradiction that no longer holds. The chart's optional terdutServer.replicas sample value moves 1 -> 2 alongside it. image.tag must be v0.36.0 or newer for any of this to hold -- stated in both the CRD field's doc comment and the chart value's comment, not enforced in code, same stance the chart takes on every other version-coupled assumption. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
88172ade29 |
Set the chart's placeholder version to 0.3.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 1m53s
Release / test (push) Successful in 1m45s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 5m27s
Release / scan-image (push) Successful in 2s
|
||
|
|
a0ea13955e |
TerdutTeam: mint and surface a real invite link (spec.invite)
The actual fix for the human-onboarding gap niklas/terdut-server#23 found -- not a terdut-server change at all. A team-scoped credential is already owner-equivalent for POST/GET/DELETE /api/teams/{teamID}/invites (requireTeamOwner's synthetic-membership mechanism, ratified not accidental per that repo's SERVICE-ACCOUNTS.md), and invite redemption bypasses signup_mode entirely -- this TerdutTeam controller just never grew a feature to use either fact. New spec.invite{enabled, role (member|owner, default member), maxUses (1-100, default 1)} and status.inviteSecretRef. The Secret lives in the TerdutTeam's OWN namespace, not the operator's: unlike status.credentialsSecretRef (a durable, high-privilege credential, kept operator-side per DESIGN.md §6), an invite is bounded and limited-use, meant for this namespace's own human operators to read and hand out -- same precedent as TerdutAlertSource's status.webhookURLSecretRef, same- namespace and OwnerReference'd so deleting the TerdutTeam garbage-collects it automatically. internal/controller/terdutteam_invite.go: mints on first spec.invite.enabled, refreshes a day ahead of terdut-server's fixed 7-day TTL (reading the Secret's own stored expiresAt, no extra server round-trip per reconcile), revokes server-side and deletes the Secret when flipped back to false. A lost invite Secret is silently re-minted rather than treated as unrecoverable the way TerdutAlertSource's webhook key is -- nothing external holds a durable dependency on one specific invite link staying stable, it's read once by one human and handed out. New tdclient.Invite/CreateInvite/RevokeInvite. New envtest coverage: mint into the team's own namespace, refresh-before-expiry, revoke-on-disable (internal/controller/terdutteam_controller_test.go's new "spec.invite" Describe block), plus the fake server growing invite support (terdutserver_controller_test.go) -- its handleTeamSubPath dispatcher was split further (deadman switches into their own handleDeadmanSubPath, matching the existing handleIntegrationSubPath precedent) to stay under golangci-lint's gocyclo threshold with the new route added. examples/demo updated to prove this end to end: 02-team-platform.yaml turns on spec.invite; run-demo.sh's bootstrap_login/join_demo_teams (the psql signup_mode flip + a direct team_members INSERT) are replaced by redeem_platform_invite (reads status.inviteSecretRef, a real POST /api/signup with the invite token) and join_payments_team (POST /api/teams/{teamID}/members using Payments' own credential and alice's user id resolved via GET /api/users, deliberately not given its own spec.invite, so the demo shows both onboarding paths this feature unlocks) -- zero kubectl exec/psql calls remain anywhere in the script. README.md's "First login" section rewritten to match; it no longer documents the admin-token curl call that 403s against current terdut-server (niklas/terdut-server#23). Depends on niklas/terdut-server#24 (the callerMayManageServiceAccount fix for terdut-operator#3) being released before this is deployed for real -- not required to build or test this change itself, since the envtest fake never modeled that authorization gap to begin with. |
||
|
|
375b5ed2f7 |
Set the chart's placeholder version to 0.2.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 57s
CI / test (push) Successful in 2m13s
Release / test (push) Successful in 1m41s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m10s
Release / scan-image (push) Successful in 3s
make helm-package passes --version and --app-version from the tag, so these fields decide nothing about what is published -- but a tree heading for v0.2.0 that still says 0.1.2 tells its reader something false. Same as |
||
|
|
6a699d4341 |
Let TerdutServer customize its pod, and never manage its own ingress
spec.pod (api/v1alpha1/terdutserver_types.go): annotations, nodeSelector,
tolerations, affinity, topologySpreadConstraints, resources, pod and
container securityContext, serviceAccountName, extraEnv/extraEnvFrom,
extraVolumes/extraVolumeMounts, imagePullSecrets, and an optional
disruptionBudget. All direct corev1 passthrough -- no wrapper types buy
anything for any of these, matching how CloudNativePG and the Zalando
postgres-operator both expose the same knobs, and matching this repo's
own SweeperSpec precedent ("wrap only when a round-trip through a
different type buys something"). affinity is pure user-supplied
passthrough, not a toggle-plus-generated-default the way a multi-replica
cluster operator's pod anti-affinity usually is: this operator never
auto-generates one, since spec.replicas above 1 isn't a supported
topology (the sweeper/notifier singleton constraint). Considered and
declined for this round: priorityClassName, pod labels beyond
annotations, and a HorizontalPodAutoscaler -- the last of those would
directly contradict the singleton constraint above.
disruptionBudget is the one field here that isn't a plain PodTemplateSpec
knob: when set, the controller now reconciles a PodDisruptionBudget
selecting the TerdutServer's own pods (new terdutserver_pdb.go); clearing
it deletes any it previously created. New RBAC marker on
poddisruptionbudgets to match.
Driven by a public-release pass: looking past this project's own use case
at what a mature, general-purpose operator CRD exposes here (researched
against Zalando postgres-operator and CloudNativePG specifically), not
just the fields this install happened to need.
Separately, and found while answering a question about exposing
TerdutServer through Istio instead of Gateway API: spec.networking's own
doc comment quietly promised a Gateway API HTTPRoute this operator would
build eventually ("a near-term follow-up, not deferred"). That promise is
wrong for a public release -- an operator managing someone's ingress
mechanism for them is a worse default than not touching it at all, and a
surprise HTTPRoute appearing once that follow-up eventually landed would
have been exactly backwards for an Istio (or plain-Ingress, or
intentionally-unexposed) install. Made the non-goal explicit and
permanent instead (DESIGN.md §1), removed the dead `gatewayListener`
field it was the only consumer of (zero runtime call sites anywhere --
setting it already had no effect, so this is a schema cleanup, not a
behavior change), and corrected ROADMAP.md's framing. hostname/servicePort
stay: both are live (TERDUT_PUBLIC_URL, container/Service port), this
operator just never acts on hostname for exposure. Added
examples/networking (Gateway API HTTPRoute, Istio VirtualService) showing
how to expose the plain ClusterIP Service the operator already creates --
outside the operator itself, as illustrations, not as something
examples/demo applies automatically.
No new terdut-server version requirement: both changes are CRD/controller-
only, nothing about the API this operator's bootstrap flow depends on
changed.
|
||
|
|
738c210505 |
Set the chart's placeholder version to 0.1.2
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree heading
for v0.1.2 that still says 0.1.1 tells its reader something false. Same
as
|
||
|
|
0e3118d6aa |
Set the chart's placeholder version to 0.1.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 54s
CI / test (push) Successful in 3m50s
Release / test (push) Successful in 1m44s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 5m47s
Release / scan-image (push) Successful in 5s
Cosmetic: `make helm-package` passes --version/--app-version from the
tag (da48814's own comment), so this field decides nothing about what
gets published. Still done so the tree doesn't say 0.1.0 while heading
for a v0.1.1 release. Cites
|
||
|
|
b4ccdb09d5 |
Stage 5: installer chart + release infra, kind e2e pass through the chart
Release / test (push) Successful in 2m48s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m3s
CI / test (push) Successful in 2m5s
Release / chart (push) Successful in 4s
Release / image (push) Successful in 7m6s
Release / scan-image (push) Failing after 33s
Chart (charts/terdut-operator) generated via kubebuilder's own helm/v2-alpha plugin from config/'s kustomize output -- CRDs + manager Deployment/RBAC come from the same markers every other stage already generates, one source of truth. Hand-added on top: the optional terdutServer values block (DESIGN.md §10's "helm install and get a server" path, off by default) and the release-skill plumbing -- .release.conf, release-vars/helm-lint/push/ helm-package/helm-push/release Makefile targets, .gitea/workflows/release.yaml (test -> image/chart -> scan-image) -- mirroring terdut-server's own shape (registry/namespace convention, multi-arch buildx push, trivy/govulncheck/ gitleaks scans). ci.yaml gains security and chart jobs to match. Two real issues caught while wiring this, fixed before either shipped: - Dockerfile's builder stage didn't pin --platform=$BUILDPLATFORM, which would have made a multi-arch release build fail outright on this org's runners (no binfmt registration) -- same fix terdut-server's own Dockerfile already needed for the same reason. - govulncheck found one real, reachable finding: google.golang.org/grpc v1.82.1 (transitive via controller-runtime's otel exporter), fixed by bumping to v1.83.1. Full golden-path kind e2e pass, this time through `helm install` rather than raw kustomize: TerdutServer (real terdut-server v0.33.0 image) -> TerdutTeam -> one of each child kind, each confirmed Ready and then independently confirmed against terdut-server's own API from inside the cluster (not just the operator's own status). Deleted every CR in reverse order and confirmed server-side cleanup the same independent way for all three child kinds, the team, and the server. No new bugs found -- Stage 1's own kind pass already caught what a real cluster catches that envtest can't. Also dropped the kubebuilder helm plugin's default .github/workflows/ scaffold, same as Stage 0 already did for the main scaffold: this org runs on Gitea, not GitHub. Not done here, deliberately: an actual tagged release. release-preflight found no terdut-operator/ entry under Ryuvia/charts yet to bump -- that one-time wrapper bootstrap is a decision about deploying this operator for real, not a side effect of finishing this stage. make fmt lint test helm-lint build all clean. |