49 Commits

Author SHA1 Message Date
Niklas Ye d50248531c Set the chart's placeholder version to 0.4.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 2m10s
Release / test (push) Successful in 7m54s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m25s
Release / scan-image (push) Successful in 36s
v0.4.0
2026-10-03 16:22:15 +02:00
Niklas Ye 4007f54279 Default TerdutServer.spec.replicas to 2 and switch to RollingUpdate
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Has been cancelled
Mirrors charts/terdut-server's own deployment.yaml change: v0.36.0 put
the sweeper, the notifier and the migration runner each behind a
Postgres advisory lock, and gave incident creation its own conflict
resolution, so the Recreate strategy and replicas-stays-at-1 guidance
this controller carried (explicitly tracking that chart's comment)
are no longer load-bearing.

spec.replicas' +kubebuilder:default moves 1 -> 2 (config/crd/bases and
the chart's CRD template regenerated via controller-gen and
kubebuilder's helm plugin respectively, then hand-verified identical
to the generator's own output rather than trusting a bulk regen --
the plugin's --output-dir charts writes a fresh charts/chart scaffold
rather than updating charts/terdut-operator in place, so only the
diff was taken, not the whole tree). terdutserver_deployment.go's
same-value fallback (reachable only for a TerdutServer stored before
this default existed) moves with it, and its Strategy changes from
Recreate to RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicas: 2, already
zero-downtime.

DESIGN.md's three places asserting multi-replica isn't a supported
topology (the illustrative spec.replicas YAML, spec.pod.affinity's
rationale, and the HPA deferred-feature note) are corrected to match;
the HPA note now gives its own standing reason (no scaling metric or
bounds decided yet) rather than a contradiction that no longer holds.

The chart's optional terdutServer.replicas sample value moves 1 -> 2
alongside it. image.tag must be v0.36.0 or newer for any of this to
hold -- stated in both the CRD field's doc comment and the chart
value's comment, not enforced in code, same stance the chart takes on
every other version-coupled assumption.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:35:18 +02:00
Niklas Ye 88172ade29 Set the chart's placeholder version to 0.3.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 1m53s
Release / test (push) Successful in 1m45s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 5m27s
Release / scan-image (push) Successful in 2s
v0.3.0
2026-10-02 22:19:18 +02:00
niklas 478ae6284a Merge pull request 'TerdutTeam: mint and surface a real invite link (spec.invite)' (#5) from terdutteam-invite-minting into main
CI / test (push) Has been cancelled
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
2026-10-02 20:17:21 +00:00
niklas fcc80b32d5 Merge pull request 'examples/demo: add run-demo.sh, an automated kind-cluster demo' (#4) from examples-demo/run-demo-script into main
CI / chart (push) Successful in 2s
CI / security (push) Successful in 1m31s
CI / test (push) Successful in 3m30s
2026-10-02 20:13:36 +00:00
Niklas Ye 4aa4f17c42 examples/demo: bump terdut-server to v0.34.0 (the service-account fix)
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m18s
CI / test (pull_request) Successful in 3m40s
Required for this demo to actually exercise the fix for
niklas/terdut-operator#3 -- v0.33.2 still has the authorization gap this
demo hit live (callerMayManageServiceAccount had no branch letting an
instance-scoped account adopt a team-scoped account's key).
2026-10-02 22:13:26 +02:00
Niklas Ye a0ea13955e TerdutTeam: mint and surface a real invite link (spec.invite)
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 52s
CI / test (pull_request) Successful in 2m43s
The actual fix for the human-onboarding gap niklas/terdut-server#23 found --
not a terdut-server change at all. A team-scoped credential is already
owner-equivalent for POST/GET/DELETE /api/teams/{teamID}/invites
(requireTeamOwner's synthetic-membership mechanism, ratified not
accidental per that repo's SERVICE-ACCOUNTS.md), and invite redemption
bypasses signup_mode entirely -- this TerdutTeam controller just never
grew a feature to use either fact.

New spec.invite{enabled, role (member|owner, default member), maxUses
(1-100, default 1)} and status.inviteSecretRef. The Secret lives in the
TerdutTeam's OWN namespace, not the operator's: unlike
status.credentialsSecretRef (a durable, high-privilege credential, kept
operator-side per DESIGN.md §6), an invite is bounded and limited-use,
meant for this namespace's own human operators to read and hand out --
same precedent as TerdutAlertSource's status.webhookURLSecretRef, same-
namespace and OwnerReference'd so deleting the TerdutTeam garbage-collects
it automatically.

internal/controller/terdutteam_invite.go: mints on first spec.invite.enabled,
refreshes a day ahead of terdut-server's fixed 7-day TTL (reading the
Secret's own stored expiresAt, no extra server round-trip per reconcile),
revokes server-side and deletes the Secret when flipped back to false. A
lost invite Secret is silently re-minted rather than treated as
unrecoverable the way TerdutAlertSource's webhook key is -- nothing
external holds a durable dependency on one specific invite link staying
stable, it's read once by one human and handed out.

New tdclient.Invite/CreateInvite/RevokeInvite. New envtest coverage: mint
into the team's own namespace, refresh-before-expiry, revoke-on-disable
(internal/controller/terdutteam_controller_test.go's new "spec.invite"
Describe block), plus the fake server growing invite support
(terdutserver_controller_test.go) -- its handleTeamSubPath dispatcher was
split further (deadman switches into their own handleDeadmanSubPath,
matching the existing handleIntegrationSubPath precedent) to stay under
golangci-lint's gocyclo threshold with the new route added.

examples/demo updated to prove this end to end: 02-team-platform.yaml
turns on spec.invite; run-demo.sh's bootstrap_login/join_demo_teams (the
psql signup_mode flip + a direct team_members INSERT) are replaced by
redeem_platform_invite (reads status.inviteSecretRef, a real POST
/api/signup with the invite token) and join_payments_team (POST
/api/teams/{teamID}/members using Payments' own credential and alice's
user id resolved via GET /api/users, deliberately not given its own
spec.invite, so the demo shows both onboarding paths this feature
unlocks) -- zero kubectl exec/psql calls remain anywhere in the script.
README.md's "First login" section rewritten to match; it no longer
documents the admin-token curl call that 403s against current
terdut-server (niklas/terdut-server#23).

Depends on niklas/terdut-server#24 (the callerMayManageServiceAccount fix
for terdut-operator#3) being released before this is deployed for real --
not required to build or test this change itself, since the envtest fake
never modeled that authorization gap to begin with.
2026-10-02 22:04:56 +02:00
Niklas Ye 2a08a8cd8e examples/demo: add run-demo.sh, an automated kind-cluster demo
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 1m5s
CI / test (pull_request) Successful in 2m49s
One script, two modes (run-demo.sh / run-demo.sh --teardown), that takes a
fresh empty kind cluster all the way to a working demo: creates the
cluster if needed, helm-installs this chart, applies every CRD kind in
this directory, waits for all nine objects to go Ready, then does what
the README's own first-login section cannot (see niklas/terdut-server#23
and niklas/terdut-operator#3 -- no service-account credential this
operator holds can ever call /api/admin/settings or POST /api/users) by
reaching into the demo's own throwaway Postgres directly: flips
signup_mode to open, signs alice up for real over the ordinary signup
endpoint, and joins her to both Platform and Payments (open signup always
creates its own new team, never joins an existing one by name, so
without this she'd have a working login that can't see a single incident
this demo fires -- /api/incidents and /api/alerts are both scoped to the
caller's own team memberships). Finishes by port-forwarding the service
and firing fire-alerts.sh at both teams, so a fresh run already has
visible incidents waiting in the web UI.

Verified end to end against a real kind cluster, including a second,
genuinely-fresh run that hit niklas/terdut-operator#3 live (terdutteam-
platform wedged in the 403 retry loop that issue describes) -- confirmed
the script itself fails cleanly on that (clear FAILED message, correct
exit code, no orphaned port-forward) rather than hanging or leaving a
mess, which is the most this script can do about a bug in the operator
it's driving.
2026-10-02 21:23:44 +02:00
Niklas Ye 375b5ed2f7 Set the chart's placeholder version to 0.2.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 57s
CI / test (push) Successful in 2m13s
Release / test (push) Successful in 1m41s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m10s
Release / scan-image (push) Successful in 3s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree
heading for v0.2.0 that still says 0.1.2 tells its reader something
false. Same as 738c210 and 0e3118d before it.
v0.2.0
2026-10-02 18:51:00 +02:00
Niklas Ye 6a699d4341 Let TerdutServer customize its pod, and never manage its own ingress
spec.pod (api/v1alpha1/terdutserver_types.go): annotations, nodeSelector,
tolerations, affinity, topologySpreadConstraints, resources, pod and
container securityContext, serviceAccountName, extraEnv/extraEnvFrom,
extraVolumes/extraVolumeMounts, imagePullSecrets, and an optional
disruptionBudget. All direct corev1 passthrough -- no wrapper types buy
anything for any of these, matching how CloudNativePG and the Zalando
postgres-operator both expose the same knobs, and matching this repo's
own SweeperSpec precedent ("wrap only when a round-trip through a
different type buys something"). affinity is pure user-supplied
passthrough, not a toggle-plus-generated-default the way a multi-replica
cluster operator's pod anti-affinity usually is: this operator never
auto-generates one, since spec.replicas above 1 isn't a supported
topology (the sweeper/notifier singleton constraint). Considered and
declined for this round: priorityClassName, pod labels beyond
annotations, and a HorizontalPodAutoscaler -- the last of those would
directly contradict the singleton constraint above.

disruptionBudget is the one field here that isn't a plain PodTemplateSpec
knob: when set, the controller now reconciles a PodDisruptionBudget
selecting the TerdutServer's own pods (new terdutserver_pdb.go); clearing
it deletes any it previously created. New RBAC marker on
poddisruptionbudgets to match.

Driven by a public-release pass: looking past this project's own use case
at what a mature, general-purpose operator CRD exposes here (researched
against Zalando postgres-operator and CloudNativePG specifically), not
just the fields this install happened to need.

Separately, and found while answering a question about exposing
TerdutServer through Istio instead of Gateway API: spec.networking's own
doc comment quietly promised a Gateway API HTTPRoute this operator would
build eventually ("a near-term follow-up, not deferred"). That promise is
wrong for a public release -- an operator managing someone's ingress
mechanism for them is a worse default than not touching it at all, and a
surprise HTTPRoute appearing once that follow-up eventually landed would
have been exactly backwards for an Istio (or plain-Ingress, or
intentionally-unexposed) install. Made the non-goal explicit and
permanent instead (DESIGN.md §1), removed the dead `gatewayListener`
field it was the only consumer of (zero runtime call sites anywhere --
setting it already had no effect, so this is a schema cleanup, not a
behavior change), and corrected ROADMAP.md's framing. hostname/servicePort
stay: both are live (TERDUT_PUBLIC_URL, container/Service port), this
operator just never acts on hostname for exposure. Added
examples/networking (Gateway API HTTPRoute, Istio VirtualService) showing
how to expose the plain ClusterIP Service the operator already creates --
outside the operator itself, as illustrations, not as something
examples/demo applies automatically.

No new terdut-server version requirement: both changes are CRD/controller-
only, nothing about the API this operator's bootstrap flow depends on
changed.
2026-10-02 18:50:38 +02:00
Niklas Ye d9315322fc examples/demo: fix two real bugs this exact demo just hit live
CI / chart (push) Successful in 1s
CI / security (push) Successful in 3m24s
CI / test (push) Successful in 10m45s
1. Renamed every object this demo creates (TerdutServer, Postgres
   Secret/Deployment/Service) from terdut-demo[-postgres] to
   terdut-operator-demo[-postgres]. The user applied this kit into the
   already-live "terdut-demo" namespace -- the real operator exercise
   from earlier in this repo's own history -- and this demo's own
   TerdutServer/Postgres objects shared that exact name. The TerdutServer
   apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
   while the live object already had postgresClusterRef violates "exactly
   one of" and the API server refused it), and the real Postgres Service
   was never touched (confirmed live: still Zalando's own spilo selector,
   endpoint still the real StatefulSet pod) -- but the Postgres Secret and
   Deployment, having no such protection, were created as brand new,
   extra, crash-looping objects sitting right next to the real ones.
   Prefixing every name this demo creates means a repeat of this exact
   mistake no longer collides with anything, documented directly in
   README.md now.

2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
   (added responding to a PodSecurity "restricted" warning) took
   CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
   entrypoint needs to chown/chmod the data directory before it drops
   privileges itself -- confirmed in a real crashed pod's logs: `chmod:
   /var/run/postgresql: Operation not permitted`. kubectl apply
   --dry-run=server, which is as far as this got verified before, only
   checks admission policy; it was never actually booted. Removed the
   capability drop and verified for real this time: applied just
   00-postgres.yaml alone into a disposable namespace, waited for the pod
   to go Ready, read its logs ("database system is ready to accept
   connections"), then deleted that namespace.
2026-10-02 13:25:47 +02:00
Niklas Ye 5b45cf72e1 Bump go.opentelemetry.io/otel to v1.45.0: v1.44.0 carries GO-2026-6505
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m22s
CI / test (push) Successful in 5m0s
Release / test (push) Successful in 5m26s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 7m22s
Release / scan-image (push) Successful in 34s
Exporter config logging may leak endpoint URLs in info logs
(otlptrace/otlptracegrpc/sdk, transitively through grpc's own otel
instrumentation -- all indirect in go.mod, nothing imports these by
name). govulncheck flagged it reachable through real call chains
(tdclient.Client.DeleteIntegration, cmd/main.go's own init), caught by
ci.yaml's security job while cutting v0.1.2 (run 897) -- same pattern as
terdut-server's own da48814 for grpc's CVE-2026-84445.

go.opentelemetry.io/otel, /metric, /sdk, /sdk/metric, /trace,
/exporters/otlp/otlptrace, /exporters/otlp/otlptrace/otlptracegrpc all
moved 1.44.0 -> 1.45.0 together, plus go-logr/logr's own patch bump and
proto/otlp + genproto that go mod tidy pulled along with them. Verified:
go build, full test suite (71.7% coverage unchanged), golangci-lint,
helm-lint, and govulncheck itself now reporting zero reachable
vulnerabilities.
v0.1.2
2026-10-02 12:59:40 +02:00
Niklas Ye 738c210505 Set the chart's placeholder version to 0.1.2
CI / chart (push) Successful in 1s
CI / security (push) Failing after 1m4s
CI / test (push) Successful in 2m11s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree heading
for v0.1.2 that still says 0.1.1 tells its reader something false. Same
as 0e3118d before it.
2026-10-02 12:56:16 +02:00
Niklas Ye 46ba0e8d5c CLAUDE.md: the wrapper-chart one-time step is done, not still pending
CI / chart (push) Successful in 1s
CI / test (push) Has been cancelled
CI / security (push) Has been cancelled
Stale since this repo's actual first release (v0.1.1, 2026-10-01) already
did it -- Ryuvia/charts/terdut-operator already exists and already pins
v0.1.1. Caught while about to repeat the same wrong assumption for this
release.
2026-10-02 12:55:23 +02:00
Niklas Ye ae97d28444 Add wait-for-postgres init container to the generated Deployment
CI / chart (push) Successful in 1s
CI / security (push) Failing after 57s
CI / test (push) Successful in 2m0s
This operator's own Deployment template crash-looped a few times against
a from-scratch postgres-operator cluster still doing initdb and Patroni
leader election -- exactly the gap examples/demo's own README just
documented for it. terdut-server's ping-retry budget on startup
(internal/db/db.go in that repo) is sized for a much shorter, different
race (NetworkPolicy propagation, a few seconds), not genuine first-time
cluster creation, so it exhausted and the process exited before ever
binding its HTTP port -- a startupProbe cannot fix that, since the crash
happens before there is anything to probe. Same root cause and same fix
as charts/terdut-server's own deployment.yaml template as of that repo's
v0.33.2.

waitForPostgresContainer reuses dbEnv unchanged: both of
resolveDatabaseEnv's two paths (DSN, postgresClusterRef) put
TERDUT_DB_DSN first, so it's already exactly what pg_isready needs, and
pg_isready needs no credentials, so dbEnv's optional PGPASSWORD riding
along too is harmless rather than load-bearing.

Covered by the existing envtest suite (asserts on Containers[0], the main
container, unaffected by adding InitContainers) -- `make test` passes
unchanged, 71.7% coverage on internal/controller. Updates examples/demo's
own README, which no longer needs to warn about this.
2026-10-02 10:11:55 +02:00
Niklas Ye ca2cd2c645 Add examples/demo: one of every CRD, plus a script to fire alerts at it
CI / chart (push) Successful in 1s
CI / security (push) Failing after 1m7s
CI / test (push) Successful in 2m37s
A self-contained demo kit: a TerdutServer against a throwaway, bare
Postgres (bring-your-own DSN -- simplest path to stand up from nothing,
ROADMAP.md Stage 1's own note), two TerdutTeams, and each team's own
TerdutEscalationRule/TerdutDeadmanSwitch/TerdutAlertSource, so every CRD
this operator manages is exercised together rather than in isolation the
way config/samples' one-of-each already does.

fire-alerts.sh sends terdut-server's own amPayload/amAlert shape (read
from internal/api/alertmanager.go in that repo, not guessed from its
docs) at whichever TerdutAlertSource's generated webhook Secret it reads
the key out of -- high-cpu/disk-full/pod-crash scenarios to open and
resolve incidents, and a heartbeat scenario matching each team's dead
man's switch matcher, so stopping it demonstrates the switch noticing
silence on its own.

Verified server-side (kubectl apply --dry-run=server -k examples/demo)
against this operator's own dev cluster, which already has these CRDs
installed: every object validates. The one warning that cluster's
"restricted" PodSecurity raises (postgres:17-alpine's entrypoint needs to
start as root before it drops privileges itself) is noted inline in
00-postgres.yaml rather than worked around -- not a real production
pattern, and this Postgres exists only to be thrown away with the rest of
the demo namespace.

README.md walks through: applying, watching status, why a few early
CrashLoopBackOff restarts on terdut-demo itself are expected (this
operator's Deployment template has no wait-for-postgres init container
yet, unlike charts/terdut-server's chart as of v0.33.2), reaching the web
UI (port-forward -- spec.networking.hostname is accepted but nothing
creates an HTTPRoute for it yet), turning on open signup with the
operator's own generated admin token since the bootstrap-created account
has no password, firing alerts, and tearing down.
2026-10-02 10:03:49 +02:00
Niklas Ye a75b23c4ad DESIGN.md: record the missing OIDC trustEmail field, found exercising a real second install
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 2m14s
Found while standing up terdut-demo (Ryuvia/charts#275), a second real
TerdutServer against the same Authentik provider as production: OIDCSpec
has no trustEmail override, so a demo install copying production's OIDC
config otherwise verbatim silently runs with the wrong default for it.
Not fixed here -- recorded in §13 as a real, found gap, not a decision,
same as the mid-life teamRef note already there.
2026-10-01 19:39:20 +02:00
Niklas Ye 4ab04d29a8 DESIGN.md: document operator mode and what it deliberately doesn't lock
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m3s
CI / test (push) Successful in 2m22s
No operator-mode section existed here before -- terdut-server's own
README.md documents the feature, but this repo's design doc never
mentioned it. Added as §6 point 7, confirmed against source
(internal/api/middleware.go's OperatorModeBlock, router.go's opMode
wrapper): it blocks human writes to exactly the resources this
operator's CRDs manage (team identity, OIDC-group binding, escalation,
dead man's switches, integrations), and nothing else -- team membership,
invites, and the on-call schedule/rota stay human-editable regardless,
confirmed from the router rather than assumed from the README's prose
alone.
2026-10-01 19:16:53 +02:00
Niklas Ye 0e3118d6aa Set the chart's placeholder version to 0.1.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 54s
CI / test (push) Successful in 3m50s
Release / test (push) Successful in 1m44s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 5m47s
Release / scan-image (push) Successful in 5s
Cosmetic: `make helm-package` passes --version/--app-version from the
tag (da48814's own comment), so this field decides nothing about what
gets published. Still done so the tree doesn't say 0.1.0 while heading
for a v0.1.1 release. Cites da48814 (the fix this version actually is).
v0.1.1
2026-10-01 15:18:24 +02:00
Niklas Ye da488146bc Bump grpc to v1.83.2: v1.83.1 itself carries CVE-2026-84445
v0.1.0's own trivy scan found it: google.golang.org/grpc v1.83.1 (bumped
in b4ccdb0 to clear an unrelated, earlier grpc CVE the same scan would
otherwise have flagged) has its own high-severity denial-of-service CVE,
fixed one patch later at v1.83.2. Pure bad timing between the two fixes,
not a different root cause. govulncheck and the trivy scan both ran
clean against this version locally before tagging.
2026-10-01 15:18:11 +02:00
Niklas Ye b4ccdb09d5 Stage 5: installer chart + release infra, kind e2e pass through the chart
Release / test (push) Successful in 2m48s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m3s
CI / test (push) Successful in 2m5s
Release / chart (push) Successful in 4s
Release / image (push) Successful in 7m6s
Release / scan-image (push) Failing after 33s
Chart (charts/terdut-operator) generated via kubebuilder's own helm/v2-alpha
plugin from config/'s kustomize output -- CRDs + manager Deployment/RBAC
come from the same markers every other stage already generates, one source
of truth. Hand-added on top: the optional terdutServer values block
(DESIGN.md §10's "helm install and get a server" path, off by default) and
the release-skill plumbing -- .release.conf, release-vars/helm-lint/push/
helm-package/helm-push/release Makefile targets, .gitea/workflows/release.yaml
(test -> image/chart -> scan-image) -- mirroring terdut-server's own shape
(registry/namespace convention, multi-arch buildx push, trivy/govulncheck/
gitleaks scans). ci.yaml gains security and chart jobs to match.

Two real issues caught while wiring this, fixed before either shipped:
- Dockerfile's builder stage didn't pin --platform=$BUILDPLATFORM, which
  would have made a multi-arch release build fail outright on this org's
  runners (no binfmt registration) -- same fix terdut-server's own
  Dockerfile already needed for the same reason.
- govulncheck found one real, reachable finding: google.golang.org/grpc
  v1.82.1 (transitive via controller-runtime's otel exporter), fixed by
  bumping to v1.83.1.

Full golden-path kind e2e pass, this time through `helm install` rather than
raw kustomize: TerdutServer (real terdut-server v0.33.0 image) -> TerdutTeam
-> one of each child kind, each confirmed Ready and then independently
confirmed against terdut-server's own API from inside the cluster (not just
the operator's own status). Deleted every CR in reverse order and confirmed
server-side cleanup the same independent way for all three child kinds, the
team, and the server. No new bugs found -- Stage 1's own kind pass already
caught what a real cluster catches that envtest can't.

Also dropped the kubebuilder helm plugin's default .github/workflows/
scaffold, same as Stage 0 already did for the main scaffold: this org runs
on Gitea, not GitHub.

Not done here, deliberately: an actual tagged release. release-preflight
found no terdut-operator/ entry under Ryuvia/charts yet to bump -- that
one-time wrapper bootstrap is a decision about deploying this operator for
real, not a side effect of finishing this stage.

make fmt lint test helm-lint build all clean.
v0.1.0
2026-10-01 14:47:10 +02:00
Niklas Ye 048f4448c4 Stage 4: TerdutAlertSource
CI / test (push) Successful in 1m34s
Covers webhook Secret generation/ownership (DESIGN.md §4.5, §7), the
WebhookSecretLost fail-closed condition, and the kind-change
delete-and-recreate rotation path.

Idempotent-create here is deliberately neither adopt-on-409
(Team/service-account) nor list-and-match-by-name (TerdutDeadmanSwitch):
terdut-server shows the webhook key exactly once, at creation, and never
again, so no server-side lookup could ever recover it after a crash.
Instead the generated webhook Secret itself -- written immediately after
the POST, before status is ever touched -- is this CR's only durable
record that a create already succeeded; found with status.integrationID
still unset on a later reconcile, it's read back directly rather than
POSTing a second, orphaned integration. Found missing with
status.integrationID *set* instead, that's the already-designed
WebhookSecretLost case: fail closed, not self-healed, since the key is
genuinely gone and recreating it would rotate a live webhook URL with no
spec change to explain why.

Renaming (PATCH) never touches the key, so it's applied unconditionally
every reconcile, same as the escalation policy's whole-policy PUT. A
spec.kind change is the one case with no in-place update verb at all:
DELETE the old integration, delete the stale webhook Secret, then run the
same create path fresh -- fires a Warning event since this breaks whatever
still sends to the old URL.

Also: fakeTerdutServer grows POST/PATCH/DELETE .../integrations routes
behind a new handleIntegrationSubPath, split out of handleTeamSubPath to
stay under gocyclo's threshold; three goconst-flagged test literals
("does-not-exist", "unready") and one unparam-flagged test helper
parameter (bootstrapReadyTerdutServer's always-"default" namespace) get
shared/removed now that a fourth same-shaped caller made the repetition
concrete enough for the linter to flag.

DESIGN.md §13 gains one honest gap found while grounding this stage, not
introduced by it: no child CRD specially detects a mid-life teamRef
change; all three always resolve spec.teamRef fresh and trust the
already-stored server-side id remains valid there.

make fmt lint test build all clean; internal/controller envtest coverage
holds at 71.6%.
2026-10-01 14:13:19 +02:00
Niklas Ye b0d50e305a ROADMAP.md: mark Stage 3 done, fix stale dead man's switch reconciliation note
CI / test (push) Successful in 1m35s
2026-10-01 13:51:22 +02:00
Niklas Ye fb9e6a38dc Stage 3: TerdutEscalationRule + TerdutDeadmanSwitch
CI / test (push) Has been cancelled
Both child CRDs resolve their own teamRef -> TerdutTeam.status via the new
shared resolveTeamAndClient helper (childref.go), never chaining up to
TerdutServer (DESIGN.md §5) -- TerdutTeam.status.serverEndpoint, added in
this same stage, is what makes that literally true.

TerdutEscalationRule: one PUT /api/teams/{id}/escalation per reconcile
(an upsert server-side, confirmed against source), resolving each "user"
target's username to a user_id via GET /api/users first and reporting
Ready: False, reason: UnknownUser if it doesn't resolve. No DELETE exists
for this resource, so its delete path PUTs an empty policy as the closest
available undo.

TerdutDeadmanSwitch: real create/update-in-place/delete, using
terdut-server v0.33.0's PUT (added specifically for this operator). No
unique-name constraint server-side, so idempotent-create here is
GET-list-and-match-by-name rather than adopt-on-409.

Extends tdclient with User/GetUserByUsername, the escalation request types
+ SetEscalation, and DeadmanSwitch + its CRUD methods. Also folds
ConditionTeamReady into the single shared ConditionReady constant, since
both were literally "Ready" and Stage 3 would otherwise have needed a
third same-valued constant.

internal/controller/terdutserver_controller_test.go's fakeTerdutServer
grows GET /api/users, PUT .../escalation, and the full dead man's switch
collection/item routes, replacing the old parseTeamPath/handleTeamByID
pair with a more general parseTeamSubPath/handleTeamSubPath dispatcher
that still covers every existing Stage 1/2 route unchanged.

make fmt lint test build all clean; envtest coverage for
internal/controller: 50.5% -> 71.7%.
2026-10-01 13:50:08 +02:00
Niklas Ye fef60caf06 DESIGN.md: ground Stage 3 against real source before writing any code
CI / test (push) Successful in 1m34s
Three real findings, same discipline as Stages 1/2:

- Dead man's switches gained PUT update-in-place in terdut-server
  v0.33.0 (internal/api/teams.go's handleUpdateTeamDeadman, whose own
  doc comment names terdut-operator as the reason it was added) --
  §5's table still described delete-and-recreate, written before that
  landed. Also: no unique-name constraint server-side at all, so this
  resource's idempotent-create step is GET-list-and-match-by-name, not
  adopt-on-409 the way Team/service-accounts work.
- §4.3's username->user_id resolution needs an endpoint: GET /api/users,
  confirmed open to any authenticated caller (router.go's own "readable
  by anyone signed in"), so the team-scoped credential already in hand
  is enough -- no new server-side capability needed here, unlike
  TEAM-LOOKUP.md's gap.
- §5's "a child never needs to chain up to TerdutServer" claim wasn't
  actually true as written -- a child still needs the server's URL to
  make any call, and the only way to get one was reading TerdutServer
  directly. Fixed at the root: TerdutTeam.status now carries
  serverEndpoint too (resolved once, by TerdutTeam's own controller,
  same reconcile as teamID/credentialsSecretRef), so the claim holds
  literally and child controllers need no terdutservers RBAC at all.
2026-10-01 11:23:20 +02:00
Niklas Ye c9c52af2f7 Stage 2: TerdutTeam (create, mint team credential, rename/oidc-groups, delete)
CI / test (push) Successful in 1m41s
Implements ROADMAP.md Stage 2 against terdut-server's now-real
GET /api/teams?name= (TEAM-LOOKUP.md, landed in terdut-server just
before this commit) -- without it, the adopt-on-409 pattern this
controller depends on for team creation had no server-side lookup to
call, the same gap TerdutServer's own bootstrap flow hit and fixed in
Stage 1.

- api/v1alpha1: TerdutTeamSpec per DESIGN.md §4.2 (serverRef, displayName,
  oidc). status.credentialsSecretRef drops namespace for key, matching
  the fix already applied to TerdutServer's.
- internal/controller:
  - terdutteam_controller.go: resolves serverRef (same-namespace by
    default; cross-namespace gated by the target TerdutServer's
    spec.allowedTeams, DESIGN.md §4.6), waits for that TerdutServer to be
    Bootstrapped (no cross-controller RPC -- reads its
    status.credentialsSecretRef directly, DESIGN.md §5), creates the team
    and mints its team-scoped credential using the TerdutServer's
    instance-scoped one, then applies rename/oidc-groups with the
    team-scoped credential every reconcile (both are idempotent PUTs of
    the whole resource -- applied unconditionally rather than diffed
    against a stored last-applied value, same "cheap because it's small"
    reasoning §5 already gives the escalation policy's whole-policy PUT).
  - terdutteam_allowedteams.go: the §4.6 consent check in isolation from
    any client, unit-tested directly against hand-built inputs.
  - terdutteam_bootstrap.go: create-or-adopt-on-409 for the team itself
    (via TEAM-LOOKUP.md) and for its team-scoped service account (via the
    same GET-by-name+mint-new-key pattern Stage 1 already uses for the
    instance account).
  - secrets.go: extracted TerdutServer's write/read-credential-Secret
    helpers into free functions, now shared by both controllers rather
    than duplicated.
  - Finalizer deletes the team server-side (owner-gated, needs the
    team-scoped credential -- confirmed against source that an
    instance-scoped one does not satisfy requireTeamOwner, same finding
    as TEAM-LOOKUP.md's) and cleans up its credentials Secret. A team
    created but never fully reconciled to Ready (no team-scoped
    credential ever minted) is left orphaned server-side on delete -- a
    known, documented limitation (terdut-server has no delete path that
    doesn't require owner-equivalent access), not a silent gap.
- internal/tdclient: Team type, CreateTeam, GetTeamByName, RenameTeam,
  DeleteTeam, SetTeamOIDCGroups, CreateTeamServiceAccount -- matching
  terdut-server's real handlers' shapes field-for-field, same as Stage
  1's client additions.
- Tests: envtest covering the happy path, both not-ready reasons
  (ServerRefNotFound, WaitingForServer), cross-namespace allow/deny
  (default-closed and explicit All), both adopt-on-409 paths (team
  itself, team-scoped service account), and deletion. Extended the shared
  fakeTerdutServer (Stage 1) with team endpoints rather than writing a
  second, separately-drifting fake. 72.8%/30.6% coverage, 0 lint issues.

Verified locally: make fmt lint test build all clean.
2026-10-01 11:09:40 +02:00
Niklas Ye 72c979e3e8 DESIGN.md: record the now-resolved Team-lookup gap, fix credentialsSecretRef shape
CI / test (push) Successful in 1m34s
§5's Team row now cites GET /api/teams?name= (terdut-server's
TEAM-LOOKUP.md, landed today) -- without it the idempotent-create general
rule's claim that every resource here has a real lookup to adopt-on-409
through wasn't actually true for Team specifically, confirmed by tracing
it before writing any TerdutTeam code, same as Stage 1's bootstrap flow.

Also: TerdutTeam.status.credentialsSecretRef drops namespace for key,
matching the TerdutServer fix from Stage 1 -- same reasoning, missed
there originally.
2026-10-01 10:56:05 +02:00
Niklas Ye b263b48510 ROADMAP.md: mark Stage 1's kind e2e pass done, with what it found
CI / test (push) Successful in 1m39s
2026-10-01 10:15:52 +02:00
Niklas Ye 1c45b7e80b Stage 1: fix two real bugs the kind e2e pass caught, neither envtest could
CI / test (push) Has been cancelled
Ran a full kind end-to-end pass per ROADMAP.md's open item: real kind
cluster, real disposable Postgres, the real terdut-server v0.33.0 image,
the operator built into a real image and deployed as a real Pod (not
`go run` against the cluster -- that was tried first and correctly failed
on cluster DNS not resolving from outside the cluster network, which is
expected, not a bug).

Result: TerdutServer went Ready, the generated credentials Secret held a
real tdsa_-prefixed service-account key, and that key successfully
authenticated and exercised its real intended capability against the
actual server (GET/POST /api/teams -> 200/201) -- confirmed from
terdut-server's own access log, not just our side. Stage 1's actual goal
(ROADMAP.md) is proven, not just asserted.

Two real bugs surfaced that no envtest suite could have caught, since
envtest's client bypasses RBAC entirely:

- .dockerignore's `!**/*.go` doesn't work under podman (the scaffold's
  own comment already named this exact gotcha, buildah/containers#6417,
  and pointed at the fix) -- `docker build` was silently building from an
  empty source tree ("package cmd/main.go is not in std") until this was
  pinned down. Fixed by re-including cmd/api/internal by name, as that
  comment suggested doing if this happened.
- The controller had no RBAC for events.k8s.io (the new events API
  GetEventRecorder uses, unlike the deprecated GetEventRecorderFor) --
  every Event emission failed server-side ("Server rejected event (will
  not retry!)"), silently, since event-recording failure doesn't fail
  reconciliation. Reconciliation itself was never affected, but DESIGN.md
  §12's observability goal (every externally-visible action emits an
  Event) silently wasn't being met in any real deployment. Added
  +kubebuilder:rbac for events.k8s.io/events (create, patch); confirmed
  fixed by restarting the operator and checking `kubectl describe
  terdutserver` actually shows the Event afterward, not just that the log
  line stopped.

Also noted, not fixed here (a different repo's bug): terdut-server's own
GET /api/me 500s for a service-account caller rather than a clean 4xx --
that endpoint assumes a human user in context. Worth a terdut-server
issue, not an operator concern.
2026-10-01 10:15:26 +02:00
Niklas Ye 8064876cb1 Stage 1: TerdutServer full lifecycle (Deployment, Service, both database
CI / test (push) Successful in 1m46s
paths, self-registration bootstrap)

Replaces the bring-your-own-only Stage 1 (commit 1be7cf2) wholesale, per
the redesign in the previous two commits: the operator creates every
server it manages, so self-registration (DESIGN.md §6) is the only
bootstrap path, and Deployment/Service/database management builds
together with it (ROADMAP.md Stage 1) rather than behind a separate
later stage.

Grounded in terdut-server's actual chart (charts/terdut-server/templates/
deployment.yaml, values.yaml), not reconstructed from DESIGN.md's
illustrative YAML alone -- env var names, the password-via-PGPASSWORD
convention, the Recreate deployment strategy, /healthz probes, and the
TERDUT_OPERATOR_MODE=true decision (always on here, unlike the chart's
default-off: every write this operator's own future controllers make
goes through a service account already) all match that source exactly.

- api/v1alpha1: full TerdutServerSpec (image, replicas, networking,
  database, sweeper, deadman, notify, oidc, passwordLogin, allowedTeams).
  spec.database is a oneOf (dsn xor postgresClusterRef) via CEL
  XValidation. No spec.credentialsSecretRef -- removed entirely in the
  prior redesign commit, not carried forward.
- internal/controller:
  - terdutserver_deployment.go: Deployment + Service via CreateOrUpdate,
    owned (OwnerReference), env built field-for-field against the chart.
  - terdutserver_database.go: both §8 paths. The Zalando path resolves
    the postgresql.acid.zalan.do CR by convention (database/role both
    "terdut", matching every DESIGN.md example) and only ever confirms
    its generated credentials Secret exists -- never reads the value,
    same "wire a secretKeyRef, don't read it" posture the DSN path takes.
    classifyClusterGetError is its own function specifically so the
    CRD-not-installed case (meta.IsNoMatchError) is unit-testable without
    a real client.
  - terdutserver_bootstrap.go: self-registration, checkpointed against
    both real crash windows (DESIGN.md §6 point 1) -- an admin-key
    checkpoint Secret, and adopt-via-GET+mint-new-key on a 409 from
    creating the service account. BootstrapStateLost is its own error
    type so Reconcile can route it to a condition instead of an infinite
    retry.
  - terdutserver_controller.go: ties it together -- finalizer add, DB
    resolution, Deployment/Service reconcile, wait for a ready replica,
    bootstrap, Ready/Bootstrapped/DatabaseReady conditions. Finalizer on
    delete only removes the generated Secrets: terdut-server's API can't
    delete a user or service account, only revoke keys, so there's
    nothing server-side to undo.
- internal/tdclient: added Bootstrap, CreateInstanceServiceAccount,
  GetServiceAccountByName, CreateServiceAccountKey, matching
  terdut-server's real handlers' request/response shapes (internal/api/
  users.go, service_accounts.go in that repo) field-for-field.
- Tests: envtest suite covering the full DSN-path lifecycle end to end
  (finalizer -> Deployment/Service -> simulated readiness -> real
  bootstrap against an httptest.Server fake), the adopt-on-409 recovery
  path, BootstrapStateLost, both Zalando outcomes (cluster not found;
  cluster + Secret found -> real DSN -> Ready), and deletion. A minimal
  test-only stub of the Zalando CRD (internal/controller/testdata) lets
  envtest create fixture objects without a real postgres-operator
  installed. 74.0%/44.7% coverage, 0 lint issues.
- Two things scoped down from §8's full ambition, called out in code and
  ROADMAP.md rather than silently dropped: no live watch on the
  Zalando-generated Secret for rotation (periodic resync notices
  eventually, not immediately), no Gateway API HTTPRoute creation from
  spec.networking (would add a new dependency; nothing about proving
  bootstrap works depends on external ingress existing). Both are
  near-term follow-ups.

Verified locally: make fmt lint test build all clean.
2026-10-01 09:11:56 +02:00
Niklas Ye fc68ee7256 ROADMAP.md: merge Stage 1 + old Stage 5, renumber
CI / test (push) Successful in 1m33s
Follows DESIGN.md's redesign (previous commit): with no hand-deployed
server to prove the simpler CRDs against, there's no reason left to defer
TerdutServer's Deployment/Service/database management behind a separate
later stage. Stage 1 now covers TerdutServer's full lifecycle --
Deployment, Service, both Postgres paths from §8 at once (bring-your-own
DSN and Zalando, per the user's call, not sequenced), bootstrap,
credentials -- built together, since bootstrap only has something to
bootstrap once the Deployment exists.

Old Stage 5 (TerdutServer absorbs Deployment/Service/bootstrap) is gone,
folded into Stage 1. Old Stage 6 (installer chart + release) renumbers to
Stage 5. Stages 2-4 (TerdutTeam, EscalationRule+DeadmanSwitch,
AlertSource) are unchanged in content, renumbering only where old Stage 5
disappears from ahead of them.

Explicitly supersedes the Stage 1 shipped before this redesign (commit
1be7cf2): that TerdutServerSpec/Status/controller/tests implemented the
now-removed bring-your-own path and get replaced wholesale when Stage 1
is actually implemented next, not extended.
2026-10-01 08:50:24 +02:00
Niklas Ye f1fd64a567 DESIGN.md: the operator only ever creates servers, never adopts one
CI / test (push) Has been cancelled
Removes the premise Stage 1's bring-your-own credential design was built
on. Confirmed with the user directly: this operator creates and owns
every TerdutServer it manages; there is no hand-deployed or chart-deployed
install it's expected to target or migrate.

- §1: states this explicitly -- the root the rest of this commit hangs off.
- §4.1: spec.credentialsSecretRef (bring-your-own input) removed entirely,
  not kept as unused flexibility. status.credentialsSecretRef stays as
  pure output.
- §6: self-registration is now the *only* bootstrap path, not one of two --
  and, since it's now load-bearing rather than a fallback with an easy
  escape hatch, closed the two real crash windows in it properly rather
  than leaving them as theoretical gaps: a checkpoint Secret for the raw
  admin key between /api/bootstrap and minting the service account, and
  adopt-on-409 (§5's general rule) if a prior interrupted attempt already
  got that far. A checkpoint lost after being used crosses into the same
  fail-closed territory §5's webhook-Secret-loss rule already established
  -- same recovery (delete and recreate), not a new, one-off workaround.
- §10: dropped the migrate-an-existing-install narrative and the
  chart-Job-vs-operator bootstrap race question entirely -- both
  presupposed an install the operator might adopt or race against, which
  doesn't exist. Kept the installer-chart framing on its own.
- §13: dropped the now-stale "Helm chart migration execution" deferred item.

§8 (Postgres) needed no change -- it already described both the DSN and
Zalando paths as co-equal, full-design detail, with no sequencing between
them to remove.
2026-10-01 08:49:04 +02:00
Niklas Ye 1be7cf2b7f Stage 1: TerdutServer, bring-your-own bootstrap credentials
CI / test (push) Successful in 1m41s
Implements the narrowed Stage 1 scope from ROADMAP.md, against the
bootstrap-flow fix from DESIGN.md §4.1/§6 (the earlier self-registration
flow couldn't work unauthenticated against terdut-server's real
AuthMiddleware -- see that commit for the full trace).

- api/v1alpha1: TerdutServer with spec.endpoint + spec.credentialsSecretRef
  + spec.allowedTeams (image/replicas/networking/database deferred to
  Stage 5, per DESIGN.md's own narrowing). SecretKeyRef has no namespace
  field -- always the operator's own, by construction.
- internal/controller: TerdutServerReconciler implements exactly the
  bring-your-own path -- adopt spec.credentialsSecretRef if the Secret
  exists and has data under the given key, probe GET /api/version as a
  reachability check, set Ready/Bootstrapped conditions accordingly.
  Self-registration (the /api/bootstrap race) is not implemented; unset
  spec.credentialsSecretRef reports Ready: False, reason:
  CredentialsSecretRefRequired, not an attempt at a flow that would fail
  unauthenticated anyway. No finalizer: this stage creates nothing
  server-side and adopts rather than generates its Secret, so there's
  nothing to clean up on delete yet.
- internal/tdclient: minimal terdut-server API client (Version only, the
  one call this stage needs), styled after terdut-tui's own
  internal/api/client.go per terdut/CLAUDE.md's mirroring convention.
- Tests: envtest suite covering all four not-ready paths plus the happy
  path (fake terdut-server via httptest.Server, per DESIGN.md §11), and a
  focused unit suite for tdclient. 75.6%/82.4% coverage.
- Event recording uses the new events.k8s.io/v1 recorder API
  (mgr.GetEventRecorder), not the deprecated GetEventRecorderFor --
  caught by golangci-lint's staticcheck before it shipped.

Verified locally: make fmt lint test build all clean, 0 lint issues, all
specs pass.
2026-09-30 22:30:22 +02:00
Niklas Ye ffc2e6441e ROADMAP.md Stage 1: scope down to bring-your-own only, defer self-registration to Stage 5
CI / test (push) Successful in 1m23s
Stage 1's own setup (chart bootstraps before the CR exists) never
exercises the self-registration fallback, and the narrowed spec has
nowhere to put the username/email /api/bootstrap needs anyway. Matches
DESIGN.md's §6 rewrite (bring-your-own is now the primary path, not an
equal alternative).
2026-09-30 22:22:09 +02:00
Niklas Ye 5f93a530fa DESIGN.md §4.1/§6: drop namespace from credentialsSecretRef, add key
CI / test (push) Has been cancelled
It's always the operator's own namespace by construction now (§6), never
anything else, so there was nothing for the field to vary -- key varies
instead (fixed 'token' when self-generated, whatever a human chose when
adopted from spec.credentialsSecretRef).
2026-09-30 22:21:28 +02:00
Niklas Ye 7f439605c4 DESIGN.md §4.1/§6: fix the bootstrap self-registration deadlock
CI / test (push) Has been cancelled
Traced the actual flow against terdut-server's real source before writing
any Stage 1 controller code, rather than trusting this section's own prior
description of it:

- internal/api/middleware.go's AuthMiddleware hard-rejects with 401 any
  request carrying neither a Bearer token nor a session cookie, before
  handleListServiceAccounts' own (more permissive) internal check ever
  runs. So "on 403, self-lookup via GET /api/service-accounts?name=" --
  this section's described fallback -- cannot work unauthenticated; an
  earlier draft of this section assumed otherwise.
- That only actually matters in the rare case where this TerdutServer's
  own controller loses the /api/bootstrap race... except Stage 1's own
  setup (ROADMAP.md) guarantees it loses every time: terdut-server is
  deployed via its existing chart, which runs its own bootstrap Job,
  before the TerdutServer CR or its controller exist at all. The
  self-registration flow was never going to complete for the one scenario
  Stage 1 actually exercises.

Fix: spec.credentialsSecretRef (§4.1), bring-your-own -- a human mints an
instance-scoped service account once, manually, with their own admin
session, and hands the controller that Secret directly. This is now the
primary, expected path; self-registration on a genuinely fresh install
(where this controller might actually win the race) stays as the
fallback it was always meant to be, not the only path.

Also corrected: this section's opening paragraph still said "v1-blocking,
not v1-shippable" pending SERVICE-ACCOUNTS.md landing -- confirmed shipped
(internal/api/service_accounts.go, migration 014) since Stage 0's work on
this repo; stale framing removed.
2026-09-30 22:20:31 +02:00
Niklas Ye dd955bbf1a CI: update comment -- second MTU fix resolves the github.com timeout
CI / test (push) Successful in 1m17s
Confirmed via the actual job log (run 857, job 1931), not just the exit
code: setup-envtest fetches envtest-v1.37.0-linux-amd64.tar.gz from
github.com in ~4s now, where it previously TLS-handshake-timed-out every
time. golangci-lint's own git-clone-to-github.com (the confound from the
first fix attempt) also went through fine in the same run.

Stage 0 is now fully green end to end: fmt, lint, test (envtest included).
2026-09-30 21:41:29 +02:00
Niklas Ye 0e89816ad4 CI: revert temporary test-only isolation; MTU fix ruled out
CI / test (push) Successful in 7m59s
Clean result, isolated from lint's own unrelated github.com flakiness:
post-MTU-fix (Ryuvia/charts#272), `make test` alone hits the exact same
"TLS handshake timeout" fetching envtest-v1.37.0-linux-amd64.tar.gz as
before the fix. Byte-for-byte identical error. The dind sidecar's MTU
mismatch was real (measured 1450 vs 1500 per the other session's report)
but it was not (solely) the cause of this specific failure.

Separately and incidentally: golangci-lint's own custom-gcl build also
does a plain `git clone https://github.com/...` and that is now failing
too (2/2, ~2.5min hang then generic exit 128) where it briefly succeeded
in an earlier pre-fix run -- noted in the comment but not chased further
here; worth someone's attention if it keeps recurring, since it'll block
`lint` regardless of the envtest question.
2026-09-30 21:26:29 +02:00
Niklas Ye a55489c7a9 CI: temporarily isolate make test from lint to probe the envtest fetch alone
CI / test (push) Failing after 3m23s
lint is currently blocked by an unrelated github.com git-clone failure
(golangci-lint's custom-gcl build), which means the last two runs never
reached setup-envtest -- the step the act-runner MTU fix (Ryuvia/charts#272)
was meant to affect. Narrowing to `make test` alone to get a clean signal;
will revert to `make fmt lint test` right after.
2026-09-30 21:22:10 +02:00
Niklas Ye a241007135 CI: revert to container-based test job; correct the NetworkPolicy claim
CI / test (push) Failing after 2m42s
Host-mode (previous commit) fails earlier and differently: "go: command not
found" -- the runner host has no Go, so that path is dead.

Checked Ryuvia/charts' act-runner/templates/networkpolicy.yaml directly
rather than assuming: it's the only NetworkPolicy in the cluster, and it is
explicitly deny-ingress only -- its own comment states egress is
deliberately untouched, "CI pulls from registries and package indexes that
are not enumerable here" (issue #128). So my earlier claim that this needs
"allowlisting github.com on the runner's NetworkPolicy" was wrong: there is
no in-repo egress rule governing this at all. Whatever blocks github.com
from the dind bridge is outside anything Ryuvia/charts or Ryuvia/k8s
expresses in a Kubernetes object -- back to container-based (matching every
other Go job in this org) as the known-good shape, with `make test` left
red on the envtest fetch until that's actually found.
2026-09-30 19:49:23 +02:00
Niklas Ye 372fbe0660 CI: try running test job on host, not in a container
CI / test (push) Failing after 1s
Confirmed the hard way (run 852, attempt 2): setup-envtest v0.25 fetches the
envtest kube-apiserver/etcd tarball from github.com's release CDN, not the
legacy GCS kubebuilder-tools bucket (that bucket 403s now for any object --
no fallback there for k8s 1.37 either). github.com is unreachable from this
job's container the same way terdut-server's ci.yaml already documents for
get.helm.sh -- TLS handshake timeout.

Dropping `container:` on this job is the same fix terdut-server's `chart` job
already uses for that exact class of problem (it reaches get.helm.sh only by
running on the host). Unproven for a Go job specifically -- no workflow in
this org has run Go outside a container before, so this also bets the runner
host has Go installed. If it fails on a missing `go` instead of the envtest
fetch, that bet was wrong and the real fix is allowlisting github.com's
release CDN on the runner's NetworkPolicy instead (Ryuvia/charts or
Ryuvia/k8s, outside this repo).
2026-09-30 19:41:08 +02:00
Niklas Ye e118a8e70d Ignore coverage output (cover.out) from make test
CI / test (push) Failing after 7m4s
2026-09-30 19:22:28 +02:00
Niklas Ye ba253b7bf7 Add CI, repo CLAUDE.md, and finish Stage 0
- .gitea/workflows/ci.yaml: fmt/lint/test, same no-actions/checkout-and-manual-clone
  shape as terdut-server's ci.yaml, and the same reasoning for why (Node/ES2022
  incompatibility on the runner image). No chart/security jobs yet -- nothing for
  either to check until Stage 6 / real controller code exists.
- CLAUDE.md: Checks + Release sections, matching the sibling repos' convention from
  the workspace-level CLAUDE.md ("each repo has its own CLAUDE.md... read it before
  working in that repo"). Release is explicitly marked not-wired-yet rather than
  copying terdut-server's, since there's no chart to release against until Stage 6.
- ROADMAP.md: moved the .release.conf bullet out of Stage 0 (it names a HELM_CHART
  this repo doesn't have yet) -- it was already duplicated into Stage 6, which is
  where it actually belongs.

Stage 0 done: `make fmt lint test` verified green locally. Real open question the CI
workflow's comments flag rather than assume past: whether storage.googleapis.com
(envtest's binary source) is reachable from this Gitea runner's container network the
way proxy.golang.org is -- terdut-server's own ci.yaml notes get.helm.sh/github.com are
not. Only running the workflow for real will confirm; the comment names the fallback
(move the job out of `container:`, like terdut-server's chart job) if it isn't.
2026-09-30 19:22:19 +02:00
Niklas Ye c97571c4c4 Scaffold project with Kubebuilder v4
kubebuilder init --domain ryuvia.com --repo git.ryuvia.com/niklas/terdut-operator
(--license none, no per-file header boilerplate -- terdut-server's source carries
none either). Go 1.26.0/controller-runtime v0.25.0/controller-tools v0.22.0, whatever
the current kubebuilder CLI (v4.16.0) scaffolds -- not pinned back to terdut-server's
go 1.25.9, since this is a separate module with its own toolchain.

Verified locally: build, vet, fmt all clean; `make lint` (golangci-lint, fetched into
bin/) 0 issues; `make test` (controller-gen + setup-envtest, fetched into bin/,
downloads real envtest binaries from storage.googleapis.com) passes.

Dropped kubebuilder's default .github/workflows/* -- this org runs on Gitea, not
GitHub; ci.yaml (next commit) is the only CI this repo gets.
2026-09-30 19:22:09 +02:00
Niklas Ye 1feffd791a Housekeeping: gitignore, and finish the webhook-Secret-loss/RBAC fix
- Add .gitignore (build artifacts, editor swapfiles, envtest testbin).
- Remove the stray .DESIGN.md.swp that was sitting untracked in the repo.
- Carries the DESIGN.md §5/§9/§13 edits from the secret-loss discussion:
  fail-closed (not self-healed) TerdutAlertSource webhook Secret loss, and
  the corrected RBAC section (the webhook Secret lives in the CR's tenant
  namespace, not the operator's own namespace as an earlier draft claimed).
2026-09-30 19:16:37 +02:00
Niklas Ye ee39b8e668 Add build roadmap
Stages the operator's implementation: TerdutServer stays bootstrap/credentials-only
(no Deployment/Service takeover) until Stage 5, so every earlier stage targets a
hand-deployed terdut-server in a disposable dev namespace instead of forcing the
chart-migration decision (§10) up front.
2026-09-30 19:16:21 +02:00
Niklas Ye 94989e2c87 Rework §6 bootstrap/credentials against confirmed server behavior
/api/bootstrap is single-shot per install (gated on COUNT(*) FROM
users, confirmed against internal/api/users.go and the chart's
bootstrap-job.yaml), not per identity — the two-identity bootstrap
plan and the delete-Secret-to-rotate runbook this section described
don't work against that. Rewrites §6 points 1/5/6 around a dedicated,
repeatable service-account credential instead (proposed server-side in
terdut-server's new SERVICE-ACCOUNTS.md), notes in §9 that Secret
mirroring is RBAC-sound but still hands out a server-admin-equivalent
credential per consenting namespace, and flags in §10 that chart-vs-
operator bootstrap ownership blocks §6 and needs deciding first.

Updates §13 to mark the service-account type as v1-blocking rather
than a someday improvement, and adds a version-discovery endpoint to
the same list (both this operator and terdut-tui currently detect
server capability by route-probing).
2026-09-29 20:35:15 +02:00
Niklas Ye 5f728a556b Switched from referencegrant the ligther parentRef 2026-09-29 16:19:06 +02:00
Niklas Ye ef5d8fcb5d First draft for design 2026-09-29 15:16:15 +02:00