27 Commits

Author SHA1 Message Date
Niklas Ye 9b59e4bdbc Update golang.org/x/net to v0.60.0
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m16s
CI / test (pull_request) Successful in 7m4s
govulncheck reports five vulnerabilities (GO-2026-6603, 6610, 6611, 6612,
6617) in x/net v0.58.0 that the code reaches through net/http's HTTP/2
client. v0.60.0 fixes them; the other x/ modules move with it.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 22:05:15 +02:00
Niklas Ye e2c4475867 Build and scan with Go 1.26.9
govulncheck in the security job reports ten standard-library
vulnerabilities (net/http, mime/multipart, crypto/tls), all fixed in
1.26.9. The workflows pinned golang:1.26.6-bookworm.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 15:17:03 +02:00
Niklas Ye e1103f2b7d Authenticate with a seeded operator key; fold escalation and switches into TerdutTeam
Credentials: the TerdutServer controller generates <name>-operator-key in
the server's own namespace (owned by it) and hands it to the pods as
TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it
at every start. A replaced Secret rolls the pods. The bootstrap handshake,
the checkpoint Secret, per-team service accounts and credentials Secrets,
BootstrapStateLost and credentials.deletionPolicy are gone.

CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule
and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[]
on the team (matched by name, extras removed); team invites are removed.
A team is created under the identity <namespace>/<name> (external_id), so a
retry, a lost status or a deleted team heal by repeating the same call, and
a display name owned by another team is TeamNameTaken instead of an
adoption. The server resolves escalation usernames (UnknownUser condition).
OIDC claim names and trustEmail are spec fields.

Fixes: query values are URL-escaped; every delete treats 404 as success;
deleting a team no longer depends on allowedTeams consent; a switch or
integration deleted on the server is recreated; unnamed switches take the
CR's name.

Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces
and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo
(run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays
cluster-wide, now stated in DESIGN.md section 9.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:22 +02:00
Niklas Ye b0a431f2a4 Set the chart's placeholder version to 0.5.0
CI / chart (push) Successful in 2s
CI / security (push) Successful in 53s
Release / chart (push) Successful in 3s
CI / test (push) Successful in 2m19s
Release / test (push) Successful in 1m41s
Release / image (push) Successful in 6m30s
Release / scan-image (push) Successful in 4s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with d502485 (0.4.0) and 88172ad (0.3.0) before it, because a
tree heading for v0.5.0 that still says 0.4.0 tells its reader
something false.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 21:10:27 +02:00
Niklas Ye 62664c93ff Keep the instance credential across a TerdutServer delete, and adopt it on recreate
Deleting a TerdutServer removed the credential Secrets but never touched the
database, so a recreated one found a server that was already bootstrapped and
no key for it: /api/bootstrap answered 403 and the operator stopped at
BootstrapStateLost, whose message and DESIGN.md both said "delete and
recreate". That is how the terdut-demo install on the cluster got stuck on
2026-10-03: Helm's cleanupOnFail deleted its TerdutServer after a failed
upgrade, the recreate found the bootstrapped database, and it sat at Ready:
False for five days until the database was reset by hand. Recreating cannot
fix it, because the finalizer clears Secrets and the database is not its to
reset, so "a fresh create starts clean" was only ever true when the database
went with it.

spec.credentials.deletionPolicy is Retain by default: the finalizer keeps the
instance credential Secret (Delete removes it, as before). The bootstrap
checkpoint is always removed. Before calling /api/bootstrap, reconcile now
looks for the retained Secret and asks the server for the operator's own
service account with its token. Accepted: adopt it and skip bootstrap.
Rejected with 401/403: the Secret outlived a database reset, so ignore it and
bootstrap like a first install, which replaces it. Any other error retries.
terdut-server's own tests already call that endpoint with an instance-scoped
key, so the permission is not new.

BootstrapStateLost is still the answer when the server is bootstrapped and no
credential it accepts survives, but its message now names the Secret to
restore and says that recreating does not clear the database. DESIGN.md §6
says the same, and the chart passes the setting through as
terdutServer.credentials.deletionPolicy.

A retained Secret of a TerdutServer that is gone for good is an orphan to
delete by hand. It is inert: nothing adopts it unless the server accepts the
token.

Checked on the kind demo with a locally built image against the real
terdut-server v0.43.0: deleting the TerdutServer kept the Secret, recreating it
reached Ready with the same credential (identical hash) and both TerdutTeams
came back Ready with their original ids. The controller specs cover adoption,
a rejected token after a reset, the bootstrapped-and-rejected failure, and
both deletion policies.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 21:09:56 +02:00
niklas 6572f63157 Merge pull request 'examples/demo: run terdut-server v0.43.0, with alerts from two clusters' (#6) from demo-matches-v0.4.0-crd into main
CI / chart (push) Successful in 2s
CI / security (push) Successful in 51s
CI / test (push) Successful in 2m26s
Reviewed-on: #6
2026-10-08 17:10:06 +00:00
Niklas Ye 50ce5bcec0 examples/demo: run terdut-server v0.43.0, with alerts from two clusters
CI / test (pull_request) Successful in 6m35s
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m6s
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the
web UI since: the queue and incident layouts, the rota and escalation
pages, the theme toggle, and the cluster chip, filter and page titles
(v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the
floor, which is what the replicas setting actually depends on.

fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus
external label plus `cluster` in Alertmanager's group_by (terdut-server's
README, "Several clusters, one team"). It goes on the alert's labels and
groupLabels, and into the group key and the fingerprint, so the same alert in
two clusters is two incidents and not one. Unset, the payload is exactly what
it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in
both, so the queue has a chip and a filter to show.

run-demo.sh also failed on its second run, though it says it is safe to
re-run: it expected HTTP 409 when alice already exists, but a spent invite
is answered with 403 "invite link is not usable" before the username is ever
checked. It now tries to log alice in first and skips the signup if that works.

Checked on the kind cluster: the server rolled to v0.43.0, every CR became
Ready and Adopted (server, both teams, both escalation rules, both dead man's
switches, both alert sources), and /api/incidents/clusters,
/api/incidents?cluster=prod-us and the incident titles came back as expected.
No operator code changed, so this needs no operator release.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 18:55:02 +02:00
Niklas Ye 822c80dda6 examples/demo: split the Ready wait so escalation rules wait on alice too
wait_for_ready waited for every demo object at once, including
terdutescalationrule-platform, which names alice as a level-1 target --
but alice does not exist yet at that point in main(): she is created by
redeem_platform_invite, which ran after wait_for_ready. terdut-server
resolves every named username at reconcile time, not just when an
escalation actually fires, so that CR could never reach Ready before
alice did, and main() had no step in between to create her.

Split into wait_for_objects (the shared loop, now taking its object list
as arguments) plus two callers: wait_for_teams_ready, covering just the
server and the two teams redeem_platform_invite/join_payments_team
need, run before alice exists; wait_for_remaining_ready, covering the
escalation rules, dead man's switches and alert sources, run after.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 18:15:42 +02:00
Niklas Ye 0ee7ede648 examples/demo: match v0.4.0's new spec.replicas default and bump terdut-server
replicas: 1 and tag: v0.34.0 were both correct when written, but the CRD's
own default moved to 2 in v0.4.0 (same release this demo is meant to show
off), and v0.34.0 predates v0.36.0's advisory locks that make a second
replica safe instead of racing the first. Left as-is, the demo would have
been the one place in this repo demonstrating the exact unsafe combination
the CRD's own doc comment warns against: more than one replica against an
image that doesn't guard the sweeper/notifier/migration-runner singletons.

replicas is now stated explicitly as 2 rather than dropped to pick up the
default silently, matching every other field in this file's own habit of
spelling out what it depends on. tag moves to v0.36.0 specifically -- the
first version where the lock landed -- with the comment keeping v0.34.0's
original reasoning (the service-account race fix) alongside the new one,
since v0.36.0 still carries that fix forward.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 18:13:09 +02:00
Niklas Ye d50248531c Set the chart's placeholder version to 0.4.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 2m10s
Release / test (push) Successful in 7m54s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m25s
Release / scan-image (push) Successful in 36s
2026-10-03 16:22:15 +02:00
Niklas Ye 4007f54279 Default TerdutServer.spec.replicas to 2 and switch to RollingUpdate
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Has been cancelled
Mirrors charts/terdut-server's own deployment.yaml change: v0.36.0 put
the sweeper, the notifier and the migration runner each behind a
Postgres advisory lock, and gave incident creation its own conflict
resolution, so the Recreate strategy and replicas-stays-at-1 guidance
this controller carried (explicitly tracking that chart's comment)
are no longer load-bearing.

spec.replicas' +kubebuilder:default moves 1 -> 2 (config/crd/bases and
the chart's CRD template regenerated via controller-gen and
kubebuilder's helm plugin respectively, then hand-verified identical
to the generator's own output rather than trusting a bulk regen --
the plugin's --output-dir charts writes a fresh charts/chart scaffold
rather than updating charts/terdut-operator in place, so only the
diff was taken, not the whole tree). terdutserver_deployment.go's
same-value fallback (reachable only for a TerdutServer stored before
this default existed) moves with it, and its Strategy changes from
Recreate to RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicas: 2, already
zero-downtime.

DESIGN.md's three places asserting multi-replica isn't a supported
topology (the illustrative spec.replicas YAML, spec.pod.affinity's
rationale, and the HPA deferred-feature note) are corrected to match;
the HPA note now gives its own standing reason (no scaling metric or
bounds decided yet) rather than a contradiction that no longer holds.

The chart's optional terdutServer.replicas sample value moves 1 -> 2
alongside it. image.tag must be v0.36.0 or newer for any of this to
hold -- stated in both the CRD field's doc comment and the chart
value's comment, not enforced in code, same stance the chart takes on
every other version-coupled assumption.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:35:18 +02:00
Niklas Ye 88172ade29 Set the chart's placeholder version to 0.3.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 1m53s
Release / test (push) Successful in 1m45s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 5m27s
Release / scan-image (push) Successful in 2s
2026-10-02 22:19:18 +02:00
niklas 478ae6284a Merge pull request 'TerdutTeam: mint and surface a real invite link (spec.invite)' (#5) from terdutteam-invite-minting into main
CI / test (push) Has been cancelled
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
2026-10-02 20:17:21 +00:00
niklas fcc80b32d5 Merge pull request 'examples/demo: add run-demo.sh, an automated kind-cluster demo' (#4) from examples-demo/run-demo-script into main
CI / chart (push) Successful in 2s
CI / security (push) Successful in 1m31s
CI / test (push) Successful in 3m30s
2026-10-02 20:13:36 +00:00
Niklas Ye 4aa4f17c42 examples/demo: bump terdut-server to v0.34.0 (the service-account fix)
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m18s
CI / test (pull_request) Successful in 3m40s
Required for this demo to actually exercise the fix for
niklas/terdut-operator#3 -- v0.33.2 still has the authorization gap this
demo hit live (callerMayManageServiceAccount had no branch letting an
instance-scoped account adopt a team-scoped account's key).
2026-10-02 22:13:26 +02:00
Niklas Ye a0ea13955e TerdutTeam: mint and surface a real invite link (spec.invite)
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 52s
CI / test (pull_request) Successful in 2m43s
The actual fix for the human-onboarding gap niklas/terdut-server#23 found --
not a terdut-server change at all. A team-scoped credential is already
owner-equivalent for POST/GET/DELETE /api/teams/{teamID}/invites
(requireTeamOwner's synthetic-membership mechanism, ratified not
accidental per that repo's SERVICE-ACCOUNTS.md), and invite redemption
bypasses signup_mode entirely -- this TerdutTeam controller just never
grew a feature to use either fact.

New spec.invite{enabled, role (member|owner, default member), maxUses
(1-100, default 1)} and status.inviteSecretRef. The Secret lives in the
TerdutTeam's OWN namespace, not the operator's: unlike
status.credentialsSecretRef (a durable, high-privilege credential, kept
operator-side per DESIGN.md §6), an invite is bounded and limited-use,
meant for this namespace's own human operators to read and hand out --
same precedent as TerdutAlertSource's status.webhookURLSecretRef, same-
namespace and OwnerReference'd so deleting the TerdutTeam garbage-collects
it automatically.

internal/controller/terdutteam_invite.go: mints on first spec.invite.enabled,
refreshes a day ahead of terdut-server's fixed 7-day TTL (reading the
Secret's own stored expiresAt, no extra server round-trip per reconcile),
revokes server-side and deletes the Secret when flipped back to false. A
lost invite Secret is silently re-minted rather than treated as
unrecoverable the way TerdutAlertSource's webhook key is -- nothing
external holds a durable dependency on one specific invite link staying
stable, it's read once by one human and handed out.

New tdclient.Invite/CreateInvite/RevokeInvite. New envtest coverage: mint
into the team's own namespace, refresh-before-expiry, revoke-on-disable
(internal/controller/terdutteam_controller_test.go's new "spec.invite"
Describe block), plus the fake server growing invite support
(terdutserver_controller_test.go) -- its handleTeamSubPath dispatcher was
split further (deadman switches into their own handleDeadmanSubPath,
matching the existing handleIntegrationSubPath precedent) to stay under
golangci-lint's gocyclo threshold with the new route added.

examples/demo updated to prove this end to end: 02-team-platform.yaml
turns on spec.invite; run-demo.sh's bootstrap_login/join_demo_teams (the
psql signup_mode flip + a direct team_members INSERT) are replaced by
redeem_platform_invite (reads status.inviteSecretRef, a real POST
/api/signup with the invite token) and join_payments_team (POST
/api/teams/{teamID}/members using Payments' own credential and alice's
user id resolved via GET /api/users, deliberately not given its own
spec.invite, so the demo shows both onboarding paths this feature
unlocks) -- zero kubectl exec/psql calls remain anywhere in the script.
README.md's "First login" section rewritten to match; it no longer
documents the admin-token curl call that 403s against current
terdut-server (niklas/terdut-server#23).

Depends on niklas/terdut-server#24 (the callerMayManageServiceAccount fix
for terdut-operator#3) being released before this is deployed for real --
not required to build or test this change itself, since the envtest fake
never modeled that authorization gap to begin with.
2026-10-02 22:04:56 +02:00
Niklas Ye 2a08a8cd8e examples/demo: add run-demo.sh, an automated kind-cluster demo
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 1m5s
CI / test (pull_request) Successful in 2m49s
One script, two modes (run-demo.sh / run-demo.sh --teardown), that takes a
fresh empty kind cluster all the way to a working demo: creates the
cluster if needed, helm-installs this chart, applies every CRD kind in
this directory, waits for all nine objects to go Ready, then does what
the README's own first-login section cannot (see niklas/terdut-server#23
and niklas/terdut-operator#3 -- no service-account credential this
operator holds can ever call /api/admin/settings or POST /api/users) by
reaching into the demo's own throwaway Postgres directly: flips
signup_mode to open, signs alice up for real over the ordinary signup
endpoint, and joins her to both Platform and Payments (open signup always
creates its own new team, never joins an existing one by name, so
without this she'd have a working login that can't see a single incident
this demo fires -- /api/incidents and /api/alerts are both scoped to the
caller's own team memberships). Finishes by port-forwarding the service
and firing fire-alerts.sh at both teams, so a fresh run already has
visible incidents waiting in the web UI.

Verified end to end against a real kind cluster, including a second,
genuinely-fresh run that hit niklas/terdut-operator#3 live (terdutteam-
platform wedged in the 403 retry loop that issue describes) -- confirmed
the script itself fails cleanly on that (clear FAILED message, correct
exit code, no orphaned port-forward) rather than hanging or leaving a
mess, which is the most this script can do about a bug in the operator
it's driving.
2026-10-02 21:23:44 +02:00
Niklas Ye 375b5ed2f7 Set the chart's placeholder version to 0.2.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 57s
CI / test (push) Successful in 2m13s
Release / test (push) Successful in 1m41s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m10s
Release / scan-image (push) Successful in 3s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree
heading for v0.2.0 that still says 0.1.2 tells its reader something
false. Same as 738c210 and 0e3118d before it.
2026-10-02 18:51:00 +02:00
Niklas Ye 6a699d4341 Let TerdutServer customize its pod, and never manage its own ingress
spec.pod (api/v1alpha1/terdutserver_types.go): annotations, nodeSelector,
tolerations, affinity, topologySpreadConstraints, resources, pod and
container securityContext, serviceAccountName, extraEnv/extraEnvFrom,
extraVolumes/extraVolumeMounts, imagePullSecrets, and an optional
disruptionBudget. All direct corev1 passthrough -- no wrapper types buy
anything for any of these, matching how CloudNativePG and the Zalando
postgres-operator both expose the same knobs, and matching this repo's
own SweeperSpec precedent ("wrap only when a round-trip through a
different type buys something"). affinity is pure user-supplied
passthrough, not a toggle-plus-generated-default the way a multi-replica
cluster operator's pod anti-affinity usually is: this operator never
auto-generates one, since spec.replicas above 1 isn't a supported
topology (the sweeper/notifier singleton constraint). Considered and
declined for this round: priorityClassName, pod labels beyond
annotations, and a HorizontalPodAutoscaler -- the last of those would
directly contradict the singleton constraint above.

disruptionBudget is the one field here that isn't a plain PodTemplateSpec
knob: when set, the controller now reconciles a PodDisruptionBudget
selecting the TerdutServer's own pods (new terdutserver_pdb.go); clearing
it deletes any it previously created. New RBAC marker on
poddisruptionbudgets to match.

Driven by a public-release pass: looking past this project's own use case
at what a mature, general-purpose operator CRD exposes here (researched
against Zalando postgres-operator and CloudNativePG specifically), not
just the fields this install happened to need.

Separately, and found while answering a question about exposing
TerdutServer through Istio instead of Gateway API: spec.networking's own
doc comment quietly promised a Gateway API HTTPRoute this operator would
build eventually ("a near-term follow-up, not deferred"). That promise is
wrong for a public release -- an operator managing someone's ingress
mechanism for them is a worse default than not touching it at all, and a
surprise HTTPRoute appearing once that follow-up eventually landed would
have been exactly backwards for an Istio (or plain-Ingress, or
intentionally-unexposed) install. Made the non-goal explicit and
permanent instead (DESIGN.md §1), removed the dead `gatewayListener`
field it was the only consumer of (zero runtime call sites anywhere --
setting it already had no effect, so this is a schema cleanup, not a
behavior change), and corrected ROADMAP.md's framing. hostname/servicePort
stay: both are live (TERDUT_PUBLIC_URL, container/Service port), this
operator just never acts on hostname for exposure. Added
examples/networking (Gateway API HTTPRoute, Istio VirtualService) showing
how to expose the plain ClusterIP Service the operator already creates --
outside the operator itself, as illustrations, not as something
examples/demo applies automatically.

No new terdut-server version requirement: both changes are CRD/controller-
only, nothing about the API this operator's bootstrap flow depends on
changed.
2026-10-02 18:50:38 +02:00
Niklas Ye d9315322fc examples/demo: fix two real bugs this exact demo just hit live
CI / chart (push) Successful in 1s
CI / security (push) Successful in 3m24s
CI / test (push) Successful in 10m45s
1. Renamed every object this demo creates (TerdutServer, Postgres
   Secret/Deployment/Service) from terdut-demo[-postgres] to
   terdut-operator-demo[-postgres]. The user applied this kit into the
   already-live "terdut-demo" namespace -- the real operator exercise
   from earlier in this repo's own history -- and this demo's own
   TerdutServer/Postgres objects shared that exact name. The TerdutServer
   apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
   while the live object already had postgresClusterRef violates "exactly
   one of" and the API server refused it), and the real Postgres Service
   was never touched (confirmed live: still Zalando's own spilo selector,
   endpoint still the real StatefulSet pod) -- but the Postgres Secret and
   Deployment, having no such protection, were created as brand new,
   extra, crash-looping objects sitting right next to the real ones.
   Prefixing every name this demo creates means a repeat of this exact
   mistake no longer collides with anything, documented directly in
   README.md now.

2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
   (added responding to a PodSecurity "restricted" warning) took
   CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
   entrypoint needs to chown/chmod the data directory before it drops
   privileges itself -- confirmed in a real crashed pod's logs: `chmod:
   /var/run/postgresql: Operation not permitted`. kubectl apply
   --dry-run=server, which is as far as this got verified before, only
   checks admission policy; it was never actually booted. Removed the
   capability drop and verified for real this time: applied just
   00-postgres.yaml alone into a disposable namespace, waited for the pod
   to go Ready, read its logs ("database system is ready to accept
   connections"), then deleted that namespace.
2026-10-02 13:25:47 +02:00
Niklas Ye 5b45cf72e1 Bump go.opentelemetry.io/otel to v1.45.0: v1.44.0 carries GO-2026-6505
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m22s
CI / test (push) Successful in 5m0s
Release / test (push) Successful in 5m26s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 7m22s
Release / scan-image (push) Successful in 34s
Exporter config logging may leak endpoint URLs in info logs
(otlptrace/otlptracegrpc/sdk, transitively through grpc's own otel
instrumentation -- all indirect in go.mod, nothing imports these by
name). govulncheck flagged it reachable through real call chains
(tdclient.Client.DeleteIntegration, cmd/main.go's own init), caught by
ci.yaml's security job while cutting v0.1.2 (run 897) -- same pattern as
terdut-server's own da48814 for grpc's CVE-2026-84445.

go.opentelemetry.io/otel, /metric, /sdk, /sdk/metric, /trace,
/exporters/otlp/otlptrace, /exporters/otlp/otlptrace/otlptracegrpc all
moved 1.44.0 -> 1.45.0 together, plus go-logr/logr's own patch bump and
proto/otlp + genproto that go mod tidy pulled along with them. Verified:
go build, full test suite (71.7% coverage unchanged), golangci-lint,
helm-lint, and govulncheck itself now reporting zero reachable
vulnerabilities.
2026-10-02 12:59:40 +02:00
Niklas Ye 738c210505 Set the chart's placeholder version to 0.1.2
CI / chart (push) Successful in 1s
CI / security (push) Failing after 1m4s
CI / test (push) Successful in 2m11s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree heading
for v0.1.2 that still says 0.1.1 tells its reader something false. Same
as 0e3118d before it.
2026-10-02 12:56:16 +02:00
Niklas Ye 46ba0e8d5c CLAUDE.md: the wrapper-chart one-time step is done, not still pending
CI / chart (push) Successful in 1s
CI / test (push) Has been cancelled
CI / security (push) Has been cancelled
Stale since this repo's actual first release (v0.1.1, 2026-10-01) already
did it -- Ryuvia/charts/terdut-operator already exists and already pins
v0.1.1. Caught while about to repeat the same wrong assumption for this
release.
2026-10-02 12:55:23 +02:00
Niklas Ye ae97d28444 Add wait-for-postgres init container to the generated Deployment
CI / chart (push) Successful in 1s
CI / security (push) Failing after 57s
CI / test (push) Successful in 2m0s
This operator's own Deployment template crash-looped a few times against
a from-scratch postgres-operator cluster still doing initdb and Patroni
leader election -- exactly the gap examples/demo's own README just
documented for it. terdut-server's ping-retry budget on startup
(internal/db/db.go in that repo) is sized for a much shorter, different
race (NetworkPolicy propagation, a few seconds), not genuine first-time
cluster creation, so it exhausted and the process exited before ever
binding its HTTP port -- a startupProbe cannot fix that, since the crash
happens before there is anything to probe. Same root cause and same fix
as charts/terdut-server's own deployment.yaml template as of that repo's
v0.33.2.

waitForPostgresContainer reuses dbEnv unchanged: both of
resolveDatabaseEnv's two paths (DSN, postgresClusterRef) put
TERDUT_DB_DSN first, so it's already exactly what pg_isready needs, and
pg_isready needs no credentials, so dbEnv's optional PGPASSWORD riding
along too is harmless rather than load-bearing.

Covered by the existing envtest suite (asserts on Containers[0], the main
container, unaffected by adding InitContainers) -- `make test` passes
unchanged, 71.7% coverage on internal/controller. Updates examples/demo's
own README, which no longer needs to warn about this.
2026-10-02 10:11:55 +02:00
Niklas Ye ca2cd2c645 Add examples/demo: one of every CRD, plus a script to fire alerts at it
CI / chart (push) Successful in 1s
CI / security (push) Failing after 1m7s
CI / test (push) Successful in 2m37s
A self-contained demo kit: a TerdutServer against a throwaway, bare
Postgres (bring-your-own DSN -- simplest path to stand up from nothing,
ROADMAP.md Stage 1's own note), two TerdutTeams, and each team's own
TerdutEscalationRule/TerdutDeadmanSwitch/TerdutAlertSource, so every CRD
this operator manages is exercised together rather than in isolation the
way config/samples' one-of-each already does.

fire-alerts.sh sends terdut-server's own amPayload/amAlert shape (read
from internal/api/alertmanager.go in that repo, not guessed from its
docs) at whichever TerdutAlertSource's generated webhook Secret it reads
the key out of -- high-cpu/disk-full/pod-crash scenarios to open and
resolve incidents, and a heartbeat scenario matching each team's dead
man's switch matcher, so stopping it demonstrates the switch noticing
silence on its own.

Verified server-side (kubectl apply --dry-run=server -k examples/demo)
against this operator's own dev cluster, which already has these CRDs
installed: every object validates. The one warning that cluster's
"restricted" PodSecurity raises (postgres:17-alpine's entrypoint needs to
start as root before it drops privileges itself) is noted inline in
00-postgres.yaml rather than worked around -- not a real production
pattern, and this Postgres exists only to be thrown away with the rest of
the demo namespace.

README.md walks through: applying, watching status, why a few early
CrashLoopBackOff restarts on terdut-demo itself are expected (this
operator's Deployment template has no wait-for-postgres init container
yet, unlike charts/terdut-server's chart as of v0.33.2), reaching the web
UI (port-forward -- spec.networking.hostname is accepted but nothing
creates an HTTPRoute for it yet), turning on open signup with the
operator's own generated admin token since the bootstrap-created account
has no password, firing alerts, and tearing down.
2026-10-02 10:03:49 +02:00
Niklas Ye a75b23c4ad DESIGN.md: record the missing OIDC trustEmail field, found exercising a real second install
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 2m14s
Found while standing up terdut-demo (Ryuvia/charts#275), a second real
TerdutServer against the same Authentik provider as production: OIDCSpec
has no trustEmail override, so a demo install copying production's OIDC
config otherwise verbatim silently runs with the wrong default for it.
Not fixed here -- recorded in §13 as a real, found gap, not a decision,
same as the mid-life teamRef note already there.
2026-10-01 19:39:20 +02:00
Niklas Ye 4ab04d29a8 DESIGN.md: document operator mode and what it deliberately doesn't lock
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m3s
CI / test (push) Successful in 2m22s
No operator-mode section existed here before -- terdut-server's own
README.md documents the feature, but this repo's design doc never
mentioned it. Added as §6 point 7, confirmed against source
(internal/api/middleware.go's OperatorModeBlock, router.go's opMode
wrapper): it blocks human writes to exactly the resources this
operator's CRDs manage (team identity, OIDC-group binding, escalation,
dead man's switches, integrations), and nothing else -- team membership,
invites, and the on-call schedule/rota stay human-editable regardless,
confirmed from the router rather than assumed from the README's prose
alone.
2026-10-01 19:16:53 +02:00
97 changed files with 11549 additions and 6547 deletions
-35
View File
@@ -1,35 +0,0 @@
{
"name": "Kubebuilder DevContainer",
"image": "golang:1.26",
"features": {
"ghcr.io/devcontainers/features/docker-in-docker:2": {
"moby": false,
"dockerDefaultAddressPool": "base=172.30.0.0/16,size=24"
},
"ghcr.io/devcontainers/features/git:1": {},
"ghcr.io/devcontainers/features/common-utils:2": {
"upgradePackages": true
}
},
"runArgs": ["--privileged", "--init"],
"customizations": {
"vscode": {
"settings": {
"terminal.integrated.shell.linux": "/bin/bash"
},
"extensions": [
"ms-kubernetes-tools.vscode-kubernetes-tools",
"ms-azuretools.vscode-docker"
]
}
},
"remoteEnv": {
"GO111MODULE": "on"
},
"onCreateCommand": "bash .devcontainer/post-install.sh"
}
-153
View File
@@ -1,153 +0,0 @@
#!/bin/bash
set -euo pipefail
echo "===================================="
echo "Kubebuilder DevContainer Setup"
echo "===================================="
# Verify running as root (required for installing to /usr/local/bin and /etc)
if [ "$(id -u)" -ne 0 ]; then
echo "ERROR: This script must be run as root"
exit 1
fi
echo ""
echo "Detecting system architecture..."
# Detect architecture using uname
MACHINE=$(uname -m)
case "${MACHINE}" in
x86_64)
ARCH="amd64"
;;
aarch64|arm64)
ARCH="arm64"
;;
*)
echo "WARNING: Unsupported architecture ${MACHINE}, defaulting to amd64"
ARCH="amd64"
;;
esac
echo "Architecture: ${ARCH}"
echo ""
echo "------------------------------------"
echo "Setting up bash completion..."
echo "------------------------------------"
BASH_COMPLETIONS_DIR="/usr/share/bash-completion/completions"
# Enable bash-completion in root's .bashrc (devcontainer runs as root)
if ! grep -q "source /usr/share/bash-completion/bash_completion" ~/.bashrc 2>/dev/null; then
echo 'source /usr/share/bash-completion/bash_completion' >> ~/.bashrc
echo "Added bash-completion to .bashrc"
fi
echo ""
echo "------------------------------------"
echo "Installing development tools..."
echo "------------------------------------"
# Install kind
if ! command -v kind &> /dev/null; then
echo "Installing kind..."
curl -Lo /usr/local/bin/kind "https://kind.sigs.k8s.io/dl/latest/kind-linux-${ARCH}"
chmod +x /usr/local/bin/kind
echo "kind installed successfully"
fi
# Generate kind bash completion
if command -v kind &> /dev/null; then
if kind completion bash > "${BASH_COMPLETIONS_DIR}/kind" 2>/dev/null; then
echo "kind completion installed"
else
echo "WARNING: Failed to generate kind completion"
fi
fi
# Install kubebuilder
if ! command -v kubebuilder &> /dev/null; then
echo "Installing kubebuilder..."
curl -Lo /usr/local/bin/kubebuilder "https://go.kubebuilder.io/dl/latest/linux/${ARCH}"
chmod +x /usr/local/bin/kubebuilder
echo "kubebuilder installed successfully"
fi
# Generate kubebuilder bash completion
if command -v kubebuilder &> /dev/null; then
if kubebuilder completion bash > "${BASH_COMPLETIONS_DIR}/kubebuilder" 2>/dev/null; then
echo "kubebuilder completion installed"
else
echo "WARNING: Failed to generate kubebuilder completion"
fi
fi
# Install kubectl
if ! command -v kubectl &> /dev/null; then
echo "Installing kubectl..."
KUBECTL_VERSION=$(curl -Ls https://dl.k8s.io/release/stable.txt)
curl -Lo /usr/local/bin/kubectl "https://dl.k8s.io/release/${KUBECTL_VERSION}/bin/linux/${ARCH}/kubectl"
chmod +x /usr/local/bin/kubectl
echo "kubectl installed successfully"
fi
# Generate kubectl bash completion
if command -v kubectl &> /dev/null; then
if kubectl completion bash > "${BASH_COMPLETIONS_DIR}/kubectl" 2>/dev/null; then
echo "kubectl completion installed"
else
echo "WARNING: Failed to generate kubectl completion"
fi
fi
# Generate Docker bash completion
if command -v docker &> /dev/null; then
if docker completion bash > "${BASH_COMPLETIONS_DIR}/docker" 2>/dev/null; then
echo "docker completion installed"
else
echo "WARNING: Failed to generate docker completion"
fi
fi
echo ""
echo "------------------------------------"
echo "Configuring Docker environment..."
echo "------------------------------------"
# Wait for Docker to be ready
echo "Waiting for Docker to be ready..."
for i in {1..30}; do
if docker info >/dev/null 2>&1; then
echo "Docker is ready"
break
fi
if [ "$i" -eq 30 ]; then
echo "WARNING: Docker not ready after 30s"
fi
sleep 1
done
# Create kind network (ignore if already exists)
if ! docker network inspect kind >/dev/null 2>&1; then
if docker network create kind >/dev/null 2>&1; then
echo "Created kind network"
else
echo "WARNING: Failed to create kind network (may already exist)"
fi
fi
echo ""
echo "------------------------------------"
echo "Verifying installations..."
echo "------------------------------------"
kind version
kubebuilder version
kubectl version --client
docker --version
go version
echo ""
echo "===================================="
echo "DevContainer ready!"
echo "===================================="
echo "All development tools installed successfully."
echo "You can now start building Kubernetes operators."
+2 -2
View File
@@ -55,7 +55,7 @@ jobs:
test:
runs-on: ubuntu-latest
container:
image: golang:1.26.6-bookworm
image: golang:1.26.9-bookworm
volumes:
- go-mod-cache:/go/pkg/mod
- go-build-cache:/root/.cache/go-build
@@ -85,7 +85,7 @@ jobs:
security:
runs-on: ubuntu-latest
container:
image: golang:1.26.6-bookworm
image: golang:1.26.9-bookworm
volumes:
- go-mod-cache:/go/pkg/mod
- go-build-cache:/root/.cache/go-build
+1 -1
View File
@@ -26,7 +26,7 @@ jobs:
test:
runs-on: ubuntu-latest
container:
image: golang:1.26.6-bookworm
image: golang:1.26.9-bookworm
volumes:
- go-mod-cache:/go/pkg/mod
- go-build-cache:/root/.cache/go-build
-320
View File
@@ -1,320 +0,0 @@
# terdut-operator - AI Agent Guide
## Project Structure
**Single-group layout (default):**
```
cmd/main.go Manager entry (registers controllers/webhooks)
api/<version>/*_types.go CRD schemas (+kubebuilder markers)
api/<version>/zz_generated.* Auto-generated (DO NOT EDIT)
internal/controller/* Reconciliation logic
internal/webhook/* Validation/defaulting (if present)
config/crd/bases/* Generated CRDs (DO NOT EDIT)
config/rbac/role.yaml Generated RBAC (DO NOT EDIT)
config/samples/* Example CRs (edit these)
Makefile Build/test/deploy commands
PROJECT Kubebuilder metadata Auto-generated (DO NOT EDIT)
```
**Multi-group layout** (for projects with multiple API groups):
```
api/<group>/<version>/*_types.go CRD schemas by group
internal/controller/<group>/* Controllers by group
internal/webhook/<group>/<version>/* Webhooks by group and version (if present)
```
Multi-group layout organizes APIs by group name (e.g., `batch`, `apps`). Check the `PROJECT` file for `multigroup: true`.
**To convert to multi-group layout:**
1. Run: `kubebuilder edit --multigroup=true`
2. Move APIs: `mkdir -p api/<group> && mv api/<version> api/<group>/`
3. Move controllers: `mkdir -p internal/controller/<group> && mv internal/controller/*.go internal/controller/<group>/`
4. Move webhooks (if present): `mkdir -p internal/webhook/<group> && mv internal/webhook/<version> internal/webhook/<group>/`
5. Update import paths in all files
6. Fix `path` in `PROJECT` file for each resource
7. Update test suite CRD paths (add one more `..` to relative paths)
## Critical Rules
### Never Edit These (Auto-Generated)
- `config/crd/bases/*.yaml` - from `make manifests`
- `config/rbac/role.yaml` - from `make manifests`
- `config/webhook/manifests.yaml` - from `make manifests`
- `**/zz_generated.*.go` - from `make generate`
- `PROJECT` - from `kubebuilder [OPTIONS]`
### Never Remove Scaffold Markers
Do NOT delete `// +kubebuilder:scaffold:*` comments. CLI injects code at these markers.
### Keep Project Structure
Do not move files around. The CLI expects files in specific locations.
### Always Use CLI Commands
Always use `kubebuilder create api` and `kubebuilder create webhook` to scaffold. Do NOT create files manually.
### E2E Tests Require an Isolated Kind Cluster
The e2e tests are designed to validate the solution in an isolated environment (similar to GitHub Actions CI).
Ensure you run them against a dedicated [Kind](https://kind.sigs.k8s.io/) cluster (not your “real” dev/prod cluster).
## After Making Changes
**After editing `*_types.go` or markers:**
```
make manifests # Regenerate CRDs/RBAC from markers
make generate # Regenerate DeepCopy methods
```
**After editing `*.go` files:**
```
make lint-fix # Auto-fix code style
make test # Run unit tests
```
## CLI Commands Cheat Sheet
### Create API (your own types)
```bash
kubebuilder create api --group <group> --version <version> --kind <Kind>
```
### Deploy Image Plugin (scaffold to deploy/manage ANY container image)
Generate a controller that deploys and manages a container image (nginx, redis, memcached, your app, etc.):
```bash
# Example: deploying memcached
kubebuilder create api --group example.com --version v1alpha1 --kind Memcached \
--image=memcached:alpine \
--plugins=deploy-image.go.kubebuilder.io/v1-alpha
```
Scaffolds good-practice code: reconciliation logic, status conditions, finalizers, RBAC. Use as a reference implementation.
### Create Webhooks
```bash
# Validation + defaulting
kubebuilder create webhook --group <group> --version <version> --kind <Kind> \
--defaulting --programmatic-validation
# Conversion webhook (for multi-version APIs)
kubebuilder create webhook --group <group> --version v1 --kind <Kind> \
--conversion --spoke v2
```
### Controller for Core Kubernetes Types
```bash
# Watch Pods
kubebuilder create api --group core --version v1 --kind Pod \
--controller=true --resource=false
# Watch Deployments
kubebuilder create api --group apps --version v1 --kind Deployment \
--controller=true --resource=false
```
### Controller for External Types (e.g., from other operators)
Watch resources from external APIs (cert-manager, Argo CD, Istio, etc.):
```bash
# Example: watching cert-manager Certificate resources
kubebuilder create api \
--group cert-manager --version v1 --kind Certificate \
--controller=true --resource=false \
--external-api-path=github.com/cert-manager/cert-manager/pkg/apis/certmanager/v1 \
--external-api-domain=io \
--external-api-module=github.com/cert-manager/cert-manager
```
**Note:** Use `--external-api-module=<module>@<version>` only if you need a specific version. Otherwise, omit `@<version>` to use what's in go.mod.
### Webhook for External Types
```bash
# Example: validating external resources
kubebuilder create webhook \
--group cert-manager --version v1 --kind Issuer \
--defaulting \
--external-api-path=github.com/cert-manager/cert-manager/pkg/apis/certmanager/v1 \
--external-api-domain=io \
--external-api-module=github.com/cert-manager/cert-manager
```
## Testing & Development
```bash
make test # Run unit tests (uses envtest: real K8s API + etcd)
make run # Run locally (uses current kubeconfig context)
```
Tests use **Ginkgo + Gomega** (BDD style). Check `suite_test.go` for setup.
## Deployment Workflow
```bash
# 1. Regenerate manifests
make manifests generate
# 2. Build & deploy
export IMG=<registry>/<project>:tag
make docker-build docker-push IMG=$IMG # Or: kind load docker-image $IMG --name <cluster>
make deploy IMG=$IMG
# 3. Test
kubectl apply -k config/samples/
# 4. Debug
kubectl logs -n <project>-system deployment/<project>-controller-manager -c manager -f
```
### API Design
**Key markers for** `api/<version>/*_types.go`:
```go
// +kubebuilder:object:root=true
// +kubebuilder:subresource:status
// +kubebuilder:resource:scope=Namespaced
// +kubebuilder:printcolumn:name="Status",type=string,JSONPath=".status.conditions[?(@.type=='Ready')].status"
// On fields:
// +kubebuilder:validation:Required
// +kubebuilder:validation:Minimum=1
// +kubebuilder:validation:MaxLength=100
// +kubebuilder:validation:Pattern="^[a-z]+$"
// +kubebuilder:default="value"
```
- **Use** `metav1.Condition` for status (not custom string fields)
- **Use predefined types**: `metav1.Time` instead of `string` for dates
- **Follow K8s API conventions**: Standard field names (`spec`, `status`, `metadata`)
### Controller Design
**RBAC markers in** `internal/controller/*_controller.go`:
```go
// +kubebuilder:rbac:groups=mygroup.example.com,resources=mykinds,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=mygroup.example.com,resources=mykinds/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=mygroup.example.com,resources=mykinds/finalizers,verbs=update
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
// +kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete
```
**Implementation rules:**
- **Idempotent reconciliation**: Safe to run multiple times
- **Re-fetch before updates**: `r.Get(ctx, req.NamespacedName, obj)` before `r.Update` to avoid conflicts
- **Structured logging**: `log := log.FromContext(ctx); log.Info("msg", "key", val)`
- **Owner references**: Enable automatic garbage collection (`SetControllerReference`)
- **Watch secondary resources**: Use `.Owns()` or `.Watches()`, not just `RequeueAfter`
- **Finalizers**: Clean up external resources (buckets, VMs, DNS entries)
### Logging
**Follow Kubernetes logging message style guidelines:**
- Start from a capital letter
- Do not end the message with a period
- Active voice: subject present (`"Deployment could not create Pod"`) or omitted (`"Could not create Pod"`)
- Past tense: `"Could not delete Pod"` not `"Cannot delete Pod"`
- Specify object type: `"Deleted Pod"` not `"Deleted"`
- Balanced key-value pairs
```go
log.Info("Starting reconciliation")
log.Info("Created Deployment", "name", deploy.Name)
log.Error(err, "Failed to create Pod", "name", name)
```
**Reference:** https://github.com/kubernetes/community/blob/master/contributors/devel/sig-instrumentation/logging.md#message-style-guidelines
### Webhooks
- **Create all types together**: `--defaulting --programmatic-validation --conversion`
- **When`--force`is used**: Backup custom logic first, then restore after scaffolding
- **For multi-version APIs**: Use hub-and-spoke pattern (`--conversion --spoke v2`)
- Hub version: Usually oldest stable version (v1)
- Spoke versions: Newer versions that convert to/from hub (v2, v3)
- Example: `--group crew --version v1 --kind Captain --conversion --spoke v2` (v1 is hub, v2 is spoke)
### Learning from Examples
The **deploy-image plugin** scaffolds a complete controller following good practices. Use it as a reference implementation:
```bash
kubebuilder create api --group example --version v1alpha1 --kind MyApp \
--image=<your-image> --plugins=deploy-image.go.kubebuilder.io/v1-alpha
```
Generated code includes: status conditions (`metav1.Condition`), finalizers, owner references, events, idempotent reconciliation.
## Distribution Options
### Option 1: YAML Bundle (Kustomize)
```bash
# Generate dist/install.yaml from Kustomize manifests
make build-installer IMG=<registry>/<project>:tag
```
**Key points:**
- The `dist/install.yaml` is generated from Kustomize manifests (CRDs, RBAC, Deployment)
- Commit this file to your repository for easy distribution
- Users only need `kubectl` to install (no additional tools required)
**Example:** Users install with a single command:
```bash
kubectl apply -f https://raw.githubusercontent.com/<org>/<repo>/<tag>/dist/install.yaml
```
### Option 2: Helm Chart
```bash
kubebuilder edit --plugins=helm/v2-alpha # Generates dist/chart/ (default)
kubebuilder edit --plugins=helm/v2-alpha --output-dir=charts # Generates charts/chart/
```
**For development:**
```bash
make helm-deploy IMG=<registry>/<project>:<tag> # Deploy manager via Helm
make helm-deploy IMG=$IMG HELM_EXTRA_ARGS="--set ..." # Deploy with custom values
make helm-status # Show release status
make helm-uninstall # Remove release
make helm-history # View release history
make helm-rollback # Rollback to previous version
```
**For end users/production:**
```bash
helm install my-release ./<output-dir>/chart/ --namespace <ns> --create-namespace
```
**Important:** If you add webhooks or modify manifests after initial chart generation:
1. Backup any customizations in `<output-dir>/chart/values.yaml` and `<output-dir>/chart/manager/manager.yaml`
2. Re-run: `kubebuilder edit --plugins=helm/v2-alpha --force` (use same `--output-dir` if customized)
3. Manually restore your custom values from the backup
### Publish Container Image
```bash
export IMG=<registry>/<project>:<version>
make docker-build docker-push IMG=$IMG
```
## References
### Essential Reading
- **Kubebuilder Book**: https://book.kubebuilder.io (comprehensive guide)
- **controller-runtime FAQ**: https://github.com/kubernetes-sigs/controller-runtime/blob/main/FAQ.md (common patterns and questions)
- **Good Practices**: https://book.kubebuilder.io/reference/good-practices.html (why reconciliation is idempotent, status conditions, etc.)
- **Logging Conventions**: https://github.com/kubernetes/community/blob/master/contributors/devel/sig-instrumentation/logging.md#message-style-guidelines (message style, verbosity levels)
### API Design & Implementation
- **API Conventions**: https://github.com/kubernetes/community/blob/master/contributors/devel/sig-architecture/api-conventions.md
- **Operator Pattern**: https://kubernetes.io/docs/concepts/extend-kubernetes/operator/
- **Markers Reference**: https://book.kubebuilder.io/reference/markers.html
### Tools & Libraries
- **controller-runtime**: https://github.com/kubernetes-sigs/controller-runtime
- **controller-tools**: https://github.com/kubernetes-sigs/controller-tools
- **Kubebuilder Repo**: https://github.com/kubernetes-sigs/kubebuilder
+11 -16
View File
@@ -1,8 +1,8 @@
# terdut-operator
Kubebuilder/controller-runtime operator for terdut-server. See `DESIGN.md` for the
settled design (CRD catalog, reconciliation semantics, bootstrap/auth, RBAC) and
`ROADMAP.md` for the staged build plan this repo is following. `README.md` stays the
design (start with its revision section: credentials and the CRD catalog) and
`ROADMAP.md` for status and deferred work. `README.md` stays the
short pitch.
## Checks
@@ -16,15 +16,12 @@ required beyond Go itself and network access to `proxy.golang.org`/`storage.goog
`security` runs `make security-go`/`security-secrets` (govulncheck/gitleaks), same as
terdut-server's own `security` job.
The kubebuilder-scaffolded `make test-e2e` (a disposable, generic smoke test) is separate
from the real golden-path `kind` e2e pass ROADMAP.md's Stage 5 describes (create every CRD
kind, verify against terdut-server's own API, delete, verify gone) — the latter is a
manual pass run and recorded in ROADMAP.md, not a CI job, matching Stage 1-4's own
precedent of validating against a real cluster outside CI.
There is no CI end-to-end job: the golden-path `kind` pass is manual, via
`examples/demo/run-demo.sh` (see ROADMAP.md).
## Release
Wired as of Stage 5 (ROADMAP.md): `.release.conf`, `make release-vars`/`helm-lint`/
Wired: `.release.conf`, `make release-vars`/`helm-lint`/
`push`/`helm-package`/`helm-push`/`release`, and `.gitea/workflows/release.yaml`
(`test` → `image`/`chart` → `scan-image`) all follow terdut-server's established shape —
see that repo's Makefile/`.release.conf` for the shared reasoning, not restated here.
@@ -36,11 +33,9 @@ charts --force` after `config/` changes, then re-review — `--force` does not t
importantly the optional `terdutServer` block, DESIGN.md §10). It installs the operator +
CRDs + RBAC, and optionally one `TerdutServer` CR (`terdutServer.enabled`, off by default).
**One manual step the release skill's own automation does not cover**: `release-preflight`
expects an existing `terdut-operator/` entry under `Ryuvia/charts` to bump on release
(steps 8-10 of the skill). There is no such entry yet — this repo's first-ever release
can publish its own image and chart (the `test`/`image`/`chart`/`scan-image` jobs), but
the wrapper-chart bump and PR will fail until someone creates that initial wrapper entry
in `Ryuvia/charts` by hand, the same one-time step every other onboarded repo already had
done for it before its own first release. That's a deliberate decision to deploy this
operator for real, not something to do as a side effect of finishing this stage.
That one-time manual step — hand-creating the initial `terdut-operator/` wrapper entry
under `Ryuvia/charts`, since `chart-bump` only ever bumps an existing one — is done. It
happened during this repo's actual first release (v0.1.1, 2026-10-01; v0.1.0 published but
never deployed anywhere, after its own `scan-image` found a CVE in the grpc version it had
just bumped to). Every release since bumps that wrapper entry like any other onboarded
repo's.
+214 -535
View File
@@ -1,138 +1,83 @@
# terdut-operator design
This is the design reference for implementing terdut-operator. It exists so
implementation can start from settled decisions instead of re-litigating them
mid-PR. It is deliberately more detailed than the README; the README stays as
the short pitch and now points here instead of carrying open questions.
Written against terdut-server as of the Postgres-only, per-team-resources
version (teams, escalation policies, dead man's switches and integrations are
all rows scoped to a team, managed over `internal/api/*` — not env/config-file
driven; see that repo's `charts/terdut-server` for the current deploy story
this operator supersedes).
The design of terdut-operator as built. terdut-server is the system of record; this operator
makes one install of it — the server, its teams, their escalation ladders, dead man's switches
and alert-source integrations — describable as Kubernetes objects and manageable through gitops.
## 1. Goals & non-goals
**Goal:** let a terdut-server install — the server itself, its teams, their
escalation policies, dead man's switches and alert-source integrations — be
fully described as Kubernetes objects and managed through gitops, following
**Goal:** a terdut-server install fully described as Kubernetes objects, following
controller-runtime / Kubebuilder conventions.
**The operator creates and owns every `TerdutServer` it manages. It never
adopts a pre-existing, independently-deployed terdut-server** — whether
deployed by hand or by `charts/terdut-server`. There is no migration path
from an existing chart-based install, and none is planned (§10): starting
with the operator means applying a fresh `TerdutServer` CR, not converting
one. This is the root a few things downstream hang off of — notably §6's
bootstrap flow, which only has to handle the operator bootstrapping a server
it just created, never a server something else already bootstrapped first.
**The operator creates and owns every `TerdutServer` it manages. It never adopts a pre-existing,
independently-deployed terdut-server**, whether deployed by hand or by `charts/terdut-server`.
There is no migration path from a chart-based install (§10): starting with the operator means
applying a fresh `TerdutServer`.
**Non-goals (v1):**
- Not a general-purpose Postgres operator. It *consumes* a database that
either the Zalando `postgres-operator` or something else already provides.
- Not managing Alertmanager itself, or the routing rules that decide which
alerts reach which integration webhook — only the terdut-server side
(creating the integration and handing back its URL/key).
- Not OLM packaging. Plain Kubebuilder manifests + Helm chart for
install, matching how terdut-server itself ships.
- Cross-namespace references are limited to exactly one edge:
`TerdutTeam.spec.serverRef` may name a `TerdutServer` in a different
namespace, gated by that `TerdutServer`'s own `spec.allowedTeams` consent
field (§4.1, §4.2) — this is the multi-tenant shape the operator exists
for (one platform team owns a `TerdutServer`; other teams self-service a
`TerdutTeam` against it without needing write access to the server's
namespace). Every other reference (`teamRef` on the escalation
rule/dead-man-switch/alert-source CRDs) stays same-namespace-as-its-`TerdutTeam`
only, in v1 — those manage a specific team's own resources and are
expected to live alongside it.
- No validating/mutating admission webhooks in v1. CEL validation rules on
the CRDs (OpenAPI `x-kubernetes-validations`) cover what they can; anything
that needs a live look at another object (e.g. "does this teamRef exist")
is a status condition, not an admission rejection — keeps v1 to a
controller-only deployment with no cert-manager/webhook dependency.
- Not a Postgres operator. It consumes a database that the Zalando `postgres-operator` or
something else provides (§8).
- Not managing Alertmanager or its routing, only the terdut-server side (creating the
integration and handing back its URL/key).
- Not OLM packaging: plain Kubebuilder manifests and a Helm chart, like terdut-server.
- Cross-namespace references are limited to one edge: `TerdutTeam.spec.serverRef` may name a
`TerdutServer` in another namespace, gated by that server's `spec.allowedTeams` (§4.6). A
`TerdutAlertSource` lives beside its `TerdutTeam`.
- No admission webhooks. CEL validation covers what it can; anything needing a live look at
another object is a status condition, not an admission rejection.
- **Never manages external exposure for a `TerdutServer`** (Ingress, HTTPRoute, VirtualService) —
a permanent non-goal. The operator creates a plain `ClusterIP` Service; see `examples/networking`.
## 2. The README's open questions, resolved
## 2. Credentials and CRD shape (the 2026-10 redesign)
> How do team-crd connect with server-crd?
What the design is, and why — each point replaced something heavier:
Explicit `spec.serverRef: {name, namespace}` on `TerdutTeam` — same reasoning
as before (explicit, greppable, trivially validated), but **`namespace` is
deliberately part of the reference**: one team can run and own a
`TerdutServer`, and other teams — in their own namespaces, without any write
access to the server-owning team's namespace — self-service a `TerdutTeam`
against it. `namespace` defaults to the `TerdutTeam`'s own namespace when
omitted, so the common single-tenant case (`serverRef: {name: terdut}`) is
unchanged.
- **No bootstrap handshake.** The operator generates a key into a Secret named
`<TerdutServer>-operator-key`, in the TerdutServer's own namespace and owned by
it (a pod can only mount Secrets of its own namespace, and an owner reference
replaces the old finalizer and `credentials.deletionPolicy`). The Deployment
hands it to terdut-server as `TERDUT_OPERATOR_KEY`; the server creates or
re-keys its instance-scoped service account `terdut-operator` from it at every
start. A replaced Secret rolls the pods (key-hash annotation). `/api/bootstrap`
stays free for the first human administrator. Gone with it: the checkpoint
Secret, `BootstrapStateLost`, `DatabaseReady`/`Bootstrapped` conditions.
- **One credential per server.** An instance-scoped account acts as owner of every
team's configuration (not a member, so it reads no incidents). The per-team
service accounts, Secrets and `status.credentialsSecretRef`/`serverEndpoint` on
`TerdutTeam` are gone.
- **Team identity is `external_id`.** The operator creates a team with
`external_id: <namespace>/<name>` of its CR; the server returns the existing team
for a known id (200) instead of creating one, so crash recovery, a lost status
and a deleted team all heal by repeating the same call, and a display name that
belongs to another team is a 409 (`TeamNameTaken`) instead of an adoption.
- **Three CRDs.** `TerdutServer`, `TerdutTeam` and `TerdutAlertSource`.
`TerdutEscalationRule` and `TerdutDeadmanSwitch` are `spec.escalation` and
`spec.deadmanSwitches[]` (matched by name, unique per team server-side; ones not
listed are removed) on the team: they were one-to-one children with the team's
lifecycle, and folding them removes `teamRef`, the two-rules-clobber footgun and
three controllers. `TerdutAlertSource` stays separate because it owns a webhook
Secret in its own namespace.
- **Usernames are resolved by the server** (`PUT .../escalation` accepts
`username`), so an unknown user is a `UnknownUser` condition, not a list-and-match.
- **Invites are removed.** Membership is not modelled; people get in through the
server's own signup/OIDC.
- **Env mirrors the chart's where it matters:** OIDC claim names and `trustEmail`
are spec fields; `TERDUT_DEADMAN_*` no longer exist on the server.
This cross-namespace edge needs the target namespace's explicit consent —
otherwise any namespace in the cluster could point a `TerdutTeam` at
someone else's `TerdutServer` and have the operator provision a team on
its behalf, which is a namespace-boundary violation, not a gitops
convenience. (This consent gate is about which `TerdutTeam`s the operator
will act on, not about credential exposure — no terdut-server credential
is ever placed in a `TerdutTeam`'s own namespace regardless of this
setting; see §6.) Kubernetes has two established patterns for this kind
of consent, and Gateway API itself uses both, for two different
relationships:
## 3. CRD catalog
- **`ReferenceGrant`** (used for a Route reaching into an arbitrary
Service/Secret): a separate object, living in the *target* namespace,
enumerating exact `{fromNamespace, fromKind} → {toKind, toName}` pairs.
No wildcard, no selector — every permitted namespace is spelled out.
- **An inline field on the parent** (used for a `ListenerSet` attaching to
a shared `Gateway`, GA since Gateway API v1.5): the parent carries
`spec.allowedListeners.namespaces: {from: None|Same|All|Selector,
selector}` directly, no separate CRD.
Group `terdut.ryuvia.com`, version `v1alpha1`; module `git.ryuvia.com/niklas/terdut-operator`.
`TerdutTeam` attaching to a shared `TerdutServer` is structurally the
second case, not the first — a bounded set of expected children attaching
to a parent they were deliberately made shareable, not an arbitrary
backend reference — so this design follows the `ListenerSet` precedent:
`TerdutServer.spec.allowedTeams` (§4.1), no extra CRD. §4.2 covers how
`TerdutTeam` resolves against it. No other reference in this design
(`teamRef` on the child CRDs) crosses a namespace boundary, so this is the
only place cross-namespace consent is needed at all (§1, §5, §6, §9).
> How does escalationrules, switches and alertsources connect to a team?
Explicit `spec.teamRef: {name}` on each of `TerdutEscalationRule`,
`TerdutDeadmanSwitch`, `TerdutAlertSource` — same reasoning, and it mirrors
terdut-server's own data model, where every one of these rows carries a
`team_id` foreign key already. A matcher/selector on `TerdutTeam` would be
inventing a second source of truth for an association the server already
models as a plain reference.
> Support for both postgres-operator (Zalando) and bring-your-own, how do we
> design that to be user friendly?
`TerdutServer.spec.database` is a oneOf, mirroring the chart's existing
`database.dsn` / `database.passwordSecret` contract (see §8):
- `dsn` + `passwordSecretRef` — bring-your-own, exactly today's chart inputs.
- `postgresClusterRef` — points at a Zalando `postgresql.acid.zalan.do` CR;
the operator derives the DSN and resolves the generated credentials Secret
itself (see §8). CloudNativePG support is a natural follow-up using the
same shape and is called out as deferred (§13), not designed in detail now.
## 3. API group, versions, CRD catalog
- Group: `terdut.ryuvia.com`, version: `v1alpha1` (matches the `ryuvia.com`
domain terdut-server already uses; bump to `v1beta1`/`v1` per the normal
Kubernetes API graduation criteria once the shapes below have proven
stable against a real install).
- Module: `git.ryuvia.com/niklas/terdut-operator`, scaffolded with
Kubebuilder (controller-runtime), matching terdut-server's Go toolchain
and house style.
| Kind | Scope | Purpose |
|---|---|---|
| `TerdutServer` | Namespaced | One terdut-server install: Deployment, Service, database wiring, bootstrap, operator credentials, cross-namespace team consent. |
| `TerdutTeam` | Namespaced | One team on a `TerdutServer`, possibly in another namespace: name, OIDC group mapping. |
| `TerdutEscalationRule` | Namespaced | A team's escalation policy (levels, targets, repeat). |
| `TerdutDeadmanSwitch` | Namespaced | One dead man's switch on a team. |
| `TerdutAlertSource` | Namespaced | One alert-ingest integration on a team (currently: Alertmanager webhook). |
| Kind | Purpose |
|---|---|
| `TerdutServer` | One terdut-server install: Deployment, Service, database wiring, operator key, cross-namespace team consent. |
| `TerdutTeam` | One team on a server, possibly in another namespace: name, OIDC groups, escalation ladder, dead man's switches. |
| `TerdutAlertSource` | One alert-ingest integration on a team; owns a webhook Secret. |
## 4. Per-CRD spec
(The escalation and dead man's switch specs that used to be §4.3 and §4.4 are part of §4.2; the section numbers are kept because code comments cite them.)
### 4.1 `TerdutServer`
```yaml
@@ -145,11 +90,10 @@ spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: v0.9.3
replicas: 1 # terdut-server is not horizontally-scale-tested; keep the field, default 1
replicas: 2 # default since terdut-server v0.36.0's advisory locks; see TerdutServerSpec.Replicas
networking:
hostname: terdut.example.com
servicePort: 8080
gatewayListener: "" # same semantics as chart's networking.listener
database:
dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef
passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password}
@@ -158,10 +102,6 @@ spec:
sweeper:
staleAfter: 6h
archiveAfter: 168h
deadman:
matchers: "alertname=Watchdog"
timeout: 15m
severity: critical
notify:
ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
fallbackTopic: ""
@@ -176,6 +116,8 @@ spec:
scopes: "openid profile email"
allowedGroups: []
adminGroup: ""
usernameClaim: preferred_username # also emailClaim, groupsClaim
trustEmail: false
sessionMaxAge: 12h
passwordLogin: true
# Consent for TerdutTeams in OTHER namespaces to set serverRef at this
@@ -188,11 +130,24 @@ spec:
# selector: # required, and only meaningful, when from: Selector
# matchLabels:
# terdut.ryuvia.com/allowed: "true"
# Pod-level customization of the Deployment -- all optional, direct corev1
# passthrough throughout (see PodSpec's own doc comment). A representative
# subset:
pod:
resources:
requests: {cpu: 100m, memory: 128Mi}
limits: {memory: 256Mi}
tolerations:
- key: dedicated
operator: Equal
value: terdut
effect: NoSchedule
disruptionBudget:
minAvailable: 1 # mutually exclusive with maxUnavailable
status:
conditions: [...] # Ready, DatabaseReady, Bootstrapped
conditions: [...] # Ready
observedGeneration: 3
serviceName: terdut
credentialsSecretRef: {name: terdut.platform-oncall-instance-credentials, key: token} # see §6; pure output -- generated by the controller's own self-registration flow, always in the OPERATOR's namespace (always, implicitly -- not stored here, since it's never anything else), under a fixed key ("token").
credentialsSecretRef: {name: terdut-operator-key, key: token} # see §6; pure output, in this TerdutServer's own namespace
```
Field-for-field this is the chart's `values.yaml` reshaped as a spec — not
@@ -214,93 +169,60 @@ per-name allowlist (no "and only these teams") — namespace-level consent is
the right granularity here, same reasoning as `ListenerSet`: the namespace
is the tenancy boundary, not the object.
`spec.pod` is pod-level customization of the Deployment, all optional and
directly reusing corev1 types wherever corev1 already models the knob
exactly (`tolerations`, `affinity`, `topologySpreadConstraints`,
`resources`, `securityContext`/`containerSecurityContext`, `extraEnv`/
`extraEnvFrom`, `extraVolumes`/`extraVolumeMounts`, `imagePullSecrets`) —
no custom wrapper buys anything for any of these, matching how
CloudNativePG and the Zalando postgres-operator both expose the same
knobs. `affinity` is pure user-supplied passthrough, not a
toggle-plus-generated-default the way a multi-replica-aware operator's
pod anti-affinity typically is: even though `replicas` now defaults to 2
(terdut-server v0.36.0's advisory locks made that safe, §4.1's own
illustrative YAML comment), this operator still never auto-generates
affinity of its own. `spec.pod.disruptionBudget` is the one field here that isn't a
straight PodTemplateSpec knob — when set, the controller reconciles a
`PodDisruptionBudget` selecting this `TerdutServer`'s pods; clearing it
deletes any it previously created (§7). `minAvailable`/`maxUnavailable`
are mutually exclusive, `+kubebuilder:validation:XValidation`-guarded the
same way as `spec.database`'s own `dsn`/`postgresClusterRef` rule.
### 4.2 `TerdutTeam`
```yaml
spec:
serverRef:
name: terdut
namespace: platform-oncall # optional; defaults to this TerdutTeam's own namespace.
# Cross-namespace requires that TerdutServer's spec.allowedTeams
# (§4.1) to admit this namespace — otherwise Ready: False, reason: RefNotPermitted.
displayName: "Platform" # -> POST /api/teams {"name": ...}; server assigns the ID
oidc:
memberGroup: "terdut-platform-members"
ownerGroup: "terdut-platform-owners"
serverRef: {name: terdut, namespace: platform-oncall} # namespace optional; cross-namespace needs allowedTeams
displayName: Platform
oidc: {memberGroup: terdut-platform-members, ownerGroup: terdut-platform-owners}
escalation:
repeatCount: 2
fallbackTopic: platform-oncall
levels:
- timeout: 5m
targets: [{kind: oncall}, {kind: user, username: alice}]
- timeout: 15m
targets: [{kind: oncall}]
deadmanSwitches:
- {name: watchdog, matcher: "alertname=Watchdog,cluster=prod", timeout: 15m, severity: critical}
status:
conditions: [...]
teamID: 42 # the server-side ID; needed by every child object's controller
credentialsSecretRef: {name: platform-oncall.platform-team-credentials, key: token} # see §6; this team's own scoped key, always in the OPERATOR's own namespace (not stored here — same reasoning as TerdutServer's, §4.1)
serverEndpoint: "http://terdut.platform-oncall.svc:8080" # resolved once, here, from spec.serverRef -- see §5: this is what actually makes "a child never needs to chain up to TerdutServer" true, not just a stated intent. Set alongside teamID/credentialsSecretRef, same reconcile.
conditions: [...] # Ready; reasons: ServerRefNotFound, RefNotPermitted, WaitingForServer,
# TeamNameTaken, UnknownUser, InvalidSpec, Adopted
teamID: 42
observedGeneration: 1
```
Team *membership* (which users belong, `team_members`) is explicitly **not**
modeled as a CRD field in v1: terdut-server already manages membership via
OIDC group sync at login for SSO installs, and manual membership for
password-login installs is a people-management action, not infrastructure —
forcing it through gitops would mean a human's team change goes through a PR
review. Flagged in §13 as revisitable if a real gitops-membership need shows up.
The team is created under the identity `<namespace>/<name>` of its CR (`external_id`), so a retry,
a lost status or a team deleted behind the operator's back all heal by repeating the same call; a
`displayName` that belongs to a different team is `TeamNameTaken`, never an adoption. Every
reconcile applies, in order: create-or-find, rename, OIDC groups, the escalation ladder (one PUT
of the whole ladder; absent `spec.escalation` clears it), and the switches. Switches are matched
by name (unique per team on the server): missing ones are created, changed ones updated in
place, and any not listed are deleted, since in operator mode nobody else can add one. Durations
are validated by a CRD pattern and re-checked (`InvalidSpec`). A username the server does not
know is `UnknownUser` until that person exists.
### 4.3 `TerdutEscalationRule`
```yaml
spec:
teamRef: {name: platform-team}
repeatCount: 2
fallbackTopic: "platform-oncall"
levels:
- timeout: 5m
targets:
- kind: oncall # "oncall" or "user"
- kind: user
username: alice # resolved to a user ID by the controller at apply time
- timeout: 15m
targets:
- kind: oncall
status:
conditions: [...]
observedGeneration: 1
```
One `TerdutEscalationRule` per team — the server itself models a policy as
one row (`escalation_policies`) with an owned list of levels, so a
one-CRD-to-one-policy mapping (not one-CRD-per-level) matches the server's
own aggregate and lets the whole thing be reconciled with the single
`PUT /api/teams/{teamID}/escalation` the API actually exposes (see §5). Not
enforced at admission (no webhooks in v1, §1) — two `TerdutEscalationRule`s
naming the same team would both `PUT` it and clobber each other every
reconcile; a footgun worth this one sentence, not a technical guard.
A `username` target is resolved to the `user_id` `PUT /api/teams/{teamID}/escalation`
actually requires (confirmed against source: `escalationTargetJSON.UserID
*int64`, no username field at all) via `GET /api/users` — confirmed open to
any authenticated caller, not gated by team membership or admin
(`internal/api/router.go`'s own comment: "readable by anyone signed in"), so
the team-scoped credential this controller already holds is enough; no
extra RBAC-equivalent server-side needed. Re-resolved every reconcile rather
than cached, in case a username is renamed. An unresolvable username is
`Ready: False, reason: UnknownUser`, naming which one.
### 4.4 `TerdutDeadmanSwitch`
```yaml
spec:
teamRef: {name: platform-team}
name: "prod-watchdog"
matcher: "alertname=Watchdog,cluster=prod"
timeout: 15m
severity: critical
status:
conditions: [...]
switchID: 7
```
No unique-name constraint server-side (§5's table row, confirmed against
source) — the controller's own idempotent-create step is a `GET`-list and
name match, not a conflict to recover from. `name` is optional, same as the
API: left empty, terdut-server derives it from `matcher`'s own canonical
form, and that's what the lookup matches against too.
Team *membership* is not modelled: the server manages it through OIDC group sync and its own UI.
### 4.5 `TerdutAlertSource`
@@ -317,7 +239,7 @@ status:
The integration key is shown by the API exactly once, at creation
(`Integration.Key`/`URL` in terdut-server's own model) — never re-readable,
same shape as the bootstrap admin key. The controller writes it straight into
a one-shot value. The controller writes it straight into
a generated, owner-referenced Secret on the create it caused and never logs
or stores it anywhere else; the CR's `status` carries only the Secret
reference, matching how e.g. cert-manager's `Certificate` exposes
@@ -337,8 +259,7 @@ of both patterns).
`serverRef` into this `TerdutServer`. Same-namespace `TerdutTeam`s are
always allowed regardless of this field.
- `from: Same` — equivalent to `None` in effect (same-namespace is already
unrestricted) but kept for parity with the upstream enum and to make the
policy self-documenting in a diff.
unrestricted); kept for parity with the upstream enum.
- `from: All` — any namespace in the cluster may reference in. Appropriate
for a genuinely shared, cluster-wide `TerdutServer`; the audit trail is
"check `allowedTeams` plus who has RBAC to create a `TerdutTeam`
@@ -361,230 +282,64 @@ of both patterns).
it used to authorize finds itself no longer permitted, flips
`Ready: False, reason: RefNotPermitted`, and — deliberately — does
**not** delete the team server-side on revocation alone; it stops
reconciling further changes until access is restored or the `TerdutTeam`
CR itself is deleted (whose finalizer still needs the credentials
Secret described in §6 to clean up, so blocking *new* changes rather
than forcing an immediate, possibly credential-less deletion is the
safer failure mode).
reconciling further changes until access is restored. Deleting the
`TerdutTeam` CR still deletes the team: consent gates what the operator
starts acting on, not whether it may clean up after itself.
## 5. Reconciliation semantics
terdut-server's REST surface (`internal/api/router.go`) does not give every
resource a full update verb, so reconciliation strategy is per-resource:
| Resource | Verbs available | Strategy |
| Resource | Server verbs | Strategy |
|---|---|---|
| Team | POST create, PUT rename, DELETE, PUT oidc-groups, GET by name | Real update-in-place: diff spec vs. last-applied, PUT the changed pieces. `GET /api/teams?name=` (terdut-server's `TEAM-LOOKUP.md`, landed 2026-10-01) is what makes the idempotent-create general rule below actually true for Team — confirmed by checking: until that endpoint existed, an instance-scoped service account had no way to recover a team's id after a 409, unlike every other resource in this table, where the adopt-on-conflict rule had a real lookup to call. |
| Escalation policy | GET/PUT whole-policy | Update-in-place: PUT the full desired policy every reconcile that finds drift; cheap because whole-policy is small and already loaded whole server-side. |
| Dead man's switch | POST create, **PUT update-in-place** (added in terdut-server v0.33.0), DELETE, GET-list | Real update-in-place, same shape as Team/Escalation: PUT the whole switch every reconcile once its id is known. No unique-name constraint server-side (confirmed against source — `handleCreateTeamDeadman` has no conflict handling at all, unlike Team/service-account creation), so idempotent-create here can't rely on a 409 to adopt from: before POSTing, `GET /api/teams/{teamID}/deadman/switches` and match by `name` first: found → adopt its id; not found → POST. `handleUpdateTeamDeadman`'s own doc comment (terdut-server) confirms the motivation directly: "Added alongside create/delete so an automated caller (terdut-operator) can reconcile a spec change without deleting and recreating the switch, which would otherwise ... needlessly rotate its id for no reason a reconciler's diff should ever manufacture." An earlier draft of this row, written before that endpoint existed, described delete-and-recreate; corrected here, confirmed against source rather than left stale. |
| Integration (alert source) | POST create, PATCH rename, DELETE | Rename via PATCH; any other spec change (kind) is delete-and-recreate, which **rotates the webhook key** — called out loudly in the CRD's field docs and in a `Warning` event, since it breaks whatever sends to the old URL/key until the new Secret is picked up. (See the general rule below for what happens if the Secret is lost with *no* spec change.) |
| Team | POST (idempotent on `external_id`), PUT rename, PUT oidc-groups, DELETE | Repeat the idempotent POST, then PUT the rest, every reconcile. Delete runs in a finalizer; the server refuses (409) while the team has open incidents, which is retried. |
| Escalation ladder | PUT whole ladder | PUT the full desired ladder every reconcile; the server resolves usernames. |
| Dead man's switch | list, POST, PUT, DELETE; names unique per team | Diff by name against the list. |
| Integration (alert source) | POST, PATCH rename, DELETE | Rename via PATCH; a kind change is delete-and-recreate, which **rotates the webhook key** (Warning event). An integration deleted on the server is recreated with a new key (Warning event). |
General rules for every controller:
- **Idempotent create**: before POSTing, check `status.<serverSideID>` is
unset; if the server already has a same-named object from a previous
partial reconcile (e.g. after a crash between POST and status-write), treat
a 409/name-conflict as "adopt" — GET-by-name and populate status, rather
than erroring forever. terdut-server's list endpoints in each of these
areas return objects by name, so this is a straightforward correlation.
- **Periodic resync** in addition to watch-triggered reconciles (Kubebuilder
default `RequeueAfter` on success, e.g. every 5–10 minutes) to catch drift
from **someone changing state directly against the server's API/UI**,
since gitops correctness means the CR wins, not "first write wins".
- **Finalizers** on every CRD that has a server-side counterpart, so deletion
calls the corresponding DELETE before the Kubernetes object disappears.
Failure to delete server-side (e.g. server unreachable) blocks finalizer
removal and surfaces as a `Degraded` condition + event, rather than
silently orphaning a row.
- **Generated Secrets holding unrecoverable server-issued material are
watched, and their loss is fail-closed, not self-healed.** Currently this
is just `TerdutAlertSource`'s webhook Secret (§4.5): the controller adds
it to its `Owns()` watches, not just the CR. If it disappears while
`status.integrationID` is still set, the controller does **not** attempt
to recreate it — the key is genuinely gone (§4.5: never stored anywhere
but that one Secret), so silently minting a replacement would rotate a
live production webhook URL with no corresponding spec change to explain
why. Instead it flips `Ready: False, reason: WebhookSecretLost` and fires
a `Warning` event telling the operator to delete and recreate the
`TerdutAlertSource`. No new mechanism is needed for recovery: deleting the
CR runs the existing finalizer (DELETE the still-live integration
server-side, above), and recreating it runs the existing idempotent-create
path (this same section) — a fresh POST, a new key, a new Secret. This is
deliberately the same recovery motion as the kind-change rotation above,
just human-triggered instead of spec-triggered.
- **Owner chain for status resolution, not API calls**: `TerdutTeam`'s
controller does not call any other controller; every child CRD's
controller independently resolves its own `teamRef` → `TerdutTeam.status`
for the `teamID`, `credentialsSecretRef`, and `serverEndpoint` it needs to
call the API (§6) — it never needs to chain further up to `TerdutServer`
at all, not just as a stated intent but literally: `serverEndpoint` is
resolved once, by `TerdutTeam`'s own controller, and stored in its status
specifically so no child ever needs its own `TerdutServer` RBAC (`get`/
`list`/`watch` on `terdutservers`) to find out where to send a request —
the team's own scoped credential plus that one status field is everything
a child resource's controller requires. If the referenced `TerdutTeam`
isn't `Ready` yet (which includes not having a `credentialsSecretRef` set),
the child requeues with backoff and reports `Ready: False, reason:
WaitingForTeam` — no cross-controller RPC.
- **Cross-namespace `serverRef` is re-checked every reconcile, not just at
creation**: `TerdutTeam`'s controller reads the target `TerdutServer`'s
`spec.allowedTeams` (and, under `Selector`, a `Get` on its own `Namespace`
object for labels) on every pass before touching a cross-namespace
`TerdutServer` — revocation (§4.6) takes effect on the team's very next
reconcile, not just when the CR is first applied.
General rules:
- **Periodic resync** (5 minutes) besides watch-triggered reconciles, to catch someone changing
state directly against the server: the CR wins.
- **Finalizers** on `TerdutTeam` and `TerdutAlertSource` call the server's DELETE first. A 404 is
success; any other failure blocks removal and surfaces as an event rather than orphaning a row.
`TerdutServer` needs none: its Secret is owned and its database is never touched.
- **A webhook Secret that is lost fails closed.** The key is never re-readable from the server, so
the controller does not mint a replacement for a live URL with no spec change to explain it:
`Ready: False, reason: WebhookSecretLost`; delete and recreate the `TerdutAlertSource`.
- **Children resolve through the team.** A `TerdutAlertSource` finds its `TerdutTeam` (same
namespace), requires it Ready, then reads the server named by the team's `serverRef` and that
server's operator key. A `TerdutTeam` re-reconciles when its `TerdutServer` changes.
- **Cross-namespace consent is re-checked every reconcile** (§4.6), so revocation takes effect on
the team's next pass.
## 6. Bootstrap & authentication to terdut-server's API
## 6. Authentication to terdut-server's API
terdut-server's scoped service-account credential type
(`terdut-server`'s `SERVICE-ACCOUNTS.md`) has shipped — confirmed against
source: `internal/api/service_accounts.go`, migration
`014_service_accounts.sql`, and `internal/api/router.go` wiring it in under
`AuthMiddleware`. This section is no longer blocked on it; the "v1-blocking"
framing here was accurate when this section was first written and is stale
now.
The operator talks to the server over HTTP with one bearer key per `TerdutServer`:
**`GET /api/service-accounts?name=` is not an unauthenticated lookup —
confirmed against `internal/api/middleware.go`'s `AuthMiddleware`, which
hard-rejects any request carrying neither a Bearer token nor a session
cookie with `401` before any handler ever runs.** This would matter a great
deal if the operator's `/api/bootstrap` call could ever lose a race to
something else bootstrapping the same server first — a credential-less
loser would have no authenticated way to recover. It doesn't matter here,
by construction (§1): **the operator only ever calls `/api/bootstrap`
against a `TerdutServer` it just created**, so there is nothing else in a
position to race it. An earlier draft of this section added a
`spec.credentialsSecretRef` bring-your-own input specifically to work around
that race, for a world where the operator might adopt a server something
else had already bootstrapped. That world doesn't exist (§1), so the field
was removed rather than kept as unused flexibility — self-registration
(point 1, below) is simply the only path, not one of two.
1. The `TerdutServer` controller creates the Secret `<name>-operator-key` (data key `token`,
`tdsa_` + 48 hex characters) in the server's own namespace, owned by the `TerdutServer`. It
is created once and never overwritten while it exists: the running server was seeded with it.
2. The Deployment hands it to the pods as `TERDUT_OPERATOR_KEY` through a `secretKeyRef` (the key
never appears in the pod spec), plus an annotation holding a hash of it so a replaced Secret
rolls the pods.
3. At every start the server creates or re-keys its instance-scoped service account
`terdut-operator` from that value. An instance-scoped account acts as owner of every team's
configuration but is not a member of any team, so it reads no incidents and is never an
administrator.
4. `status.credentialsSecretRef` points at the Secret; `TerdutTeam` and `TerdutAlertSource`
controllers read it from the server's namespace.
**Every credential the operator holds — the one instance-scoped key per
`TerdutServer`, and one team-scoped key per `TerdutTeam` — lives in a Secret
in the *operator's own* namespace, never in the namespace of the CR it
authenticates for.** Reconciliation happens entirely inside the operator's
controller loop, which is a single Deployment/ServiceAccount already
watching every namespace it's granted (§9); nothing about calling
terdut-server's API on a CR's behalf requires the credential to be
physically located near that CR, and no CR owner (human or otherwise) ever
needs to see, hold, or have RBAC to read a terdut-server credential. This is
a straight simplification of an earlier draft of this section, which mirrored
a shared credential into each consenting namespace instead — that version
conflated "the CR's owner never needs to see this" (true, and preserved
here) with "so the credential must live in the CR's namespace" (a
non-sequitur once you don't need to grant *anyone else* namespace-local
read access). Dropping that assumption also removes an entire class of
complexity: no on-demand mirroring, no garbage-collecting an orphaned copy
when `allowedTeams` narrows, no "OwnerReferences can't cross namespaces so
track it in status instead" workaround — none of that machinery is needed
when nothing ever crosses into a tenant namespace in the first place.
1. **First reconcile, confirmed against source**
(`internal/api/users.go`'s `handleBootstrap`): `/api/bootstrap` is
single-shot *per install*, gated on `SELECT COUNT(*) FROM users` — once
non-zero, every call `403`s regardless of identity. There are two crash
windows between "get an admin key" and "have a lasting, usable
credential" — getting from `/api/bootstrap`'s key to a minted
service-account key, and getting from that key to a persisted Secret —
and both get a checkpoint rather than being left as a theoretical gap,
the same rigor §5's general idempotent-create rule already applies
elsewhere:
- `status.credentialsSecretRef` already set: done, nothing to do.
- Otherwise, check for an intermediate
`<namespace>.<name>-bootstrap-admin` Secret in
the operator's own namespace first. If it exists, its key is a still-
valid admin credential from an earlier, interrupted attempt — skip
`/api/bootstrap` entirely and reuse it. If not, call `/api/bootstrap`
once the Deployment this `TerdutServer` created has a ready replica;
on `201`, immediately checkpoint its response's raw admin key
(`{"user": ..., "api_key": {"key": "<raw>", ...}}`) into that Secret
before doing anything else with it. A `403` with neither
`status.credentialsSecretRef` nor this checkpoint Secret present is
the one genuinely pathological case left (the checkpoint deleted out
from under a reconcile already past this point) — handled the same
way the design already handles unrecoverable server-issued material
elsewhere (§5's webhook-Secret-loss rule): fail closed,
`Ready: False, reason: BootstrapStateLost`, with the same recovery as
that case, delete and recreate the `TerdutServer` (its finalizer tears
down the Deployment/database-backing and server-side rows; a fresh
create starts clean) — not a workaround peculiar to this one path.
- With an admin key in hand (fresh or checkpointed): `POST
/api/service-accounts {name: "terdut-operator", scope: "instance"}`.
A `409` here means a prior attempt got this far before being
interrupted — adopt rather than error, per §5's general rule:
`GET /api/service-accounts?name=terdut-operator` (authenticated with
the checkpointed admin key, not an unauthenticated lookup) to find its
id, then `POST /api/service-accounts/{id}/keys` to mint a fresh key —
an orphaned first key some interrupted attempt minted and never used
is inert, not a cleanup obligation.
2. The resulting instance-scoped key — not the checkpointed admin key, which
is deleted once this step succeeds — is written to a generated Secret in
the **operator's own namespace** (e.g.
`<serverRef.namespace>.<serverRef.name>-instance-credentials`,
under a fixed data key, `token`), referenced back from
`TerdutServer.status.credentialsSecretRef: {name, key}` (§4.1). No
`OwnerReference` (those can't cross namespaces, and this Secret doesn't
share a namespace with the `TerdutServer` that caused it); the `TerdutServer`'s
finalizer deletes this Secret directly as part of its own teardown,
the same way it already has to clean up the server-side resources it
created (§5's general finalizer rule extends naturally to this Secret).
3. When a `TerdutTeam` first becomes `Ready` (its `serverRef` resolved,
`allowedTeams` satisfied if cross-namespace), its controller uses the
`TerdutServer`'s instance-scoped credential (read from the operator's own
namespace, resolved via the owner chain in §5) to mint a **team-scoped**
service account for itself: `POST /api/service-accounts` with
`scope: team, teamID: <status.teamID>`. The resulting key is written to
its own generated Secret, again in the **operator's own namespace**
(e.g. `<teamNamespace>.<teamName>-team-credentials`), referenced from
`TerdutTeam.status.credentialsSecretRef` (§4.2). Same finalizer pattern as
point 2: the `TerdutTeam`'s finalizer deletes this Secret as part of its
own teardown.
4. Every child controller (`TerdutEscalationRule`, `TerdutDeadmanSwitch`,
`TerdutAlertSource`) reads its team's `credentialsSecretRef` — resolved
through its `teamRef` → `TerdutTeam.status` (§5) — and never touches the
instance-scoped credential at all. Since child CRDs stay same-namespace-
as-their-`TerdutTeam` in v1 (§1), and the credential itself lives in the
operator's namespace regardless of where the `TerdutTeam` or its children
are, this works identically whether the `TerdutTeam` is same-namespace or
cross-namespace relative to its `TerdutServer` — there is no separate
cross-namespace case to handle here at all, unlike the mirroring design
this replaced.
5. **Blast radius**: a team-scoped key can only touch its own `team_id`'s
escalation policy, dead-man switches, integrations, schedule and OIDC
group bindings server-side (enforced by terdut-server itself, per
`SERVICE-ACCOUNTS.md`) — compromising one such Secret (e.g. a bug that
leaks operator-namespace Secrets, or an overly broad RBAC grant on that
one namespace) exposes exactly one team, never the whole server. This is
the real fix for what an earlier draft of this section called out as its
weak point (every mirrored copy being server-admin-equivalent); it falls
out of team-scoped credentials existing at all, independent of where
they're stored — the operator-private storage described above closes the
RBAC-footprint half of the problem, team scoping closes the credential-
privilege half.
6. **Rotation**: `POST /api/service-accounts/{id}/keys` mints a new key on
the existing account without recreating it; the old key is revoked via
`DELETE /api/service-accounts/{id}/keys/{keyID}`; the operator's local
Secret is updated in place. No DB-level workaround, no re-triggering a
single-shot endpoint that can't fire twice (which is what made rotation
unworkable under the old `/api/bootstrap`-only design).
There is no bootstrap handshake: `/api/bootstrap` stays free for the first human administrator.
Deleting the `TerdutServer` deletes the key with it; a recreated one gets a new key and the server
re-seeds on its next start. Rotating by hand means deleting the Secret: the next reconcile makes
a new one and rolls the pods.
## 7. Ownership, status, garbage collection
- Every generated object that lives in the *same* namespace as the CR that
caused it (Deployment, Service, webhook Secret) carries a standard
`metav1.OwnerReference` — GC handles these, no finalizer needed. The two
credential Secrets from §6 are the one exception: they live in the
operator's own namespace regardless of where their owning CR lives, so
`OwnerReference` doesn't apply (cross-namespace) and cleanup instead runs
through that CR's finalizer directly, alongside the server-side DELETE
it already has to issue (§5).
- Status conditions follow the standard `metav1.Condition` shape with at
least `Ready` on every kind, plus kind-specific ones (`TerdutServer`:
`DatabaseReady`, `Bootstrapped`; children: `Synced`).
- `status.observedGeneration` on every kind, bumped only after a successful
reconcile against that generation's spec — the standard way a client
(or `kubectl wait`) tells "applied" from "seen".
- No cluster-scoped aggregation object (e.g. no cluster-wide "all servers"
status) in v1 — `kubectl get terdutservers -A` is the aggregate view.
- Everything the controllers generate in a CR's namespace carries an `OwnerReference` (Deployment,
Service, PodDisruptionBudget, the operator key Secret, the webhook Secret), so GC cleans up
and no finalizer is needed. The PDB exists only while `spec.pod.disruptionBudget` is set; the
controller deletes it itself when the field is cleared.
- Every kind has a `Ready` condition and `status.observedGeneration`.
- No cluster-scoped aggregate object: `kubectl get terdutservers -A` is the overview.
## 8. Postgres integration
@@ -605,85 +360,43 @@ documented and tested operationally:
(`<user>.<cluster>.credentials.postgresql.acid.zalan.do`) the same way
the chart's comment already documents, and wires it in as
`PGPASSWORD` the same way.
- Watches that Secret (not just the `postgresql` CR) so a credential
rotation triggers a requeue — the chart today requires a manual pod
restart for this; the operator can at least detect and report it via a
condition even if restarting on rotation is left as a §13 follow-up
rather than done automatically (a rolling restart on credential change
is a behavior change worth its own design pass, not folded in here).
- Does not watch that Secret: a rotated credential is noticed at the next
5-minute resync, and restarting pods on rotation is a §13 follow-up.
- Requires read RBAC on `postgresql.acid.zalan.do` (optional CRD — the
operator's ClusterRole/Role should not hard-fail if the CRD isn't
installed and a given `TerdutServer` uses BYO DSN instead).
## 9. RBAC
- The operator's own ServiceAccount needs, per namespace it's granted:
`get/list/watch/create/update/patch/delete` on `Deployments`, `Services`
it owns, and `get/list/watch` on `postgresql.acid.zalan.do` (optional,
degrade gracefully if absent per §8), plus cluster-wide `get/list` on
`Namespace` (labels only, for `allowedTeams: {from: Selector}` evaluation
— §4.1, §4.6).
- **Two different `Secret` scopes, not one — corrected from an earlier draft
of this section.** That earlier draft said `Secret` access was "scoped to
the operator's own namespace only... nowhere else," reasoning that with
every credential held privately in the operator's own namespace (§6)
there was no legitimate reason to touch a `Secret` anywhere else. That was
wrong once §4.5 existed:
- The §6 credential Secrets (one instance-scoped key per `TerdutServer`,
one team-scoped key per `TerdutTeam`) do live in, and are only ever
touched from, the operator's own namespace —
`get/list/watch/create/update/patch/delete` there, nowhere else. The
rest of the original reasoning stands for *these* Secrets specifically:
no human or team's own RBAC is ever granted access to a terdut-server
credential by this design, and the operator itself never needs
cross-namespace access to reach them.
- The §4.5 webhook Secret is different: it's owned by and lives beside
its `TerdutAlertSource`, in that CR's own tenant namespace, not the
operator's. The per-namespace `Role` already granted for
`Deployments`/`Services` in each watched namespace (below) must carry
the same `Secret` verbs there too, or the controller cannot create,
watch, or even detect the loss of (§5) that Secret at all.
- This necessarily widens the operator's footprint in each watched
tenant namespace to "any `Secret` in that namespace," not just the ones
it created — Kubernetes RBAC has no owner-scoped grant finer than the
namespace itself, and the design already accepts this same granularity
for Deployments/Services there. Flagged as an accepted trade-off, not a
silent gap (§13).
- No cluster-scoped resources are created by this operator (namespaced CRDs
only, per §1) — a `Role` + `RoleBinding` per watched (tenant) namespace is
sufficient for Deployments/Services/the webhook `Secret`/the optional
Zalando CRD, plus a separate `Role` + `RoleBinding` in the operator's own
namespace for the §6 credential Secrets; a `ClusterRole` is only needed
for watching CRDs across all namespaces (the normal Kubebuilder
multi-tenant-operator default) — none of the `Secret` access above needs
to be cluster-scoped.
- terdut-server's own RBAC is unaffected — the operator talks to it purely
over HTTP with service-account API keys (§6), never via the Kubernetes
API for app-level state.
- The operator needs `get/list/watch/create/update/patch/delete` on `Deployments`, `Services`,
`PodDisruptionBudgets` and `Secrets` it owns, `get/list/watch` on `postgresql.acid.zalan.do`
(optional; degraded gracefully when the CRD is absent), and `get/list/watch` on `Namespaces`
(labels only, for `allowedTeams: {from: Selector}`).
- Secrets live in the `TerdutServer`'s namespace (the operator key) and in each
`TerdutAlertSource`'s namespace (the webhook Secret), so the operator needs Secret access in
every tenant namespace. Kubernetes RBAC has no owner-scoped grant finer than the namespace.
- **As built (2026-10): broader than the above.** The shipped default is a
`ClusterRole` with full verbs on `Secrets` in every namespace (the chart's
`rbac.namespaced: true` gives a `Role` in the release namespace only, which
cannot serve tenant namespaces), and the manager's cache is not restricted
to watched namespaces. A per-namespace `Role` split (a Role and RoleBinding per watched namespace, with the cache
restricted to them) is the intended end state, not implemented, and the decision (2026-10) is to stay
cluster-wide for now: treat this operator as able to read every Secret in the cluster.
- terdut-server's own RBAC is unaffected: the operator uses only its HTTP API (§6), never the
Kubernetes API for app-level state.
## 10. Relationship to `charts/terdut-server`
The chart's Deployment/Service/bootstrap-job templates are redundant once
`TerdutServer` exists — running both would mean two controllers (Helm and
this operator) reconciling the same Deployment, which is exactly the
conflict Kubernetes operators exist to avoid. The chart is repurposed into
an **installer chart**: it installs the operator + CRDs (and optionally one
`TerdutServer` CR from `values.yaml`, for users who want "helm install and
get a server" without hand-writing a CR) rather than templating the
Deployment directly.
The server chart's Deployment, Service and bootstrap Job are redundant once `TerdutServer`
exists: running both would have two controllers reconciling the same Deployment. The operator's
chart (`charts/terdut-operator`) installs the operator, CRDs and RBAC, and optionally one
`TerdutServer` from `values.yaml` (`terdutServer.enabled`). There is no migration from a
chart-based install and none is planned (§1); whatever the server chart deployed stays a separate
install until someone deletes it.
**No migration path from an existing chart-based install, and none is
planned (§1).** An earlier draft of this section spent most of its length on
one — `helm template` the current release's `values.yaml` into an equivalent
`TerdutServer` CR, uninstall or shrink the old release, and a whole
sub-question about who gets to call `/api/bootstrap` first, the chart's Job
or the operator — all of which presupposed the operator might end up
managing a server the chart had already deployed and bootstrapped. §1 rules
that out: the operator only ever manages servers it created itself, so
there's nothing to migrate and no bootstrap race to settle (§6 covers why
that race doesn't exist either). Adopting the operator means applying a
fresh `TerdutServer` CR; whatever the chart deployed before stays exactly
what it was, a separate install, until someone deletes it.
The operator's Deployment builder mirrors the chart's env block for the knobs both expose (the
chart's `deployment.yaml` and `buildEnv` in `internal/controller/terdutserver_deployment.go`);
a new server setting is added in `config.go`, the chart, and the operator, in that order.
## 11. Testing strategy
@@ -694,61 +407,27 @@ what it was, a separate install, until someone deletes it.
request/response shapes (already well-documented in
`internal/api/*_test.go` on the server side) — no real Postgres or real
terdut-server binary needed for controller unit tests.
- A smaller number of true end-to-end tests (`kind` cluster + real
terdut-server image + real Postgres) covering the golden path per CRD:
create `TerdutServer` → `TerdutTeam` → one of each child kind → verify via
terdut-server's own API that the objects exist with the right shape →
delete the CR → verify the server-side object is gone.
- The golden path (`kind` cluster + real terdut-server image + real Postgres:
create `TerdutServer` → `TerdutTeam` → `TerdutAlertSource`, verify through
terdut-server's own API, delete, verify it is gone) is a manual pass via
`examples/demo/run-demo.sh`, not a CI job.
## 12. Observability
- Standard controller-runtime metrics (reconcile duration/error counts) are
enough for v1 — no custom metrics.
- Every externally-visible action (bootstrap, key rotation-needed, delete-
and-recreate on the no-PUT resources, adopt-on-conflict) emits a
- Every externally-visible action (a team created or renamed, an integration recreated, a delete
that failed) emits a
Kubernetes `Event` on the CR, since that's what shows up in `kubectl
describe` and gitops tooling (Argo CD/Flux) surfaces without extra wiring.
## 13. Deferred / explicitly out of scope for this design
## 13. Deferred / out of scope
- **A real scoped service-account/token type in terdut-server — not merely
deferred, this is v1-blocking for §6 as written** (verified: without it,
§6's bootstrap flow has no working credential-rotation path and no clean
answer to the chart-vs-operator bootstrap race; see §6 points 1, 5, 6 and
§10). Sequence this server-side change *before* implementing the
`TerdutServer` controller's bootstrap logic, not after.
- **A version-discovery endpoint on terdut-server** (e.g. `GET /api/version`).
Neither this operator nor terdut-tui has one today — both independently
detect capability by probing specific routes (terdut-tui via `GET
/api/teams` 404-checking; this operator would otherwise need to invent
its own equivalent probe). An unattended reconciler is more exposed to a
silent breaking API change than an interactive TUI a human is watching;
raising this alongside the service-account request rather than inventing
another route-probe here.
- CloudNativePG support — same `spec.database` shape as Zalando should
extend to it, but the concrete field/Secret-naming conventions need their
own look.
- Cross-namespace `teamRef` on the child CRDs (`TerdutEscalationRule`,
`TerdutDeadmanSwitch`, `TerdutAlertSource`) — only `TerdutTeam.serverRef`
crosses namespaces in v1 (§2, §4.2, §4.6); these stay same-namespace as
their `TerdutTeam` until a real need for splitting them out shows up.
- Narrower-than-namespace RBAC for the §4.5 webhook Secret (Kubernetes RBAC
has no owner-scoped grant below the namespace itself, per §9) — revisit
if the widened per-tenant-namespace `Secret` access proves too broad in
practice.
- A mid-life `spec.teamRef` change on a child CRD (`TerdutEscalationRule`,
`TerdutDeadmanSwitch`, `TerdutAlertSource`) isn't specially detected —
noticed while grounding Stage 4 against source, not newly introduced by
it: all three controllers always resolve `spec.teamRef` fresh every
reconcile and act against whatever `TerdutTeam` that currently names,
trusting the server-side id already stored in `status` remains valid
there. There's no server-side verb that could move an existing
integration/policy/switch to a different team in place regardless, so
retargeting one onto a live child isn't a supported operation in v1 —
delete and recreate the CR instead.
- Gitops-managed team *membership* (see §4.2).
- CloudNativePG support, alongside the Zalando `postgresClusterRef` (same `spec.database` shape).
- Gitops-managed team membership (§4.2).
- Per-namespace RBAC with a restricted cache (§9).
- Automatic Deployment restart on upstream Postgres credential rotation.
- Admission webhooks / CEL-only validation limits (e.g. verifying a
`teamRef` exists at admission time rather than surfacing it as a status
condition after the fact).
- `spec.pod.priorityClassName`, pod labels beyond annotations, and an HPA for `TerdutServer`.
- Admission webhooks beyond CEL (for example, checking a `teamRef` exists at admission time).
- A shared API types module or generated client for terdut-server, terdut-operator and terdut-tui.
- OLM packaging.
+1 -33
View File
@@ -61,38 +61,7 @@ vet: ## Run go vet against code.
.PHONY: test
test: manifests generate fmt vet setup-envtest ## Run tests.
KUBEBUILDER_ASSETS="$(shell "$(ENVTEST)" use $(ENVTEST_K8S_VERSION) --bin-dir "$(LOCALBIN)" -p path)" go test $$(go list ./... | grep -v /e2e) -coverprofile cover.out
# TODO(user): To use a different vendor for e2e tests, modify the setup under 'tests/e2e'.
# The default setup assumes Kind is pre-installed and builds/loads the Manager Docker image locally.
# kubectl kuberc is disabled by default for test isolation; enable with:
# - KUBECTL_KUBERC=true
# CertManager is installed by default; skip with:
# - CERT_MANAGER_INSTALL_SKIP=true
KIND_CLUSTER ?= terdut-operator-test-e2e
.PHONY: setup-test-e2e
setup-test-e2e: ## Set up a Kind cluster for e2e tests if it does not exist
@command -v $(KIND) >/dev/null 2>&1 || { \
echo "Kind is not installed. Please install Kind manually."; \
exit 1; \
}
@case "$$($(KIND) get clusters)" in \
*"$(KIND_CLUSTER)"*) \
echo "Kind cluster '$(KIND_CLUSTER)' already exists. Skipping creation." ;; \
*) \
echo "Creating Kind cluster '$(KIND_CLUSTER)'..."; \
$(KIND) create cluster --name $(KIND_CLUSTER) ;; \
esac
.PHONY: test-e2e
test-e2e: setup-test-e2e manifests generate fmt vet ## Run the e2e tests. Expected an isolated environment using Kind.
KIND=$(KIND) KIND_CLUSTER=$(KIND_CLUSTER) go test -tags=e2e ./test/e2e/ -v -ginkgo.v
$(MAKE) cleanup-test-e2e
.PHONY: cleanup-test-e2e
cleanup-test-e2e: ## Tear down the Kind cluster used for e2e tests
@$(KIND) delete cluster --name $(KIND_CLUSTER)
KUBEBUILDER_ASSETS="$(shell "$(ENVTEST)" use $(ENVTEST_K8S_VERSION) --bin-dir "$(LOCALBIN)" -p path)" go test $$(go list ./...) -coverprofile cover.out
.PHONY: lint
lint: golangci-lint ## Run golangci-lint linter
@@ -186,7 +155,6 @@ $(LOCALBIN):
## Tool Binaries
KUBECTL ?= kubectl
KIND ?= kind
KUSTOMIZE ?= $(LOCALBIN)/kustomize
CONTROLLER_GEN ?= $(LOCALBIN)/controller-gen
ENVTEST ?= $(LOCALBIN)/setup-envtest
-18
View File
@@ -31,24 +31,6 @@ resources:
kind: TerdutTeam
path: git.ryuvia.com/niklas/terdut-operator/api/v1alpha1
version: v1alpha1
- api:
crdVersion: v1
namespaced: true
controller: true
domain: ryuvia.com
group: terdut
kind: TerdutEscalationRule
path: git.ryuvia.com/niklas/terdut-operator/api/v1alpha1
version: v1alpha1
- api:
crdVersion: v1
namespaced: true
controller: true
domain: ryuvia.com
group: terdut
kind: TerdutDeadmanSwitch
path: git.ryuvia.com/niklas/terdut-operator/api/v1alpha1
version: v1alpha1
- api:
crdVersion: v1
namespaced: true
+23 -28
View File
@@ -1,36 +1,31 @@
# Terdut operator
Aims to expose most config as CRD's, so end users can self-service over gitops.
Exposes terdut-server's configuration as Kubernetes objects, so teams can self-service it over
gitops. See [DESIGN.md](./DESIGN.md) for the design (read its revision section first: it is the
current shape of credentials and the CRD catalog), and [examples/demo](./examples/demo) for a
working install.
See [DESIGN.md](./DESIGN.md) for the full design: CRD catalog and specs,
reconciliation semantics, bootstrap/auth, Postgres integration, RBAC, and the
relationship to `charts/terdut-server`. This README stays a short pitch; the
open questions it used to carry are now resolved decisions there (§2).
## CRDs
## CRD's
### TerdutServer
A terdut-server install: Deployment, Service, database wiring (a DSN, or a Zalando
`postgresClusterRef`), and an operator key Secret it hands to the server so the operator can
authenticate. `allowedTeams` consents to `TerdutTeam`s in other namespaces.
### terdutServers
Creates a server — Deployment, Service, database wiring, bootstrap, operator
credentials, and `allowedTeams` consent for cross-namespace teams. See
DESIGN.md §4.1, §4.6.
### TerdutTeam
One team on a server, possibly in another namespace (`serverRef`, gated by that server's
`allowedTeams`). It carries the team's whole configuration:
- `displayName` and OIDC group bindings
- `escalation` — the escalation ladder
- `deadmanSwitches` — dead man's switches, by name (switches not listed are removed)
### terdutTeams
- team name
- oidc groups
- `serverRef` — explicit reference to its `TerdutServer`, may be in a
different namespace (one team owns the server, others self-service a
team against it), gated by that `TerdutServer`'s own `allowedTeams`
field (DESIGN.md §2, §4.1, §4.2, §4.6)
### TerdutAlertSource
An alert-ingest integration on a team (`teamRef`). The server shows the webhook key once; it is
surfaced only through a generated Secret next to the object, never set explicitly.
### terdutEscalationrules
- rule
- `teamRef` — explicit reference to its `TerdutTeam` (DESIGN.md §2, §4.3)
## Demo
### terdutDeadmansswitches
- rule
- `teamRef` (DESIGN.md §4.4)
### terdutAlertSources
- `teamRef` (DESIGN.md §4.5)
- URL/key are generated by the server at creation and surfaced only via a
generated Secret, never set explicitly
[`examples/demo`](./examples/demo) wires one of each together — two teams, each with an
escalation ladder, a dead man's switch and an alert source — plus a script that fires synthetic
Alertmanager webhooks at it, so you can watch incidents open, escalate and resolve without a real
Alertmanager.
+20 -233
View File
@@ -1,238 +1,25 @@
# terdut-operator build roadmap
# terdut-operator status and deferred work
This is the staging plan for implementing the operator against `DESIGN.md`'s
settled decisions. It exists for the same reason `DESIGN.md` and
`SERVICE-ACCOUNTS.md` do: so each stage starts from an agreed sequencing
instead of re-litigating "what do we build first" mid-PR.
The staged build plan this file used to hold (scaffolding, TerdutServer, TerdutTeam, the child
kinds, the installer chart and first release) is done and shipped; the history is in git. The
design it implemented is in `DESIGN.md`, and its revision section at the top is the current
shape of credentials and the CRD catalog.
## Sequencing call this roadmap makes
## Validation
**The operator creates and owns every `TerdutServer` it manages — it never
adopts one deployed independently, by hand or by `charts/terdut-server`**
(`DESIGN.md` §1). An earlier version of this roadmap staged `TerdutServer`'s
Deployment/Service/bootstrap takeover separately (old Stage 5), behind a
hand-deployed server the simpler CRDs could be proven against first.
That staging existed only because a credential-less operator couldn't
`/api/bootstrap` its way into a server something else had already
bootstrapped (`DESIGN.md` §6's original gap). With no server to adopt at
all, that split has nothing left to justify it: `TerdutServer` now builds
its full lifecycle — Deployment, Service, database wiring, bootstrap,
credentials — in one stage, Stage 1, since bootstrap only has something to
bootstrap once the Deployment exists.
There is no CI end-to-end job. The golden path (create every kind against a real terdut-server on
`kind`, check the server's own API, delete, check it is gone) is a manual pass, using
`examples/demo/run-demo.sh`. It has not been re-run since the credential and CRD redesign
(2026-10): do that before the next release.
## Stage 0 — Scaffolding & CI
## Deferred
- `go.mod` (`git.ryuvia.com/niklas/terdut-operator`) + Kubebuilder v4
scaffold (`cmd/main.go`, `config/`, `Makefile`, `PROJECT`), matching
terdut-server's Go toolchain and house style (§3). Kubebuilder's own
scaffolded `Makefile` already wires `manifests`/`generate`
(`controller-gen`) and `setup-envtest` into `test`, and `golangci-lint`
into `lint`, all fetched on demand into `bin/` — no separate install
step needed beyond what `make test`/`make lint` already do.
- `.release.conf` deliberately **not** added yet: it names a `HELM_CHART`
this repo doesn't have until Stage 5. Adding it now would either be a
stub that lies about what's releasable or dead config nobody can run —
it lands in Stage 5, alongside the chart it describes.
- Gitea Actions CI calling `fmt lint test`, mirroring terdut-server's
`ci.yaml` convention (its `CLAUDE.md`: "a green gate here and a green
pipeline are the same code, not two descriptions of it") minus the
`chart`/`security` jobs, which need a chart (Stage 5) and real controller
code (Stage 1+) respectively to have anything to check.
- Drop kubebuilder's default `.github/workflows/*` scaffold — this org
runs on Gitea, not GitHub; `.gitea/workflows/ci.yaml` is the only CI this
repo has.
- Housekeeping: drop the stray `.DESIGN.md.swp` (leftover vim swapfile,
shouldn't be committed); correct `DESIGN.md` §6/§13's "v1-blocking, not
v1-shippable" language — the service-account feature it was blocking on
has since shipped in terdut-server.
**Done when:** CI is green on an otherwise-empty scaffold.
## Stage 1 — `TerdutServer`, full lifecycle
Supersedes the Stage 1 shipped before this redesign (commit `1be7cf2`)
outright — that `TerdutServerSpec`/`Status`/controller/tests implemented the
now-removed bring-your-own path and get replaced wholesale, not extended.
New commits build forward over the old ones; no git history rewrite.
- Full §4.1 spec: `image`, `replicas`, `networking`, `database`, `sweeper`,
`deadman`, `notify`, `oidc`, `passwordLogin`, `allowedTeams`, all together
— no narrowing, since bootstrap needs the Deployment it's narrowed away
from in the version this replaces.
- Controller manages the Deployment + Service, both Postgres paths from §8
at once (bring-your-own DSN *and* the Zalando `postgres-operator`
`postgresClusterRef` integration — not sequenced, per the user's call),
and bootstrap/credentials per §6's self-registration flow: `/api/bootstrap`
once the Deployment has a ready replica, checkpoint the admin key, mint
the instance-scoped service account, generated credentials Secret in the
operator's own namespace, `status.credentialsSecretRef`. Finalizer cleans
up that Secret (and the checkpoint, if one's still there) on delete —
there's no server-side row to clean up alongside it: terdut-server's API
has no way to delete a user or a service account, only to revoke
individual keys, so there's nothing to undo there regardless.
- RBAC: read-only watch on `postgresql.acid.zalan.do`, degrading gracefully
if that CRD isn't installed (§8, §9).
- Shipped, scoped down from §8's full ambition in two ways, both called out
in code rather than silently dropped: no live watch on the Zalando-
generated credentials Secret for rotation (relies on the periodic resync
to notice eventually, higher latency than a watch); no Gateway API
`HTTPRoute` creation from `spec.networking.hostname`/`gatewayListener`
(needs the Gateway API types as a new dependency, and nothing about
proving a `TerdutServer` boots and bootstraps a real server depends on
external ingress existing). Both are near-term follow-ups, not deferred
to a later stage.
- `envtest` covering Deployment/Service reconciliation and both database
paths — the Zalando path needs that CRD's schema vendored into the test
environment (there's no real `postgres-operator` controller in `envtest`,
only the CRD shape to create fixture objects against) — plus the
self-registration flow against an `httptest.Server` fake of
`/api/bootstrap` and `/api/service-accounts` (§11), including the
adopt-on-409 recovery path and the one fail-closed case
(`BootstrapStateLost`), not just the happy path.
- **Done, 2026-10-01**: a real `kind` end-to-end pass (bring-your-own DSN,
real terdut-server `v0.33.0` image, operator built into a real image and
deployed as a real Pod, not `go run` against the cluster). `TerdutServer`
went `Ready`; the generated credential authenticated and exercised its
real capability against the actual server
(`GET`/`POST /api/teams` → `200`/`201`, confirmed from terdut-server's own
access log). Caught two real bugs no `envtest` suite could have (its
client bypasses RBAC): `.dockerignore`'s `!**/*.go` not working under
podman, and missing RBAC for `events.k8s.io` (the new events API
`GetEventRecorder` uses) — both fixed. The Zalando path can additionally
be validated for real against the org's own cluster later, where
`postgres-operator` already runs, rather than only in a disposable `kind`
stand-in — not done in this pass.
## Stage 2 — `TerdutTeam`
- §4.2: `serverRef` resolution, real update-in-place (POST create / PUT
rename / PUT oidc-groups), team-scoped service-account minting once
`Ready` (§6 point 3), finalizer that DELETEs the team server-side and its
credential Secret.
- First place the "every child resolves its own `teamRef` →
`TerdutTeam.status`, never chains up to `TerdutServer`" pattern (§5) gets
proven end to end.
## Stage 3 — `TerdutEscalationRule` + `TerdutDeadmanSwitch`
- Built together: both stay same-namespace-as-their-`TerdutTeam` (§1), so
neither exercises cross-namespace complexity, but together they cover the
two different reconciliation shapes §5's table calls out — whole-policy
PUT-upsert for the escalation policy (no separate create step at all),
real create/update-in-place/delete for the dead man's switch (PUT added
in terdut-server `v0.33.0` specifically for this operator) — against the
same shared create/finalizer/resync scaffolding Stage 2 already built.
- New shared `resolveTeamAndClient` helper (`childref.go`) implements §5's
"every child resolves its own `teamRef` → `TerdutTeam.status`, never
chains up to `TerdutServer`" rule once, for both controllers —
`TerdutTeam.status.serverEndpoint`, added in this stage, is what makes
that literally true rather than just a stated intent.
- `TerdutEscalationRule` resolves each "user" target's username to a
user_id via `GET /api/users` (confirmed open to any authenticated
caller) and reports `Ready: False, reason: UnknownUser` if it doesn't
resolve. No `DELETE` exists for this resource, so its delete path `PUT`s
an empty policy as the closest available undo.
- `TerdutDeadmanSwitch` has no unique-name constraint server-side, so its
idempotent-create is `GET`-list-and-match-by-name rather than
adopt-on-409 (unlike every other resource in this operator).
- **Done, 2026-10-01**: `envtest` coverage for both controllers' happy
path, `TeamRefNotFound`/`WaitingForTeam`, `UnknownUser`, list-and-match
adoption, update-in-place on spec drift, and deletion. `make fmt lint
test build` all clean; `internal/controller` envtest coverage
50.5% → 71.7%. No `kind` e2e pass for this stage — Stage 1's already
proved the real-cluster mechanics (RBAC, image, bootstrap) these two
controllers reuse unchanged, and neither introduces a new mechanism that
pass would exercise differently (same reasoning Stage 2 used to skip
one).
## Stage 4 — `TerdutAlertSource`
- Last of the children on purpose: it has the subtlest failure mode of the
four. Covers webhook Secret generation/ownership (§4.5, §7), the
`WebhookSecretLost` fail-closed condition + `Warning` event (§5, added
2026-09-30), and the kind-change delete-and-recreate rotation path — all
easier to get right with the other three controllers' patterns already
in place to build on.
- Idempotent-create here is neither adopt-on-409 (Team/service-account) nor
list-and-match-by-name (`TerdutDeadmanSwitch`): terdut-server shows the
webhook key exactly once, at creation, and never again, so no server-side
lookup could ever recover it after a crash. The generated webhook Secret
itself — written immediately after the POST, before `status` is ever
touched — is this CR's only durable record that a create already
succeeded; found on a later reconcile with `status.integrationID` still
unset, it's read back directly rather than POSTing again. Found missing
with `status.integrationID` *set*, that's the already-designed
`WebhookSecretLost` fail-closed case instead.
- **Done, 2026-10-01**: `envtest` coverage for the happy path, rename
(PATCH, no key rotation), a kind change (delete-and-recreate, new id and
key), crash recovery between POST and the Secret write, `WebhookSecretLost`,
`TeamRefNotFound`/`WaitingForTeam`, and deletion. `make fmt lint test
build` all clean; `internal/controller` envtest coverage holds at 71.6%.
No `kind` e2e pass for this stage, same reasoning as Stage 3 (reuses
Stage 1's already-proven real-cluster mechanics unchanged).
## Stage 5 — Installer chart + real release
- Package CRDs + the operator's own Deployment/RBAC into the installer
chart §10 describes; wire `.release.conf`/release-vars the same way
terdut-server does; run it through the `release` skill for a real first
cut.
- Full `kind` end-to-end test per §11: create `TerdutServer` → `TerdutTeam`
→ one of each child kind → verify via terdut-server's own API that each
object exists with the right shape → delete the CR → verify the
server-side object is gone.
- Chart built via kubebuilder's own `helm/v2-alpha` plugin from `config/`'s
kustomize output (`charts/terdut-operator`), not hand-rolled -- CRDs +
manager Deployment/RBAC come from the same markers/manifests every other
stage already generates, so there's exactly one source of truth for
them. Hand-added on top: the optional `terdutServer` values block (§10's
"helm install and get a server" path), `.release.conf`, and the
`release-vars`/`helm-lint`/`push`/`helm-package`/`helm-push`/`release`
Makefile targets `.gitea/workflows/release.yaml` calls, mirroring
terdut-server's own shape end to end (same registry/namespace
convention, same multi-arch buildx push, same trivy/govulncheck/gitleaks
scans). Also fixed while wiring this: the Dockerfile's builder stage
didn't pin `--platform=$BUILDPLATFORM`, which would have made a
multi-arch release build fail outright on this org's runners (no binfmt
registration) -- caught before it ever shipped, not discovered mid-release;
and govulncheck surfaced one real, reachable finding (`google.golang.org/grpc`
v1.82.1, transitive via controller-runtime's otel exporter), fixed by
bumping to v1.83.1.
- **Done, 2026-10-01**: the full golden-path pass above, run for real
against a `kind` cluster, installed via `helm install` (not raw
kustomize/kubectl apply -- the first time the chart itself, not just
`config/`, was exercised): `TerdutServer` (real terdut-server `v0.33.0`
image, bring-your-own DSN against a throwaway in-cluster Postgres) →
`TerdutTeam` → one `TerdutEscalationRule` + `TerdutDeadmanSwitch` +
`TerdutAlertSource`, each confirmed `Ready` and then confirmed a second
way, independent of the operator's own status: a `curl` pod inside the
cluster, authenticated with the generated team credential, hit
terdut-server's real API directly (`GET /api/teams/{id}/escalation`,
`.../deadman/switches`, `.../integrations`) and got back exactly the
policy/switch/integration each spec declared. Deleting every CR in
reverse order was verified the same way: the escalation policy came back
empty (its only available "undo"), the switch and the integration were
both gone from their list endpoints, the team no longer resolved by
name, and the Deployment/Service/every generated Secret were gone from
the cluster. No new bugs found this pass -- Stage 1's own kind e2e pass
already caught the two issues (`events.k8s.io` RBAC, the podman
`.dockerignore` fix) a real cluster catches and `envtest` can't, and
nothing since has touched that surface.
- Not done in this pass, deliberately: an actual tagged release. `make
release-vars`/`helm-lint`/`push`/`helm-package`/`helm-push` all work
locally and `.gitea/workflows/release.yaml` is wired, but
`release-preflight` found there is no `terdut-operator/` entry under
`Ryuvia/charts` yet to bump -- every other onboarded repo had that
one-time wrapper-chart bootstrap done for it before its own first
release, and this one doesn't, since deploying this operator for real is
a decision for whoever runs the cluster, not a side effect of finishing
this stage. Cutting the first real release (and creating that wrapper
entry) is therefore the next action, not yet taken.
## Deferred (§13, unchanged by this roadmap)
Cross-namespace `allowedTeams` exercised against a real second namespace,
CloudNativePG support, narrower-than-namespace Secret RBAC for the webhook
Secret, gitops-managed team membership, automatic Deployment restart on
upstream Postgres credential rotation, admission webhooks/CEL-only
validation limits, OLM packaging.
- Narrower Secret RBAC: per-namespace Roles and a restricted cache (today a ClusterRole with
Secret access cluster-wide, DESIGN.md §9).
- CloudNativePG support alongside the Zalando `postgresClusterRef`.
- Gitops-managed team membership.
- Automatic Deployment restart on upstream Postgres credential rotation.
- Admission webhooks beyond CEL validation.
- OLM packaging.
- A shared API types module (or generated client) between terdut-server, terdut-operator and
terdut-tui, so contract drift is a compile error and not a manual mirror.
-98
View File
@@ -1,98 +0,0 @@
package v1alpha1
import (
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
)
// TerdutDeadmanSwitchSpec defines the desired state of TerdutDeadmanSwitch.
//
// One per switch (DESIGN.md §4.4). Reconciled with real update-in-place
// (terdut-server v0.33.0 added PUT specifically for this, §5) -- but with no
// unique-name constraint server-side, idempotent-create here means
// GET-list-and-match-by-name, not adopt-on-409.
type TerdutDeadmanSwitchSpec struct {
// +required
TeamRef TerdutTeamRef `json:"teamRef"`
// name is optional, same as the API: left empty, terdut-server derives
// it from matcher's own canonical form, and that's what the
// idempotent-create lookup matches against too.
// +optional
Name string `json:"name,omitempty"`
// matcher names the alerts this switch watches, e.g.
// "alertname=Watchdog,cluster=prod". One matcher per switch -- add
// another TerdutDeadmanSwitch instead of separating with ";"
// (terdut-server's own restriction, mirrored here so a bad spec is
// rejected at apply time).
// +required
// +kubebuilder:validation:MinLength=1
// +kubebuilder:validation:XValidation:rule="!self.contains(';')",message="one matcher per switch: add another TerdutDeadmanSwitch instead of separating with ;"
Matcher string `json:"matcher"`
// timeout is a Go duration string, e.g. "15m".
// +required
// +kubebuilder:validation:MinLength=1
Timeout string `json:"timeout"`
// +kubebuilder:validation:Enum=critical;error;warning;info
// +kubebuilder:default=critical
// +optional
Severity string `json:"severity,omitempty"`
}
// TerdutDeadmanSwitchStatus defines the observed state of TerdutDeadmanSwitch.
type TerdutDeadmanSwitchStatus struct {
// +listType=map
// +listMapKey=type
// +optional
Conditions []metav1.Condition `json:"conditions,omitempty"`
// switchID is the server-side id.
// +optional
SwitchID int64 `json:"switchID,omitempty"`
// +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"`
}
// +kubebuilder:object:root=true
// +kubebuilder:subresource:status
// +kubebuilder:printcolumn:name="Team",type=string,JSONPath=`.spec.teamRef.name`
// +kubebuilder:printcolumn:name="SwitchID",type=integer,JSONPath=`.status.switchID`
// +kubebuilder:printcolumn:name="Ready",type=string,JSONPath=`.status.conditions[?(@.type=="Ready")].status`
// +kubebuilder:printcolumn:name="Reason",type=string,JSONPath=`.status.conditions[?(@.type=="Ready")].reason`
// TerdutDeadmanSwitch is the Schema for the terdutdeadmanswitches API
type TerdutDeadmanSwitch struct {
metav1.TypeMeta `json:",inline"`
// metadata is a standard object metadata
// +optional
metav1.ObjectMeta `json:"metadata,omitzero"`
// spec defines the desired state of TerdutDeadmanSwitch
// +required
Spec TerdutDeadmanSwitchSpec `json:"spec"`
// status defines the observed state of TerdutDeadmanSwitch
// +optional
Status TerdutDeadmanSwitchStatus `json:"status,omitzero"`
}
// +kubebuilder:object:root=true
// TerdutDeadmanSwitchList contains a list of TerdutDeadmanSwitch
type TerdutDeadmanSwitchList struct {
metav1.TypeMeta `json:",inline"`
metav1.ListMeta `json:"metadata,omitzero"`
Items []TerdutDeadmanSwitch `json:"items"`
}
func init() {
SchemeBuilder.Register(func(s *runtime.Scheme) error {
s.AddKnownTypes(SchemeGroupVersion, &TerdutDeadmanSwitch{}, &TerdutDeadmanSwitchList{})
return nil
})
}
-137
View File
@@ -1,137 +0,0 @@
package v1alpha1
import (
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
)
// TerdutTeamRef names the TerdutTeam this resource belongs to. Always
// same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
// crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
type TerdutTeamRef struct {
// +kubebuilder:validation:MinLength=1
Name string `json:"name"`
}
// EscalationTargetKind is who one rung of the ladder pages.
// +kubebuilder:validation:Enum=oncall;user
type EscalationTargetKind string
const (
EscalationTargetOncall EscalationTargetKind = "oncall"
EscalationTargetUser EscalationTargetKind = "user"
)
// EscalationTarget is one page within a level. username is required iff
// kind is "user" (terdut-server's own validation, internal/api/escalation.go's
// handleSetEscalation -- mirrored here as a CEL rule so a bad spec is
// rejected at apply time, not discovered on the next failed PUT).
// +kubebuilder:validation:XValidation:rule="self.kind != 'user' || has(self.username)",message="username is required when kind is user"
// +kubebuilder:validation:XValidation:rule="self.kind != 'oncall' || !has(self.username)",message="username must not be set when kind is oncall"
type EscalationTarget struct {
// +required
Kind EscalationTargetKind `json:"kind"`
// +optional
Username string `json:"username,omitempty"`
}
// EscalationLevel is one rung of the ladder: how long to wait, and who to
// page if nobody's acknowledged by then.
type EscalationLevel struct {
// timeout is a Go duration string, e.g. "5m".
// +required
// +kubebuilder:validation:MinLength=1
Timeout string `json:"timeout"`
// +required
// +kubebuilder:validation:MinItems=1
Targets []EscalationTarget `json:"targets"`
}
// TerdutEscalationRuleSpec defines the desired state of TerdutEscalationRule.
//
// One per team (DESIGN.md §4.3) -- terdut-server models a policy as one row
// with an owned list of levels, reconciled with a single whole-policy PUT.
// Not enforced at admission if two CRs name the same team (no webhooks in
// v1, §1); they would simply clobber each other every reconcile.
type TerdutEscalationRuleSpec struct {
// +required
TeamRef TerdutTeamRef `json:"teamRef"`
// +kubebuilder:validation:Minimum=0
// +kubebuilder:validation:Maximum=10
// +optional
RepeatCount int64 `json:"repeatCount,omitempty"`
// +optional
FallbackTopic string `json:"fallbackTopic,omitempty"`
// +required
// +kubebuilder:validation:MinItems=1
Levels []EscalationLevel `json:"levels"`
}
// Condition reasons shared by TerdutEscalationRule and TerdutDeadmanSwitch
// (both resolve a teamRef the same way, DESIGN.md §5).
const (
// ReasonTeamRefNotFound: spec.teamRef names no TerdutTeam (yet).
ReasonTeamRefNotFound = "TeamRefNotFound"
// ReasonWaitingForTeam: the referenced TerdutTeam exists but isn't
// Ready yet (no status.credentialsSecretRef to read).
ReasonWaitingForTeam = "WaitingForTeam"
// ReasonUnknownUser: an escalation target's username doesn't resolve to
// any user server-side (TerdutEscalationRule only).
ReasonUnknownUser = "UnknownUser"
// ReasonChildAdopted: the happy path, shared by both child kinds.
ReasonChildAdopted = "Adopted"
)
// TerdutEscalationRuleStatus defines the observed state of TerdutEscalationRule.
type TerdutEscalationRuleStatus struct {
// +listType=map
// +listMapKey=type
// +optional
Conditions []metav1.Condition `json:"conditions,omitempty"`
// +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"`
}
// +kubebuilder:object:root=true
// +kubebuilder:subresource:status
// +kubebuilder:printcolumn:name="Team",type=string,JSONPath=`.spec.teamRef.name`
// +kubebuilder:printcolumn:name="Ready",type=string,JSONPath=`.status.conditions[?(@.type=="Ready")].status`
// +kubebuilder:printcolumn:name="Reason",type=string,JSONPath=`.status.conditions[?(@.type=="Ready")].reason`
// TerdutEscalationRule is the Schema for the terdutescalationrules API
type TerdutEscalationRule struct {
metav1.TypeMeta `json:",inline"`
// metadata is a standard object metadata
// +optional
metav1.ObjectMeta `json:"metadata,omitzero"`
// spec defines the desired state of TerdutEscalationRule
// +required
Spec TerdutEscalationRuleSpec `json:"spec"`
// status defines the observed state of TerdutEscalationRule
// +optional
Status TerdutEscalationRuleStatus `json:"status,omitzero"`
}
// +kubebuilder:object:root=true
// TerdutEscalationRuleList contains a list of TerdutEscalationRule
type TerdutEscalationRuleList struct {
metav1.TypeMeta `json:",inline"`
metav1.ListMeta `json:"metadata,omitzero"`
Items []TerdutEscalationRule `json:"items"`
}
func init() {
SchemeBuilder.Register(func(s *runtime.Scheme) error {
s.AddKnownTypes(SchemeGroupVersion, &TerdutEscalationRule{}, &TerdutEscalationRuleList{})
return nil
})
}
+157 -76
View File
@@ -1,16 +1,17 @@
package v1alpha1
import (
corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/util/intstr"
)
// SecretKeyRef names one data key inside a Secret. Every use of this type in
// TerdutServerSpec resolves in the TerdutServer's own namespace (it's wired
// straight into the Deployment's pod spec as a secretKeyRef env source,
// which Kubernetes itself only allows same-namespace) -- unlike the
// generated credentials Secret (DESIGN.md §6), which always lives in the
// operator's own namespace and is never referenced through this type.
// SecretKeyRef names one data key inside a Secret in the TerdutServer's own
// namespace. Every use of this type is wired into the Deployment's pod spec as
// a secretKeyRef env source, which Kubernetes only allows same-namespace --
// including status.credentialsSecretRef, the operator key the controller
// generates there.
type SecretKeyRef struct {
// name is the Secret's name.
// +kubebuilder:validation:MinLength=1
@@ -29,33 +30,28 @@ type ImageSpec struct {
Tag string `json:"tag"`
}
// NetworkingSpec is how this TerdutServer is reached from outside the
// cluster.
//
// hostname/gatewayListener describe the intended Gateway API HTTPRoute
// (matching charts/terdut-server's own templates/httpproxy.yaml, despite its
// name — that chart carries a Gateway API HTTPRoute, not a Contour
// HTTPProxy), but creating that HTTPRoute isn't implemented yet: it needs
// the Gateway API types as a new dependency, and nothing about proving a
// TerdutServer boots and bootstraps a real server depends on external
// ingress existing. Tracked as a near-term follow-up, not deferred to a
// later ROADMAP.md stage the way Deployment/database/bootstrap once were.
// NetworkingSpec configures terdut-server itself and the plain ClusterIP
// Service the operator creates in front of it. It does not expose
// TerdutServer outside the cluster in any way, and never will (DESIGN.md
// §1 -- a permanent non-goal, not a staged one): exposing it is entirely up
// to whoever deploys it -- a Gateway API HTTPRoute, a plain Ingress, an
// Istio VirtualService, or nothing at all if it should stay cluster-
// internal. See examples/networking for worked examples against the
// Service this creates.
type NetworkingSpec struct {
// hostname the HTTPRoute will carry once it exists.
// hostname is terdut-server's own public URL (TERDUT_PUBLIC_URL) --
// used for absolute links terdut-server generates itself
// (notifications, OIDC redirect URIs), not read by this operator for
// anything ingress-related. Set it to whatever hostname your own
// exposure mechanism, if any, actually serves this on.
// +optional
Hostname string `json:"hostname,omitempty"`
// servicePort is both the Service's port and the HTTPRoute's backend
// port once it exists. Defaults to 8080, matching the chart's own
// service.port default.
// servicePort is both the container's port and the ClusterIP Service's
// port. Defaults to 8080, matching the chart's own service.port default.
// +kubebuilder:default=8080
// +optional
ServicePort int32 `json:"servicePort,omitempty"`
// gatewayListener is the HTTPRoute's sectionName once it exists. Empty
// attaches to every matching listener, including plaintext HTTP.
// +optional
GatewayListener string `json:"gatewayListener,omitempty"`
}
// PostgresClusterRef names a Zalando postgres-operator `postgresql` CR
@@ -107,18 +103,6 @@ type SweeperSpec struct {
ArchiveAfter string `json:"archiveAfter,omitempty"`
}
// DeadmanSpec controls dead man's switch alerts. Matchers/Timeout/Severity
// map straight to TERDUT_DEADMAN_MATCHERS/TERDUT_DEADMAN_TIMEOUT/
// TERDUT_DEADMAN_SEVERITY.
type DeadmanSpec struct {
// +optional
Matchers string `json:"matchers,omitempty"`
// +optional
Timeout string `json:"timeout,omitempty"`
// +optional
Severity string `json:"severity,omitempty"`
}
// NotifySpec controls push notifications via ntfy. Empty ntfyURL disables
// notifications entirely (matches the chart's own default).
type NotifySpec struct {
@@ -134,11 +118,7 @@ type NotifySpec struct {
TokenSecretRef *SecretKeyRef `json:"tokenSecretRef,omitempty"`
}
// OIDCSpec controls single sign-on. Fields the chart also exposes but
// DESIGN.md's spec doesn't (usernameClaim, emailClaim, groupsClaim,
// trustEmail) use terdut-server's own defaults
// (preferred_username/email/groups/false) rather than being added here
// speculatively.
// OIDCSpec controls single sign-on.
type OIDCSpec struct {
// +optional
Enabled bool `json:"enabled,omitempty"`
@@ -161,6 +141,21 @@ type OIDCSpec struct {
// +kubebuilder:default="12h"
// +optional
SessionMaxAge string `json:"sessionMaxAge,omitempty"`
// usernameClaim, emailClaim and groupsClaim name the ID token claims read.
// +kubebuilder:default="preferred_username"
// +optional
UsernameClaim string `json:"usernameClaim,omitempty"`
// +kubebuilder:default="email"
// +optional
EmailClaim string `json:"emailClaim,omitempty"`
// +kubebuilder:default="groups"
// +optional
GroupsClaim string `json:"groupsClaim,omitempty"`
// trustEmail links a sign-in to an existing local user by email even when
// the provider does not vouch the address is verified (Authentik reports
// email_verified false unless told otherwise).
// +optional
TrustEmail bool `json:"trustEmail,omitempty"`
}
// AllowedTeamsNamespaces gates which namespaces a TerdutTeam may resolve a
@@ -190,6 +185,109 @@ type AllowedTeams struct {
Namespaces AllowedTeamsNamespaces `json:"namespaces,omitempty"`
}
// PodDisruptionBudgetSpec configures an optional PodDisruptionBudget for
// this TerdutServer's Deployment. Exactly one of minAvailable or
// maxUnavailable may be set, matching policyv1.PodDisruptionBudgetSpec's own
// upstream rule (both wrap intstr.IntOrString unchanged here -- this is pure
// passthrough, not reshaped) and mirroring DatabaseSpec's own
// dsn/postgresClusterRef mutual-exclusion pattern. Clearing this field
// deletes any PodDisruptionBudget the controller previously created for this
// TerdutServer (DESIGN.md §7).
// +kubebuilder:validation:XValidation:rule="(has(self.minAvailable) ? 1 : 0) + (has(self.maxUnavailable) ? 1 : 0) == 1",message="exactly one of minAvailable or maxUnavailable must be set"
type PodDisruptionBudgetSpec struct {
// minAvailable -- mutually exclusive with maxUnavailable.
// +optional
MinAvailable *intstr.IntOrString `json:"minAvailable,omitempty"`
// maxUnavailable -- mutually exclusive with minAvailable.
// +optional
MaxUnavailable *intstr.IntOrString `json:"maxUnavailable,omitempty"`
}
// PodSpec is pod-level customization of the Deployment this TerdutServer
// creates. Fields here directly reuse corev1 types wherever corev1 already
// models the knob exactly, rather than wrapping (unlike SecretKeyRef's own
// "wrap only when a round-trip through a different type buys something"
// standard would suggest at first glance -- none of these do: Tolerations,
// Affinity, TopologySpreadConstraints, Resources, SecurityContext, EnvVar,
// EnvFromSource, Volume, VolumeMount and LocalObjectReference are all passed
// straight through to the pod template with no added semantics, matching how
// CloudNativePG and the Zalando postgres-operator both expose the same
// knobs).
type PodSpec struct {
// annotations are merged onto the pod template's own metadata.
// Operator-managed labels (labelsFor) are never touched by this field.
// +optional
Annotations map[string]string `json:"annotations,omitempty"`
// +optional
NodeSelector map[string]string `json:"nodeSelector,omitempty"`
// +optional
Tolerations []corev1.Toleration `json:"tolerations,omitempty"`
// affinity covers node affinity, pod affinity and pod anti-affinity in
// one field -- even though replicas now defaults to 2 (see
// TerdutServerSpec.Replicas's own doc comment), this operator still
// never generates a default anti-affinity of its own the way a
// multi-replica-aware operator typically would, so this stays pure
// user-supplied passthrough, not a toggle-plus-generated-default.
// +optional
Affinity *corev1.Affinity `json:"affinity,omitempty"`
// +optional
TopologySpreadConstraints []corev1.TopologySpreadConstraint `json:"topologySpreadConstraints,omitempty"`
// resources applied to the main terdut-server container. Unset today --
// this field closes a pre-existing gap, not a behavior change for
// anyone not setting it.
// +optional
Resources corev1.ResourceRequirements `json:"resources,omitempty"`
// securityContext is pod-level.
// +optional
SecurityContext *corev1.PodSecurityContext `json:"securityContext,omitempty"`
// containerSecurityContext applies to the main terdut-server container
// only -- not wait-for-postgres, which runs a stock postgres image this
// operator doesn't control the entrypoint of. No implicit defaults are
// merged underneath it.
// +optional
ContainerSecurityContext *corev1.SecurityContext `json:"containerSecurityContext,omitempty"`
// serviceAccountName. Defaults to "default", same as any pod that
// doesn't set it.
// +optional
ServiceAccountName string `json:"serviceAccountName,omitempty"`
// extraEnv is appended after the fixed env vars buildEnv produces.
// +optional
ExtraEnv []corev1.EnvVar `json:"extraEnv,omitempty"`
// +optional
ExtraEnvFrom []corev1.EnvFromSource `json:"extraEnvFrom,omitempty"`
// extraVolumes are added to the pod spec; pair with extraVolumeMounts to
// actually mount one on the main container.
// +optional
ExtraVolumes []corev1.Volume `json:"extraVolumes,omitempty"`
// extraVolumeMounts are added to the main terdut-server container only
// -- not wait-for-postgres.
// +optional
ExtraVolumeMounts []corev1.VolumeMount `json:"extraVolumeMounts,omitempty"`
// +optional
ImagePullSecrets []corev1.LocalObjectReference `json:"imagePullSecrets,omitempty"`
// disruptionBudget, when set, causes the controller to reconcile a
// policyv1.PodDisruptionBudget selecting this TerdutServer's pods.
// Removing this field deletes any PodDisruptionBudget the controller
// previously created.
// +optional
DisruptionBudget *PodDisruptionBudgetSpec `json:"disruptionBudget,omitempty"`
}
// TerdutServerSpec defines the desired state of TerdutServer.
//
// The operator creates and owns every TerdutServer it manages (DESIGN.md
@@ -201,10 +299,14 @@ type TerdutServerSpec struct {
// +required
Image ImageSpec `json:"image"`
// replicas. terdut-server is not horizontally-scale-tested; keep this
// at its default of 1 unless you've verified otherwise -- the sweeper
// and the notifier are unsynchronised singletons.
// +kubebuilder:default=1
// replicas. Defaults to 2: terdut-server v0.36.0 put the sweeper, the
// notifier and the migration runner each behind a Postgres advisory
// lock, and gave incident creation its own conflict resolution, so
// more than one replica no longer double-pages, races a migration, or
// drops a webhook payload. image.tag must be v0.36.0 or newer for
// that to hold -- an older terdut-server has none of these guards,
// and this field does not check the tag for you.
// +kubebuilder:default=2
// +optional
Replicas int32 `json:"replicas,omitempty"`
@@ -217,9 +319,6 @@ type TerdutServerSpec struct {
// +optional
Sweeper SweeperSpec `json:"sweeper,omitempty"`
// +optional
Deadman DeadmanSpec `json:"deadman,omitempty"`
// +optional
Notify NotifySpec `json:"notify,omitempty"`
@@ -237,6 +336,11 @@ type TerdutServerSpec struct {
// this CRD's schema doesn't need a breaking change to grow it later.
// +optional
AllowedTeams AllowedTeams `json:"allowedTeams,omitempty"`
// pod is pod-level customization of the Deployment this TerdutServer
// creates (DESIGN.md §4.1).
// +optional
Pod PodSpec `json:"pod,omitempty"`
}
// Condition types this controller sets on TerdutServer.
@@ -244,15 +348,6 @@ const (
// ConditionReady is the standard top-level condition every CRD carries
// (DESIGN.md §7).
ConditionReady = "Ready"
// ConditionDatabaseReady reflects whether the configured database is
// usable -- for postgresClusterRef, whether the Zalando CR and its
// generated credentials Secret both resolved; for a plain dsn, always
// true once set (DESIGN.md §8: "no connectivity check beyond what the
// Deployment's own readiness probe already gives").
ConditionDatabaseReady = "DatabaseReady"
// ConditionBootstrapped reflects whether a working credential has been
// acquired via self-registration (DESIGN.md §6).
ConditionBootstrapped = "Bootstrapped"
)
// Condition reasons this controller sets.
@@ -270,13 +365,6 @@ const (
// instead -- but a TerdutServer that explicitly asks for it still needs
// to say clearly that it can't be satisfied).
ReasonPostgresOperatorCRDNotInstalled = "PostgresOperatorCRDNotInstalled"
// ReasonBootstrapStateLost: a checkpointed admin credential
// (DESIGN.md §6) was lost after being used but before the lasting
// credential it was for could be persisted -- the one genuinely
// pathological case in the self-registration flow. Fail-closed, same
// recovery as DESIGN.md §5's webhook-Secret-loss rule: delete and
// recreate this TerdutServer.
ReasonBootstrapStateLost = "BootstrapStateLost"
// ReasonAdopted: the happy path. A working credential is in hand, the
// Deployment has a ready replica, and the database (if postgresClusterRef)
// resolved.
@@ -297,16 +385,9 @@ type TerdutServerStatus struct {
// +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"`
// serviceName is the Service this controller created for the
// Deployment, so other objects can reference it without recomputing the
// naming convention.
// +optional
ServiceName string `json:"serviceName,omitempty"`
// credentialsSecretRef is the generated instance-scoped credential
// (DESIGN.md §6) -- pure output, always in the operator's own
// namespace, under a fixed data key ("token"). Set only once
// Bootstrapped is True.
// credentialsSecretRef is the operator key this controller generated for
// the server (TERDUT_OPERATOR_KEY): pure output, in the TerdutServer's own
// namespace and owned by it, under the data key "token".
// +optional
CredentialsSecretRef *SecretKeyRef `json:"credentialsSecretRef,omitempty"`
}
+125 -17
View File
@@ -33,16 +33,117 @@ type TerdutTeamSpec struct {
// +required
ServerRef TerdutServerRef `json:"serverRef"`
// displayName is this team's name, both in terdut-server's own data
// (POST /api/teams {"name": ...}) and as the identity POST /api/teams
// and GET /api/teams?name= correlate on (DESIGN.md §5's idempotent-
// create rule, via TEAM-LOOKUP.md).
// displayName is this team's name on the server. It can be changed freely:
// the team is found by the CR's own identity (<namespace>/<name>, sent as
// external_id), not by this name.
// +required
// +kubebuilder:validation:MinLength=1
DisplayName string `json:"displayName"`
// +optional
OIDC TerdutTeamOIDC `json:"oidc,omitempty"`
// escalation is this team's escalation ladder. Omitted, the team has none
// (the server's plain reminder behaviour applies).
// +optional
Escalation *EscalationSpec `json:"escalation,omitempty"`
// deadmanSwitches are this team's dead man's switches, by name. Switches on
// the server that are not listed here are removed: in operator mode this
// list is the whole truth.
// +listType=map
// +listMapKey=name
// +kubebuilder:validation:MaxItems=50
// +optional
DeadmanSwitches []DeadmanSwitchSpec `json:"deadmanSwitches,omitempty"`
}
// EscalationTargetKind is who one rung of the ladder pages.
// +kubebuilder:validation:Enum=oncall;user
type EscalationTargetKind string
const (
EscalationTargetOncall EscalationTargetKind = "oncall"
EscalationTargetUser EscalationTargetKind = "user"
)
// EscalationTarget is one page within a level. username is required iff kind is
// "user".
// +kubebuilder:validation:XValidation:rule="self.kind != 'user' || has(self.username)",message="username is required when kind is user"
// +kubebuilder:validation:XValidation:rule="self.kind != 'oncall' || !has(self.username)",message="username must not be set when kind is oncall"
type EscalationTarget struct {
// +required
Kind EscalationTargetKind `json:"kind"`
// +kubebuilder:validation:MaxLength=255
// +optional
Username string `json:"username,omitempty"`
}
// EscalationLevel is one rung of the ladder: how long to wait, and who to page
// if nobody has acknowledged by then.
type EscalationLevel struct {
// timeout is a Go duration string, e.g. "5m".
// +required
// +kubebuilder:validation:MaxLength=32
// +kubebuilder:validation:Pattern=`^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$`
Timeout string `json:"timeout"`
// +required
// +kubebuilder:validation:MinItems=1
// +kubebuilder:validation:MaxItems=20
Targets []EscalationTarget `json:"targets"`
}
// EscalationSpec is a team's escalation ladder.
type EscalationSpec struct {
// +kubebuilder:validation:Minimum=0
// +kubebuilder:validation:Maximum=10
// +optional
RepeatCount int64 `json:"repeatCount,omitempty"`
// +optional
FallbackTopic string `json:"fallbackTopic,omitempty"`
// +required
// +kubebuilder:validation:MinItems=1
// +kubebuilder:validation:MaxItems=10
Levels []EscalationLevel `json:"levels"`
}
// DeadmanSwitchSpec is one dead man's switch: the absence of an alert matching
// matcher for longer than timeout opens an incident.
type DeadmanSwitchSpec struct {
// name identifies the switch within the team.
// +required
// +kubebuilder:validation:MinLength=1
// +kubebuilder:validation:MaxLength=100
Name string `json:"name"`
// matcher names the alerts this switch watches, e.g.
// "alertname=Watchdog,cluster=prod". One matcher per switch.
// +required
// +kubebuilder:validation:MinLength=1
// +kubebuilder:validation:MaxLength=512
// +kubebuilder:validation:XValidation:rule="!self.contains(';')",message="one matcher per switch: add another entry instead of separating with ;"
Matcher string `json:"matcher"`
// timeout is a Go duration string, e.g. "15m".
// +required
// +kubebuilder:validation:MaxLength=32
// +kubebuilder:validation:Pattern=`^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$`
Timeout string `json:"timeout"`
// +kubebuilder:validation:Enum=critical;error;warning;info
// +kubebuilder:default=critical
// +optional
Severity string `json:"severity,omitempty"`
}
// TerdutTeamRef names the TerdutTeam a TerdutAlertSource belongs to. Always
// same-namespace as the CR itself.
type TerdutTeamRef struct {
// +kubebuilder:validation:MinLength=1
Name string `json:"name"`
}
// Condition reasons this controller sets.
@@ -58,10 +159,30 @@ const (
// DESIGN.md §5's "every child requeues with backoff, no cross-
// controller RPC" rule.
ReasonWaitingForServer = "WaitingForServer"
// ReasonTeamNameTaken: the server already has a team with spec.displayName
// that belongs to a different TerdutTeam (or to a person). The name is global
// to the server, so the operator waits for one of them to change.
ReasonTeamNameTaken = "TeamNameTaken"
// ReasonUnknownUser: an escalation target's username matches no user on the
// server (yet).
ReasonUnknownUser = "UnknownUser"
// ReasonInvalidSpec: a duration in the spec does not parse.
ReasonInvalidSpec = "InvalidSpec"
// ReasonTeamAdopted: the happy path.
ReasonTeamAdopted = "Adopted"
)
// Condition reasons shared by the resources that hang off a TerdutTeam
// (currently TerdutAlertSource).
const (
// ReasonTeamRefNotFound: spec.teamRef names no TerdutTeam (yet).
ReasonTeamRefNotFound = "TeamRefNotFound"
// ReasonWaitingForTeam: the referenced TerdutTeam exists but is not Ready.
ReasonWaitingForTeam = "WaitingForTeam"
// ReasonChildAdopted: the happy path.
ReasonChildAdopted = "Adopted"
)
// TerdutTeamStatus defines the observed state of TerdutTeam.
type TerdutTeamStatus struct {
// +listType=map
@@ -74,19 +195,6 @@ type TerdutTeamStatus struct {
// +optional
TeamID int64 `json:"teamID,omitempty"`
// credentialsSecretRef is this team's own scoped credential
// (DESIGN.md §6 point 3) -- pure output, always in the operator's own
// namespace, under a fixed data key ("token").
// +optional
CredentialsSecretRef *SecretKeyRef `json:"credentialsSecretRef,omitempty"`
// serverEndpoint is the resolved TerdutServer's base URL, resolved once
// here so no child controller (TerdutEscalationRule, TerdutDeadmanSwitch,
// TerdutAlertSource) ever needs its own RBAC on terdutservers just to
// find out where to send a request (DESIGN.md §5).
// +optional
ServerEndpoint string `json:"serverEndpoint,omitempty"`
// +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"`
}
+162 -212
View File
@@ -5,8 +5,10 @@
package v1alpha1
import (
corev1 "k8s.io/api/core/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/util/intstr"
)
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
@@ -71,16 +73,16 @@ func (in *DatabaseSpec) DeepCopy() *DatabaseSpec {
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *DeadmanSpec) DeepCopyInto(out *DeadmanSpec) {
func (in *DeadmanSwitchSpec) DeepCopyInto(out *DeadmanSwitchSpec) {
*out = *in
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new DeadmanSpec.
func (in *DeadmanSpec) DeepCopy() *DeadmanSpec {
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new DeadmanSwitchSpec.
func (in *DeadmanSwitchSpec) DeepCopy() *DeadmanSwitchSpec {
if in == nil {
return nil
}
out := new(DeadmanSpec)
out := new(DeadmanSwitchSpec)
in.DeepCopyInto(out)
return out
}
@@ -105,6 +107,28 @@ func (in *EscalationLevel) DeepCopy() *EscalationLevel {
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *EscalationSpec) DeepCopyInto(out *EscalationSpec) {
*out = *in
if in.Levels != nil {
in, out := &in.Levels, &out.Levels
*out = make([]EscalationLevel, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new EscalationSpec.
func (in *EscalationSpec) DeepCopy() *EscalationSpec {
if in == nil {
return nil
}
out := new(EscalationSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *EscalationTarget) DeepCopyInto(out *EscalationTarget) {
*out = *in
@@ -210,6 +234,128 @@ func (in *OIDCSpec) DeepCopy() *OIDCSpec {
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *PodDisruptionBudgetSpec) DeepCopyInto(out *PodDisruptionBudgetSpec) {
*out = *in
if in.MinAvailable != nil {
in, out := &in.MinAvailable, &out.MinAvailable
*out = new(intstr.IntOrString)
**out = **in
}
if in.MaxUnavailable != nil {
in, out := &in.MaxUnavailable, &out.MaxUnavailable
*out = new(intstr.IntOrString)
**out = **in
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PodDisruptionBudgetSpec.
func (in *PodDisruptionBudgetSpec) DeepCopy() *PodDisruptionBudgetSpec {
if in == nil {
return nil
}
out := new(PodDisruptionBudgetSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *PodSpec) DeepCopyInto(out *PodSpec) {
*out = *in
if in.Annotations != nil {
in, out := &in.Annotations, &out.Annotations
*out = make(map[string]string, len(*in))
for key, val := range *in {
(*out)[key] = val
}
}
if in.NodeSelector != nil {
in, out := &in.NodeSelector, &out.NodeSelector
*out = make(map[string]string, len(*in))
for key, val := range *in {
(*out)[key] = val
}
}
if in.Tolerations != nil {
in, out := &in.Tolerations, &out.Tolerations
*out = make([]corev1.Toleration, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.Affinity != nil {
in, out := &in.Affinity, &out.Affinity
*out = new(corev1.Affinity)
(*in).DeepCopyInto(*out)
}
if in.TopologySpreadConstraints != nil {
in, out := &in.TopologySpreadConstraints, &out.TopologySpreadConstraints
*out = make([]corev1.TopologySpreadConstraint, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
in.Resources.DeepCopyInto(&out.Resources)
if in.SecurityContext != nil {
in, out := &in.SecurityContext, &out.SecurityContext
*out = new(corev1.PodSecurityContext)
(*in).DeepCopyInto(*out)
}
if in.ContainerSecurityContext != nil {
in, out := &in.ContainerSecurityContext, &out.ContainerSecurityContext
*out = new(corev1.SecurityContext)
(*in).DeepCopyInto(*out)
}
if in.ExtraEnv != nil {
in, out := &in.ExtraEnv, &out.ExtraEnv
*out = make([]corev1.EnvVar, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ExtraEnvFrom != nil {
in, out := &in.ExtraEnvFrom, &out.ExtraEnvFrom
*out = make([]corev1.EnvFromSource, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ExtraVolumes != nil {
in, out := &in.ExtraVolumes, &out.ExtraVolumes
*out = make([]corev1.Volume, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ExtraVolumeMounts != nil {
in, out := &in.ExtraVolumeMounts, &out.ExtraVolumeMounts
*out = make([]corev1.VolumeMount, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ImagePullSecrets != nil {
in, out := &in.ImagePullSecrets, &out.ImagePullSecrets
*out = make([]corev1.LocalObjectReference, len(*in))
copy(*out, *in)
}
if in.DisruptionBudget != nil {
in, out := &in.DisruptionBudget, &out.DisruptionBudget
*out = new(PodDisruptionBudgetSpec)
(*in).DeepCopyInto(*out)
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PodSpec.
func (in *PodSpec) DeepCopy() *PodSpec {
if in == nil {
return nil
}
out := new(PodSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *PostgresClusterRef) DeepCopyInto(out *PostgresClusterRef) {
*out = *in
@@ -357,207 +503,6 @@ func (in *TerdutAlertSourceStatus) DeepCopy() *TerdutAlertSourceStatus {
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutDeadmanSwitch) DeepCopyInto(out *TerdutDeadmanSwitch) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ObjectMeta.DeepCopyInto(&out.ObjectMeta)
out.Spec = in.Spec
in.Status.DeepCopyInto(&out.Status)
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutDeadmanSwitch.
func (in *TerdutDeadmanSwitch) DeepCopy() *TerdutDeadmanSwitch {
if in == nil {
return nil
}
out := new(TerdutDeadmanSwitch)
in.DeepCopyInto(out)
return out
}
// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object.
func (in *TerdutDeadmanSwitch) DeepCopyObject() runtime.Object {
if c := in.DeepCopy(); c != nil {
return c
}
return nil
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutDeadmanSwitchList) DeepCopyInto(out *TerdutDeadmanSwitchList) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ListMeta.DeepCopyInto(&out.ListMeta)
if in.Items != nil {
in, out := &in.Items, &out.Items
*out = make([]TerdutDeadmanSwitch, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutDeadmanSwitchList.
func (in *TerdutDeadmanSwitchList) DeepCopy() *TerdutDeadmanSwitchList {
if in == nil {
return nil
}
out := new(TerdutDeadmanSwitchList)
in.DeepCopyInto(out)
return out
}
// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object.
func (in *TerdutDeadmanSwitchList) DeepCopyObject() runtime.Object {
if c := in.DeepCopy(); c != nil {
return c
}
return nil
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutDeadmanSwitchSpec) DeepCopyInto(out *TerdutDeadmanSwitchSpec) {
*out = *in
out.TeamRef = in.TeamRef
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutDeadmanSwitchSpec.
func (in *TerdutDeadmanSwitchSpec) DeepCopy() *TerdutDeadmanSwitchSpec {
if in == nil {
return nil
}
out := new(TerdutDeadmanSwitchSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutDeadmanSwitchStatus) DeepCopyInto(out *TerdutDeadmanSwitchStatus) {
*out = *in
if in.Conditions != nil {
in, out := &in.Conditions, &out.Conditions
*out = make([]v1.Condition, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutDeadmanSwitchStatus.
func (in *TerdutDeadmanSwitchStatus) DeepCopy() *TerdutDeadmanSwitchStatus {
if in == nil {
return nil
}
out := new(TerdutDeadmanSwitchStatus)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutEscalationRule) DeepCopyInto(out *TerdutEscalationRule) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ObjectMeta.DeepCopyInto(&out.ObjectMeta)
in.Spec.DeepCopyInto(&out.Spec)
in.Status.DeepCopyInto(&out.Status)
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutEscalationRule.
func (in *TerdutEscalationRule) DeepCopy() *TerdutEscalationRule {
if in == nil {
return nil
}
out := new(TerdutEscalationRule)
in.DeepCopyInto(out)
return out
}
// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object.
func (in *TerdutEscalationRule) DeepCopyObject() runtime.Object {
if c := in.DeepCopy(); c != nil {
return c
}
return nil
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutEscalationRuleList) DeepCopyInto(out *TerdutEscalationRuleList) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ListMeta.DeepCopyInto(&out.ListMeta)
if in.Items != nil {
in, out := &in.Items, &out.Items
*out = make([]TerdutEscalationRule, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutEscalationRuleList.
func (in *TerdutEscalationRuleList) DeepCopy() *TerdutEscalationRuleList {
if in == nil {
return nil
}
out := new(TerdutEscalationRuleList)
in.DeepCopyInto(out)
return out
}
// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object.
func (in *TerdutEscalationRuleList) DeepCopyObject() runtime.Object {
if c := in.DeepCopy(); c != nil {
return c
}
return nil
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutEscalationRuleSpec) DeepCopyInto(out *TerdutEscalationRuleSpec) {
*out = *in
out.TeamRef = in.TeamRef
if in.Levels != nil {
in, out := &in.Levels, &out.Levels
*out = make([]EscalationLevel, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutEscalationRuleSpec.
func (in *TerdutEscalationRuleSpec) DeepCopy() *TerdutEscalationRuleSpec {
if in == nil {
return nil
}
out := new(TerdutEscalationRuleSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutEscalationRuleStatus) DeepCopyInto(out *TerdutEscalationRuleStatus) {
*out = *in
if in.Conditions != nil {
in, out := &in.Conditions, &out.Conditions
*out = make([]v1.Condition, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutEscalationRuleStatus.
func (in *TerdutEscalationRuleStatus) DeepCopy() *TerdutEscalationRuleStatus {
if in == nil {
return nil
}
out := new(TerdutEscalationRuleStatus)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutServer) DeepCopyInto(out *TerdutServer) {
*out = *in
@@ -639,10 +584,10 @@ func (in *TerdutServerSpec) DeepCopyInto(out *TerdutServerSpec) {
out.Networking = in.Networking
in.Database.DeepCopyInto(&out.Database)
out.Sweeper = in.Sweeper
out.Deadman = in.Deadman
in.Notify.DeepCopyInto(&out.Notify)
in.OIDC.DeepCopyInto(&out.OIDC)
in.AllowedTeams.DeepCopyInto(&out.AllowedTeams)
in.Pod.DeepCopyInto(&out.Pod)
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutServerSpec.
@@ -687,7 +632,7 @@ func (in *TerdutTeam) DeepCopyInto(out *TerdutTeam) {
*out = *in
out.TypeMeta = in.TypeMeta
in.ObjectMeta.DeepCopyInto(&out.ObjectMeta)
out.Spec = in.Spec
in.Spec.DeepCopyInto(&out.Spec)
in.Status.DeepCopyInto(&out.Status)
}
@@ -776,6 +721,16 @@ func (in *TerdutTeamSpec) DeepCopyInto(out *TerdutTeamSpec) {
*out = *in
out.ServerRef = in.ServerRef
out.OIDC = in.OIDC
if in.Escalation != nil {
in, out := &in.Escalation, &out.Escalation
*out = new(EscalationSpec)
(*in).DeepCopyInto(*out)
}
if in.DeadmanSwitches != nil {
in, out := &in.DeadmanSwitches, &out.DeadmanSwitches
*out = make([]DeadmanSwitchSpec, len(*in))
copy(*out, *in)
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamSpec.
@@ -798,11 +753,6 @@ func (in *TerdutTeamStatus) DeepCopyInto(out *TerdutTeamStatus) {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.CredentialsSecretRef != nil {
in, out := &in.CredentialsSecretRef, &out.CredentialsSecretRef
*out = new(SecretKeyRef)
**out = **in
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamStatus.
+2 -2
View File
@@ -6,8 +6,8 @@ type: application
# These fields decide nothing: `make helm-package` passes --version and
# --app-version from the release tag (same reasoning as terdut-server's own
# chart). They're for whoever reads the tree before a tag exists.
version: 0.1.1
appVersion: "v0.1.1"
version: 0.5.0
appVersion: "v0.5.0"
keywords:
- kubernetes
@@ -1,184 +0,0 @@
{{- if .Values.crd.enabled }}
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
annotations:
{{- if .Values.crd.keep }}
"helm.sh/resource-policy": keep
{{- end }}
controller-gen.kubebuilder.io/version: v0.22.0
name: terdutdeadmanswitches.terdut.ryuvia.com
spec:
group: terdut.ryuvia.com
names:
kind: TerdutDeadmanSwitch
listKind: TerdutDeadmanSwitchList
plural: terdutdeadmanswitches
singular: terdutdeadmanswitch
scope: Namespaced
versions:
- additionalPrinterColumns:
- jsonPath: .spec.teamRef.name
name: Team
type: string
- jsonPath: .status.switchID
name: SwitchID
type: integer
- jsonPath: .status.conditions[?(@.type=="Ready")].status
name: Ready
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].reason
name: Reason
type: string
name: v1alpha1
schema:
openAPIV3Schema:
description: TerdutDeadmanSwitch is the Schema for the terdutdeadmanswitches
API
properties:
apiVersion:
description: |-
APIVersion defines the versioned schema of this representation of an object.
Servers should convert recognized schemas to the latest internal value, and
may reject unrecognized values.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources
type: string
kind:
description: |-
Kind is a string value representing the REST resource this object represents.
Servers may infer this from the endpoint the client submits requests to.
Cannot be updated.
In CamelCase.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds
type: string
metadata:
type: object
spec:
description: spec defines the desired state of TerdutDeadmanSwitch
properties:
matcher:
description: |-
matcher names the alerts this switch watches, e.g.
"alertname=Watchdog,cluster=prod". One matcher per switch -- add
another TerdutDeadmanSwitch instead of separating with ";"
(terdut-server's own restriction, mirrored here so a bad spec is
rejected at apply time).
minLength: 1
type: string
x-kubernetes-validations:
- message: 'one matcher per switch: add another TerdutDeadmanSwitch
instead of separating with ;'
rule: '!self.contains('';'')'
name:
description: |-
name is optional, same as the API: left empty, terdut-server derives
it from matcher's own canonical form, and that's what the
idempotent-create lookup matches against too.
type: string
severity:
default: critical
enum:
- critical
- error
- warning
- info
type: string
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
properties:
name:
minLength: 1
type: string
required:
- name
type: object
timeout:
description: timeout is a Go duration string, e.g. "15m".
minLength: 1
type: string
required:
- matcher
- teamRef
- timeout
type: object
status:
description: status defines the observed state of TerdutDeadmanSwitch
properties:
conditions:
items:
description: Condition contains details for one aspect of the current
state of this API Resource.
properties:
lastTransitionTime:
description: |-
lastTransitionTime is the last time the condition transitioned from one status to another.
This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable.
format: date-time
type: string
message:
description: |-
message is a human readable message indicating details about the transition.
This may be an empty string.
maxLength: 32768
type: string
observedGeneration:
description: |-
observedGeneration represents the .metadata.generation that the condition was set based upon.
For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date
with respect to the current state of the instance.
format: int64
minimum: 0
type: integer
reason:
description: |-
reason contains a programmatic identifier indicating the reason for the condition's last transition.
Producers of specific condition types may define expected values and meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
maxLength: 1024
minLength: 1
pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$
type: string
status:
description: status of the condition, one of True, False, Unknown.
enum:
- "True"
- "False"
- Unknown
type: string
type:
description: type of condition in CamelCase or in foo.example.com/CamelCase.
maxLength: 316
pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$
type: string
required:
- lastTransitionTime
- message
- reason
- status
- type
type: object
type: array
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
observedGeneration:
format: int64
type: integer
switchID:
description: switchID is the server-side id.
format: int64
type: integer
type: object
required:
- spec
type: object
served: true
storage: true
subresources:
status: {}
{{- end }}
@@ -1,195 +0,0 @@
{{- if .Values.crd.enabled }}
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
annotations:
{{- if .Values.crd.keep }}
"helm.sh/resource-policy": keep
{{- end }}
controller-gen.kubebuilder.io/version: v0.22.0
name: terdutescalationrules.terdut.ryuvia.com
spec:
group: terdut.ryuvia.com
names:
kind: TerdutEscalationRule
listKind: TerdutEscalationRuleList
plural: terdutescalationrules
singular: terdutescalationrule
scope: Namespaced
versions:
- additionalPrinterColumns:
- jsonPath: .spec.teamRef.name
name: Team
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].status
name: Ready
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].reason
name: Reason
type: string
name: v1alpha1
schema:
openAPIV3Schema:
description: TerdutEscalationRule is the Schema for the terdutescalationrules
API
properties:
apiVersion:
description: |-
APIVersion defines the versioned schema of this representation of an object.
Servers should convert recognized schemas to the latest internal value, and
may reject unrecognized values.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources
type: string
kind:
description: |-
Kind is a string value representing the REST resource this object represents.
Servers may infer this from the endpoint the client submits requests to.
Cannot be updated.
In CamelCase.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds
type: string
metadata:
type: object
spec:
description: spec defines the desired state of TerdutEscalationRule
properties:
fallbackTopic:
type: string
levels:
items:
description: |-
EscalationLevel is one rung of the ladder: how long to wait, and who to
page if nobody's acknowledged by then.
properties:
targets:
items:
description: |-
EscalationTarget is one page within a level. username is required iff
kind is "user" (terdut-server's own validation, internal/api/escalation.go's
handleSetEscalation -- mirrored here as a CEL rule so a bad spec is
rejected at apply time, not discovered on the next failed PUT).
properties:
kind:
description: EscalationTargetKind is who one rung of the
ladder pages.
enum:
- oncall
- user
type: string
username:
type: string
required:
- kind
type: object
x-kubernetes-validations:
- message: username is required when kind is user
rule: self.kind != 'user' || has(self.username)
- message: username must not be set when kind is oncall
rule: self.kind != 'oncall' || !has(self.username)
minItems: 1
type: array
timeout:
description: timeout is a Go duration string, e.g. "5m".
minLength: 1
type: string
required:
- targets
- timeout
type: object
minItems: 1
type: array
repeatCount:
format: int64
maximum: 10
minimum: 0
type: integer
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
properties:
name:
minLength: 1
type: string
required:
- name
type: object
required:
- levels
- teamRef
type: object
status:
description: status defines the observed state of TerdutEscalationRule
properties:
conditions:
items:
description: Condition contains details for one aspect of the current
state of this API Resource.
properties:
lastTransitionTime:
description: |-
lastTransitionTime is the last time the condition transitioned from one status to another.
This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable.
format: date-time
type: string
message:
description: |-
message is a human readable message indicating details about the transition.
This may be an empty string.
maxLength: 32768
type: string
observedGeneration:
description: |-
observedGeneration represents the .metadata.generation that the condition was set based upon.
For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date
with respect to the current state of the instance.
format: int64
minimum: 0
type: integer
reason:
description: |-
reason contains a programmatic identifier indicating the reason for the condition's last transition.
Producers of specific condition types may define expected values and meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
maxLength: 1024
minLength: 1
pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$
type: string
status:
description: status of the condition, one of True, False, Unknown.
enum:
- "True"
- "False"
- Unknown
type: string
type:
description: type of condition in CamelCase or in foo.example.com/CamelCase.
maxLength: 316
pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$
type: string
required:
- lastTransitionTime
- message
- reason
- status
- type
type: object
type: array
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
observedGeneration:
format: int64
type: integer
type: object
required:
- spec
type: object
served: true
storage: true
subresources:
status: {}
{{- end }}
File diff suppressed because it is too large Load Diff
@@ -55,14 +55,122 @@ spec:
spec:
description: spec defines the desired state of TerdutTeam
properties:
deadmanSwitches:
description: |-
deadmanSwitches are this team's dead man's switches, by name. Switches on
the server that are not listed here are removed: in operator mode this
list is the whole truth.
items:
description: |-
DeadmanSwitchSpec is one dead man's switch: the absence of an alert matching
matcher for longer than timeout opens an incident.
properties:
matcher:
description: |-
matcher names the alerts this switch watches, e.g.
"alertname=Watchdog,cluster=prod". One matcher per switch.
maxLength: 512
minLength: 1
type: string
x-kubernetes-validations:
- message: 'one matcher per switch: add another entry instead
of separating with ;'
rule: '!self.contains('';'')'
name:
description: name identifies the switch within the team.
maxLength: 100
minLength: 1
type: string
severity:
default: critical
enum:
- critical
- error
- warning
- info
type: string
timeout:
description: timeout is a Go duration string, e.g. "15m".
maxLength: 32
pattern: ^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$
type: string
required:
- matcher
- name
- timeout
type: object
maxItems: 50
type: array
x-kubernetes-list-map-keys:
- name
x-kubernetes-list-type: map
displayName:
description: |-
displayName is this team's name, both in terdut-server's own data
(POST /api/teams {"name": ...}) and as the identity POST /api/teams
and GET /api/teams?name= correlate on (DESIGN.md §5's idempotent-
create rule, via TEAM-LOOKUP.md).
displayName is this team's name on the server. It can be changed freely:
the team is found by the CR's own identity (<namespace>/<name>, sent as
external_id), not by this name.
minLength: 1
type: string
escalation:
description: |-
escalation is this team's escalation ladder. Omitted, the team has none
(the server's plain reminder behaviour applies).
properties:
fallbackTopic:
type: string
levels:
items:
description: |-
EscalationLevel is one rung of the ladder: how long to wait, and who to page
if nobody has acknowledged by then.
properties:
targets:
items:
description: |-
EscalationTarget is one page within a level. username is required iff kind is
"user".
properties:
kind:
description: EscalationTargetKind is who one rung
of the ladder pages.
enum:
- oncall
- user
type: string
username:
maxLength: 255
type: string
required:
- kind
type: object
x-kubernetes-validations:
- message: username is required when kind is user
rule: self.kind != 'user' || has(self.username)
- message: username must not be set when kind is oncall
rule: self.kind != 'oncall' || !has(self.username)
maxItems: 20
minItems: 1
type: array
timeout:
description: timeout is a Go duration string, e.g. "5m".
maxLength: 32
pattern: ^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$
type: string
required:
- targets
- timeout
type: object
maxItems: 10
minItems: 1
type: array
repeatCount:
format: int64
maximum: 10
minimum: 0
type: integer
required:
- levels
type: object
oidc:
description: |-
TerdutTeamOIDC binds which identity-provider groups grant membership and
@@ -152,35 +260,9 @@ spec:
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
credentialsSecretRef:
description: |-
credentialsSecretRef is this team's own scoped credential
(DESIGN.md §6 point 3) -- pure output, always in the operator's own
namespace, under a fixed data key ("token").
properties:
key:
description: key is the data key inside the Secret holding the
raw value.
minLength: 1
type: string
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- key
- name
type: object
observedGeneration:
format: int64
type: integer
serverEndpoint:
description: |-
serverEndpoint is the resolved TerdutServer's base URL, resolved once
here so no child controller (TerdutEscalationRule, TerdutDeadmanSwitch,
TerdutAlertSource) ever needs its own RBAC on terdutservers just to
find out where to send a request (DESIGN.md §5).
type: string
teamID:
description: |-
teamID is the server-side id -- needed by every child object's
@@ -58,12 +58,22 @@ rules:
verbs:
- create
- patch
- apiGroups:
- policy
resources:
- poddisruptionbudgets
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutalertsources
- terdutdeadmanswitches
- terdutescalationrules
- terdutservers
- terdutteams
verbs:
@@ -78,8 +88,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources/finalizers
- terdutdeadmanswitches/finalizers
- terdutescalationrules/finalizers
- terdutservers/finalizers
- terdutteams/finalizers
verbs:
@@ -88,8 +96,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources/status
- terdutdeadmanswitches/status
- terdutescalationrules/status
- terdutservers/status
- terdutteams/status
verbs:
@@ -1,31 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutdeadmanswitch-admin-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- '*'
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
{{- end }}
@@ -1,37 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutdeadmanswitch-editor-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
{{- end }}
@@ -1,33 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutdeadmanswitch-viewer-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- get
- list
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
{{- end }}
@@ -1,31 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutescalationrule-admin-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- '*'
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
{{- end }}
@@ -1,37 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutescalationrule-editor-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
{{- end }}
@@ -1,33 +0,0 @@
{{- if .Values.rbac.helpers.enabled }}
apiVersion: rbac.authorization.k8s.io/v1
{{- if .Values.rbac.namespaced }}
kind: Role
{{- else }}
kind: ClusterRole
{{- end }}
metadata:
{{- if .Values.rbac.namespaced }}
namespace: {{ .Release.Namespace }}
{{- end }}
labels:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/name: {{ include "terdut-operator.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/instance: {{ .Release.Name }}
name: {{ include "terdut-operator.resourceName" (dict "suffix" "terdutescalationrule-viewer-role" "context" $) }}
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- get
- list
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
{{- end }}
@@ -34,10 +34,6 @@ spec:
sweeper:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with .Values.terdutServer.deadman }}
deadman:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with .Values.terdutServer.notify }}
notify:
{{- toYaml . | nindent 4 }}
@@ -51,4 +47,8 @@ spec:
allowedTeams:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with .Values.terdutServer.pod }}
pod:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- end }}
+27 -8
View File
@@ -233,12 +233,19 @@ terdutServer:
## Required when terdutServer.enabled.
# tag: ""
replicas: 1
## Safe above 1 since terdut-server v0.36.0 (image.tag above must be that or
## newer): the sweeper, notifier and migration runner are each behind a
## Postgres advisory lock, and incident creation resolves its own insert
## conflict, matching this CRD's own spec.replicas default.
replicas: 2
networking:
## Required when terdutServer.enabled -- the hostname a future
## HTTPRoute will carry (see NetworkingSpec's own doc comment: creating
## that HTTPRoute isn't implemented yet).
## Required when terdutServer.enabled -- terdut-server's own public
## URL, used for absolute links it generates itself (notifications,
## OIDC redirect URIs). This chart/operator never creates any
## ingress/HTTPRoute for it -- see examples/networking in the repo for
## how to expose the Service this chart's TerdutServer CR causes to
## be created, if you want to expose it at all.
# hostname: ""
servicePort: 8080
@@ -259,10 +266,6 @@ terdutServer:
# sweeper:
# staleAfter: 6h
# archiveAfter: 168h
# deadman:
# matchers: "alertname=Watchdog"
# timeout: 15m
# severity: critical
# notify: {}
# oidc: {}
@@ -270,3 +273,19 @@ terdutServer:
# allowedTeams: {}
## Pod-level customization of the Deployment this TerdutServer creates --
## see api/v1alpha1/terdutserver_types.go's PodSpec for the full shape
## (nodeSelector, tolerations, affinity, topologySpreadConstraints,
## securityContext, containerSecurityContext, serviceAccountName,
## extraEnv/extraEnvFrom, extraVolumes/extraVolumeMounts,
## imagePullSecrets, disruptionBudget). All optional; resources is the
## one most installs will want to set, since the container otherwise
## runs with no requests/limits at all:
# pod:
# resources:
# requests:
# cpu: 100m
# memory: 128Mi
# limits:
# memory: 256Mi
+6 -36
View File
@@ -166,53 +166,23 @@ func main() {
os.Exit(1)
}
// POD_NAMESPACE is the operator's own namespace, via the Deployment's
// downward API (config/manager/manager.yaml) — every credentials Secret
// TerdutServerReconciler reads or writes lives here, never in a
// TerdutServer's own namespace (DESIGN.md §6). Falling back to "default"
// keeps `go run` usable for local development against a real cluster;
// production always sets it.
operatorNamespace := os.Getenv("POD_NAMESPACE")
if operatorNamespace == "" {
operatorNamespace = "default"
}
if err := (&controller.TerdutServerReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutserver")
os.Exit(1)
}
if err := (&controller.TerdutTeamReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutteam")
os.Exit(1)
}
if err := (&controller.TerdutEscalationRuleReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutescalationrule")
os.Exit(1)
}
if err := (&controller.TerdutDeadmanSwitchReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutdeadmanswitch")
os.Exit(1)
}
if err := (&controller.TerdutAlertSourceReconciler{
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
OperatorNamespace: operatorNamespace,
Client: mgr.GetClient(),
Scheme: mgr.GetScheme(),
}).SetupWithManager(mgr); err != nil {
setupLog.Error(err, "Failed to create controller", "controller", "terdutalertsource")
os.Exit(1)
@@ -76,9 +76,8 @@ spec:
type: string
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
TerdutTeamRef names the TerdutTeam a TerdutAlertSource belongs to. Always
same-namespace as the CR itself.
properties:
name:
minLength: 1
@@ -1,180 +0,0 @@
---
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
annotations:
controller-gen.kubebuilder.io/version: v0.22.0
name: terdutdeadmanswitches.terdut.ryuvia.com
spec:
group: terdut.ryuvia.com
names:
kind: TerdutDeadmanSwitch
listKind: TerdutDeadmanSwitchList
plural: terdutdeadmanswitches
singular: terdutdeadmanswitch
scope: Namespaced
versions:
- additionalPrinterColumns:
- jsonPath: .spec.teamRef.name
name: Team
type: string
- jsonPath: .status.switchID
name: SwitchID
type: integer
- jsonPath: .status.conditions[?(@.type=="Ready")].status
name: Ready
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].reason
name: Reason
type: string
name: v1alpha1
schema:
openAPIV3Schema:
description: TerdutDeadmanSwitch is the Schema for the terdutdeadmanswitches
API
properties:
apiVersion:
description: |-
APIVersion defines the versioned schema of this representation of an object.
Servers should convert recognized schemas to the latest internal value, and
may reject unrecognized values.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources
type: string
kind:
description: |-
Kind is a string value representing the REST resource this object represents.
Servers may infer this from the endpoint the client submits requests to.
Cannot be updated.
In CamelCase.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds
type: string
metadata:
type: object
spec:
description: spec defines the desired state of TerdutDeadmanSwitch
properties:
matcher:
description: |-
matcher names the alerts this switch watches, e.g.
"alertname=Watchdog,cluster=prod". One matcher per switch -- add
another TerdutDeadmanSwitch instead of separating with ";"
(terdut-server's own restriction, mirrored here so a bad spec is
rejected at apply time).
minLength: 1
type: string
x-kubernetes-validations:
- message: 'one matcher per switch: add another TerdutDeadmanSwitch
instead of separating with ;'
rule: '!self.contains('';'')'
name:
description: |-
name is optional, same as the API: left empty, terdut-server derives
it from matcher's own canonical form, and that's what the
idempotent-create lookup matches against too.
type: string
severity:
default: critical
enum:
- critical
- error
- warning
- info
type: string
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
properties:
name:
minLength: 1
type: string
required:
- name
type: object
timeout:
description: timeout is a Go duration string, e.g. "15m".
minLength: 1
type: string
required:
- matcher
- teamRef
- timeout
type: object
status:
description: status defines the observed state of TerdutDeadmanSwitch
properties:
conditions:
items:
description: Condition contains details for one aspect of the current
state of this API Resource.
properties:
lastTransitionTime:
description: |-
lastTransitionTime is the last time the condition transitioned from one status to another.
This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable.
format: date-time
type: string
message:
description: |-
message is a human readable message indicating details about the transition.
This may be an empty string.
maxLength: 32768
type: string
observedGeneration:
description: |-
observedGeneration represents the .metadata.generation that the condition was set based upon.
For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date
with respect to the current state of the instance.
format: int64
minimum: 0
type: integer
reason:
description: |-
reason contains a programmatic identifier indicating the reason for the condition's last transition.
Producers of specific condition types may define expected values and meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
maxLength: 1024
minLength: 1
pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$
type: string
status:
description: status of the condition, one of True, False, Unknown.
enum:
- "True"
- "False"
- Unknown
type: string
type:
description: type of condition in CamelCase or in foo.example.com/CamelCase.
maxLength: 316
pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$
type: string
required:
- lastTransitionTime
- message
- reason
- status
- type
type: object
type: array
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
observedGeneration:
format: int64
type: integer
switchID:
description: switchID is the server-side id.
format: int64
type: integer
type: object
required:
- spec
type: object
served: true
storage: true
subresources:
status: {}
@@ -1,191 +0,0 @@
---
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
annotations:
controller-gen.kubebuilder.io/version: v0.22.0
name: terdutescalationrules.terdut.ryuvia.com
spec:
group: terdut.ryuvia.com
names:
kind: TerdutEscalationRule
listKind: TerdutEscalationRuleList
plural: terdutescalationrules
singular: terdutescalationrule
scope: Namespaced
versions:
- additionalPrinterColumns:
- jsonPath: .spec.teamRef.name
name: Team
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].status
name: Ready
type: string
- jsonPath: .status.conditions[?(@.type=="Ready")].reason
name: Reason
type: string
name: v1alpha1
schema:
openAPIV3Schema:
description: TerdutEscalationRule is the Schema for the terdutescalationrules
API
properties:
apiVersion:
description: |-
APIVersion defines the versioned schema of this representation of an object.
Servers should convert recognized schemas to the latest internal value, and
may reject unrecognized values.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources
type: string
kind:
description: |-
Kind is a string value representing the REST resource this object represents.
Servers may infer this from the endpoint the client submits requests to.
Cannot be updated.
In CamelCase.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds
type: string
metadata:
type: object
spec:
description: spec defines the desired state of TerdutEscalationRule
properties:
fallbackTopic:
type: string
levels:
items:
description: |-
EscalationLevel is one rung of the ladder: how long to wait, and who to
page if nobody's acknowledged by then.
properties:
targets:
items:
description: |-
EscalationTarget is one page within a level. username is required iff
kind is "user" (terdut-server's own validation, internal/api/escalation.go's
handleSetEscalation -- mirrored here as a CEL rule so a bad spec is
rejected at apply time, not discovered on the next failed PUT).
properties:
kind:
description: EscalationTargetKind is who one rung of the
ladder pages.
enum:
- oncall
- user
type: string
username:
type: string
required:
- kind
type: object
x-kubernetes-validations:
- message: username is required when kind is user
rule: self.kind != 'user' || has(self.username)
- message: username must not be set when kind is oncall
rule: self.kind != 'oncall' || !has(self.username)
minItems: 1
type: array
timeout:
description: timeout is a Go duration string, e.g. "5m".
minLength: 1
type: string
required:
- targets
- timeout
type: object
minItems: 1
type: array
repeatCount:
format: int64
maximum: 10
minimum: 0
type: integer
teamRef:
description: |-
TerdutTeamRef names the TerdutTeam this resource belongs to. Always
same-namespace as the CR itself (DESIGN.md §1: only TerdutTeam.spec.serverRef
crosses namespaces in v1) -- no namespace field, unlike TerdutServerRef.
properties:
name:
minLength: 1
type: string
required:
- name
type: object
required:
- levels
- teamRef
type: object
status:
description: status defines the observed state of TerdutEscalationRule
properties:
conditions:
items:
description: Condition contains details for one aspect of the current
state of this API Resource.
properties:
lastTransitionTime:
description: |-
lastTransitionTime is the last time the condition transitioned from one status to another.
This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable.
format: date-time
type: string
message:
description: |-
message is a human readable message indicating details about the transition.
This may be an empty string.
maxLength: 32768
type: string
observedGeneration:
description: |-
observedGeneration represents the .metadata.generation that the condition was set based upon.
For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date
with respect to the current state of the instance.
format: int64
minimum: 0
type: integer
reason:
description: |-
reason contains a programmatic identifier indicating the reason for the condition's last transition.
Producers of specific condition types may define expected values and meanings for this field,
and whether the values are considered a guaranteed API.
The value should be a CamelCase string.
This field may not be empty.
maxLength: 1024
minLength: 1
pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$
type: string
status:
description: status of the condition, one of True, False, Unknown.
enum:
- "True"
- "False"
- Unknown
type: string
type:
description: type of condition in CamelCase or in foo.example.com/CamelCase.
maxLength: 316
pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$
type: string
required:
- lastTransitionTime
- message
- reason
- status
- type
type: object
type: array
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
observedGeneration:
format: int64
type: integer
type: object
required:
- spec
type: object
served: true
storage: true
subresources:
status: {}
File diff suppressed because it is too large Load Diff
@@ -52,14 +52,122 @@ spec:
spec:
description: spec defines the desired state of TerdutTeam
properties:
deadmanSwitches:
description: |-
deadmanSwitches are this team's dead man's switches, by name. Switches on
the server that are not listed here are removed: in operator mode this
list is the whole truth.
items:
description: |-
DeadmanSwitchSpec is one dead man's switch: the absence of an alert matching
matcher for longer than timeout opens an incident.
properties:
matcher:
description: |-
matcher names the alerts this switch watches, e.g.
"alertname=Watchdog,cluster=prod". One matcher per switch.
maxLength: 512
minLength: 1
type: string
x-kubernetes-validations:
- message: 'one matcher per switch: add another entry instead
of separating with ;'
rule: '!self.contains('';'')'
name:
description: name identifies the switch within the team.
maxLength: 100
minLength: 1
type: string
severity:
default: critical
enum:
- critical
- error
- warning
- info
type: string
timeout:
description: timeout is a Go duration string, e.g. "15m".
maxLength: 32
pattern: ^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$
type: string
required:
- matcher
- name
- timeout
type: object
maxItems: 50
type: array
x-kubernetes-list-map-keys:
- name
x-kubernetes-list-type: map
displayName:
description: |-
displayName is this team's name, both in terdut-server's own data
(POST /api/teams {"name": ...}) and as the identity POST /api/teams
and GET /api/teams?name= correlate on (DESIGN.md §5's idempotent-
create rule, via TEAM-LOOKUP.md).
displayName is this team's name on the server. It can be changed freely:
the team is found by the CR's own identity (<namespace>/<name>, sent as
external_id), not by this name.
minLength: 1
type: string
escalation:
description: |-
escalation is this team's escalation ladder. Omitted, the team has none
(the server's plain reminder behaviour applies).
properties:
fallbackTopic:
type: string
levels:
items:
description: |-
EscalationLevel is one rung of the ladder: how long to wait, and who to page
if nobody has acknowledged by then.
properties:
targets:
items:
description: |-
EscalationTarget is one page within a level. username is required iff kind is
"user".
properties:
kind:
description: EscalationTargetKind is who one rung
of the ladder pages.
enum:
- oncall
- user
type: string
username:
maxLength: 255
type: string
required:
- kind
type: object
x-kubernetes-validations:
- message: username is required when kind is user
rule: self.kind != 'user' || has(self.username)
- message: username must not be set when kind is oncall
rule: self.kind != 'oncall' || !has(self.username)
maxItems: 20
minItems: 1
type: array
timeout:
description: timeout is a Go duration string, e.g. "5m".
maxLength: 32
pattern: ^([0-9]+(\.[0-9]+)?(ns|us|µs|ms|s|m|h))+$
type: string
required:
- targets
- timeout
type: object
maxItems: 10
minItems: 1
type: array
repeatCount:
format: int64
maximum: 10
minimum: 0
type: integer
required:
- levels
type: object
oidc:
description: |-
TerdutTeamOIDC binds which identity-provider groups grant membership and
@@ -149,35 +257,9 @@ spec:
x-kubernetes-list-map-keys:
- type
x-kubernetes-list-type: map
credentialsSecretRef:
description: |-
credentialsSecretRef is this team's own scoped credential
(DESIGN.md §6 point 3) -- pure output, always in the operator's own
namespace, under a fixed data key ("token").
properties:
key:
description: key is the data key inside the Secret holding the
raw value.
minLength: 1
type: string
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- key
- name
type: object
observedGeneration:
format: int64
type: integer
serverEndpoint:
description: |-
serverEndpoint is the resolved TerdutServer's base URL, resolved once
here so no child controller (TerdutEscalationRule, TerdutDeadmanSwitch,
TerdutAlertSource) ever needs its own RBAC on terdutservers just to
find out where to send a request (DESIGN.md §5).
type: string
teamID:
description: |-
teamID is the server-side id -- needed by every child object's
-2
View File
@@ -4,8 +4,6 @@
resources:
- bases/terdut.ryuvia.com_terdutservers.yaml
- bases/terdut.ryuvia.com_terdutteams.yaml
- bases/terdut.ryuvia.com_terdutescalationrules.yaml
- bases/terdut.ryuvia.com_terdutdeadmanswitches.yaml
- bases/terdut.ryuvia.com_terdutalertsources.yaml
# +kubebuilder:scaffold:crdkustomizeresource
-6
View File
@@ -23,14 +23,8 @@ resources:
#- ../webhook
# [CERTMANAGER] To enable cert-manager, uncomment all sections with 'CERTMANAGER'. 'WEBHOOK' components are required.
#- ../certmanager
# [PROMETHEUS] To enable prometheus monitor, uncomment all sections with 'PROMETHEUS'.
#- ../prometheus
# [METRICS] Expose the controller manager metrics service.
- metrics_service.yaml
# [NETWORK POLICY] Control ingress to metrics and webhook ports.
# Allow metrics traffic from pods in namespaces labeled 'metrics: enabled'.
# Allow webhook traffic from all sources.
#- ../network-policy
# Uncomment the patches line if you enable Metrics
patches:
@@ -1,26 +0,0 @@
# Allow metrics traffic from pods in namespaces labeled 'metrics: enabled'.
# Add this label to namespaces whose pods should scrape metrics.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: allow-metrics-traffic
namespace: system
spec:
podSelector:
matchLabels:
control-plane: controller-manager
app.kubernetes.io/name: terdut-operator
policyTypes:
- Ingress
ingress:
# Allow pods in namespaces labeled 'metrics: enabled' to scrape metrics.
- from:
- namespaceSelector:
matchLabels:
metrics: enabled # Only from namespaces with this label
ports:
- port: 8443
protocol: TCP
-2
View File
@@ -1,2 +0,0 @@
resources:
- allow-metrics-traffic.yaml
-11
View File
@@ -1,11 +0,0 @@
resources:
- monitor.yaml
# [PROMETHEUS-WITH-CERTS] The following patch configures the ServiceMonitor in ../prometheus
# to securely reference certificates created and managed by cert-manager.
# Additionally, ensure that you uncomment the [METRICS WITH CERTMANAGER] patch under config/default/kustomization.yaml
# to mount the "metrics-server-cert" secret in the Manager Deployment.
#patches:
# - path: monitor_tls_patch.yaml
# target:
# kind: ServiceMonitor
-27
View File
@@ -1,27 +0,0 @@
# Prometheus Monitor Service (Metrics)
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
labels:
control-plane: controller-manager
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: controller-manager-metrics-monitor
namespace: system
spec:
endpoints:
- path: /metrics
port: https # Ensure this is the name of the port that exposes HTTPS metrics
scheme: https
bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token
tlsConfig:
# TODO(user): The option insecureSkipVerify: true is not recommended for production since it disables
# certificate verification, exposing the system to potential man-in-the-middle attacks.
# For production environments, it is recommended to use cert-manager for automatic TLS certificate management.
# To apply this configuration, enable cert-manager and use the patch located at config/prometheus/servicemonitor_tls_patch.yaml,
# which securely references the certificate from the 'metrics-server-cert' secret.
insecureSkipVerify: true
selector:
matchLabels:
control-plane: controller-manager
app.kubernetes.io/name: terdut-operator
-19
View File
@@ -1,19 +0,0 @@
# Patch for Prometheus ServiceMonitor to enable secure TLS configuration
# using certificates managed by cert-manager
- op: replace
path: /spec/endpoints/0/tlsConfig
value:
# SERVICE_NAME and SERVICE_NAMESPACE will be substituted by kustomize
serverName: SERVICE_NAME.SERVICE_NAMESPACE.svc
insecureSkipVerify: false
ca:
secret:
name: metrics-server-cert
key: ca.crt
cert:
secret:
name: metrics-server-cert
key: tls.crt
keySecret:
name: metrics-server-cert
key: tls.key
-6
View File
@@ -25,12 +25,6 @@ resources:
- terdutalertsource_admin_role.yaml
- terdutalertsource_editor_role.yaml
- terdutalertsource_viewer_role.yaml
- terdutdeadmanswitch_admin_role.yaml
- terdutdeadmanswitch_editor_role.yaml
- terdutdeadmanswitch_viewer_role.yaml
- terdutescalationrule_admin_role.yaml
- terdutescalationrule_editor_role.yaml
- terdutescalationrule_viewer_role.yaml
- terdutteam_admin_role.yaml
- terdutteam_editor_role.yaml
- terdutteam_viewer_role.yaml
+12 -6
View File
@@ -52,12 +52,22 @@ rules:
verbs:
- create
- patch
- apiGroups:
- policy
resources:
- poddisruptionbudgets
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutalertsources
- terdutdeadmanswitches
- terdutescalationrules
- terdutservers
- terdutteams
verbs:
@@ -72,8 +82,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources/finalizers
- terdutdeadmanswitches/finalizers
- terdutescalationrules/finalizers
- terdutservers/finalizers
- terdutteams/finalizers
verbs:
@@ -82,8 +90,6 @@ rules:
- terdut.ryuvia.com
resources:
- terdutalertsources/status
- terdutdeadmanswitches/status
- terdutescalationrules/status
- terdutservers/status
- terdutteams/status
verbs:
@@ -1,27 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants full permissions ('*') over terdut.ryuvia.com.
# This role is intended for users authorized to modify roles and bindings within the cluster,
# enabling them to delegate specific permissions to other users or groups as needed.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutdeadmanswitch-admin-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- '*'
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
@@ -1,33 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants permissions to create, update, and delete resources within the terdut.ryuvia.com.
# This role is intended for users who need to manage these resources
# but should not control RBAC or manage permissions for others.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutdeadmanswitch-editor-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
@@ -1,29 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants read-only access to terdut.ryuvia.com resources.
# This role is intended for users who need visibility into these resources
# without permissions to modify them. It is ideal for monitoring purposes and limited-access viewing.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutdeadmanswitch-viewer-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches
verbs:
- get
- list
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutdeadmanswitches/status
verbs:
- get
@@ -1,27 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants full permissions ('*') over terdut.ryuvia.com.
# This role is intended for users authorized to modify roles and bindings within the cluster,
# enabling them to delegate specific permissions to other users or groups as needed.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutescalationrule-admin-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- '*'
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
@@ -1,33 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants permissions to create, update, and delete resources within the terdut.ryuvia.com.
# This role is intended for users who need to manage these resources
# but should not control RBAC or manage permissions for others.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutescalationrule-editor-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
@@ -1,29 +0,0 @@
# This rule is not used by the project terdut-operator itself.
# It is provided to allow the cluster admin to help manage permissions for users.
#
# Grants read-only access to terdut.ryuvia.com resources.
# This role is intended for users who need visibility into these resources
# without permissions to modify them. It is ideal for monitoring purposes and limited-access viewing.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutescalationrule-viewer-role
rules:
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules
verbs:
- get
- list
- watch
- apiGroups:
- terdut.ryuvia.com
resources:
- terdutescalationrules/status
verbs:
- get
-8
View File
@@ -1,8 +0,0 @@
## Append samples of your project ##
resources:
- terdut_v1alpha1_terdutserver.yaml
- terdut_v1alpha1_terdutteam.yaml
- terdut_v1alpha1_terdutescalationrule.yaml
- terdut_v1alpha1_terdutdeadmanswitch.yaml
- terdut_v1alpha1_terdutalertsource.yaml
# +kubebuilder:scaffold:manifestskustomizesamples
@@ -1,19 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutAlertSource
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutalertsource-sample
spec:
teamRef:
name: terdutteam-sample
# kind defaults to "alertmanager" -- the only value terdut-server
# supports today. Changing it after this object exists rotates the
# webhook key (DESIGN.md §5): the old integration is deleted and a new
# one created, which breaks whatever still sends to the old URL.
kind: alertmanager
# name is this source's own display name server-side, distinct from this
# object's own metadata.name above -- renaming it is safe and never
# rotates the key.
name: prod-alertmanager
@@ -1,15 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutDeadmanSwitch
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutdeadmanswitch-sample
spec:
teamRef:
name: terdutteam-sample
# name is optional -- left empty, terdut-server derives it from matcher's
# own canonical form (DESIGN.md §4.4).
matcher: "alertname=Watchdog"
timeout: 15m
severity: critical
@@ -1,25 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutEscalationRule
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutescalationrule-sample
spec:
# One per team (DESIGN.md §4.3) -- a second TerdutEscalationRule naming
# the same teamRef would simply clobber this one every reconcile, since
# there's no admission-time check for it in v1.
teamRef:
name: terdutteam-sample
repeatCount: 2
fallbackTopic: platform-fallback
levels:
# username is required iff kind is "user", and rejected otherwise --
# enforced at apply time via CEL (api/v1alpha1/terdutescalationrule_types.go).
- timeout: 5m
targets:
- kind: user
username: alice
- timeout: 10m
targets:
- kind: oncall
@@ -1,33 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutserver-sample
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
tag: v0.20.0
replicas: 1
networking:
hostname: terdut.example.com
servicePort: 8080
# Bring-your-own DSN (simplest path, no external CRD dependency). For the
# Zalando postgres-operator path instead, use:
# database:
# postgresClusterRef:
# name: terdut-postgres
database:
dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require"
passwordSecretRef:
name: terdut-postgres-password
key: password
sweeper:
staleAfter: 6h
archiveAfter: 168h
deadman:
matchers: "alertname=Watchdog"
timeout: 15m
severity: critical
passwordLogin: true
@@ -1,20 +0,0 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
labels:
app.kubernetes.io/name: terdut-operator
app.kubernetes.io/managed-by: kustomize
name: terdutteam-sample
spec:
# serverRef.namespace is optional, defaulting to this TerdutTeam's own
# namespace (the common case). Set it only to reference a TerdutServer in
# a different namespace -- which needs that TerdutServer's own
# spec.allowedTeams to admit this namespace (DESIGN.md §4.6), otherwise
# this reports Ready: False, reason: RefNotPermitted.
serverRef:
name: terdutserver-sample
displayName: Platform
# oidc is optional -- omit entirely for a password-login-only install.
# oidc:
# memberGroup: terdut-platform-members
# ownerGroup: terdut-platform-owners
+96
View File
@@ -0,0 +1,96 @@
# Demo-only Postgres: a bare Deployment+Service+Secret, not the Zalando
# postgres-operator path (DatabaseSpec.postgresClusterRef, DESIGN.md §8).
# Bring-your-own DSN is the simpler of the two paths to stand up from
# nothing (ROADMAP.md Stage 1's own note), which is all this needs to be.
#
# emptyDir, one replica, a password sitting in a plaintext Secret below --
# none of that is how you'd run Postgres for real. It exists only so
# 01-server.yaml has something to talk to. Throw the whole demo namespace
# away when you're done; nothing here is meant to survive that.
apiVersion: v1
kind: Secret
metadata:
name: terdut-operator-demo-postgres
type: Opaque
stringData:
password: demo-not-a-real-password
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: terdut-operator-demo-postgres
labels:
app: terdut-operator-demo-postgres
spec:
replicas: 1
# Recreate, not RollingUpdate: emptyDir means a new pod starts with an
# empty database anyway, and two Postgres pods would never agree on one
# emptyDir each.
strategy:
type: Recreate
selector:
matchLabels:
app: terdut-operator-demo-postgres
template:
metadata:
labels:
app: terdut-operator-demo-postgres
spec:
containers:
- name: postgres
image: postgres:17-alpine
# Partial, deliberately: the official image's entrypoint needs to
# start as root to chown/chmod the data directory before it drops
# privileges itself (gosu, to the postgres user) -- forcing
# runAsNonRoot would just refuse to start the container, and
# dropping all capabilities (an earlier version of this file did)
# takes CAP_CHOWN/CAP_FOWNER away from that same root user, which
# is a different way of breaking the identical startup step:
# confirmed the hard way, as `chmod: /var/run/postgresql:
# Operation not permitted` in a real pod's logs, not caught by
# `kubectl apply --dry-run=server` -- that only checks admission
# policy, never whether the container actually boots. A
# "restricted" PodSecurity namespace warns on the remaining gap
# (no runAsNonRoot) rather than blocking, which is an acceptable
# tradeoff for Postgres that exists only to be thrown away with
# the rest of this demo.
securityContext:
allowPrivilegeEscalation: false
seccompProfile:
type: RuntimeDefault
ports:
- name: postgres
containerPort: 5432
env:
- name: POSTGRES_USER
value: terdut
- name: POSTGRES_DB
value: terdut
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: terdut-operator-demo-postgres
key: password
volumeMounts:
- name: data
mountPath: /var/lib/postgresql/data
subPath: pgdata
readinessProbe:
exec:
command: ["pg_isready", "-U", "terdut"]
initialDelaySeconds: 5
volumes:
- name: data
emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
name: terdut-operator-demo-postgres
spec:
selector:
app: terdut-operator-demo-postgres
ports:
- name: postgres
port: 5432
targetPort: postgres
+46
View File
@@ -0,0 +1,46 @@
# The one TerdutServer this whole demo runs against. Everything else in
# this directory (teams with their escalation and dead man's switches,
# alert sources) references it by name.
#
# The operator never creates any ingress/HTTPRoute for this TerdutServer --
# that's a permanent non-goal (DESIGN.md §1, NetworkingSpec's own doc
# comment), not a missing feature. This demo reaches it only by
# port-forwarding its Service, same name as this object (see README.md);
# see ../networking for worked examples of exposing it yourself instead.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
name: terdut-operator-demo
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
# v0.36.0 is the floor now that replicas below is 2 (this demo pins
# the current release, v0.43.0, so it shows the current web UI too): that
# release put the sweeper, the notifier and the migration runner each
# behind a Postgres advisory lock, and gave incident creation its own
# conflict resolution, which is what makes a second replica safe
# instead of racing the first. (Still carries v0.34.0's fix too --
# callerMayManageServiceAccount, so an instance-scoped service account
# can adopt/rotate a key on a team-scoped account it didn't just create
# in the same call -- without which terdutteam-* can wedge permanently
# on the crash-window race this demo hit live, niklas/terdut-operator#3.)
tag: v0.43.0
# Matches this CRD's own spec.replicas default (v0.4.0) -- stated
# explicitly, like every other field in this file, rather than left to
# the default. RollingUpdate follows automatically; this operator does
# not expose Strategy as a spec field.
replicas: 2
networking:
hostname: terdut-operator-demo.example
servicePort: 8080
database:
dsn: "postgres://terdut@terdut-operator-demo-postgres:5432/terdut?sslmode=disable"
passwordSecretRef:
name: terdut-operator-demo-postgres
key: password
sweeper:
staleAfter: 6h
archiveAfter: 168h
# No oidc block: password login only, so there's nothing external to
# register a redirect URI with before this demo can sign in.
passwordLogin: true
+45
View File
@@ -0,0 +1,45 @@
# Two teams (this one and 03-team-payments.yaml) so the demo shows
# per-team isolation -- separate incident lists, separate escalation
# ladders, separate alert sources -- rather than one team standing in for
# everything. A team carries its own escalation ladder and dead man's
# switches; alert sources are separate objects (04-alertsource-*.yaml)
# because each owns a webhook Secret.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
name: terdutteam-platform
spec:
# serverRef.namespace omitted: both this and terdut-operator-demo (01-server.yaml)
# live in whatever namespace you apply this directory into, which is the
# common case and needs no allowedTeams consent on the TerdutServer side
# (DESIGN.md §4.1, §4.6).
serverRef:
name: terdut-operator-demo
displayName: Platform
# No oidc block: this demo is password-login only (01-server.yaml).
# The whole ladder, replaced as one unit. username is required iff kind is
# "user" (CEL validation at apply time). The server resolves it, so a
# username it does not know yet leaves this team at Reason: UnknownUser
# until that person exists -- run-demo.sh creates alice for exactly that.
escalation:
repeatCount: 2
fallbackTopic: platform-fallback
levels:
- timeout: 5m
targets:
- kind: user
username: alice
- timeout: 10m
targets:
- kind: oncall
# Dead man's switches, by name. fire-alerts.sh's "heartbeat" scenario sends
# a matching alert; stop sending it and terdut-server itself opens an
# incident once `timeout` passes with no heartbeat. A switch on the server
# that is not listed here is removed.
deadmanSwitches:
- name: platform-watchdog
matcher: "alertname=PlatformWatchdog"
timeout: 15m
severity: critical
+20
View File
@@ -0,0 +1,20 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
name: terdutteam-payments
spec:
serverRef:
name: terdut-operator-demo
displayName: Payments
escalation:
repeatCount: 1
fallbackTopic: payments-fallback
levels:
- timeout: 5m
targets:
- kind: oncall
deadmanSwitches:
- name: payments-watchdog
matcher: "alertname=PaymentsWatchdog"
timeout: 15m
severity: critical
@@ -0,0 +1,18 @@
# The webhook URL/key fire-alerts.sh sends to, for the Platform team.
# terdut-server shows the key exactly once, at creation, and never again
# (DESIGN.md §4.5) -- this object's status.webhookURLSecretRef names the
# generated Secret holding it (keys "url" and "key"), which is what
# fire-alerts.sh reads. See README.md before applying this: it's the one
# object in this directory whose Secret you can't just re-read if you
# miss it.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutAlertSource
metadata:
name: terdutalertsource-platform
spec:
teamRef:
name: terdutteam-platform
# kind defaults to "alertmanager" -- the only value terdut-server
# supports today.
kind: alertmanager
name: platform-demo-alertmanager
@@ -0,0 +1,9 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutAlertSource
metadata:
name: terdutalertsource-payments
spec:
teamRef:
name: terdutteam-payments
kind: alertmanager
name: payments-demo-alertmanager
+159
View File
@@ -0,0 +1,159 @@
# Demo
Every CRD this operator reconciles, wired into one working install: one
`TerdutServer`, two `TerdutTeam`s (Platform and Payments), each carrying its
own escalation ladder and dead man's switches, and a `TerdutAlertSource` per
team, plus a script that fires synthetic Alertmanager webhooks at it so you can
watch real incidents appear, escalate and resolve.
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
Postgres with `emptyDir` storage and a password committed in this
directory. Throw the whole namespace away when you're done.
**Want this fully automated instead of walking through it by hand?**
`./run-demo.sh` does everything below itself, against a fresh (or
already-set-up) `kind` cluster — creates the cluster, installs the
operator, applies every CR here, creates `alice` as the first user, and fires a
few alerts. `./run-demo.sh --help` for the knobs, `./run-demo.sh
--teardown` to tear it back down. The rest of this file is the manual
walkthrough it automates.
## Prerequisites
- The operator and its CRDs installed and running (`make install
deploy IMG=...`, or `charts/terdut-operator` — see this repo's own
README.md/DESIGN.md), pointed at a cluster you're fine creating
throwaway resources in. A `kind` cluster is the easy choice.
- `kubectl`, `jq`, `curl` on your path.
**Apply this into a namespace of its own.** Every object name in this
directory is prefixed `terdut-operator-demo` specifically so applying it
by mistake into some other namespace that already has unrelated objects
doesn't collide with them -- but that only helps if this directory's own
objects don't collide with *each other* across two applies. Applying it
twice into two different namespaces is fine; applying it a second time
into a namespace that already has something else named `terdut-demo` (a
real install from following `terdut-operator`'s own repo along, say) is
exactly the mistake this prefix exists to avoid, and it only works if you
don't override these names yourself.
## Apply it
```sh
kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .
```
Listed and applied in dependency order (server → team → alert source), but
you don't have to preserve that order yourself: every controller here
re-queues and waits rather than failing when a ref isn't resolvable yet
(`kubectl describe` shows `Reason: WaitingForServer` / `WaitingForTeam` while
that settles). Platform stays at `Reason: UnknownUser` until `alice` exists:
its escalation ladder names her.
Watch it converge:
```sh
kubectl get terdutservers,terdutteams,terdutalertsources \
-n terdut-operator-demo
```
The operator's own Deployment template now carries a `wait-for-postgres`
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
so `terdut-operator-demo`'s pod should come up clean even against this brand-new
Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
Once `terdut-operator-demo`'s own `Ready` condition is `True`, everything downstream
of it should settle within a reconcile interval or two.
## See the web UI
The operator never creates any external exposure for a `TerdutServer` --
that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this
demo just reaches it the simplest way there is:
```sh
kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
```
and open http://localhost:8080. See `../networking` for worked examples of
exposing it for real (Gateway API, Istio, or a plain `Ingress`) instead.
### First login
The operator authenticates with a key of its own (the `TerdutServer`'s
`<name>-operator-key` Secret, handed to the server as `TERDUT_OPERATOR_KEY`) and
never creates a user. A person gets in the way anyone does on a fresh
terdut-server: `/api/bootstrap` creates the first user, an administrator,
while no user exists yet:
```sh
curl -sS -X POST http://localhost:8080/api/bootstrap -H 'Content-Type: application/json' \
-d '{"username":"alice","email":"alice@example.com","password":"a-long-demo-password"}'
```
The response carries an API key, shown once. An administrator can manage any
team, so add `alice` to both teams with it (`POST /api/teams/{teamID}/members`;
`kubectl get terdutteam -o jsonpath='{.status.teamID}'` gives the ids), or just
use the UI's team pages. `run-demo.sh` does exactly this for you.
## Fire some alerts
In another terminal, with the port-forward above still running:
```sh
export NAMESPACE=terdut-operator-demo
./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash
# Watch it open an incident, escalate per the escalation ladders in
# 02/03-team-*.yaml, and show up on the Platform/Payments team's own incident list.
./fire-alerts.sh platform high-cpu resolve
```
`fire-alerts.sh -h` (or any bad argument) prints the full scenario list.
Each `(team, scenario)` pair is one stable fingerprint, so firing the same
one twice updates the same alert (a real re-fire) and `resolve` closes
exactly that one.
Set `CLUSTER` to send the alert as if it came from one of several clusters:
```sh
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-us ./fire-alerts.sh platform high-cpu # a second incident, not a join
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
```
It stands in for a Prometheus external label plus `cluster` in Alertmanager's
`group_by` (terdut-server's README, "Several clusters, one team"): the web UI
then shows the cluster chip on each incident and a cluster filter in the
queue. `CLUSTER` is part of the fingerprint, so resolve with the same value you
fired with. `./run-demo.sh` fires its alerts across `prod-eu` and `prod-us`.
### Dead man's switches
The `deadmanSwitches` in `02-team-platform.yaml` / `03-team-payments.yaml` expect a heartbeat
alert on a 15-minute timeout:
```sh
./fire-alerts.sh platform heartbeat
```
Keep sending that (e.g. a `watch -n 60`) and nothing happens — that's the
point. Stop sending it and, 15 minutes after the last one, terdut-server
opens a `critical` incident on its own, with no webhook involved: proof
the switch is watching for silence, not for a signal.
## Tear down
```sh
kubectl delete namespace terdut-operator-demo
```
The operator's finalizers delete each team on the server (with its
escalation, switches and integrations) before this namespace's objects
actually disappear — give it a few seconds past the `kubectl delete`
returning. A team with open incidents is not deleted until they are resolved.
+150
View File
@@ -0,0 +1,150 @@
#!/usr/bin/env bash
# Sends a synthetic Alertmanager v4 webhook payload at the "platform" or
# "payments" demo team's TerdutAlertSource, so terdut-server opens (or
# resolves) an incident exactly the way it would for a real Alertmanager.
#
# The payload shape here is amPayload/amAlert, read straight out of
# terdut-server's own internal/api/alertmanager.go rather than guessed from
# its docs -- version/status/groupKey/groupLabels, and alerts[] carrying
# status/labels/annotations/startsAt/endsAt/generatorURL/fingerprint.
#
# Why this reads the webhook key out of a kubectl Secret instead of using
# the "url" key already in it: that URL is built from spec.networking.hostname
# (TERDUT_PUBLIC_URL), and nothing in this demo stands up real ingress for
# it (01-server.yaml's own comment) -- so it resolves nowhere. The key
# alone, against whatever you've actually port-forwarded BASE_URL to below,
# is the one part of that URL still usable here.
#
# Usage:
# [CLUSTER=prod-eu] ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
#
# CLUSTER stands in for a Prometheus externalLabel plus `cluster` in
# Alertmanager's group_by (terdut-server's README, "Several clusters, one
# team"): it is put on the alert's labels and on groupLabels, so the incident
# carries it and the web UI shows the cluster chip and the queue's cluster
# filter. It is also part of the group key and the fingerprint, which is what
# keeps the same alert in two clusters from joining one incident. Unset, the
# alert is sent exactly as before.
#
# Prerequisites: kubectl context pointed at the demo namespace, jq, curl,
# and (in another terminal) a running:
# kubectl port-forward svc/terdut-operator-demo 8080:8080
set -euo pipefail
NAMESPACE="${NAMESPACE:-}"
CLUSTER="${CLUSTER:-}"
BASE_URL="${BASE_URL:-http://localhost:8080}"
usage() {
cat >&2 <<'EOF'
usage: fire-alerts.sh <platform|payments> <scenario> [resolve]
scenarios:
high-cpu warning -- CPU usage above 90% for 10 minutes
disk-full critical -- disk usage above 95%
pod-crash error -- a pod crash-looping
heartbeat critical -- the team's dead man's switch heartbeat
(matches the matcher in the deadmanSwitches in 02/03-team-*.yaml -- send this
repeatedly to keep the switch alive, or stop sending it and
watch terdut-server open an incident on its own once
`timeout` passes with no heartbeat. "resolve" is not a valid
third argument for this scenario: a heartbeat is only ever
firing.)
env vars:
NAMESPACE kubectl -n for reading the webhook Secret (required)
BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080)
CLUSTER optional cluster name, e.g. prod-eu: sent as a `cluster` label and
group label, so the UI shows where the incident came from
EOF
exit 1
}
[ $# -ge 2 ] || usage
team="$1" scenario="$2" verb="${3:-fire}"
[ -n "$NAMESPACE" ] || { echo "fire-alerts.sh: set NAMESPACE" >&2; exit 1; }
case "$team" in
platform|payments) ;;
*) usage ;;
esac
case "$scenario" in
high-cpu) alertname=TerdutDemoHighCPU severity=warning summary="CPU usage above 90% for 10 minutes" ;;
disk-full) alertname=TerdutDemoDiskFull severity=critical summary="Disk usage above 95% on /data" ;;
pod-crash) alertname=TerdutDemoPodCrashLooping severity=error summary="Pod web-7f8b9 is crash-looping (5 restarts in 10m)" ;;
heartbeat)
# Must match 02-team-platform.yaml / 03-team-payments.yaml's own
# matcher exactly -- that's what makes this a heartbeat rather than a
# third ordinary alert.
case "$team" in
platform) alertname=PlatformWatchdog ;;
payments) alertname=PaymentsWatchdog ;;
esac
severity=critical summary="demo heartbeat"
[ "$verb" = fire ] || { echo "fire-alerts.sh: heartbeat is only ever fired, never resolved -- just stop sending it" >&2; exit 1; }
;;
*) usage ;;
esac
case "$verb" in
fire) status=firing ;;
resolve) status=resolved ;;
*) usage ;;
esac
secret_name="terdutalertsource-${team}-terdut-webhook"
key="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.key}' | base64 -d)"
[ -n "$key" ] || { echo "fire-alerts.sh: empty key read from Secret $secret_name -- has 04/05-alertsource-*.yaml reconciled yet?" >&2; exit 1; }
# Stable per (team, scenario) so a resolve targets the same alert a fire
# created: terdut-server correlates on (team_id, fingerprint), not on
# anything else in the payload. Real Alertmanager computes this from the
# alert's label set; a fixed string plays the same role here.
fingerprint="demo-${team}-${scenario}${CLUSTER:+-$CLUSTER}"
now="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
if [ "$status" = firing ]; then
ends_at="0001-01-01T00:00:00Z" # Alertmanager's own "not resolved" zero value
else
ends_at="$now"
fi
payload="$(jq -n \
--arg status "$status" \
--arg groupKey "demo:${team}:${scenario}${CLUSTER:+:$CLUSTER}" \
--arg cluster "$CLUSTER" \
--arg alertname "$alertname" \
--arg team "$team" \
--arg severity "$severity" \
--arg summary "$summary" \
--arg startsAt "$now" \
--arg endsAt "$ends_at" \
--arg fingerprint "$fingerprint" \
'{
version: "4",
status: $status,
groupKey: $groupKey,
groupLabels: ({ alertname: $alertname, team: $team } + (if $cluster != "" then { cluster: $cluster } else {} end)),
alerts: [{
status: $status,
labels: ({ alertname: $alertname, severity: $severity, team: $team, instance: "demo" } + (if $cluster != "" then { cluster: $cluster } else {} end)),
annotations: { summary: $summary },
startsAt: $startsAt,
endsAt: $endsAt,
generatorURL: "https://example.com/demo",
fingerprint: $fingerprint
}]
}')"
url="${BASE_URL}/api/integrations/${key}/alertmanager"
echo "POST $url (team=$team scenario=$scenario status=$status${CLUSTER:+ cluster=$CLUSTER})" >&2
code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \
-X POST "$url" -H 'Content-Type: application/json' -d "$payload")"
echo "-> HTTP $code" >&2
cat /tmp/fire-alerts-response.json >&2
echo >&2
if [ "$code" != "200" ]; then
exit 1
fi
+13
View File
@@ -0,0 +1,13 @@
## kubectl apply -n <your-demo-namespace> -k examples/demo
##
## Listed in apply order even though kustomize itself doesn't need that --
## a human reading this file top-to-bottom should see the same dependency
## order the controllers themselves require (server before team, team
## before everything that teamRefs it).
resources:
- 00-postgres.yaml
- 01-server.yaml
- 02-team-platform.yaml
- 03-team-payments.yaml
- 04-alertsource-platform.yaml
- 05-alertsource-payments.yaml
+317
View File
@@ -0,0 +1,317 @@
#!/usr/bin/env bash
# Stands up the complete terdut demo (terdut-operator + every CRD kind this
# repo ships + a working local login + a few synthetic incidents) on a kind
# cluster, fully automated. Password login only -- this demo kit's own
# 01-server.yaml carries no oidc: block at all, so there is nothing to
# disable; OIDC is simply absent.
#
# Usage:
# ./run-demo.sh deploy the whole demo (idempotent: safe to
# re-run against a cluster that already has it)
# ./run-demo.sh --teardown delete the demo namespace (and, if
# TEARDOWN_CLUSTER=true, the kind cluster too)
# ./run-demo.sh --help
#
# All of the defaults below are overridable as environment variables.
set -euo pipefail
CLUSTER_NAME="${CLUSTER_NAME:-terdut-demo}"
NAMESPACE="${NAMESPACE:-terdut-operator-demo}"
OPERATOR_NAMESPACE="${OPERATOR_NAMESPACE:-terdut-operator-system}"
RELEASE_NAME="${RELEASE_NAME:-terdut-operator}"
ALICE_USERNAME="${ALICE_USERNAME:-alice}"
ALICE_EMAIL="${ALICE_EMAIL:-alice@terdut-demo.local}"
DEMO_PASSWORD="${DEMO_PASSWORD:-terdut-demo-1234}"
BASE_URL="${BASE_URL:-http://localhost:8080}"
LOCAL_PORT="${LOCAL_PORT:-8080}"
HELM_TIMEOUT="${HELM_TIMEOUT:-180s}"
WAIT_TIMEOUT="${WAIT_TIMEOUT:-180s}"
TEARDOWN_CLUSTER="${TEARDOWN_CLUSTER:-false}"
PF_PIDFILE="${PF_PIDFILE:-/tmp/terdut-demo-port-forward.pid}"
PF_LOGFILE="${PF_LOGFILE:-/tmp/terdut-demo-port-forward.log}"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
CHART_DIR="$(cd "$SCRIPT_DIR/../.." && pwd)/charts/terdut-operator"
# ---------------------------------------------------------------------------
# helpers
# ---------------------------------------------------------------------------
log() { printf '[run-demo] %s\n' "$*" >&2; }
die() { printf '[run-demo] FAILED: %s\n' "$*" >&2; exit 1; }
usage() {
cat <<EOF
Usage: $0 [--teardown|--help]
Deploys (or tears down) the complete terdut demo on a kind cluster.
See the top of this file for every overridable environment variable.
EOF
}
# Only kills a port-forward THIS invocation started, so a successful run
# never has its background job reaped by its own exit trap.
STARTED_PF_PID=""
cleanup_on_failure() {
local rc=$?
if [ "$rc" -ne 0 ] && [ -n "$STARTED_PF_PID" ]; then
log "run failed -- stopping the port-forward it started (pid $STARTED_PF_PID)"
kill "$STARTED_PF_PID" 2>/dev/null || true
fi
exit "$rc"
}
trap cleanup_on_failure EXIT
# ---------------------------------------------------------------------------
# steps
# ---------------------------------------------------------------------------
preflight() {
local missing=()
for bin in kubectl kind helm jq curl; do
command -v "$bin" >/dev/null 2>&1 || missing+=("$bin")
done
if [ "${#missing[@]}" -gt 0 ]; then
die "missing required tools on PATH: ${missing[*]}"
fi
}
ensure_kind_cluster() {
case "$(kind get clusters 2>/dev/null)" in
*"$CLUSTER_NAME"*)
log "kind cluster '$CLUSTER_NAME' already exists, skipping creation" ;;
*)
log "creating kind cluster '$CLUSTER_NAME'"
kind create cluster --name "$CLUSTER_NAME" ;;
esac
kubectl config use-context "kind-${CLUSTER_NAME}" >/dev/null
}
install_operator() {
log "installing terdut-operator into namespace $OPERATOR_NAMESPACE"
helm upgrade --install "$RELEASE_NAME" "$CHART_DIR" \
--namespace "$OPERATOR_NAMESPACE" --create-namespace \
--wait --timeout "$HELM_TIMEOUT" \
|| die "helm install of terdut-operator did not become ready"
}
apply_demo() {
log "creating namespace $NAMESPACE"
kubectl create namespace "$NAMESPACE" --dry-run=client -o yaml | kubectl apply -f - >/dev/null
log "applying demo CRs (00-05) into $NAMESPACE"
kubectl apply -n "$NAMESPACE" -k "$SCRIPT_DIR" >/dev/null
}
wait_for_objects() {
local obj
for obj in "$@"; do
log "waiting for $obj to become Ready"
kubectl wait --for=condition=Ready --timeout "$WAIT_TIMEOUT" -n "$NAMESPACE" "$obj" >/dev/null \
|| die "timed out waiting for $obj -- try: kubectl describe -n $NAMESPACE $obj"
done
}
# Only the server: the teams cannot reach Ready until alice exists (their
# escalation ladders name her, and terdut-server resolves usernames when the
# ladder is applied), and alice is created below, through the server itself.
wait_for_server_ready() {
wait_for_objects "terdutserver/terdut-operator-demo"
}
# Everything that was waiting on alice to exist.
wait_for_remaining_ready() {
wait_for_objects \
"terdutteam/terdutteam-platform" \
"terdutteam/terdutteam-payments" \
"terdutalertsource/terdutalertsource-platform" \
"terdutalertsource/terdutalertsource-payments"
}
start_port_forward() {
# A stale pidfile from an earlier run would otherwise collide with us on
# $LOCAL_PORT -- if that pid is still alive, stop it first.
if [ -f "$PF_PIDFILE" ]; then
local old_pid
old_pid="$(cat "$PF_PIDFILE" 2>/dev/null || true)"
if [ -n "$old_pid" ] && kill -0 "$old_pid" 2>/dev/null; then
log "stopping stale port-forward from a previous run (pid $old_pid)"
kill "$old_pid" 2>/dev/null || true
sleep 1
fi
rm -f "$PF_PIDFILE"
fi
log "starting port-forward svc/terdut-operator-demo ${LOCAL_PORT}:8080"
kubectl -n "$NAMESPACE" port-forward svc/terdut-operator-demo "${LOCAL_PORT}:8080" \
>"$PF_LOGFILE" 2>&1 &
STARTED_PF_PID=$!
echo "$STARTED_PF_PID" > "$PF_PIDFILE"
local tries=0
until curl -sf -o /dev/null "http://localhost:${LOCAL_PORT}/healthz"; do
tries=$((tries + 1))
if [ "$tries" -ge 30 ]; then
die "port-forward never became ready -- see $PF_LOGFILE"
fi
sleep 1
done
log "port-forward ready (pid $STARTED_PF_PID, log $PF_LOGFILE)"
}
# Creates alice as the install's first user and administrator through
# /api/bootstrap, which is open until a first user exists -- the operator
# authenticates with its own seeded key and never uses it, so this is how a
# person gets in. An administrator may manage any team, so alice's own key is
# enough to add her to both; the teams' own status.teamID is set as soon as
# the operator has created each team, even while they still wait on her.
create_alice_and_join_teams() {
# A re-run: she exists already (and was added to the teams the first time).
local login_code
login_code="$(curl -sS -o /dev/null -w '%{http_code}' \
-X POST "${BASE_URL}/api/login" -H 'Content-Type: application/json' \
-d "$(jq -n --arg u "$ALICE_USERNAME" --arg p "$DEMO_PASSWORD" '{username: $u, password: $p}')")"
if [ "$login_code" = "200" ]; then
log "account ${ALICE_USERNAME} already exists and can sign in, skipping creation (re-run detected)"
return 0
fi
log "creating ${ALICE_USERNAME} as the first user (administrator)"
local resp_file code alice_key
resp_file="$(mktemp)"
code="$(curl -sS -o "$resp_file" -w '%{http_code}' \
-X POST "${BASE_URL}/api/bootstrap" -H 'Content-Type: application/json' \
-d "$(jq -n --arg u "$ALICE_USERNAME" --arg e "$ALICE_EMAIL" --arg p "$DEMO_PASSWORD" \
'{username: $u, email: $e, password: $p}')")"
[ "$code" = "201" ] || die "bootstrap failed (HTTP $code): $(cat "$resp_file")"
alice_key="$(jq -r '.api_key.key' "$resp_file")"
rm -f "$resp_file"
local team team_id
for team in terdutteam-platform terdutteam-payments; do
team_id=""
local tries=0
until team_id="$(kubectl -n "$NAMESPACE" get terdutteam "$team" -o jsonpath='{.status.teamID}' 2>/dev/null)" \
&& [ -n "$team_id" ]; do
tries=$((tries + 1))
[ "$tries" -lt 60 ] || die "$team never reported status.teamID -- kubectl describe it"
sleep 1
done
local alice_id
alice_id="$(curl -sS -H "Authorization: Bearer $alice_key" "${BASE_URL}/api/users" \
| jq -r --arg u "$ALICE_USERNAME" '.[] | select(.username == $u) | .id')"
code="$(curl -sS -o /dev/null -w '%{http_code}' \
-X POST "${BASE_URL}/api/teams/${team_id}/members" \
-H "Authorization: Bearer $alice_key" -H 'Content-Type: application/json' \
-d "$(jq -n --argjson id "$alice_id" '{user_id: $id, role: "owner"}')")"
[ "$code" = "204" ] || die "adding ${ALICE_USERNAME} to ${team} failed (HTTP $code)"
log "added ${ALICE_USERNAME} to ${team}"
done
}
fire_demo_alerts() {
log "firing representative demo alerts"
# Two clusters, so the queue shows the cluster chip and offers its filter.
# high-cpu fires in both: the same alert in two clusters is two incidents.
local fire="$SCRIPT_DIR/fire-alerts.sh"
CLUSTER=prod-eu NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform high-cpu
CLUSTER=prod-us NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform high-cpu
CLUSTER=prod-us NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform disk-full
CLUSTER=prod-eu NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" payments pod-crash
}
print_summary() {
cat <<EOF
terdut demo is up.
Web UI: http://localhost:${LOCAL_PORT}
Login: ${ALICE_USERNAME} / ${DEMO_PASSWORD}
Cluster: kind-${CLUSTER_NAME}
Namespace: ${NAMESPACE}
Port-forward pid: ${STARTED_PF_PID:-$(cat "$PF_PIDFILE" 2>/dev/null || echo unknown)} (log: ${PF_LOGFILE})
stop it with: kill \$(cat ${PF_PIDFILE})
Fire more alerts:
export NAMESPACE=${NAMESPACE} BASE_URL=${BASE_URL}
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
./fire-alerts.sh payments heartbeat # send repeatedly (e.g. every minute)
# to keep a dead man's switch alive;
# stop sending it and, 15 minutes
# later, terdut-server opens a
# critical incident on its own.
Tear down:
$0 --teardown
# add TEARDOWN_CLUSTER=true to also delete the kind cluster itself
EOF
}
teardown() {
if [ -f "$PF_PIDFILE" ]; then
local pid
pid="$(cat "$PF_PIDFILE" 2>/dev/null || true)"
if [ -n "$pid" ] && kill -0 "$pid" 2>/dev/null; then
log "stopping port-forward (pid $pid)"
kill "$pid" 2>/dev/null || true
fi
rm -f "$PF_PIDFILE"
fi
log "deleting namespace $NAMESPACE"
kubectl delete namespace "$NAMESPACE" --ignore-not-found --wait=true --timeout "$WAIT_TIMEOUT"
if [ "$TEARDOWN_CLUSTER" = "true" ]; then
log "deleting kind cluster $CLUSTER_NAME"
kind delete cluster --name "$CLUSTER_NAME"
else
log "leaving kind cluster '$CLUSTER_NAME' and the operator install in place" \
"(set TEARDOWN_CLUSTER=true to also delete the cluster)"
fi
}
# ---------------------------------------------------------------------------
# main
# ---------------------------------------------------------------------------
main() {
case "${1:-}" in
--teardown)
preflight
teardown
trap - EXIT
exit 0
;;
--help|-h)
usage
trap - EXIT
exit 0
;;
"") ;;
*)
usage
die "unknown argument: $1"
;;
esac
preflight
ensure_kind_cluster
install_operator
apply_demo
wait_for_server_ready
start_port_forward
create_alice_and_join_teams
wait_for_remaining_ready
fire_demo_alerts
print_summary
# Success: leave the port-forward running, don't let the EXIT trap kill it.
trap - EXIT
}
main "$@"
+31
View File
@@ -0,0 +1,31 @@
# Exposing a TerdutServer
This operator never manages external exposure/ingress for `TerdutServer`,
in any form — a permanent non-goal (`DESIGN.md` §1), not a missing
feature. Some installs won't expose it outside the cluster at all (see
`examples/demo`, which just port-forwards); others will put it behind
whatever their cluster already uses. That choice is entirely yours, not
the operator's.
The only contract the operator gives you to build on: a plain `ClusterIP`
Service, named after the `TerdutServer` CR (same name, same namespace),
with a port named `http` (`spec.networking.servicePort`, default `8080`).
Everything here targets exactly that Service — none of it is applied by
`examples/demo`'s `kustomization.yaml`, and none of it depends on anything
the operator creates beyond that one Service.
Pick whichever matches your cluster:
- **`httproute.yaml`** — a [Gateway API](https://gateway-api.sigs.k8s.io/)
`HTTPRoute`, attached to a `Gateway` you already have.
- **`istio-virtualservice.yaml`** — an Istio `VirtualService`, attached to
a `Gateway` (Istio's own CRD, not Gateway API's) you already have.
- **A plain `Ingress`** needs no example here — it's the same idea with
one fewer layer of indirection: an `Ingress` with a single rule whose
`backend.service.name`/`port.name` point at the `TerdutServer`'s Service
and `http` port.
Remember to set `spec.networking.hostname` on the `TerdutServer` itself to
whatever hostname you expose it on — that's not read by the operator for
any of this, but terdut-server uses it for its own absolute links
(notifications, OIDC redirect URIs).
+21
View File
@@ -0,0 +1,21 @@
# Gateway API HTTPRoute exposing a TerdutServer through a Gateway you
# already have (not something this operator creates or watches -- see
# ../networking/README.md). Replace terdut-operator-demo and the Gateway
# reference/hostname with your own; terdut-operator-demo matches
# ../demo/01-server.yaml, if you're layering this onto that demo.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: terdut-operator-demo
namespace: terdut-operator-demo
spec:
parentRefs:
- name: my-gateway # an existing Gateway in this namespace (or
# namespace: gateway-ns # a different one, if the Gateway allows it)
sectionName: https # the listener to attach to, if it's picky
hostnames:
- terdut-operator-demo.example # matches networking.hostname on the CR
rules:
- backendRefs:
- name: terdut-operator-demo # the Service the operator created --
port: 8080 # same name as the TerdutServer CR
@@ -0,0 +1,23 @@
# Istio VirtualService exposing a TerdutServer through an Istio Gateway
# you already have (Istio's own Gateway CRD, not Gateway API's -- not
# something this operator creates or watches, see ../networking/README.md).
# Replace terdut-operator-demo and the gateway reference/hostname with your
# own; terdut-operator-demo matches ../demo/01-server.yaml, if you're
# layering this onto that demo.
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: terdut-operator-demo
namespace: terdut-operator-demo
spec:
hosts:
- terdut-operator-demo.example # matches networking.hostname on the CR
gateways:
- my-gateway-namespace/my-gateway # an existing istio Gateway
http:
- route:
- destination:
host: terdut-operator-demo.terdut-operator-demo.svc.cluster.local
port:
number: 8080 # the Service's port -- same name as the
# TerdutServer CR, default servicePort
+17 -17
View File
@@ -25,7 +25,7 @@ require (
github.com/felixge/httpsnoop v1.0.4 // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/fxamacker/cbor/v2 v2.9.1 // indirect
github.com/go-logr/logr v1.4.3 // indirect
github.com/go-logr/logr v1.4.4 // indirect
github.com/go-logr/stdr v1.2.2 // indirect
github.com/go-logr/zapr v1.3.0 // indirect
github.com/go-openapi/jsonpointer v1.0.0 // indirect
@@ -64,30 +64,30 @@ require (
github.com/x448/float16 v0.8.4 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect
go.opentelemetry.io/otel v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 // indirect
go.opentelemetry.io/otel/metric v1.44.0 // indirect
go.opentelemetry.io/otel/sdk v1.44.0 // indirect
go.opentelemetry.io/otel/trace v1.44.0 // indirect
go.opentelemetry.io/proto/otlp v1.10.0 // indirect
go.opentelemetry.io/otel v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 // indirect
go.opentelemetry.io/otel/metric v1.45.0 // indirect
go.opentelemetry.io/otel/sdk v1.45.0 // indirect
go.opentelemetry.io/otel/trace v1.45.0 // indirect
go.opentelemetry.io/proto/otlp v1.11.0 // indirect
go.uber.org/multierr v1.11.0 // indirect
go.uber.org/zap v1.27.1 // indirect
go.yaml.in/yaml/v2 v2.4.4 // indirect
go.yaml.in/yaml/v3 v3.0.4 // indirect
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f // indirect
golang.org/x/mod v0.38.0 // indirect
golang.org/x/net v0.58.0 // indirect
golang.org/x/mod v0.41.0 // indirect
golang.org/x/net v0.60.0 // indirect
golang.org/x/oauth2 v0.36.0 // indirect
golang.org/x/sync v0.22.0 // indirect
golang.org/x/sys v0.47.0 // indirect
golang.org/x/term v0.45.0 // indirect
golang.org/x/text v0.41.0 // indirect
golang.org/x/sync v0.23.0 // indirect
golang.org/x/sys v0.48.0 // indirect
golang.org/x/term v0.46.0 // indirect
golang.org/x/text v0.42.0 // indirect
golang.org/x/time v0.15.0 // indirect
golang.org/x/tools v0.48.0 // indirect
golang.org/x/tools v0.49.0 // indirect
gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260803160001-6ac0973c030d // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260803160001-6ac0973c030d // indirect
google.golang.org/grpc v1.83.2 // indirect
google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af // indirect
gopkg.in/evanphx/json-patch.v4 v4.13.0 // indirect
+36 -36
View File
@@ -36,8 +36,8 @@ github.com/gkampitakis/go-diff v1.3.2/go.mod h1:LLgOrpqleQe26cte8s36HTWcTmMEur6O
github.com/gkampitakis/go-snaps v0.5.15 h1:amyJrvM1D33cPHwVrjo9jQxX8g/7E2wYdZ+01KS3zGE=
github.com/gkampitakis/go-snaps v0.5.15/go.mod h1:HNpx/9GoKisdhw9AFOBT1N7DBs9DiHo/hGheFGBZ+mc=
github.com/go-logr/logr v1.2.2/go.mod h1:jdQByPbusPIv2/zmleS9BjJVeZ6kBagPoEUsqbVz/1A=
github.com/go-logr/logr v1.4.3 h1:CjnDlHq8ikf6E492q6eKboGOC0T8CDaOvkHCIg8idEI=
github.com/go-logr/logr v1.4.3/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY=
github.com/go-logr/logr v1.4.4 h1:tG4xh9yMsRCAiodLVTxyrkzSZ9+o0L1Kg/+cPVcbP/8=
github.com/go-logr/logr v1.4.4/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY=
github.com/go-logr/stdr v1.2.2 h1:hSWxHoqTgW2S2qGc0LTAI563KZ5YKYRhT3MFKZMbjag=
github.com/go-logr/stdr v1.2.2/go.mod h1:mMo/vtBO5dYbehREoey6XUKy/eSumjCCveDpRre4VKE=
github.com/go-logr/zapr v1.3.0 h1:XGdV8XW8zdwFiwOA2Dryh1gj2KRQyOOoNmBy4EplIcQ=
@@ -168,22 +168,22 @@ go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ
go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI=
go.opentelemetry.io/otel v1.44.0 h1:JjwHmHpA4iZ3wBxluu2fbbE7j4kqlE8jXyAyPXH7HqU=
go.opentelemetry.io/otel v1.44.0/go.mod h1:BMgjTHL9WPRlRjL2oZCBTL4whCGtXch2H4BhOPIAyYc=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 h1:4YsVu3B8+3qtWYYrsUYgn0OG78pN0rnNPRGX4SbokQI=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0/go.mod h1:+wnlSn0mD1ADVMe3v9Z/WIaiz6q6gL2J/ejaAmdmv80=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 h1:qazEJlUOQzhCpzQpFETGby7EdqjI1wsd0W+6Gg1SCTU=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0/go.mod h1:fOD2Yefuxixkx3ahVNf0O/PERb6r4OlbxfATVnYvzCo=
go.opentelemetry.io/otel/metric v1.44.0 h1:1w0gILTcHdr3YI+ixLyjemwrVnsMURbTZFrSYCdDdmc=
go.opentelemetry.io/otel/metric v1.44.0/go.mod h1:8O7hanEPBNgEMmybD3s2VBKcgWOCsA6tzHBPODAiquo=
go.opentelemetry.io/otel/sdk v1.44.0 h1:nHYwb9lK+fJPU/dnT6s7W7Z8itMWyqrnVfbheVYrZ58=
go.opentelemetry.io/otel/sdk v1.44.0/go.mod h1:Osuydd3Se74nqjAKxid74N5eC+jfEqfTegHRnq58oK0=
go.opentelemetry.io/otel/sdk/metric v1.44.0 h1:3LlKgI+VjbVsjNRFZJZAJ30WjXC5VkNRks6si09iEfI=
go.opentelemetry.io/otel/sdk/metric v1.44.0/go.mod h1:5B5pMARnXxKhltooO4xUuCBorl65a4EpnTalObqOigA=
go.opentelemetry.io/otel/trace v1.44.0 h1:jxF5CsGYCe74MCRx2X4g7WsY/VBKRqqpNvXlX/6gtIk=
go.opentelemetry.io/otel/trace v1.44.0/go.mod h1:oLl1jrMQAVo6v3GAggN+1VH9VIz9iUSvW53sW1Q8PIE=
go.opentelemetry.io/proto/otlp v1.10.0 h1:IQRWgT5srOCYfiWnpqUYz9CVmbO8bFmKcwYxpuCSL2g=
go.opentelemetry.io/proto/otlp v1.10.0/go.mod h1:/CV4QoCR/S9yaPj8utp3lvQPoqMtxXdzn7ozvvozVqk=
go.opentelemetry.io/otel v1.45.0 h1:pdrWmLHofpubmArBv1LgFSv1Z0Ie/ppdZzu+kUN5EeU=
go.opentelemetry.io/otel v1.45.0/go.mod h1:XZxIqPapzEYnhNSScF5DIqXhm/rYi0FzCe2XddAwZfQ=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 h1:QRefszxJmfPdjXUUm3j6iDzY03mTPXMjqErFqQ67vUg=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0/go.mod h1:Tiz03lTBVBrm7eWZBOidzEaYaJa8tjwGUGv6d8mlTyk=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 h1:fG5MCxGz8+2VtrN/WgqSpJFctVz24gpxj8CxkKmc8Ww=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0/go.mod h1:BmAYTn+3ysbRe+IU2msxmf5Rx3g6DHvex+tWI3LdhYI=
go.opentelemetry.io/otel/metric v1.45.0 h1:7Eg1uH7CJ5cXv9is6tnBe1FI6rj1nwUdbFypRm3br/M=
go.opentelemetry.io/otel/metric v1.45.0/go.mod h1:HAPbm1nd3p1PmFH7v2dR+6BjXxw+Lq4a2+pndMAm08s=
go.opentelemetry.io/otel/sdk v1.45.0 h1:4VVSMgQ83dUgW2aoX5f6JgLvHwIvzcuLnF9lUdCSpCw=
go.opentelemetry.io/otel/sdk v1.45.0/go.mod h1:Sr40LgXV7DsKMMJMKOhUWOgMWTfAaqvm2kF0g7ilwuA=
go.opentelemetry.io/otel/sdk/metric v1.45.0 h1:oVFszMfyj1Am6s24Vtc7wBb8BKLcwepJjNEYILuiE3o=
go.opentelemetry.io/otel/sdk/metric v1.45.0/go.mod h1:vUWUxDZvu1WVRj8JA8S0AdhsPrZoDpA2DdZauIh4mDA=
go.opentelemetry.io/otel/trace v1.45.0 h1:l/mP6Uv7oNO7/TblbhpbgMidxhq1uO/rPsikOyVhxag=
go.opentelemetry.io/otel/trace v1.45.0/go.mod h1:qoJJA2xNMnxRrdISU/kLtfUH2wNeQbiv+jhs/CxI8bc=
go.opentelemetry.io/proto/otlp v1.11.0 h1:5rrYs0Ykyj50sdU/JU0x8etU+LubXWb+gED6TbEdMIk=
go.opentelemetry.io/proto/otlp v1.11.0/go.mod h1:SmVizdCOAm3XBtG1g1NnOdhW6jtddT72hLMhv8VwA8E=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
go.uber.org/multierr v1.11.0 h1:blXXJkSxSSfBVBlC76pxqeO+LN3aDfLQo+309xJstO0=
@@ -196,32 +196,32 @@ go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc=
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f h1:W3F4c+6OLc6H2lb//N1q4WpJkhzJCK5J6kUi1NTVXfM=
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f/go.mod h1:J1xhfL/vlindoeF/aINzNzt2Bket5bjo9sdOYzOsU80=
golang.org/x/mod v0.38.0 h1:MECBjubtXD7yj4HrhIUcywNaGeNVUdfVnxmPajOk4yk=
golang.org/x/mod v0.38.0/go.mod h1:V6Xz0pq8TQ3dGqVQ1FVHuelZpAL0uNhSkk9ogYP3c40=
golang.org/x/net v0.58.0 h1:ynWG7rqYi4ccpTEuPZ2QGWHktVEM9DMCj9yzDE0Q7To=
golang.org/x/net v0.58.0/go.mod h1:YwCddHnFlT7eLQqVprV19OnhLGtc5xOKgE0RyqgfWAU=
golang.org/x/mod v0.41.0 h1:qJmnOUb4YB+FsEuM3HcWucdZASCPGhsX6uljO6pog0c=
golang.org/x/mod v0.41.0/go.mod h1:Ek9pY8RKWXwsWvd3rQiHYtMqkjSUV+s1Rj7j4H5Ur6o=
golang.org/x/net v0.60.0 h1:79p50tfZlm0J9YfoDsSi639qSXNGVwEzOPLCxM2FsYU=
golang.org/x/net v0.60.0/go.mod h1:2DA/G1UfVbCpQPeWTmMPGY7Cs2PkBkwu743bVX5PIVg=
golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs=
golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q=
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/term v0.45.0 h1:NwWyBmoJCbfTHpxrWoZ9C6/VxOf7ic219I8xZZFdrf0=
golang.org/x/term v0.45.0/go.mod h1:9aqxs0blBcrm/n0L9QW0aRVD+ktan8ssZromtqJC43w=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
golang.org/x/sync v0.23.0 h1:KameEIfc1IkluZyXWLn39Wd4tURc6GbCiISGiZm2bQk=
golang.org/x/sync v0.23.0/go.mod h1:sUUOizhqBxiL6pEWpqNLUiaJn1ShEbZ6BBqskPbjZm0=
golang.org/x/sys v0.48.0 h1:bbX/i/6MgT9BVLM9RT1thmxL04yeTAhbEz4SyadbXoo=
golang.org/x/sys v0.48.0/go.mod h1:hNLxWAXmnKAxqDtdwIYC4bM9oQPEecfsnNMuSxOs3og=
golang.org/x/term v0.46.0 h1:3+OXuTbaKDgwk8jTi3aSLHRlmWqHEUDUtxnbFigO4YE=
golang.org/x/term v0.46.0/go.mod h1:+K02xbkittuwc0Am4abfA3Fc+XRGXkvBXNO88NCXPoc=
golang.org/x/text v0.42.0 h1:JbOZXgfeCPU9gacVtYliJqOhD+zhrEqK4LfdpmlUZqI=
golang.org/x/text v0.42.0/go.mod h1:ojzP1Z+2QtioaF8DTtO8K5q7JWVVYwZKenzujK0Zd0E=
golang.org/x/time v0.15.0 h1:bbrp8t3bGUeFOx08pvsMYRTCVSMk89u4tKbNOZbp88U=
golang.org/x/time v0.15.0/go.mod h1:Y4YMaQmXwGQZoFaVFk4YpCt4FLQMYKZe9oeV/f4MSno=
golang.org/x/tools v0.48.0 h1:3+hClM1aLL5mjMKm5ovokw9epgRXPuu2tILgismM6RE=
golang.org/x/tools v0.48.0/go.mod h1:08xX0orndb/F7jJxGDicx061tyd5pcMto75YMAXr6lk=
golang.org/x/tools v0.49.0 h1:3NI7VXzL9+1WZD52Dx2ttoPwD5DWrFGpl9mFZDlmisI=
golang.org/x/tools v0.49.0/go.mod h1:SJNXV9DBKT0UbdttsQjbfJlAE/q+y36++zo3uL3N0Oo=
gomodules.xyz/jsonpatch/v2 v2.4.0 h1:Ci3iUJyx9UeRx7CeFN8ARgGbkESwJK+KB9lLcWxY/Zw=
gomodules.xyz/jsonpatch/v2 v2.4.0/go.mod h1:AH3dM2RI6uoBZxn3LVrfvJ3E0/9dG4cSrbuBJT4moAY=
gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4=
gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E=
google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa h1:Kjn0N0tCrDgiAFW+lGO4JZ3ck44CehvJQMAwj9QF0G8=
google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa/go.mod h1:q4lMZS6kskjT5HvCPrnnypcDPVJqT/f4nfxmkE7gryY=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa h1:mZHHdPZl0dbGHCflZgAq/Q468DWVFcU2whhB2KAo8fk=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8=
google.golang.org/genproto/googleapis/api v0.0.0-20260803160001-6ac0973c030d h1:FarXi840EJWSHYTN3ERkADbPWjl307+FGrA22KAVjjc=
google.golang.org/genproto/googleapis/api v0.0.0-20260803160001-6ac0973c030d/go.mod h1:K/+WGbmBY7aNW1HDw1fJnKYo10i0DkAX6pows00dLig=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260803160001-6ac0973c030d h1:IL4hdHzcUv2l/gcg98/Rj3FbtE6axwqslOW8SW0C+S0=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260803160001-6ac0973c030d/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8=
google.golang.org/grpc v1.83.2 h1:EManeRomTObA0BU7I8vXgg/78uE5MJ9M8B39EX2WscU=
google.golang.org/grpc v1.83.2/go.mod h1:YPI1hK3kDked6iHvgX3tR0y+nX/qpMFKhPgFsokw1S8=
google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af h1:+5/Sw3GsDNlEmu7TfklWKPdQ0Ykja5VEmq2i817+jbI=
+20 -10
View File
@@ -5,14 +5,15 @@ import (
"fmt"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta"
"sigs.k8s.io/controller-runtime/pkg/client"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// childError carries a condition reason/message, the same role teamError
// and databaseError play for their own controllers: an expected,
// childError carries a condition reason/message, the same role teamError and
// databaseError play for their own controllers: an expected,
// requeue-and-retry outcome, not a reconcile failure.
type childError struct {
reason string
@@ -21,12 +22,13 @@ type childError struct {
func (e *childError) Error() string { return e.message }
// resolveTeamAndClient implements DESIGN.md §5's "every child resolves its
// own teamRef -> TerdutTeam.status, never chains up to TerdutServer"
// rule -- shared by TerdutEscalationRule and TerdutDeadmanSwitch, which
// both need exactly this and nothing else to call terdut-server's API.
// resolveTeamAndClient finds the TerdutTeam a child (TerdutAlertSource) names,
// requires it to be Ready, and returns a client for its server authenticated
// with that server's operator key. The server is read from the team's own
// serverRef: consent was already checked when the team went Ready, and does not
// gate a resource the team owns.
func resolveTeamAndClient(
ctx context.Context, c client.Client, operatorNamespace, namespace string,
ctx context.Context, c client.Client, namespace string,
ref terdutv1alpha1.TerdutTeamRef, newClient func(string) *tdclient.Client,
) (*terdutv1alpha1.TerdutTeam, *tdclient.Client, *childError) {
var team terdutv1alpha1.TerdutTeam
@@ -39,16 +41,24 @@ func resolveTeamAndClient(
}
return nil, nil, &childError{reason: terdutv1alpha1.ReasonTeamRefNotFound, message: err.Error()}
}
if team.Status.CredentialsSecretRef == nil || team.Status.TeamID == 0 {
if team.Status.TeamID == 0 || !meta.IsStatusConditionTrue(team.Status.Conditions, terdutv1alpha1.ConditionReady) {
return nil, nil, &childError{
reason: terdutv1alpha1.ReasonWaitingForTeam,
message: fmt.Sprintf("TerdutTeam %q is not Ready yet", ref.Name),
}
}
teamKey, err := readOperatorSecret(ctx, c, operatorNamespace, team.Status.CredentialsSecretRef)
srvNS := team.Spec.ServerRef.Namespace
if srvNS == "" {
srvNS = team.Namespace
}
var srv terdutv1alpha1.TerdutServer
if err := c.Get(ctx, client.ObjectKey{Namespace: srvNS, Name: team.Spec.ServerRef.Name}, &srv); err != nil {
return nil, nil, &childError{reason: terdutv1alpha1.ReasonWaitingForTeam, message: err.Error()}
}
key, err := readOperatorKey(ctx, c, &srv)
if err != nil {
return nil, nil, &childError{reason: terdutv1alpha1.ReasonWaitingForTeam, message: err.Error()}
}
return &team, newClient(team.Status.ServerEndpoint).WithToken(teamKey), nil
return &team, newClient(serviceURL(&srv)).WithToken(key), nil
}
+375
View File
@@ -0,0 +1,375 @@
package controller
import (
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strconv"
"strings"
"sync"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// errJSONKey is the key every error body the fake writes uses, and
// deadmanSwitchesPath the literal shared by its dead man's switch routes --
// goconst would otherwise flag the repeats.
const (
errJSONKey = "error"
deadmanSwitchesPath = "/deadman/switches"
)
// fakeTerdutServer reproduces the parts of terdut-server's API the controllers
// call, with the same status codes and idempotency rules (POST /api/teams
// keyed by external_id, unique team names, unique switch names per team), so a
// controller test exercises the real contract and not a canned reply.
type fakeTerdutServer struct {
mu sync.Mutex
// expectKey, when set, makes every request without that bearer token a 401.
expectKey string
nextTeamID int64
teams map[string]int64 // name -> id
teamNames map[int64]string // id -> current name
teamExt map[string]int64 // external_id -> id
teamOIDC map[int64][2]string
teamDelete map[int64]bool // id -> true once DELETEd
// unknownUsers are usernames the escalation PUT answers 400 "unknown user" for.
unknownUsers map[string]bool
escalation map[int64]tdclient.SetEscalationRequest
nextSwitchID int64
switches map[int64]map[int64]tdclient.DeadmanSwitch // teamID -> switchID -> switch
switchDelete map[int64]bool
nextIntegrationID int64
integrations map[int64]map[int64]tdclient.Integration
integrationDelete map[int64]bool
}
func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) {
f := &fakeTerdutServer{
teams: map[string]int64{},
teamNames: map[int64]string{},
teamExt: map[string]int64{},
teamOIDC: map[int64][2]string{},
teamDelete: map[int64]bool{},
unknownUsers: map[string]bool{},
escalation: map[int64]tdclient.SetEscalationRequest{},
switches: map[int64]map[int64]tdclient.DeadmanSwitch{},
switchDelete: map[int64]bool{},
integrations: map[int64]map[int64]tdclient.Integration{},
integrationDelete: map[int64]bool{},
}
return f, httptest.NewServer(f)
}
func (f *fakeTerdutServer) ServeHTTP(w http.ResponseWriter, r *http.Request) {
f.mu.Lock()
defer f.mu.Unlock()
if f.expectKey != "" && r.Header.Get("Authorization") != "Bearer "+f.expectKey {
writeJSON(w, http.StatusUnauthorized, map[string]string{errJSONKey: "invalid or expired API key"})
return
}
switch {
case r.URL.Path == "/api/version":
writeJSON(w, http.StatusOK, map[string]string{"version": "test"})
case r.URL.Path == "/api/teams" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
ExternalID string `json:"external_id"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if id, ok := f.teamExt[req.ExternalID]; ok && req.ExternalID != "" {
writeJSON(w, http.StatusOK, tdclient.Team{ID: id, Name: f.teamNames[id]})
return
}
if _, exists := f.teams[req.Name]; exists {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a team with that name already exists"})
return
}
f.nextTeamID++
id := f.nextTeamID
f.teams[req.Name] = id
f.teamNames[id] = req.Name
if req.ExternalID != "" {
f.teamExt[req.ExternalID] = id
}
writeJSON(w, http.StatusCreated, tdclient.Team{ID: id, Name: req.Name})
default:
if id, rest, ok := parseTeamSubPath(r.URL.Path); ok {
f.handleTeamSubPath(w, r, id, rest)
return
}
w.WriteHeader(http.StatusNotFound)
}
}
// handleTeamSubPath answers everything under /api/teams/{id}. rest is whatever
// parseTeamSubPath found after "/api/teams/{id}" -- "" for the bare resource.
func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "" && r.Method == http.MethodPut:
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
oldName, exists := f.teamNames[id]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
if other, taken := f.teams[req.Name]; taken && other != id {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a team with that name already exists"})
return
}
delete(f.teams, oldName)
f.teamNames[id] = req.Name
f.teams[req.Name] = id
w.WriteHeader(http.StatusNoContent)
case rest == "" && r.Method == http.MethodDelete:
name, exists := f.teamNames[id]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.teams, name)
delete(f.teamNames, id)
f.teamDelete[id] = true
w.WriteHeader(http.StatusNoContent)
case rest == "/oidc-groups" && r.Method == http.MethodPut:
var req struct {
MemberGroup string `json:"member_group"`
OwnerGroup string `json:"owner_group"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if _, exists := f.teamNames[id]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
f.teamOIDC[id] = [2]string{req.MemberGroup, req.OwnerGroup}
w.WriteHeader(http.StatusNoContent)
case rest == "/escalation" && r.Method == http.MethodPut:
var req tdclient.SetEscalationRequest
_ = json.NewDecoder(r.Body).Decode(&req)
for _, l := range req.Levels {
for _, t := range l.Targets {
if t.Username != "" && f.unknownUsers[t.Username] {
writeJSON(w, http.StatusBadRequest, map[string]string{errJSONKey: fmt.Sprintf("unknown user %q", t.Username)})
return
}
}
}
f.escalation[id] = req
w.WriteHeader(http.StatusNoContent)
case rest == deadmanSwitchesPath || strings.HasPrefix(rest, deadmanSwitchesPath+"/"):
f.handleDeadmanSubPath(w, r, id, rest)
case rest == "/integrations" || strings.HasPrefix(rest, "/integrations/"):
f.handleIntegrationSubPath(w, r, id, rest)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleDeadmanSubPath answers GET/POST /api/teams/{id}/deadman/switches and
// PUT/DELETE .../deadman/switches/{switchID} -- split out of
// handleTeamSubPath for the same gocyclo reason as handleIntegrationSubPath.
func (f *fakeTerdutServer) handleDeadmanSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == deadmanSwitchesPath && r.Method == http.MethodGet:
existing := f.switches[id]
out := make([]tdclient.DeadmanSwitch, 0, len(existing))
for _, s := range existing {
out = append(out, s)
}
writeJSON(w, http.StatusOK, out)
case rest == deadmanSwitchesPath && r.Method == http.MethodPost:
var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req)
name := req.Name
for _, s := range f.switches[id] {
if s.Name == name {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a switch with that name already exists in this team"})
return
}
}
f.nextSwitchID++
switchID := f.nextSwitchID
sw := tdclient.DeadmanSwitch{
ID: switchID, Name: name, Matcher: req.Matcher,
TimeoutSeconds: req.TimeoutSeconds, Severity: req.Severity,
}
if f.switches[id] == nil {
f.switches[id] = map[int64]tdclient.DeadmanSwitch{}
}
f.switches[id][switchID] = sw
writeJSON(w, http.StatusCreated, sw)
case strings.HasPrefix(rest, "/deadman/switches/") && r.Method == http.MethodPut:
switchID, ok := parseTrailingID(rest, "/deadman/switches/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.switches[id][switchID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req)
name := req.Name
if name == "" {
name = f.switches[id][switchID].Name
}
f.switches[id][switchID] = tdclient.DeadmanSwitch{
ID: switchID, Name: name, Matcher: req.Matcher,
TimeoutSeconds: req.TimeoutSeconds, Severity: req.Severity,
}
w.WriteHeader(http.StatusNoContent)
case strings.HasPrefix(rest, "/deadman/switches/") && r.Method == http.MethodDelete:
switchID, ok := parseTrailingID(rest, "/deadman/switches/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.switches[id][switchID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.switches[id], switchID)
f.switchDelete[switchID] = true
w.WriteHeader(http.StatusNoContent)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleIntegrationSubPath answers POST /api/teams/{id}/integrations,
// PATCH .../integrations/{integrationID} and DELETE .../integrations/{integrationID}
// -- split out of handleTeamSubPath so that switch's own cyclomatic
// complexity stays under golangci-lint's gocyclo threshold.
func (f *fakeTerdutServer) handleIntegrationSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "/integrations" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
Kind string `json:"kind"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
f.nextIntegrationID++
integID := f.nextIntegrationID
integ := tdclient.Integration{
ID: integID, TeamID: id, Kind: req.Kind, Name: req.Name,
// Key/URL are only ever in *this* response -- never again,
// matching terdut-server's own one-time-show semantics
// (DESIGN.md §4.5) -- so what's stored for later GET/PATCH
// calls in this fake deliberately omits them too.
Key: fmt.Sprintf("webhook-key-%d", integID),
URL: fmt.Sprintf("https://terdut.example.invalid/api/integrations/webhook-key-%d/%s", integID, req.Kind),
}
if f.integrations[id] == nil {
f.integrations[id] = map[int64]tdclient.Integration{}
}
f.integrations[id][integID] = tdclient.Integration{ID: integID, TeamID: id, Kind: req.Kind, Name: req.Name}
writeJSON(w, http.StatusCreated, integ)
case strings.HasPrefix(rest, "/integrations/") && r.Method == http.MethodPatch:
integID, ok := parseTrailingID(rest, "/integrations/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
existing, exists := f.integrations[id][integID]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
existing.Name = req.Name
f.integrations[id][integID] = existing
w.WriteHeader(http.StatusNoContent)
case strings.HasPrefix(rest, "/integrations/") && r.Method == http.MethodDelete:
integID, ok := parseTrailingID(rest, "/integrations/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.integrations[id][integID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.integrations[id], integID)
f.integrationDelete[integID] = true
w.WriteHeader(http.StatusNoContent)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// deadmanSwitchFakeRequest mirrors tdclient's own (unexported)
// deadmanSwitchRequest -- the fake needs its own copy to decode the same
// wire shape without reaching across package boundaries for an internal type.
type deadmanSwitchFakeRequest struct {
Name string `json:"name,omitempty"`
Matcher string `json:"matcher"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Severity string `json:"severity"`
}
// parseTeamSubPath splits "/api/teams/{id}" from anything after it --
// "" for an exact match, "/oidc-groups", "/escalation", "/deadman/switches"
// or "/deadman/switches/{switchID}" otherwise. Doesn't itself validate the
// suffix; handleTeamSubPath's own switch does that.
func parseTeamSubPath(path string) (id int64, rest string, ok bool) {
const prefix = "/api/teams/"
if !strings.HasPrefix(path, prefix) {
return 0, "", false
}
trimmed := path[len(prefix):]
parts := strings.SplitN(trimmed, "/", 2)
parsedID, err := strconv.ParseInt(parts[0], 10, 64)
if err != nil {
return 0, "", false
}
if len(parts) == 1 {
return parsedID, "", true
}
return parsedID, "/" + parts[1], true
}
// parseTrailingID parses the numeric id after prefix within rest, e.g.
// parseTrailingID("/deadman/switches/7", "/deadman/switches/") -> 7, true.
func parseTrailingID(rest, prefix string) (id int64, ok bool) {
parsedID, err := strconv.ParseInt(strings.TrimPrefix(rest, prefix), 10, 64)
if err != nil {
return 0, false
}
return parsedID, true
}
func writeJSON(w http.ResponseWriter, status int, v any) {
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(status)
_ = json.NewEncoder(w).Encode(v)
}
-49
View File
@@ -1,49 +0,0 @@
package controller
import (
"context"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
)
// credentialsSecretDataKey is the fixed data key every generated
// credentials Secret this controller writes uses — "whatever a human chose"
// only applied to the bring-your-own input an earlier design draft had and
// removed (DESIGN.md §6); every Secret any controller in this repo
// generates itself uses this one key.
const credentialsSecretDataKey = "token"
// writeOperatorSecret creates or replaces a Secret in namespace (always the
// operator's own, DESIGN.md §6) holding one raw value under
// credentialsSecretDataKey. Shared by every controller that generates a
// credential there -- TerdutServer's instance-scoped key and the bootstrap
// admin-key checkpoint, TerdutTeam's team-scoped key.
func writeOperatorSecret(ctx context.Context, c client.Client, namespace, name, rawValue string) error {
secret := &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: namespace},
Data: map[string][]byte{credentialsSecretDataKey: []byte(rawValue)},
}
if err := c.Create(ctx, secret); err != nil {
if apierrors.IsAlreadyExists(err) {
return c.Update(ctx, secret)
}
return err
}
return nil
}
// readOperatorSecret reads one generated credential back, by the
// SecretKeyRef a controller's own status stores (always resolved in the
// operator's own namespace, never the referencing CR's -- DESIGN.md §6).
func readOperatorSecret(ctx context.Context, c client.Client, namespace string, ref *terdutv1alpha1.SecretKeyRef) (string, error) {
var secret corev1.Secret
if err := c.Get(ctx, client.ObjectKey{Namespace: namespace, Name: ref.Name}, &secret); err != nil {
return "", err
}
return string(secret.Data[ref.Key]), nil
}
@@ -2,7 +2,9 @@ package controller
import (
"context"
"errors"
"fmt"
"net/http"
"strconv"
"time"
@@ -39,9 +41,8 @@ type TerdutAlertSourceReconciler struct {
client.Client
Scheme *runtime.Scheme
OperatorNamespace string
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
}
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutalertsources,verbs=get;list;watch;create;update;patch;delete
@@ -79,7 +80,7 @@ func (r *TerdutAlertSourceReconciler) Reconcile(ctx context.Context, req ctrl.Re
return ctrl.Result{}, nil
}
team, tc, resolveErr := resolveTeamAndClient(ctx, r.Client, r.OperatorNamespace, as.Namespace, as.Spec.TeamRef, newClient)
team, tc, resolveErr := resolveTeamAndClient(ctx, r.Client, as.Namespace, as.Spec.TeamRef, newClient)
if resolveErr != nil {
return r.setNotReady(ctx, &as, resolveErr.reason, resolveErr.message, waitInterval)
}
@@ -111,7 +112,27 @@ func (r *TerdutAlertSourceReconciler) Reconcile(ctx context.Context, req ctrl.Re
return ctrl.Result{}, err
}
} else if err := tc.RenameIntegration(ctx, teamID, as.Status.IntegrationID, as.Spec.Name); err != nil {
return ctrl.Result{}, fmt.Errorf("PATCH /api/teams/%d/integrations/%d: %w", teamID, as.Status.IntegrationID, err)
se, ok := errors.AsType[*tdclient.StatusError](err)
if !ok || se.Code != http.StatusNotFound {
return ctrl.Result{}, fmt.Errorf("PATCH /api/teams/%d/integrations/%d: %w", teamID, as.Status.IntegrationID, err)
}
// Deleted on the server behind our back, so the webhook URL
// already stopped working: recreate it (a new URL, in the same
// Secret) rather than fail forever. reconcileCreate trusts the
// Secret's existence as "already created", so drop it first.
lostID := as.Status.IntegrationID
if err := r.Delete(ctx, &secret); err != nil && !apierrors.IsNotFound(err) {
return ctrl.Result{}, fmt.Errorf("deleting stale webhook Secret %s/%s: %w", as.Namespace, secretName, err)
}
as.Status.IntegrationID = 0
if err := r.reconcileCreate(ctx, &as, tc, teamID); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(&as, nil, corev1.EventTypeWarning, "IntegrationRecreated", "IntegrationRecreated",
"integration %d no longer exists on the server; created %d with a new webhook URL (Secret %s)",
lostID, as.Status.IntegrationID, secretName)
}
}
}
@@ -279,7 +300,7 @@ func (r *TerdutAlertSourceReconciler) reconcileDelete(
if as.Status.IntegrationID != 0 {
if team, tc, resolveErr := resolveTeamAndClient(
ctx, r.Client, r.OperatorNamespace, as.Namespace, as.Spec.TeamRef, newClient,
ctx, r.Client, as.Namespace, as.Spec.TeamRef, newClient,
); resolveErr == nil {
if err := tc.DeleteIntegration(ctx, team.Status.TeamID, as.Status.IntegrationID); err != nil {
if r.Recorder != nil {
@@ -34,14 +34,13 @@ var _ = Describe("TerdutAlertSource Controller", func() {
fake, fakeSrv = newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
srv = bootstrapReadyTerdutServer(ctx, uniqueName("asserver"), fakeSrv.URL)
srv = readyTerdutServer(ctx, uniqueName("asserver"))
team = readyTerdutTeam(ctx, operatorNamespace, uniqueName("asteam"), srv, fakeSrv.URL)
reconciler = &TerdutAlertSourceReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
Client: k8sClient,
Scheme: k8sClient.Scheme(),
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
}
srcName = uniqueName("alertsource")
srcKey = types.NamespacedName{Name: srcName, Namespace: operatorNamespace}
@@ -1,199 +0,0 @@
package controller
import (
"context"
"fmt"
"time"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
logf "sigs.k8s.io/controller-runtime/pkg/log"
"sigs.k8s.io/controller-runtime/pkg/recorder"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
const deadmanFinalizerName = "terdut.ryuvia.com/terdutdeadmanswitch"
// TerdutDeadmanSwitchReconciler reconciles a TerdutDeadmanSwitch object.
type TerdutDeadmanSwitchReconciler struct {
client.Client
Scheme *runtime.Scheme
OperatorNamespace string
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
}
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutdeadmanswitches,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutdeadmanswitches/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutdeadmanswitches/finalizers,verbs=update
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutteams,verbs=get;list;watch
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
func (r *TerdutDeadmanSwitchReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
log := logf.FromContext(ctx)
var sw terdutv1alpha1.TerdutDeadmanSwitch
if err := r.Get(ctx, req.NamespacedName, &sw); err != nil {
if apierrors.IsNotFound(err) {
return ctrl.Result{}, nil
}
return ctrl.Result{}, err
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
if !sw.DeletionTimestamp.IsZero() {
return r.reconcileDeadmanDelete(ctx, &sw, newClient)
}
if !controllerutil.ContainsFinalizer(&sw, deadmanFinalizerName) {
controllerutil.AddFinalizer(&sw, deadmanFinalizerName)
if err := r.Update(ctx, &sw); err != nil {
return ctrl.Result{}, err
}
return ctrl.Result{}, nil
}
team, tc, resolveErr := resolveTeamAndClient(ctx, r.Client, r.OperatorNamespace, sw.Namespace, sw.Spec.TeamRef, newClient)
if resolveErr != nil {
return r.setDeadmanNotReady(ctx, &sw, resolveErr.reason, resolveErr.message, waitInterval)
}
timeout, err := time.ParseDuration(sw.Spec.Timeout)
if err != nil {
return ctrl.Result{}, fmt.Errorf("spec.timeout %q: %w", sw.Spec.Timeout, err)
}
severity := sw.Spec.Severity
if severity == "" {
severity = "critical"
}
if sw.Status.SwitchID == 0 {
if err := r.createOrAdoptDeadmanSwitch(ctx, &sw, tc, team.Status.TeamID, timeout, severity); err != nil {
return ctrl.Result{}, err
}
} else if err := tc.UpdateDeadmanSwitch(ctx, team.Status.TeamID, sw.Status.SwitchID, sw.Spec.Name, sw.Spec.Matcher, int64(timeout.Seconds()), severity); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/deadman/switches/%d: %w", team.Status.TeamID, sw.Status.SwitchID, err)
}
meta.SetStatusCondition(&sw.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonChildAdopted,
Message: fmt.Sprintf("switch %d applied on team %d", sw.Status.SwitchID, team.Status.TeamID),
})
sw.Status.ObservedGeneration = sw.Generation
if err := r.Status().Update(ctx, &sw); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(&sw, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonChildAdopted, terdutv1alpha1.ReasonChildAdopted,
"dead man's switch applied")
}
log.Info("TerdutDeadmanSwitch applied", "name", sw.Name, "switchID", sw.Status.SwitchID)
return ctrl.Result{RequeueAfter: resyncInterval}, nil
}
// createOrAdoptDeadmanSwitch implements this resource's own idempotent-
// create shape (DESIGN.md §4.4, §5): there's no unique-name constraint
// server-side to 409 on, so this lists first and matches by name (the
// server's own derived name, when spec.name is empty) rather than adopting
// after a conflict the API would never actually raise.
func (r *TerdutDeadmanSwitchReconciler) createOrAdoptDeadmanSwitch(
ctx context.Context, sw *terdutv1alpha1.TerdutDeadmanSwitch, tc *tdclient.Client,
teamID int64, timeout time.Duration, severity string,
) error {
existing, err := tc.ListDeadmanSwitches(ctx, teamID)
if err != nil {
return fmt.Errorf("GET /api/teams/%d/deadman/switches: %w", teamID, err)
}
if sw.Spec.Name != "" {
for _, s := range existing {
if s.Name == sw.Spec.Name {
sw.Status.SwitchID = s.ID
return nil
}
}
}
created, err := tc.CreateDeadmanSwitch(ctx, teamID, sw.Spec.Name, sw.Spec.Matcher, int64(timeout.Seconds()), severity)
if err != nil {
return fmt.Errorf("POST /api/teams/%d/deadman/switches: %w", teamID, err)
}
sw.Status.SwitchID = created.ID
return nil
}
func (r *TerdutDeadmanSwitchReconciler) setDeadmanNotReady(
ctx context.Context, sw *terdutv1alpha1.TerdutDeadmanSwitch, reason, message string, d time.Duration,
) (ctrl.Result, error) {
meta.SetStatusCondition(&sw.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionFalse,
Reason: reason,
Message: message,
})
sw.Status.ObservedGeneration = sw.Generation
if err := r.Status().Update(ctx, sw); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(sw, nil, corev1.EventTypeWarning, reason, reason, message)
}
return ctrl.Result{RequeueAfter: d}, nil
}
// reconcileDeadmanDelete calls the real DELETE this resource actually has
// (unlike TerdutEscalationRule) if the team is still resolvable and a
// switch was ever created, then removes the finalizer unconditionally.
func (r *TerdutDeadmanSwitchReconciler) reconcileDeadmanDelete(
ctx context.Context, sw *terdutv1alpha1.TerdutDeadmanSwitch, newClient func(string) *tdclient.Client,
) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(sw, deadmanFinalizerName) {
return ctrl.Result{}, nil
}
if sw.Status.SwitchID != 0 {
if team, tc, resolveErr := resolveTeamAndClient(
ctx, r.Client, r.OperatorNamespace, sw.Namespace, sw.Spec.TeamRef, newClient,
); resolveErr == nil {
if err := tc.DeleteDeadmanSwitch(ctx, team.Status.TeamID, sw.Status.SwitchID); err != nil {
if r.Recorder != nil {
r.Recorder.Eventf(sw, nil, corev1.EventTypeWarning, "DeleteFailed", "DeleteFailed", err.Error())
}
return ctrl.Result{}, err
}
}
}
controllerutil.RemoveFinalizer(sw, deadmanFinalizerName)
return ctrl.Result{}, r.Update(ctx, sw)
}
// SetupWithManager sets up the controller with the Manager.
func (r *TerdutDeadmanSwitchReconciler) SetupWithManager(mgr ctrl.Manager) error {
if r.NewClient == nil {
r.NewClient = tdclient.New
}
if r.Recorder == nil {
r.Recorder = mgr.GetEventRecorder("terdutdeadmanswitch-controller")
}
return ctrl.NewControllerManagedBy(mgr).
For(&terdutv1alpha1.TerdutDeadmanSwitch{}).
Named("terdutdeadmanswitch").
Complete(r)
}
@@ -1,213 +0,0 @@
package controller
import (
"context"
"net/http/httptest"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/types"
"sigs.k8s.io/controller-runtime/pkg/reconcile"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
var _ = Describe("TerdutDeadmanSwitch Controller", func() {
const operatorNamespace = "default"
var (
reconciler *TerdutDeadmanSwitchReconciler
fake *fakeTerdutServer
fakeSrv *httptest.Server
srv *terdutv1alpha1.TerdutServer
team *terdutv1alpha1.TerdutTeam
swName string
swKey types.NamespacedName
)
BeforeEach(func(ctx SpecContext) {
fake, fakeSrv = newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
srv = bootstrapReadyTerdutServer(ctx, uniqueName("dmserver"), fakeSrv.URL)
team = readyTerdutTeam(ctx, operatorNamespace, uniqueName("dmteam"), srv, fakeSrv.URL)
reconciler = &TerdutDeadmanSwitchReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
}
swName = uniqueName("switch")
swKey = types.NamespacedName{Name: swName, Namespace: operatorNamespace}
})
AfterEach(func(ctx SpecContext) {
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
if err := k8sClient.Get(ctx, swKey, sw); err == nil {
sw.Finalizers = nil
_ = k8sClient.Update(ctx, sw)
_ = k8sClient.Delete(ctx, sw)
}
teamKey := types.NamespacedName{Name: team.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, teamKey, team); err == nil {
team.Finalizers = nil
_ = k8sClient.Update(ctx, team)
_ = k8sClient.Delete(ctx, team)
}
srvKey := types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, srvKey, srv); err == nil {
srv.Finalizers = nil
_ = k8sClient.Update(ctx, srv)
_ = k8sClient.Delete(ctx, srv)
}
})
createSwitch := func(ctx context.Context, teamRef terdutv1alpha1.TerdutTeamRef, name, matcher, timeout string) {
sw := &terdutv1alpha1.TerdutDeadmanSwitch{
ObjectMeta: metav1.ObjectMeta{Name: swName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutDeadmanSwitchSpec{
TeamRef: teamRef,
Name: name,
Matcher: matcher,
Timeout: timeout,
},
}
Expect(k8sClient.Create(ctx, sw)).To(Succeed())
}
reconcileOnce := func(ctx context.Context) {
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: swKey})
Expect(err).NotTo(HaveOccurred())
}
readyCondition := func(ctx context.Context) metav1.Condition {
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
c := meta.FindStatusCondition(sw.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(c).NotTo(BeNil())
return *c
}
sameTeamRef := func() terdutv1alpha1.TerdutTeamRef {
return terdutv1alpha1.TerdutTeamRef{Name: team.Name}
}
Describe("the happy path", func() {
It("creates the switch server-side", func(ctx SpecContext) {
createSwitch(ctx, sameTeamRef(), "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // create
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionTrue))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonChildAdopted))
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
Expect(sw.Status.SwitchID).NotTo(BeZero())
created, ok := fake.switches[team.Status.TeamID][sw.Status.SwitchID]
Expect(ok).To(BeTrue())
Expect(created.Matcher).To(Equal("alertname=Watchdog"))
Expect(created.TimeoutSeconds).To(Equal(int64(900)))
Expect(created.Severity).To(Equal("critical")) // kubebuilder default
})
})
Describe("list-and-match-by-name adoption", func() {
It("adopts an already-created switch instead of creating a duplicate", func(ctx SpecContext) {
// Simulates a prior, interrupted reconcile that got as far as
// POSTing the switch -- no 409 signal exists for this resource
// (DESIGN.md §4.4), so the recovery path is GET-list-and-match,
// not adopt-on-409.
fake.nextSwitchID = 1
fake.switches[team.Status.TeamID] = map[int64]tdclient.DeadmanSwitch{
1: {ID: 1, Name: "heartbeat", Matcher: "alertname=Watchdog", TimeoutSeconds: 900, Severity: "critical"},
}
createSwitch(ctx, sameTeamRef(), "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
Expect(sw.Status.SwitchID).To(Equal(int64(1)))
Expect(fake.switches[team.Status.TeamID]).To(HaveLen(1), "should not have created a second switch")
})
})
Describe("update-in-place on spec drift", func() {
It("PUTs the new spec rather than creating a second switch", func(ctx SpecContext) {
createSwitch(ctx, sameTeamRef(), "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
switchID := sw.Status.SwitchID
sw.Spec.Timeout = "30m"
Expect(k8sClient.Update(ctx, sw)).To(Succeed())
reconcileOnce(ctx)
Expect(fake.switches[team.Status.TeamID]).To(HaveLen(1), "update-in-place, not a second switch")
Expect(fake.switches[team.Status.TeamID][switchID].TimeoutSeconds).To(Equal(int64(1800)))
})
})
Describe("waiting on the referenced TerdutTeam", func() {
It("reports TeamRefNotFound when the TerdutTeam doesn't exist", func(ctx SpecContext) {
createSwitch(ctx, terdutv1alpha1.TerdutTeamRef{Name: testRefNotFoundName}, "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonTeamRefNotFound))
})
It("reports WaitingForTeam when the TerdutTeam exists but isn't Ready yet", func(ctx SpecContext) {
unreadyName := uniqueName("dmteam-unready")
unready := &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: unreadyName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutTeamSpec{
ServerRef: terdutv1alpha1.TerdutServerRef{Name: srv.Name},
DisplayName: testUnreadyDisplayName,
},
}
Expect(k8sClient.Create(ctx, unready)).To(Succeed())
DeferCleanup(func() { _ = k8sClient.Delete(ctx, unready) })
createSwitch(ctx, terdutv1alpha1.TerdutTeamRef{Name: unreadyName}, "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonWaitingForTeam))
})
})
Describe("deletion", func() {
It("deletes the switch server-side (the real DELETE this resource has) and removes the finalizer", func(ctx SpecContext) {
createSwitch(ctx, sameTeamRef(), "heartbeat", "alertname=Watchdog", "15m")
reconcileOnce(ctx)
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
sw := &terdutv1alpha1.TerdutDeadmanSwitch{}
Expect(k8sClient.Get(ctx, swKey, sw)).To(Succeed())
switchID := sw.Status.SwitchID
Expect(k8sClient.Delete(ctx, sw)).To(Succeed())
reconcileOnce(ctx) // runs the finalizer
Expect(fake.switchDelete[switchID]).To(BeTrue())
err := k8sClient.Get(ctx, swKey, sw)
Expect(err).To(HaveOccurred(), "the TerdutDeadmanSwitch itself should be gone once the finalizer clears")
})
})
})
@@ -1,217 +0,0 @@
package controller
import (
"context"
"fmt"
"time"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
logf "sigs.k8s.io/controller-runtime/pkg/log"
"sigs.k8s.io/controller-runtime/pkg/recorder"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// escalationFinalizerName exists for the one undo PUT /api/teams/{teamID}/escalation
// supports, on delete: an empty policy (DESIGN.md §5's general finalizer
// rule expects a server-side counterpart to be undone, but this resource's
// API is GET/PUT-only, with no DELETE at all -- a zeroed PUT is the closest
// equivalent terdut-server itself recognizes as "no policy"
// (escalationPolicy.configured(), confirmed against source: "a policy row
// with no levels is the same as no policy").
const escalationFinalizerName = "terdut.ryuvia.com/terdutescalationrule"
// TerdutEscalationRuleReconciler reconciles a TerdutEscalationRule object.
type TerdutEscalationRuleReconciler struct {
client.Client
Scheme *runtime.Scheme
OperatorNamespace string
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
}
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutescalationrules,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutescalationrules/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutescalationrules/finalizers,verbs=update
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutteams,verbs=get;list;watch
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
func (r *TerdutEscalationRuleReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
log := logf.FromContext(ctx)
var rule terdutv1alpha1.TerdutEscalationRule
if err := r.Get(ctx, req.NamespacedName, &rule); err != nil {
if apierrors.IsNotFound(err) {
return ctrl.Result{}, nil
}
return ctrl.Result{}, err
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
if !rule.DeletionTimestamp.IsZero() {
return r.reconcileEscalationDelete(ctx, &rule, newClient)
}
if !controllerutil.ContainsFinalizer(&rule, escalationFinalizerName) {
controllerutil.AddFinalizer(&rule, escalationFinalizerName)
if err := r.Update(ctx, &rule); err != nil {
return ctrl.Result{}, err
}
return ctrl.Result{}, nil
}
team, tc, resolveErr := resolveTeamAndClient(ctx, r.Client, r.OperatorNamespace, rule.Namespace, rule.Spec.TeamRef, newClient)
if resolveErr != nil {
return r.setEscalationNotReady(ctx, &rule, resolveErr.reason, resolveErr.message, waitInterval)
}
body, unknownUser, err := buildEscalationRequest(ctx, tc, rule.Spec)
if err != nil {
return ctrl.Result{}, err
}
if unknownUser != "" {
return r.setEscalationNotReady(ctx, &rule, terdutv1alpha1.ReasonUnknownUser,
fmt.Sprintf("username %q does not resolve to any user", unknownUser), waitInterval)
}
if err := tc.SetEscalation(ctx, team.Status.TeamID, body); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/escalation: %w", team.Status.TeamID, err)
}
meta.SetStatusCondition(&rule.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady, // "Ready" -- same name, shared across every CRD (DESIGN.md §7)
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonChildAdopted,
Message: fmt.Sprintf("escalation policy applied to team %d", team.Status.TeamID),
})
rule.Status.ObservedGeneration = rule.Generation
if err := r.Status().Update(ctx, &rule); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(&rule, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonChildAdopted, terdutv1alpha1.ReasonChildAdopted,
"escalation policy applied")
}
log.Info("TerdutEscalationRule applied", "name", rule.Name, "teamID", team.Status.TeamID)
return ctrl.Result{RequeueAfter: resyncInterval}, nil
}
// buildEscalationRequest resolves every "user" target's username to a
// user_id (DESIGN.md §4.3) and translates spec into the wire shape
// SetEscalation sends. Returns the first unresolvable username, if any,
// distinct from a plain error: that's an expected, reportable condition
// (ReasonUnknownUser), not a reconcile failure.
func buildEscalationRequest(
ctx context.Context, tc *tdclient.Client, spec terdutv1alpha1.TerdutEscalationRuleSpec,
) (tdclient.SetEscalationRequest, string, error) {
levels := make([]tdclient.EscalationLevelRequest, len(spec.Levels))
for i, lvl := range spec.Levels {
timeout, err := time.ParseDuration(lvl.Timeout)
if err != nil {
return tdclient.SetEscalationRequest{}, "", fmt.Errorf("spec.levels[%d].timeout %q: %w", i, lvl.Timeout, err)
}
targets := make([]tdclient.EscalationTargetRequest, len(lvl.Targets))
for j, t := range lvl.Targets {
if t.Kind == terdutv1alpha1.EscalationTargetOncall {
targets[j] = tdclient.EscalationTargetRequest{Kind: string(terdutv1alpha1.EscalationTargetOncall)}
continue
}
user, err := tc.GetUserByUsername(ctx, t.Username)
if err != nil {
return tdclient.SetEscalationRequest{}, "", fmt.Errorf("GET /api/users (resolving %q): %w", t.Username, err)
}
if user == nil {
return tdclient.SetEscalationRequest{}, t.Username, nil
}
targets[j] = tdclient.EscalationTargetRequest{Kind: string(terdutv1alpha1.EscalationTargetUser), UserID: &user.ID}
}
levels[i] = tdclient.EscalationLevelRequest{
Position: int64(i + 1),
TimeoutSeconds: int64(timeout.Seconds()),
Targets: targets,
}
}
return tdclient.SetEscalationRequest{
RepeatCount: spec.RepeatCount,
FallbackTopic: spec.FallbackTopic,
Levels: levels,
}, "", nil
}
func (r *TerdutEscalationRuleReconciler) setEscalationNotReady(
ctx context.Context, rule *terdutv1alpha1.TerdutEscalationRule, reason, message string, d time.Duration,
) (ctrl.Result, error) {
meta.SetStatusCondition(&rule.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionFalse,
Reason: reason,
Message: message,
})
rule.Status.ObservedGeneration = rule.Generation
if err := r.Status().Update(ctx, rule); err != nil {
return ctrl.Result{}, err
}
if r.Recorder != nil {
r.Recorder.Eventf(rule, nil, corev1.EventTypeWarning, reason, reason, message)
}
return ctrl.Result{RequeueAfter: d}, nil
}
// reconcileEscalationDelete PUTs an empty policy (this resource's only
// available "undo", per this file's own const comment) if the team is
// still resolvable, then removes the finalizer unconditionally -- same
// "the parent's probably going away too" reasoning TerdutTeam's own delete
// path uses for a TerdutServer that's gone.
func (r *TerdutEscalationRuleReconciler) reconcileEscalationDelete(
ctx context.Context, rule *terdutv1alpha1.TerdutEscalationRule, newClient func(string) *tdclient.Client,
) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(rule, escalationFinalizerName) {
return ctrl.Result{}, nil
}
if team, tc, resolveErr := resolveTeamAndClient(
ctx, r.Client, r.OperatorNamespace, rule.Namespace, rule.Spec.TeamRef, newClient,
); resolveErr == nil {
if err := tc.SetEscalation(ctx, team.Status.TeamID, tdclient.SetEscalationRequest{Levels: []tdclient.EscalationLevelRequest{}}); err != nil {
if r.Recorder != nil {
r.Recorder.Eventf(rule, nil, corev1.EventTypeWarning, "DeleteFailed", "DeleteFailed", err.Error())
}
return ctrl.Result{}, err
}
}
controllerutil.RemoveFinalizer(rule, escalationFinalizerName)
return ctrl.Result{}, r.Update(ctx, rule)
}
// SetupWithManager sets up the controller with the Manager.
func (r *TerdutEscalationRuleReconciler) SetupWithManager(mgr ctrl.Manager) error {
if r.NewClient == nil {
r.NewClient = tdclient.New
}
if r.Recorder == nil {
r.Recorder = mgr.GetEventRecorder("terdutescalationrule-controller")
}
return ctrl.NewControllerManagedBy(mgr).
For(&terdutv1alpha1.TerdutEscalationRule{}).
Named("terdutescalationrule").
Complete(r)
}
@@ -1,206 +0,0 @@
package controller
import (
"context"
"net/http/httptest"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/types"
"sigs.k8s.io/controller-runtime/pkg/reconcile"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
var _ = Describe("TerdutEscalationRule Controller", func() {
const operatorNamespace = "default"
var (
reconciler *TerdutEscalationRuleReconciler
fake *fakeTerdutServer
fakeSrv *httptest.Server
srv *terdutv1alpha1.TerdutServer
team *terdutv1alpha1.TerdutTeam
ruleName string
ruleKey types.NamespacedName
)
BeforeEach(func(ctx SpecContext) {
fake, fakeSrv = newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
srv = bootstrapReadyTerdutServer(ctx, uniqueName("erserver"), fakeSrv.URL)
team = readyTerdutTeam(ctx, operatorNamespace, uniqueName("erteam"), srv, fakeSrv.URL)
reconciler = &TerdutEscalationRuleReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
}
ruleName = uniqueName("escalation")
ruleKey = types.NamespacedName{Name: ruleName, Namespace: operatorNamespace}
})
AfterEach(func(ctx SpecContext) {
rule := &terdutv1alpha1.TerdutEscalationRule{}
if err := k8sClient.Get(ctx, ruleKey, rule); err == nil {
rule.Finalizers = nil
_ = k8sClient.Update(ctx, rule)
_ = k8sClient.Delete(ctx, rule)
}
// team/srv carry their own finalizers from readyTerdutTeam/
// bootstrapReadyTerdutServer -- clear them directly the same way
// terdutteam_controller_test.go's own AfterEach does, rather than
// relying on either reconciler to ever run again here.
teamKey := types.NamespacedName{Name: team.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, teamKey, team); err == nil {
team.Finalizers = nil
_ = k8sClient.Update(ctx, team)
_ = k8sClient.Delete(ctx, team)
}
srvKey := types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, srvKey, srv); err == nil {
srv.Finalizers = nil
_ = k8sClient.Update(ctx, srv)
_ = k8sClient.Delete(ctx, srv)
}
})
createRule := func(ctx context.Context, teamRef terdutv1alpha1.TerdutTeamRef, levels []terdutv1alpha1.EscalationLevel) {
rule := &terdutv1alpha1.TerdutEscalationRule{
ObjectMeta: metav1.ObjectMeta{Name: ruleName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutEscalationRuleSpec{
TeamRef: teamRef,
Levels: levels,
},
}
Expect(k8sClient.Create(ctx, rule)).To(Succeed())
}
reconcileOnce := func(ctx context.Context) {
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: ruleKey})
Expect(err).NotTo(HaveOccurred())
}
readyCondition := func(ctx context.Context) metav1.Condition {
rule := &terdutv1alpha1.TerdutEscalationRule{}
Expect(k8sClient.Get(ctx, ruleKey, rule)).To(Succeed())
c := meta.FindStatusCondition(rule.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(c).NotTo(BeNil())
return *c
}
sameTeamRef := func() terdutv1alpha1.TerdutTeamRef {
return terdutv1alpha1.TerdutTeamRef{Name: team.Name}
}
Describe("the happy path", func() {
It("resolves usernames and applies the escalation policy", func(ctx SpecContext) {
fake.seedUser("alice")
createRule(ctx, sameTeamRef(), []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{
{Kind: terdutv1alpha1.EscalationTargetUser, Username: "alice"},
}},
{Timeout: "10m", Targets: []terdutv1alpha1.EscalationTarget{
{Kind: terdutv1alpha1.EscalationTargetOncall},
}},
})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // resolve + apply
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionTrue))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonChildAdopted))
applied, ok := fake.escalation[team.Status.TeamID]
Expect(ok).To(BeTrue())
Expect(applied.Levels).To(HaveLen(2))
Expect(applied.Levels[0].Targets[0].Kind).To(Equal("user"))
Expect(applied.Levels[0].Targets[0].UserID).NotTo(BeNil())
Expect(*applied.Levels[0].Targets[0].UserID).To(Equal(fake.users["alice"]))
Expect(applied.Levels[1].Targets[0].Kind).To(Equal("oncall"))
Expect(applied.Levels[1].Targets[0].UserID).To(BeNil())
})
})
Describe("an unresolvable username", func() {
It("reports UnknownUser and never calls PUT /api/teams/{id}/escalation", func(ctx SpecContext) {
createRule(ctx, sameTeamRef(), []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{
{Kind: terdutv1alpha1.EscalationTargetUser, Username: "ghost"},
}},
})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionFalse))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonUnknownUser))
_, applied := fake.escalation[team.Status.TeamID]
Expect(applied).To(BeFalse())
})
})
Describe("waiting on the referenced TerdutTeam", func() {
It("reports TeamRefNotFound when the TerdutTeam doesn't exist", func(ctx SpecContext) {
createRule(ctx, terdutv1alpha1.TerdutTeamRef{Name: testRefNotFoundName}, []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{{Kind: terdutv1alpha1.EscalationTargetOncall}}},
})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonTeamRefNotFound))
})
It("reports WaitingForTeam when the TerdutTeam exists but isn't Ready yet", func(ctx SpecContext) {
unreadyName := uniqueName("erteam-unready")
unready := &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: unreadyName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutTeamSpec{
ServerRef: terdutv1alpha1.TerdutServerRef{Name: srv.Name},
DisplayName: testUnreadyDisplayName,
},
}
Expect(k8sClient.Create(ctx, unready)).To(Succeed())
DeferCleanup(func() { _ = k8sClient.Delete(ctx, unready) })
createRule(ctx, terdutv1alpha1.TerdutTeamRef{Name: unreadyName}, []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{{Kind: terdutv1alpha1.EscalationTargetOncall}}},
})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonWaitingForTeam))
})
})
Describe("deletion", func() {
It("PUTs an empty policy (this resource's only available undo) and removes the finalizer", func(ctx SpecContext) {
fake.seedUser("alice")
createRule(ctx, sameTeamRef(), []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{
{Kind: terdutv1alpha1.EscalationTargetUser, Username: "alice"},
}},
})
reconcileOnce(ctx)
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
Expect(fake.escalation[team.Status.TeamID].Levels).To(HaveLen(1))
rule := &terdutv1alpha1.TerdutEscalationRule{}
Expect(k8sClient.Get(ctx, ruleKey, rule)).To(Succeed())
Expect(k8sClient.Delete(ctx, rule)).To(Succeed())
reconcileOnce(ctx) // runs the finalizer
Expect(fake.escalation[team.Status.TeamID].Levels).To(BeEmpty())
err := k8sClient.Get(ctx, ruleKey, rule)
Expect(err).To(HaveOccurred(), "the TerdutEscalationRule itself should be gone once the finalizer clears")
})
})
})
@@ -1,131 +0,0 @@
package controller
import (
"context"
"errors"
"fmt"
"net/http"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// bootstrapStateLostError is DESIGN.md §6's one genuinely pathological
// case: a checkpointed admin credential was used and then lost before the
// lasting credential it was for could be persisted. Distinct from a plain
// error so Reconcile can route it to a Ready: False condition (the
// documented recovery is delete-and-recreate, not an automatic retry) rather
// than treating it as a transient reconcile failure.
type bootstrapStateLostError struct{ detail string }
func (e *bootstrapStateLostError) Error() string {
return fmt.Sprintf(
"server reports already bootstrapped, but neither status.credentialsSecretRef nor a "+
"checkpointed admin credential exist here: %s. This TerdutServer cannot recover a "+
"credential on its own; delete and recreate it", e.detail)
}
// reconcileBootstrap implements DESIGN.md §6 point 1's self-registration
// flow, checkpointed against the two real crash windows in it rather than
// leaving them as theoretical gaps. Only called once
// srv.Status.CredentialsSecretRef is nil and the Deployment has a ready
// replica.
func (r *TerdutServerReconciler) reconcileBootstrap(ctx context.Context, srv *terdutv1alpha1.TerdutServer) error {
adminKey, err := r.getOrCreateCheckpointedAdminKey(ctx, srv)
if err != nil {
return err
}
bc := r.NewClient(serviceURL(srv)).WithToken(adminKey)
instanceKey, err := r.getOrMintInstanceServiceAccountKey(ctx, bc)
if err != nil {
return err
}
credsName := credentialsSecretName(srv)
if err := writeOperatorSecret(ctx, r.Client, r.OperatorNamespace, credsName, instanceKey); err != nil {
return err
}
srv.Status.CredentialsSecretRef = &terdutv1alpha1.SecretKeyRef{Name: credsName, Key: credentialsSecretDataKey}
// Best-effort: the checkpoint has done its job. Leaving it behind on a
// delete failure here isn't a correctness problem (the next reconcile
// finds status.CredentialsSecretRef already set and never looks at the
// checkpoint again) — it would just be an unused Secret sitting around,
// cleaned up for real by the finalizer on delete.
checkpoint := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: checkpointSecretName(srv), Namespace: r.OperatorNamespace}}
_ = r.Delete(ctx, checkpoint)
return nil
}
// getOrCreateCheckpointedAdminKey returns a usable admin key: from the
// checkpoint Secret if an earlier, interrupted attempt already got one, or
// freshly from /api/bootstrap, immediately checkpointed before it's used
// for anything else.
func (r *TerdutServerReconciler) getOrCreateCheckpointedAdminKey(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (string, error) {
checkpointName := checkpointSecretName(srv)
var checkpoint corev1.Secret
err := r.Get(ctx, client.ObjectKey{Namespace: r.OperatorNamespace, Name: checkpointName}, &checkpoint)
switch {
case err == nil:
return string(checkpoint.Data[credentialsSecretDataKey]), nil
case !apierrors.IsNotFound(err):
return "", err
}
bc := r.NewClient(serviceURL(srv))
result, err := bc.Bootstrap(ctx, bootstrapUsername, bootstrapEmail)
if err != nil {
if statusErr, ok := errors.AsType[*tdclient.StatusError](err); ok && statusErr.Code == http.StatusForbidden {
// §1: this operator is the only thing that ever bootstraps a
// server it created, so a 403 here (no checkpoint, no
// status.credentialsSecretRef) means a prior reconcile already
// won this exact race and its checkpoint was lost afterward --
// the one case §6 doesn't try to paper over.
return "", &bootstrapStateLostError{detail: "/api/bootstrap returned 403"}
}
return "", fmt.Errorf("POST /api/bootstrap: %w", err)
}
if err := writeOperatorSecret(ctx, r.Client, r.OperatorNamespace, checkpointName, result.APIKey.Key); err != nil {
return "", fmt.Errorf("checkpointing admin key: %w", err)
}
return result.APIKey.Key, nil
}
// getOrMintInstanceServiceAccountKey mints the operator's own instance-
// scoped service account, or, if an earlier interrupted attempt already
// created it (409), adopts it and mints a fresh key rather than treating
// the conflict as an error (DESIGN.md §6 point 1, §5's general
// adopt-on-conflict rule).
func (r *TerdutServerReconciler) getOrMintInstanceServiceAccountKey(ctx context.Context, bc *tdclient.Client) (string, error) {
result, err := bc.CreateInstanceServiceAccount(ctx, serviceAccountName)
if err == nil {
return result.Key.Key, nil
}
statusErr, ok := errors.AsType[*tdclient.StatusError](err)
if !ok || statusErr.Code != http.StatusConflict {
return "", fmt.Errorf("POST /api/service-accounts: %w", err)
}
sa, err := bc.GetServiceAccountByName(ctx, serviceAccountName)
if err != nil {
return "", fmt.Errorf("GET /api/service-accounts?name=%s (adopting after 409): %w", serviceAccountName, err)
}
if sa == nil {
return "", fmt.Errorf("POST /api/service-accounts 409'd for %q but GET found nothing", serviceAccountName)
}
key, err := bc.CreateServiceAccountKey(ctx, sa.ID, "initial")
if err != nil {
return "", fmt.Errorf("POST /api/service-accounts/%d/keys (adopting after 409): %w", sa.ID, err)
}
return key.Key, nil
}
+16 -110
View File
@@ -2,12 +2,13 @@ package controller
import (
"context"
"errors"
"fmt"
"time"
appsv1 "k8s.io/api/apps/v1"
corev1 "k8s.io/api/core/v1"
policyv1 "k8s.io/api/policy/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
@@ -15,12 +16,10 @@ import (
"k8s.io/apimachinery/pkg/runtime/schema"
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
logf "sigs.k8s.io/controller-runtime/pkg/log"
"sigs.k8s.io/controller-runtime/pkg/recorder"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// resyncInterval is the periodic requeue on a successful reconcile (DESIGN.md
@@ -38,27 +37,6 @@ const resyncInterval = 5 * time.Minute
// than "someone edited something out of band."
const waitInterval = 15 * time.Second
// finalizerName cleans up the credentials Secret(s) this controller
// generates in the operator's own namespace on delete — the Deployment and
// Service are owned (OwnerReference, DESIGN.md §7) and need no finalizer of
// their own.
const finalizerName = "terdut.ryuvia.com/terdutserver"
// serviceAccountName is the name the operator registers itself under
// server-side (DESIGN.md §6) — a fixed, repo-wide constant, not a spec
// field: it names the automation, not anything about this one TerdutServer.
const serviceAccountName = "terdut-operator"
// bootstrapUsername/bootstrapEmail found the one human-shaped user every
// fresh install needs (terdut-server's handleBootstrap requires both).
// Nobody signs in as this user afterward — its only purpose is minting the
// admin key the controller immediately trades for a real service-account
// key — so these are fixed, not spec fields.
const (
bootstrapUsername = "terdut-operator-bootstrap"
bootstrapEmail = "bootstrap@terdut-operator.local"
)
// postgresqlGVK is the Zalando postgres-operator's CR (DESIGN.md §8).
// Resolved via unstructured rather than vendoring Zalando's own client, to
// keep this operator's dependency on it to "an optional CRD read" rather
@@ -74,21 +52,9 @@ type TerdutServerReconciler struct {
client.Client
Scheme *runtime.Scheme
// OperatorNamespace is where every credentials Secret this controller
// generates lives (DESIGN.md §6) — never the TerdutServer's own
// namespace. Set from the POD_NAMESPACE downward-API env var in
// production (cmd/main.go); tests set it directly.
OperatorNamespace string
// Recorder emits the Kubernetes Events DESIGN.md §12 asks for on every
// externally-visible outcome.
Recorder recorder.EventRecorder
// NewClient builds the terdut-server API client for a given endpoint. A
// field, not a direct tdclient.New call, so tests can substitute an
// httptest.Server's client without a real network round trip. Defaults
// to tdclient.New via SetupWithManager.
NewClient func(endpoint string) *tdclient.Client
}
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutservers,verbs=get;list;watch;create;update;patch;delete
@@ -97,6 +63,7 @@ type TerdutServerReconciler struct {
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups="",resources=services,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=policy,resources=poddisruptionbudgets,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=acid.zalan.do,resources=postgresqls,verbs=get;list;watch
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
@@ -111,39 +78,27 @@ func (r *TerdutServerReconciler) Reconcile(ctx context.Context, req ctrl.Request
return ctrl.Result{}, err
}
if !srv.DeletionTimestamp.IsZero() {
return r.reconcileDelete(ctx, &srv)
}
if !controllerutil.ContainsFinalizer(&srv, finalizerName) {
controllerutil.AddFinalizer(&srv, finalizerName)
if err := r.Update(ctx, &srv); err != nil {
return ctrl.Result{}, err
}
// The Update above re-triggers a reconcile via the watch; nothing
// further to do on this pass.
return ctrl.Result{}, nil
}
dbEnv, dbErr := r.resolveDatabaseEnv(ctx, &srv)
if dbErr != nil {
return r.setNotReady(ctx, &srv, dbErr.reason, dbErr.message, waitInterval)
}
deploy, err := r.reconcileDeployment(ctx, &srv, dbEnv)
operatorKey, err := r.reconcileOperatorKey(ctx, &srv)
if err != nil {
return ctrl.Result{}, err
}
srv.Status.CredentialsSecretRef = operatorKeyRef(&srv)
deploy, err := r.reconcileDeployment(ctx, &srv, dbEnv, operatorKeyHash(operatorKey))
if err != nil {
return ctrl.Result{}, err
}
if err := r.reconcileService(ctx, &srv); err != nil {
return ctrl.Result{}, err
}
meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionDatabaseReady,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonAdopted,
Message: "database resolved",
})
if err := r.reconcilePodDisruptionBudget(ctx, &srv); err != nil {
return ctrl.Result{}, err
}
if deploy.Status.ReadyReplicas < 1 {
return r.setNotReady(ctx, &srv,
@@ -152,28 +107,12 @@ func (r *TerdutServerReconciler) Reconcile(ctx context.Context, req ctrl.Request
waitInterval)
}
if srv.Status.CredentialsSecretRef == nil {
if err := r.reconcileBootstrap(ctx, &srv); err != nil {
if pending, ok := errors.AsType[*bootstrapStateLostError](err); ok {
return r.setNotReady(ctx, &srv, terdutv1alpha1.ReasonBootstrapStateLost, pending.Error(), waitInterval)
}
return ctrl.Result{}, err
}
}
meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionBootstrapped,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonAdopted,
Message: fmt.Sprintf("credentials in Secret %q", srv.Status.CredentialsSecretRef.Name),
})
meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonAdopted,
Message: "deployment ready, database resolved, credentials bootstrapped",
Message: "deployment ready, database resolved, operator key seeded",
})
srv.Status.ServiceName = srv.Name
srv.Status.ObservedGeneration = srv.Generation
if err := r.Status().Update(ctx, &srv); err != nil {
return ctrl.Result{}, err
@@ -210,35 +149,8 @@ func (r *TerdutServerReconciler) setNotReady(
return ctrl.Result{RequeueAfter: d}, nil
}
// reconcileDelete cleans up the credentials Secret(s) this controller
// generated in the operator's own namespace. The Deployment and Service are
// owned (OwnerReference, DESIGN.md §7) and need no attention here — normal
// GC handles them. There is no server-side "delete this install" call to
// make: bootstrap created a user and a service account, and terdut-server's
// API has no way to delete either (only to revoke individual keys), so
// there is nothing meaningful to undo there either.
func (r *TerdutServerReconciler) reconcileDelete(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(srv, finalizerName) {
return ctrl.Result{}, nil
}
for _, name := range []string{checkpointSecretName(srv), credentialsSecretName(srv)} {
sec := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: r.OperatorNamespace}}
if err := r.Delete(ctx, sec); err != nil && !apierrors.IsNotFound(err) {
return ctrl.Result{}, err
}
}
controllerutil.RemoveFinalizer(srv, finalizerName)
if err := r.Update(ctx, srv); err != nil {
return ctrl.Result{}, err
}
return ctrl.Result{}, nil
}
// SetupWithManager sets up the controller with the Manager.
func (r *TerdutServerReconciler) SetupWithManager(mgr ctrl.Manager) error {
if r.NewClient == nil {
r.NewClient = tdclient.New
}
if r.Recorder == nil {
r.Recorder = mgr.GetEventRecorder("terdutserver-controller")
}
@@ -246,20 +158,14 @@ func (r *TerdutServerReconciler) SetupWithManager(mgr ctrl.Manager) error {
For(&terdutv1alpha1.TerdutServer{}).
Owns(&appsv1.Deployment{}).
Owns(&corev1.Service{}).
Owns(&policyv1.PodDisruptionBudget{}).
Owns(&corev1.Secret{}).
Named("terdutserver").
Complete(r)
}
// --- naming ---
func checkpointSecretName(srv *terdutv1alpha1.TerdutServer) string {
return fmt.Sprintf("%s.%s-bootstrap-admin", srv.Namespace, srv.Name)
}
func credentialsSecretName(srv *terdutv1alpha1.TerdutServer) string {
return fmt.Sprintf("%s.%s-instance-credentials", srv.Namespace, srv.Name)
}
func serviceURL(srv *terdutv1alpha1.TerdutServer) string {
port := srv.Spec.Networking.ServicePort
if port == 0 {
@@ -2,459 +2,26 @@ package controller
import (
"context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strconv"
"strings"
"sync"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
appsv1 "k8s.io/api/apps/v1"
corev1 "k8s.io/api/core/v1"
policyv1 "k8s.io/api/policy/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta"
"k8s.io/apimachinery/pkg/api/resource"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
"k8s.io/apimachinery/pkg/types"
"k8s.io/apimachinery/pkg/util/intstr"
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/reconcile"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// fakeVersionString/errJSONKey are shared by every response fakeTerdutServer
// writes -- goconst would otherwise flag "test" and "error" as repeated
// literals across its handlers.
const (
fakeVersionString = "test"
errJSONKey = "error"
)
// fakeTerdutServer reproduces the exact stateful semantics of
// /api/bootstrap, /api/service-accounts and /api/version that the
// bootstrap flow depends on (DESIGN.md §6, §11: "terdut-server's REST API
// is faked with a small httptest.Server per controller test ... matching
// the real handlers' request/response shapes"), including the 403-after-
// first-success and 409-on-name-conflict behavior the adopt-on-conflict
// recovery path exists for.
type fakeTerdutServer struct {
mu sync.Mutex
bootstrapped bool
bootstrap403 bool // force every /api/bootstrap call to 403, even the first
nextID int64
accounts map[string]int64 // name -> id
keyMints map[int64]int // id -> number of keys minted so far
nextTeamID int64
teams map[string]int64 // name -> id
teamNames map[int64]string // id -> current name (renames update this)
teamOIDC map[int64][2]string
teamDelete map[int64]bool // id -> true once DELETEd, for 404-on-redelete
// users backs GET /api/users for TerdutEscalationRule's username
// resolution (DESIGN.md §4.3) -- a fixed, pre-seeded directory, since
// nothing in this controller's own flow ever creates a user.
users map[string]int64 // username -> id
// escalation backs PUT /api/teams/{id}/escalation -- an upsert
// server-side (confirmed against source), so this is just "the last
// body PUT for this team", keyed by teamID, with no separate create
// step to model.
escalation map[int64]tdclient.SetEscalationRequest
// switches/nextSwitchID/switchDelete back the dead man's switch
// endpoints -- no unique-name constraint server-side (DESIGN.md §4.4),
// so switches is keyed by id, not name, same as the real API's own
// GET-list-and-match-by-name idempotent-create shape requires.
nextSwitchID int64
switches map[int64]map[int64]tdclient.DeadmanSwitch // teamID -> switchID -> switch
switchDelete map[int64]bool // switchID -> true once DELETEd, for 404-on-redelete
// integrations/nextIntegrationID/integrationDelete back the alert
// source endpoints -- no unique-name constraint server-side either
// (DESIGN.md §4.5), same shape as switches, keyed by id.
nextIntegrationID int64
integrations map[int64]map[int64]tdclient.Integration // teamID -> integrationID -> integration
integrationDelete map[int64]bool // integrationID -> true once DELETEd, for 404-on-redelete
}
func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) {
f := &fakeTerdutServer{
accounts: map[string]int64{},
keyMints: map[int64]int{},
teams: map[string]int64{},
teamNames: map[int64]string{},
teamOIDC: map[int64][2]string{},
teamDelete: map[int64]bool{},
users: map[string]int64{},
escalation: map[int64]tdclient.SetEscalationRequest{},
switches: map[int64]map[int64]tdclient.DeadmanSwitch{},
switchDelete: map[int64]bool{},
integrations: map[int64]map[int64]tdclient.Integration{},
integrationDelete: map[int64]bool{},
}
return f, httptest.NewServer(f)
}
// seedUser registers a username the fake GET /api/users will return --
// called from test setup, before the controller under test ever runs.
// Callers read the assigned id back from f.users themselves, so this has
// nothing left to return.
func (f *fakeTerdutServer) seedUser(username string) {
f.mu.Lock()
defer f.mu.Unlock()
f.nextID++
f.users[username] = f.nextID
}
func (f *fakeTerdutServer) ServeHTTP(w http.ResponseWriter, r *http.Request) {
f.mu.Lock()
defer f.mu.Unlock()
switch {
case r.URL.Path == "/api/version":
writeJSON(w, http.StatusOK, map[string]string{"version": fakeVersionString})
case r.URL.Path == "/api/bootstrap" && r.Method == http.MethodPost:
if f.bootstrapped || f.bootstrap403 {
writeJSON(w, http.StatusForbidden, map[string]string{errJSONKey: "bootstrap already completed"})
return
}
f.bootstrapped = true
writeJSON(w, http.StatusCreated, tdclient.BootstrapResult{
APIKey: tdclient.APIKey{ID: 1, Name: "bootstrap", Key: "admin-key-raw"},
})
case r.URL.Path == "/api/service-accounts" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
Scope string `json:"scope"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if _, exists := f.accounts[req.Name]; exists {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a service account with that name already exists"})
return
}
f.nextID++
id := f.nextID
f.accounts[req.Name] = id
f.keyMints[id] = 1
writeJSON(w, http.StatusCreated, tdclient.CreateServiceAccountResult{
ServiceAccount: tdclient.ServiceAccount{ID: id, Name: req.Name, Scope: req.Scope},
Key: tdclient.APIKey{ID: 1, Name: "initial", Key: fmt.Sprintf("instance-key-%d-initial", id)},
})
case r.URL.Path == "/api/service-accounts" && r.Method == http.MethodGet:
name := r.URL.Query().Get("name")
id, exists := f.accounts[name]
if !exists {
writeJSON(w, http.StatusOK, []tdclient.ServiceAccount{})
return
}
writeJSON(w, http.StatusOK, []tdclient.ServiceAccount{{ID: id, Name: name, Scope: "instance"}})
case r.URL.Path == "/api/teams" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if _, exists := f.teams[req.Name]; exists {
writeJSON(w, http.StatusConflict, map[string]string{errJSONKey: "a team with that name already exists"})
return
}
f.nextTeamID++
id := f.nextTeamID
f.teams[req.Name] = id
f.teamNames[id] = req.Name
writeJSON(w, http.StatusCreated, tdclient.Team{ID: id, Name: req.Name})
case r.URL.Path == "/api/teams" && r.Method == http.MethodGet:
// The controller never calls this without ?name= (TEAM-LOOKUP.md's
// own lookup shape) -- the fake only needs to answer that form.
name := r.URL.Query().Get("name")
id, exists := f.teams[name]
if !exists {
writeJSON(w, http.StatusOK, []tdclient.Team{})
return
}
writeJSON(w, http.StatusOK, []tdclient.Team{{ID: id, Name: f.teamNames[id]}})
case r.URL.Path == "/api/users" && r.Method == http.MethodGet:
// No query filter -- GetUserByUsername fetches the whole list and
// matches client-side (confirmed against source: no server-side
// filter either), so the fake does the same.
users := make([]tdclient.User, 0, len(f.users))
for name, id := range f.users {
users = append(users, tdclient.User{ID: id, Username: name})
}
writeJSON(w, http.StatusOK, users)
default:
if id, name, ok := parseKeysPath(r.URL.Path); ok && r.Method == http.MethodPost {
f.keyMints[id]++
writeJSON(w, http.StatusCreated, tdclient.APIKey{
ID: int64(f.keyMints[id]), Name: name,
Key: fmt.Sprintf("instance-key-%d-mint%d", id, f.keyMints[id]),
})
return
}
if id, rest, ok := parseTeamSubPath(r.URL.Path); ok {
f.handleTeamSubPath(w, r, id, rest)
return
}
w.WriteHeader(http.StatusNotFound)
}
}
// handleTeamSubPath answers everything under /api/teams/{id}: PUT (rename),
// DELETE, PUT .../oidc-groups, PUT .../escalation, and the dead man's
// switch collection/item endpoints. rest is whatever parseTeamSubPath found
// after "/api/teams/{id}" -- "" for the bare resource.
func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "" && r.Method == http.MethodPut:
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
oldName, exists := f.teamNames[id]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.teams, oldName)
f.teamNames[id] = req.Name
f.teams[req.Name] = id
w.WriteHeader(http.StatusNoContent)
case rest == "" && r.Method == http.MethodDelete:
name, exists := f.teamNames[id]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.teams, name)
delete(f.teamNames, id)
f.teamDelete[id] = true
w.WriteHeader(http.StatusNoContent)
case rest == "/oidc-groups" && r.Method == http.MethodPut:
var req struct {
MemberGroup string `json:"member_group"`
OwnerGroup string `json:"owner_group"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
if _, exists := f.teamNames[id]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
f.teamOIDC[id] = [2]string{req.MemberGroup, req.OwnerGroup}
w.WriteHeader(http.StatusNoContent)
case rest == "/escalation" && r.Method == http.MethodPut:
var req tdclient.SetEscalationRequest
_ = json.NewDecoder(r.Body).Decode(&req)
f.escalation[id] = req
w.WriteHeader(http.StatusNoContent)
case rest == "/deadman/switches" && r.Method == http.MethodGet:
existing := f.switches[id]
out := make([]tdclient.DeadmanSwitch, 0, len(existing))
for _, s := range existing {
out = append(out, s)
}
writeJSON(w, http.StatusOK, out)
case rest == "/deadman/switches" && r.Method == http.MethodPost:
var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req)
f.nextSwitchID++
switchID := f.nextSwitchID
name := req.Name
if name == "" {
name = "derived-" + req.Matcher
}
sw := tdclient.DeadmanSwitch{
ID: switchID, Name: name, Matcher: req.Matcher,
TimeoutSeconds: req.TimeoutSeconds, Severity: req.Severity,
}
if f.switches[id] == nil {
f.switches[id] = map[int64]tdclient.DeadmanSwitch{}
}
f.switches[id][switchID] = sw
writeJSON(w, http.StatusCreated, sw)
case strings.HasPrefix(rest, "/deadman/switches/") && r.Method == http.MethodPut:
switchID, ok := parseTrailingID(rest, "/deadman/switches/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.switches[id][switchID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req)
name := req.Name
if name == "" {
name = f.switches[id][switchID].Name
}
f.switches[id][switchID] = tdclient.DeadmanSwitch{
ID: switchID, Name: name, Matcher: req.Matcher,
TimeoutSeconds: req.TimeoutSeconds, Severity: req.Severity,
}
w.WriteHeader(http.StatusNoContent)
case strings.HasPrefix(rest, "/deadman/switches/") && r.Method == http.MethodDelete:
switchID, ok := parseTrailingID(rest, "/deadman/switches/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.switches[id][switchID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.switches[id], switchID)
f.switchDelete[switchID] = true
w.WriteHeader(http.StatusNoContent)
case rest == "/integrations" || strings.HasPrefix(rest, "/integrations/"):
f.handleIntegrationSubPath(w, r, id, rest)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleIntegrationSubPath answers POST /api/teams/{id}/integrations,
// PATCH .../integrations/{integrationID} and DELETE .../integrations/{integrationID}
// -- split out of handleTeamSubPath so that switch's own cyclomatic
// complexity stays under golangci-lint's gocyclo threshold.
func (f *fakeTerdutServer) handleIntegrationSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "/integrations" && r.Method == http.MethodPost:
var req struct {
Name string `json:"name"`
Kind string `json:"kind"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
f.nextIntegrationID++
integID := f.nextIntegrationID
integ := tdclient.Integration{
ID: integID, TeamID: id, Kind: req.Kind, Name: req.Name,
// Key/URL are only ever in *this* response -- never again,
// matching terdut-server's own one-time-show semantics
// (DESIGN.md §4.5) -- so what's stored for later GET/PATCH
// calls in this fake deliberately omits them too.
Key: fmt.Sprintf("webhook-key-%d", integID),
URL: fmt.Sprintf("https://terdut.example.invalid/api/integrations/webhook-key-%d/%s", integID, req.Kind),
}
if f.integrations[id] == nil {
f.integrations[id] = map[int64]tdclient.Integration{}
}
f.integrations[id][integID] = tdclient.Integration{ID: integID, TeamID: id, Kind: req.Kind, Name: req.Name}
writeJSON(w, http.StatusCreated, integ)
case strings.HasPrefix(rest, "/integrations/") && r.Method == http.MethodPatch:
integID, ok := parseTrailingID(rest, "/integrations/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
existing, exists := f.integrations[id][integID]
if !exists {
w.WriteHeader(http.StatusNotFound)
return
}
var req struct {
Name string `json:"name"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
existing.Name = req.Name
f.integrations[id][integID] = existing
w.WriteHeader(http.StatusNoContent)
case strings.HasPrefix(rest, "/integrations/") && r.Method == http.MethodDelete:
integID, ok := parseTrailingID(rest, "/integrations/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.integrations[id][integID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.integrations[id], integID)
f.integrationDelete[integID] = true
w.WriteHeader(http.StatusNoContent)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// deadmanSwitchFakeRequest mirrors tdclient's own (unexported)
// deadmanSwitchRequest -- the fake needs its own copy to decode the same
// wire shape without reaching across package boundaries for an internal type.
type deadmanSwitchFakeRequest struct {
Name string `json:"name,omitempty"`
Matcher string `json:"matcher"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Severity string `json:"severity"`
}
// parseTeamSubPath splits "/api/teams/{id}" from anything after it --
// "" for an exact match, "/oidc-groups", "/escalation", "/deadman/switches"
// or "/deadman/switches/{switchID}" otherwise. Doesn't itself validate the
// suffix; handleTeamSubPath's own switch does that.
func parseTeamSubPath(path string) (id int64, rest string, ok bool) {
const prefix = "/api/teams/"
if !strings.HasPrefix(path, prefix) {
return 0, "", false
}
trimmed := path[len(prefix):]
parts := strings.SplitN(trimmed, "/", 2)
parsedID, err := strconv.ParseInt(parts[0], 10, 64)
if err != nil {
return 0, "", false
}
if len(parts) == 1 {
return parsedID, "", true
}
return parsedID, "/" + parts[1], true
}
// parseTrailingID parses the numeric id after prefix within rest, e.g.
// parseTrailingID("/deadman/switches/7", "/deadman/switches/") -> 7, true.
func parseTrailingID(rest, prefix string) (id int64, ok bool) {
parsedID, err := strconv.ParseInt(strings.TrimPrefix(rest, prefix), 10, 64)
if err != nil {
return 0, false
}
return parsedID, true
}
func parseKeysPath(path string) (id int64, mintName string, ok bool) {
var parsedID int64
n, err := fmt.Sscanf(path, "/api/service-accounts/%d/keys", &parsedID)
if err != nil || n != 1 {
return 0, "", false
}
return parsedID, "minted", true
}
func writeJSON(w http.ResponseWriter, status int, v any) {
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(status)
_ = json.NewEncoder(w).Encode(v)
}
var _ = Describe("TerdutServer Controller", func() {
const operatorNamespace = "default"
@@ -466,9 +33,8 @@ var _ = Describe("TerdutServer Controller", func() {
BeforeEach(func() {
reconciler = &TerdutServerReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
Client: k8sClient,
Scheme: k8sClient.Scheme(),
}
name = fmt.Sprintf("test-server-%d-%d", GinkgoRandomSeed(), GinkgoParallelProcess())
objKey = types.NamespacedName{Name: name, Namespace: operatorNamespace}
@@ -481,9 +47,7 @@ var _ = Describe("TerdutServer Controller", func() {
_ = k8sClient.Update(ctx, srv)
_ = k8sClient.Delete(ctx, srv)
}
for _, n := range []string{checkpointSecretNameFor(name), credentialsSecretNameFor(name)} {
_ = k8sClient.Delete(ctx, &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: n, Namespace: operatorNamespace}})
}
_ = k8sClient.Delete(ctx, &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: name + "-operator-key", Namespace: operatorNamespace}})
})
dsnSpec := func() terdutv1alpha1.TerdutServerSpec {
@@ -520,127 +84,117 @@ var _ = Describe("TerdutServer Controller", func() {
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
c := meta.FindStatusCondition(srv.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(c).NotTo(BeNil(), "Ready condition should always be set after a reconcile past the finalizer-add pass")
Expect(c).NotTo(BeNil(), "Ready condition should always be set after a reconcile")
return *c
}
Describe("bring-your-own DSN path", func() {
It("reaches Ready through the full lifecycle: finalizer, Deployment/Service, wait, bootstrap", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
envOf := func(ctx context.Context) map[string]corev1.EnvVar {
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
out := map[string]corev1.EnvVar{}
for _, e := range deploy.Spec.Template.Spec.Containers[0].Env {
out[e.Name] = e
}
return out
}
Describe("bring-your-own DSN path", func() {
It("creates Deployment, Service and operator key, then reaches Ready once a replica is up", func(ctx SpecContext) {
createServer(ctx, dsnSpec())
// Pass 1: adds the finalizer and returns early.
reconcileOnce(ctx)
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(srv.Finalizers).To(ContainElement(finalizerName))
// Pass 2: creates Deployment + Service, waits for readiness.
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonWaitingForDeployment))
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
Expect(deploy.Spec.Template.Spec.Containers[0].Image).To(Equal("example.invalid/terdut-server:test"))
envNames := map[string]string{}
for _, e := range deploy.Spec.Template.Spec.Containers[0].Env {
envNames[e.Name] = e.Value
}
Expect(envNames).To(HaveKeyWithValue("TERDUT_DB_DSN", testDSN))
Expect(envNames).To(HaveKeyWithValue("TERDUT_OPERATOR_MODE", "true"))
env := envOf(ctx)
Expect(env["TERDUT_DB_DSN"].Value).To(Equal(testDSN))
Expect(env["TERDUT_OPERATOR_MODE"].Value).To(Equal("true"))
Expect(env).NotTo(HaveKey("TERDUT_DEADMAN_MATCHERS"), "dead man's switches are per team now")
var svc corev1.Service
Expect(k8sClient.Get(ctx, objKey, &svc)).To(Succeed())
Expect(svc.Spec.Ports[0].Port).To(Equal(int32(8080)))
// Simulate the Deployment becoming ready (envtest has no
// kubelet/deployment-controller to do this for real).
markDeploymentReady(ctx)
// Pass 3: bootstraps for real against the fake server.
reconcileOnce(ctx)
Expect(fake.bootstrapped).To(BeTrue())
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
ready := meta.FindStatusCondition(srv.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(ready.Status).To(Equal(metav1.ConditionTrue))
Expect(ready.Reason).To(Equal(terdutv1alpha1.ReasonAdopted))
Expect(srv.Status.CredentialsSecretRef).NotTo(BeNil())
Expect(srv.Status.CredentialsSecretRef.Key).To(Equal(credentialsSecretDataKey))
var credsSecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: srv.Status.CredentialsSecretRef.Name, Namespace: operatorNamespace,
}, &credsSecret)).To(Succeed())
Expect(string(credsSecret.Data[credentialsSecretDataKey])).To(Equal("instance-key-1-initial"))
// The checkpoint is cleaned up once the lasting credential is
// written (DESIGN.md §6 point 2).
var checkpoint corev1.Secret
err := k8sClient.Get(ctx, types.NamespacedName{Name: checkpointSecretName(srv), Namespace: operatorNamespace}, &checkpoint)
Expect(err).To(HaveOccurred())
})
})
Describe("the adopt-on-409 recovery path", func() {
It("mints a fresh key instead of erroring when the service account already exists", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
fake.bootstrapped = true // an earlier attempt already bootstrapped...
fake.nextID = 1
fake.accounts[serviceAccountName] = 1 // ...and already created the service account.
fake.keyMints[1] = 1
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, dsnSpec())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service, WaitingForDeployment
markDeploymentReady(ctx)
// The earlier attempt's checkpoint survived (that's how this
// reconcile can authenticate at all to recover).
Expect(writeOperatorSecret(ctx, k8sClient, operatorNamespace, checkpointSecretName(&terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: operatorNamespace},
}), "admin-key-raw")).To(Succeed())
reconcileOnce(ctx)
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(meta.FindStatusCondition(srv.Status.Conditions, terdutv1alpha1.ConditionReady).Status).To(Equal(metav1.ConditionTrue))
var credsSecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: srv.Status.CredentialsSecretRef.Name, Namespace: operatorNamespace,
}, &credsSecret)).To(Succeed())
// Minted fresh, not the (never-seen-by-this-reconcile) "initial"
// key from the account's original creation.
Expect(string(credsSecret.Data[credentialsSecretDataKey])).To(Equal("instance-key-1-mint2"))
ready := meta.FindStatusCondition(srv.Status.Conditions, terdutv1alpha1.ConditionReady)
Expect(ready.Status).To(Equal(metav1.ConditionTrue))
Expect(ready.Reason).To(Equal(terdutv1alpha1.ReasonAdopted))
Expect(srv.Status.CredentialsSecretRef).To(Equal(&terdutv1alpha1.SecretKeyRef{
Name: name + "-operator-key", Key: operatorKeyDataKey,
}))
})
})
Describe("the BootstrapStateLost path", func() {
It("fails closed when the server reports already-bootstrapped with no checkpoint to recover from", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
fake.bootstrap403 = true
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
Describe("the operator key", func() {
keySecret := func(ctx context.Context) corev1.Secret {
var s corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: name + "-operator-key", Namespace: operatorNamespace}, &s)).To(Succeed())
return s
}
It("is generated once, owned by the TerdutServer, and wired into the pod", func(ctx SpecContext) {
createServer(ctx, dsnSpec())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service
markDeploymentReady(ctx)
reconcileOnce(ctx)
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionFalse))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonBootstrapStateLost))
s := keySecret(ctx)
value := string(s.Data[operatorKeyDataKey])
Expect(len(value)).To(BeNumerically(">=", 32), "terdut-server refuses a shorter TERDUT_OPERATOR_KEY")
Expect(value).To(HavePrefix(operatorKeyPrefix))
Expect(s.OwnerReferences).To(ContainElement(HaveField("Name", name)))
env := envOf(ctx)
Expect(env["TERDUT_OPERATOR_KEY"].ValueFrom).NotTo(BeNil())
Expect(env["TERDUT_OPERATOR_KEY"].ValueFrom.SecretKeyRef.Name).To(Equal(name + "-operator-key"))
Expect(env["TERDUT_OPERATOR_KEY"].Value).To(BeEmpty(), "the key itself must never appear in the pod spec")
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
Expect(deploy.Spec.Template.Annotations).To(HaveKeyWithValue(operatorKeyHashAnnotation, operatorKeyHash(value)))
// A later reconcile keeps the same key: the running server was seeded with it.
reconcileOnce(ctx)
Expect(string(keySecret(ctx).Data[operatorKeyDataKey])).To(Equal(value))
})
It("rolls the Deployment when the Secret is replaced", func(ctx SpecContext) {
createServer(ctx, dsnSpec())
reconcileOnce(ctx)
first := string(keySecret(ctx).Data[operatorKeyDataKey])
s := keySecret(ctx)
Expect(k8sClient.Delete(ctx, &s)).To(Succeed())
reconcileOnce(ctx)
second := string(keySecret(ctx).Data[operatorKeyDataKey])
Expect(second).NotTo(Equal(first))
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
Expect(deploy.Spec.Template.Annotations).To(HaveKeyWithValue(operatorKeyHashAnnotation, operatorKeyHash(second)))
})
})
Describe("spec.oidc claims", func() {
It("passes the claim names and trustEmail through", func(ctx SpecContext) {
spec := dsnSpec()
spec.OIDC = terdutv1alpha1.OIDCSpec{
Enabled: true, Issuer: "https://sso.example.invalid", ClientID: "terdut",
UsernameClaim: "upn", EmailClaim: "mail", GroupsClaim: "roles", TrustEmail: true,
SessionMaxAge: "12h",
}
createServer(ctx, spec)
reconcileOnce(ctx)
env := envOf(ctx)
Expect(env["TERDUT_OIDC_USERNAME_CLAIM"].Value).To(Equal("upn"))
Expect(env["TERDUT_OIDC_EMAIL_CLAIM"].Value).To(Equal("mail"))
Expect(env["TERDUT_OIDC_GROUPS_CLAIM"].Value).To(Equal("roles"))
Expect(env["TERDUT_OIDC_TRUST_EMAIL"].Value).To(Equal("true"))
})
})
@@ -655,7 +209,6 @@ var _ = Describe("TerdutServer Controller", func() {
It("waits with reason PostgresClusterNotFound when the postgresql CR doesn't exist yet", func(ctx SpecContext) {
createServer(ctx, zalandoSpec("missing-cluster"))
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonPostgresClusterNotFound))
@@ -679,13 +232,7 @@ var _ = Describe("TerdutServer Controller", func() {
Expect(k8sClient.Create(ctx, zalandoSecret)).To(Succeed())
DeferCleanup(func() { _ = k8sClient.Delete(ctx, zalandoSecret) })
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, zalandoSpec(clusterName))
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service created with resolved DB env
var deploy appsv1.Deployment
@@ -705,32 +252,80 @@ var _ = Describe("TerdutServer Controller", func() {
})
})
Describe("deletion", func() {
It("removes the credentials and checkpoint Secrets and the finalizer", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
Describe("spec.pod", func() {
It("wires pod-level customization onto the right spot on the Deployment", func(ctx SpecContext) {
spec := dsnSpec()
qty := resource.MustParse("250m")
spec.Pod = terdutv1alpha1.PodSpec{
Resources: corev1.ResourceRequirements{Requests: corev1.ResourceList{corev1.ResourceCPU: qty}},
Tolerations: []corev1.Toleration{{Key: "dedicated", Operator: corev1.TolerationOpEqual, Value: "terdut", Effect: corev1.TaintEffectNoSchedule}},
ExtraEnv: []corev1.EnvVar{{Name: "EXTRA_FLAG", Value: "on"}},
ServiceAccountName: "terdut-server-custom",
ExtraVolumes: []corev1.Volume{{Name: "extra-ca", VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}},
ExtraVolumeMounts: []corev1.VolumeMount{{Name: "extra-ca", MountPath: "/etc/extra-ca"}},
}
createServer(ctx, spec)
reconcileOnce(ctx)
createServer(ctx, dsnSpec())
reconcileOnce(ctx)
reconcileOnce(ctx)
markDeploymentReady(ctx)
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
podSpec := deploy.Spec.Template.Spec
Expect(podSpec.Tolerations).To(ConsistOf(spec.Pod.Tolerations))
Expect(podSpec.ServiceAccountName).To(Equal("terdut-server-custom"))
Expect(podSpec.Volumes).To(ConsistOf(spec.Pod.ExtraVolumes))
main := podSpec.Containers[0]
Expect(main.Name).To(Equal("terdut-server"))
Expect(main.Resources).To(Equal(spec.Pod.Resources))
Expect(main.VolumeMounts).To(ConsistOf(spec.Pod.ExtraVolumeMounts))
Expect(main.Env).To(ContainElement(corev1.EnvVar{Name: "EXTRA_FLAG", Value: "on"}))
initContainer := podSpec.InitContainers[0]
Expect(initContainer.Name).To(Equal("wait-for-postgres"))
Expect(initContainer.VolumeMounts).To(BeEmpty(), "extraVolumeMounts must not leak onto wait-for-postgres")
})
})
Describe("spec.pod.disruptionBudget", func() {
It("creates an owned PodDisruptionBudget when set, and deletes it once cleared", func(ctx SpecContext) {
spec := dsnSpec()
minAvail := intstr.FromInt32(1)
spec.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{MinAvailable: &minAvail}
createServer(ctx, spec)
reconcileOnce(ctx)
var pdb policyv1.PodDisruptionBudget
Expect(k8sClient.Get(ctx, objKey, &pdb)).To(Succeed())
Expect(pdb.Spec.Selector.MatchLabels).To(Equal(labelsFor(&terdutv1alpha1.TerdutServer{ObjectMeta: metav1.ObjectMeta{Name: name}})))
Expect(pdb.Spec.MinAvailable).To(Equal(&minAvail))
Expect(pdb.OwnerReferences).To(ContainElement(HaveField("Name", name)))
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
credsName := srv.Status.CredentialsSecretRef.Name
srv.Spec.Pod.DisruptionBudget = nil
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Delete(ctx, srv)).To(Succeed())
reconcileOnce(ctx) // runs the finalizer
err := k8sClient.Get(ctx, objKey, &policyv1.PodDisruptionBudget{})
Expect(apierrors.IsNotFound(err)).To(BeTrue(), "PodDisruptionBudget should be deleted once spec.pod.disruptionBudget is cleared")
})
err := k8sClient.Get(ctx, objKey, srv)
Expect(err).To(HaveOccurred(), "the TerdutServer itself should be gone once the finalizer clears")
It("rejects both minAvailable and maxUnavailable set together, and neither set", func(ctx SpecContext) {
bothSet := dsnSpec()
minAvail, maxUnavail := intstr.FromInt32(1), intstr.FromInt32(1)
bothSet.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{MinAvailable: &minAvail, MaxUnavailable: &maxUnavail}
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: name + "-both", Namespace: operatorNamespace},
Spec: bothSet,
})).To(HaveOccurred())
var leftover corev1.Secret
err = k8sClient.Get(ctx, types.NamespacedName{Name: credsName, Namespace: operatorNamespace}, &leftover)
Expect(err).To(HaveOccurred(), "the credentials Secret should have been cleaned up by the finalizer")
neitherSet := dsnSpec()
neitherSet.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{}
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: name + "-neither", Namespace: operatorNamespace},
Spec: neitherSet,
})).To(HaveOccurred())
})
})
@@ -743,13 +338,3 @@ var _ = Describe("TerdutServer Controller", func() {
})
})
})
// checkpointSecretNameFor/credentialsSecretNameFor let AfterEach clean up
// without needing a live TerdutServer object (it may already be gone by
// then in the deletion test).
func checkpointSecretNameFor(name string) string {
return fmt.Sprintf("default.%s-bootstrap-admin", name)
}
func credentialsSecretNameFor(name string) string {
return fmt.Sprintf("default.%s-instance-credentials", name)
}
@@ -0,0 +1,101 @@
package controller
import (
"context"
"crypto/rand"
"crypto/sha256"
"encoding/hex"
"fmt"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
)
// operatorKeyDataKey is the data key of the operator key Secret.
const operatorKeyDataKey = "token"
// operatorKeyPrefix marks the key as a service-account credential in logs, the
// way terdut-server's own generated ones are.
const operatorKeyPrefix = "tdsa_"
// operatorKeySecretName is the Secret holding a TerdutServer's operator key. It
// lives beside the Deployment because the pod mounts it (TERDUT_OPERATOR_KEY)
// and a pod can only reference Secrets of its own namespace; it is owned by the
// TerdutServer, so deleting that deletes the key and a recreated one starts
// with a fresh one the server re-seeds on its next start.
func operatorKeySecretName(srv *terdutv1alpha1.TerdutServer) string {
return srv.Name + "-operator-key"
}
// reconcileOperatorKey returns this server's operator key, generating the Secret
// on first use. An existing Secret is never overwritten: the running server was
// seeded with that value, and replacing it would lock the operator out until
// every pod restarted.
func (r *TerdutServerReconciler) reconcileOperatorKey(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (string, error) {
var secret corev1.Secret
key := client.ObjectKey{Namespace: srv.Namespace, Name: operatorKeySecretName(srv)}
err := r.Get(ctx, key, &secret)
if err == nil {
if v := string(secret.Data[operatorKeyDataKey]); v != "" {
return v, nil
}
// Present but empty or foreign: treat as ours to fill rather than
// failing every reconcile on it.
} else if !apierrors.IsNotFound(err) {
return "", err
}
raw := make([]byte, 24)
if _, err := rand.Read(raw); err != nil {
return "", fmt.Errorf("generating operator key: %w", err)
}
value := operatorKeyPrefix + hex.EncodeToString(raw)
secret = corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: key.Name, Namespace: key.Namespace},
Data: map[string][]byte{operatorKeyDataKey: []byte(value)},
}
if err := controllerutil.SetControllerReference(srv, &secret, r.Scheme); err != nil {
return "", err
}
if err := r.Create(ctx, &secret); err != nil {
if apierrors.IsAlreadyExists(err) {
// Lost a race with a stale cache read: the next reconcile finds it.
return "", fmt.Errorf("operator key Secret %s appeared during creation; retrying", key.Name)
}
return "", err
}
return value, nil
}
// operatorKeyHash is a short digest of the key, stamped on the pod template so
// a replaced Secret rolls the Deployment: the server only reads the env var at
// start.
func operatorKeyHash(value string) string {
sum := sha256.Sum256([]byte(value))
return hex.EncodeToString(sum[:8])
}
// operatorKeyRef is the reference status.credentialsSecretRef reports.
func operatorKeyRef(srv *terdutv1alpha1.TerdutServer) *terdutv1alpha1.SecretKeyRef {
return &terdutv1alpha1.SecretKeyRef{Name: operatorKeySecretName(srv), Key: operatorKeyDataKey}
}
// readOperatorKey reads a TerdutServer's key from its own namespace, for the
// controllers that call the server's API.
func readOperatorKey(ctx context.Context, c client.Client, srv *terdutv1alpha1.TerdutServer) (string, error) {
var secret corev1.Secret
if err := c.Get(ctx, client.ObjectKey{Namespace: srv.Namespace, Name: operatorKeySecretName(srv)}, &secret); err != nil {
return "", err
}
v := string(secret.Data[operatorKeyDataKey])
if v == "" {
return "", fmt.Errorf("operator key Secret %s/%s is empty", srv.Namespace, secret.Name)
}
return v, nil
}
+91 -29
View File
@@ -3,6 +3,7 @@ package controller
import (
"context"
"fmt"
"maps"
"strings"
appsv1 "k8s.io/api/apps/v1"
@@ -31,27 +32,48 @@ func labelsFor(srv *terdutv1alpha1.TerdutServer) map[string]string {
// object (not just the desired one) so the caller can check
// status.readyReplicas.
func (r *TerdutServerReconciler) reconcileDeployment(
ctx context.Context, srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar,
ctx context.Context, srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar, keyHash string,
) (*appsv1.Deployment, error) {
deploy := &appsv1.Deployment{ObjectMeta: metav1.ObjectMeta{Name: srv.Name, Namespace: srv.Namespace}}
_, err := controllerutil.CreateOrUpdate(ctx, r.Client, deploy, func() error {
replicas := srv.Spec.Replicas
if replicas == 0 {
replicas = 1
// Only reachable for a TerdutServer stored before the
// +kubebuilder:default=2 marker existed -- the API server's own
// CRD defaulting fills this in for anything created or updated
// through it, so a fresh zero value here means a pre-existing
// object that predates the default, not a deliberate "none"
// (there is no way to request zero replicas).
replicas = 2
}
labels := labelsFor(srv)
deploy.Spec.Replicas = &replicas
deploy.Spec.Selector = &metav1.LabelSelector{MatchLabels: labels}
// Recreate, not RollingUpdate: the sweeper and the notifier are
// unsynchronised singletons inside terdut-server, and two replicas
// overlapping during a rollout would both page for the same
// incident (matches the chart's own deployment.yaml comment).
deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RecreateDeploymentStrategyType}
// RollingUpdate, not Recreate: terdut-server v0.36.0 put the sweeper,
// the notifier and the migration runner each behind a Postgres
// advisory lock, and gave incident creation its own conflict
// resolution, so two replicas overlapping during a rollout no longer
// double-page, race a migration, or drop a webhook payload (matches
// the chart's own deployment.yaml comment). No explicit
// maxUnavailable/maxSurge: left at the 25%/25% default, which rounds
// to 0/1 at the default replicas: 2 -- already zero-downtime.
deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RollingUpdateDeploymentStrategyType}
pod := srv.Spec.Pod
deploy.Spec.Template = corev1.PodTemplateSpec{
ObjectMeta: metav1.ObjectMeta{Labels: labels},
// The controller's own annotation wins over a same-named user one.
ObjectMeta: metav1.ObjectMeta{Labels: labels, Annotations: podAnnotations(pod.Annotations, keyHash)},
Spec: corev1.PodSpec{
EnableServiceLinks: new(false),
EnableServiceLinks: new(false),
NodeSelector: pod.NodeSelector,
Tolerations: pod.Tolerations,
Affinity: pod.Affinity,
TopologySpreadConstraints: pod.TopologySpreadConstraints,
SecurityContext: pod.SecurityContext,
ServiceAccountName: pod.ServiceAccountName,
ImagePullSecrets: pod.ImagePullSecrets,
InitContainers: []corev1.Container{waitForPostgresContainer(dbEnv)},
Volumes: pod.ExtraVolumes,
Containers: []corev1.Container{{
Name: "terdut-server",
Image: fmt.Sprintf("%s:%s", srv.Spec.Image.Repository, srv.Spec.Image.Tag),
@@ -60,9 +82,13 @@ func (r *TerdutServerReconciler) reconcileDeployment(
ContainerPort: servicePort(srv),
Protocol: corev1.ProtocolTCP,
}},
Env: buildEnv(srv, dbEnv),
LivenessProbe: healthzProbe(),
ReadinessProbe: healthzProbe(),
Env: buildEnv(srv, dbEnv),
EnvFrom: pod.ExtraEnvFrom,
VolumeMounts: pod.ExtraVolumeMounts,
Resources: pod.Resources,
SecurityContext: pod.ContainerSecurityContext,
LivenessProbe: healthzProbe(),
ReadinessProbe: healthzProbe(),
}},
},
}
@@ -97,6 +123,17 @@ func (r *TerdutServerReconciler) reconcileService(ctx context.Context, srv *terd
return err
}
// operatorKeyHashAnnotation records which operator key the pods were started
// with, so replacing the key's Secret rolls them.
const operatorKeyHashAnnotation = "terdut.ryuvia.com/operator-key-hash"
func podAnnotations(user map[string]string, keyHash string) map[string]string {
out := make(map[string]string, len(user)+1)
maps.Copy(out, user)
out[operatorKeyHashAnnotation] = keyHash
return out
}
func servicePort(srv *terdutv1alpha1.TerdutServer) int32 {
if srv.Spec.Networking.ServicePort == 0 {
return 8080
@@ -104,6 +141,36 @@ func servicePort(srv *terdutv1alpha1.TerdutServer) int32 {
return srv.Spec.Networking.ServicePort
}
// waitForPostgresContainer blocks the main container from starting until
// Postgres accepts connections, matching charts/terdut-server's own
// deployment.yaml template as of v0.33.2 (that repo's CLAUDE.md/release
// notes) -- that chart grew this the moment this exact Deployment, created
// by this controller, crash-looped a few times against a from-scratch
// postgres-operator cluster still doing initdb and Patroni leader election:
// terdut-server's own ping-retry budget on startup (internal/db/db.go) is
// sized for a much shorter, different race (NetworkPolicy propagation, a
// few seconds), not for genuine first-time cluster creation, so it
// exhausted and the process exited before ever binding its HTTP port -- a
// startupProbe cannot help there, since the crash happens before there is
// anything to probe.
//
// Reuses dbEnv unchanged: both of resolveDatabaseEnv's paths put
// TERDUT_DB_DSN first (terdutserver_database.go), so it's already exactly
// what pg_isready needs, and pg_isready needs no credentials -- it reports
// PQPING_OK on anything that amounts to a Postgres backend answering,
// including an auth challenge -- so including dbEnv's optional PGPASSWORD
// here too is harmless rather than load-bearing.
func waitForPostgresContainer(dbEnv []corev1.EnvVar) corev1.Container {
return corev1.Container{
Name: "wait-for-postgres",
Image: "postgres:17-alpine",
Env: dbEnv,
Command: []string{"sh", "-c",
`until pg_isready -d "$TERDUT_DB_DSN"; do echo "wait-for-postgres: not ready yet, retrying in 2s"; sleep 2; done`,
},
}
}
func healthzProbe() *corev1.Probe {
return &corev1.Probe{
ProbeHandler: corev1.ProbeHandler{
@@ -120,7 +187,10 @@ func healthzProbe() *corev1.Probe {
// field-for-field (confirmed against that source, not reconstructed from
// DESIGN.md's illustrative YAML alone) — dbEnv (TERDUT_DB_DSN, optionally
// PGPASSWORD) comes from resolveDatabaseEnv, since which of §8's two paths
// produced it doesn't matter past this point.
// produced it doesn't matter past this point. spec.pod.extraEnv is appended
// last, after every fixed var -- this is the one place that owns "what env
// this container gets," so the escape hatch lives here rather than being
// appended separately in reconcileDeployment.
func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.EnvVar {
env := []corev1.EnvVar{{Name: "TERDUT_ADDR", Value: fmt.Sprintf(":%d", servicePort(srv))}}
env = append(env, dbEnv...)
@@ -131,13 +201,6 @@ func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.
corev1.EnvVar{Name: "TERDUT_ARCHIVE_AFTER", Value: sweeper.ArchiveAfter},
)
deadman := srv.Spec.Deadman
env = append(env,
corev1.EnvVar{Name: "TERDUT_DEADMAN_MATCHERS", Value: deadman.Matchers},
corev1.EnvVar{Name: "TERDUT_DEADMAN_TIMEOUT", Value: deadman.Timeout},
corev1.EnvVar{Name: "TERDUT_DEADMAN_SEVERITY", Value: deadman.Severity},
)
notify := srv.Spec.Notify
if notify.NtfyURL != "" {
env = append(env,
@@ -167,6 +230,9 @@ func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.
// here on" -- which is unconditionally true for anything this
// operator creates).
corev1.EnvVar{Name: "TERDUT_OPERATOR_MODE", Value: "true"},
// The credential this operator calls the API with; the server
// creates or re-keys its instance-scoped account from it at start.
corev1.EnvVar{Name: "TERDUT_OPERATOR_KEY", ValueFrom: secretEnvSource(operatorKeyRef(srv))},
)
if oidc := srv.Spec.OIDC; oidc.Enabled {
@@ -175,14 +241,10 @@ func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.
corev1.EnvVar{Name: "TERDUT_OIDC_CLIENT_ID", Value: oidc.ClientID},
corev1.EnvVar{Name: "TERDUT_OIDC_NAME", Value: oidc.Name},
corev1.EnvVar{Name: "TERDUT_OIDC_SCOPES", Value: oidc.Scopes},
// terdut-server's own defaults for the claims/trust-email knobs
// the chart exposes but DESIGN.md's spec doesn't (§4.1's doc
// comment on OIDCSpec) -- not configurable here, not an
// oversight.
corev1.EnvVar{Name: "TERDUT_OIDC_USERNAME_CLAIM", Value: "preferred_username"},
corev1.EnvVar{Name: "TERDUT_OIDC_EMAIL_CLAIM", Value: "email"},
corev1.EnvVar{Name: "TERDUT_OIDC_GROUPS_CLAIM", Value: "groups"},
corev1.EnvVar{Name: "TERDUT_OIDC_TRUST_EMAIL", Value: "false"},
corev1.EnvVar{Name: "TERDUT_OIDC_USERNAME_CLAIM", Value: oidc.UsernameClaim},
corev1.EnvVar{Name: "TERDUT_OIDC_EMAIL_CLAIM", Value: oidc.EmailClaim},
corev1.EnvVar{Name: "TERDUT_OIDC_GROUPS_CLAIM", Value: oidc.GroupsClaim},
corev1.EnvVar{Name: "TERDUT_OIDC_TRUST_EMAIL", Value: boolString(oidc.TrustEmail)},
corev1.EnvVar{Name: "TERDUT_OIDC_SESSION_MAX_AGE", Value: oidc.SessionMaxAge},
)
if oidc.ClientSecretRef != nil {
@@ -199,7 +261,7 @@ func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.
}
}
return env
return append(env, srv.Spec.Pod.ExtraEnv...)
}
func secretEnvSource(ref *terdutv1alpha1.SecretKeyRef) *corev1.EnvVarSource {
+38
View File
@@ -0,0 +1,38 @@
package controller
import (
"context"
policyv1 "k8s.io/api/policy/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
)
// reconcilePodDisruptionBudget creates/updates the PodDisruptionBudget
// spec.pod.disruptionBudget asks for, or deletes a previously-created one
// when the field has been cleared -- the one conditionally-created child
// object in this controller (Deployment/Service are unconditional). Owned
// by srv, same plain-OwnerReference shape as Deployment/Service (DESIGN.md
// §7): same namespace, GC handles it, no finalizer needed.
func (r *TerdutServerReconciler) reconcilePodDisruptionBudget(ctx context.Context, srv *terdutv1alpha1.TerdutServer) error {
pdb := &policyv1.PodDisruptionBudget{ObjectMeta: metav1.ObjectMeta{Name: srv.Name, Namespace: srv.Namespace}}
spec := srv.Spec.Pod.DisruptionBudget
if spec == nil {
if err := r.Delete(ctx, pdb); err != nil && !apierrors.IsNotFound(err) {
return err
}
return nil
}
_, err := controllerutil.CreateOrUpdate(ctx, r.Client, pdb, func() error {
pdb.Spec.Selector = &metav1.LabelSelector{MatchLabels: labelsFor(srv)}
pdb.Spec.MinAvailable = spec.MinAvailable
pdb.Spec.MaxUnavailable = spec.MaxUnavailable
return controllerutil.SetControllerReference(srv, pdb, r.Scheme)
})
return err
}
@@ -1,84 +0,0 @@
package controller
import (
"context"
"errors"
"fmt"
"net/http"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// createOrAdoptTeam calls POST /api/teams with the TerdutServer's
// instance-scoped credential, or, if an earlier interrupted attempt
// already created this name (409), adopts it via GET /api/teams?name=
// (TEAM-LOOKUP.md) rather than treating the conflict as an error --
// DESIGN.md §5's general adopt-on-conflict rule, the same shape
// TerdutServer's own bootstrap flow uses for minting its instance account.
func (r *TerdutTeamReconciler) createOrAdoptTeam(ctx context.Context, team *terdutv1alpha1.TerdutTeam, instanceClient *tdclient.Client) error {
created, err := instanceClient.CreateTeam(ctx, team.Spec.DisplayName)
if err == nil {
team.Status.TeamID = created.ID
return nil
}
statusErr, ok := errors.AsType[*tdclient.StatusError](err)
if !ok || statusErr.Code != http.StatusConflict {
return fmt.Errorf("POST /api/teams: %w", err)
}
found, err := instanceClient.GetTeamByName(ctx, team.Spec.DisplayName)
if err != nil {
return fmt.Errorf("GET /api/teams?name=%s (adopting after 409): %w", team.Spec.DisplayName, err)
}
if found == nil {
// Genuinely pathological, not just a narrow crash window: the name
// was taken a moment ago and isn't now. Surfaced as a plain error
// (standard requeue-with-backoff) rather than a dedicated
// condition -- there's no documented recovery to point at that
// differs from "try again".
return fmt.Errorf("POST /api/teams 409'd for %q but GET found nothing", team.Spec.DisplayName)
}
team.Status.TeamID = found.ID
return nil
}
// mintTeamCredential mints this team's own team-scoped service account,
// using the TerdutServer's instance-scoped credential (DESIGN.md §6 point
// 3: an instance-scoped caller may do this against any team). Adopts via
// GET+mint-new-key on a 409, the same pattern TerdutServer's own bootstrap
// flow uses.
func (r *TerdutTeamReconciler) mintTeamCredential(ctx context.Context, team *terdutv1alpha1.TerdutTeam, instanceClient *tdclient.Client) error {
saName := teamServiceAccountName(team)
result, err := instanceClient.CreateTeamServiceAccount(ctx, saName, team.Status.TeamID)
var key string
if err == nil {
key = result.Key.Key
} else {
statusErr, ok := errors.AsType[*tdclient.StatusError](err)
if !ok || statusErr.Code != http.StatusConflict {
return fmt.Errorf("POST /api/service-accounts (team scope): %w", err)
}
sa, err := instanceClient.GetServiceAccountByName(ctx, saName)
if err != nil {
return fmt.Errorf("GET /api/service-accounts?name=%s (adopting after 409): %w", saName, err)
}
if sa == nil {
return fmt.Errorf("POST /api/service-accounts 409'd for %q but GET found nothing", saName)
}
minted, err := instanceClient.CreateServiceAccountKey(ctx, sa.ID, "initial")
if err != nil {
return fmt.Errorf("POST /api/service-accounts/%d/keys (adopting after 409): %w", sa.ID, err)
}
key = minted.Key
}
credsName := teamCredentialsSecretName(team)
if err := writeOperatorSecret(ctx, r.Client, r.OperatorNamespace, credsName, key); err != nil {
return err
}
team.Status.CredentialsSecretRef = &terdutv1alpha1.SecretKeyRef{Name: credsName, Key: credentialsSecretDataKey}
return nil
}
+127
View File
@@ -0,0 +1,127 @@
package controller
import (
"context"
"errors"
"fmt"
"net/http"
"strings"
"time"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// reconcileEscalation applies spec.escalation as the team's whole ladder, or
// clears it when the field is absent (an empty ladder is the server's "none").
// The second return is an expected, reportable condition (a bad duration, an
// unknown user), the third a failure to retry.
func (r *TerdutTeamReconciler) reconcileEscalation(
ctx context.Context, tc *tdclient.Client, team *terdutv1alpha1.TerdutTeam,
) (*teamError, error) {
body, err := buildEscalationRequest(team.Spec.Escalation)
if err != nil {
return &teamError{reason: terdutv1alpha1.ReasonInvalidSpec, message: err.Error()}, nil
}
if err := tc.SetEscalation(ctx, team.Status.TeamID, body); err != nil {
// The server resolves usernames; one it does not know is a state to
// wait out (the person may be created later), not a failure.
if se, ok := errors.AsType[*tdclient.StatusError](err); ok &&
se.Code == http.StatusBadRequest && strings.HasPrefix(se.Message, "unknown user") {
return &teamError{reason: terdutv1alpha1.ReasonUnknownUser, message: se.Message}, nil
}
return nil, fmt.Errorf("PUT /api/teams/%d/escalation: %w", team.Status.TeamID, err)
}
return nil, nil
}
func buildEscalationRequest(spec *terdutv1alpha1.EscalationSpec) (tdclient.SetEscalationRequest, error) {
if spec == nil {
return tdclient.SetEscalationRequest{Levels: []tdclient.EscalationLevelRequest{}}, nil
}
levels := make([]tdclient.EscalationLevelRequest, len(spec.Levels))
for i, lvl := range spec.Levels {
timeout, err := time.ParseDuration(lvl.Timeout)
if err != nil || timeout <= 0 {
return tdclient.SetEscalationRequest{}, fmt.Errorf("spec.escalation.levels[%d].timeout %q is not a positive duration", i, lvl.Timeout)
}
targets := make([]tdclient.EscalationTargetRequest, len(lvl.Targets))
for j, t := range lvl.Targets {
targets[j] = tdclient.EscalationTargetRequest{Kind: string(t.Kind), Username: t.Username}
}
levels[i] = tdclient.EscalationLevelRequest{
Position: int64(i + 1),
TimeoutSeconds: int64(timeout.Seconds()),
Targets: targets,
}
}
return tdclient.SetEscalationRequest{
RepeatCount: spec.RepeatCount,
FallbackTopic: spec.FallbackTopic,
Levels: levels,
}, nil
}
// reconcileDeadmanSwitches makes the team's switches on the server exactly
// spec.deadmanSwitches, matched by name (unique per team server-side): create
// what is missing, update what differs, delete what is not listed. In operator
// mode nobody else can add one, so anything extra is leftover to remove.
func (r *TerdutTeamReconciler) reconcileDeadmanSwitches(
ctx context.Context, tc *tdclient.Client, team *terdutv1alpha1.TerdutTeam,
) (*teamError, error) {
type want struct {
spec terdutv1alpha1.DeadmanSwitchSpec
seconds int64
}
desired := make(map[string]want, len(team.Spec.DeadmanSwitches))
for _, sw := range team.Spec.DeadmanSwitches {
timeout, err := time.ParseDuration(sw.Timeout)
if err != nil || timeout <= 0 {
return &teamError{
reason: terdutv1alpha1.ReasonInvalidSpec,
message: fmt.Sprintf("spec.deadmanSwitches[%q].timeout %q is not a positive duration", sw.Name, sw.Timeout),
}, nil
}
desired[sw.Name] = want{spec: sw, seconds: int64(timeout.Seconds())}
}
existing, err := tc.ListDeadmanSwitches(ctx, team.Status.TeamID)
if err != nil {
return nil, fmt.Errorf("GET /api/teams/%d/deadman/switches: %w", team.Status.TeamID, err)
}
seen := make(map[string]bool, len(existing))
for _, have := range existing {
w, keep := desired[have.Name]
if !keep {
if err := tc.DeleteDeadmanSwitch(ctx, team.Status.TeamID, have.ID); err != nil {
return nil, fmt.Errorf("DELETE deadman switch %q: %w", have.Name, err)
}
continue
}
seen[have.Name] = true
severity := severityOrDefault(w.spec.Severity)
if have.Matcher == w.spec.Matcher && have.TimeoutSeconds == w.seconds && have.Severity == severity {
continue
}
if err := tc.UpdateDeadmanSwitch(ctx, team.Status.TeamID, have.ID, w.spec.Name, w.spec.Matcher, w.seconds, severity); err != nil {
return nil, fmt.Errorf("PUT deadman switch %q: %w", have.Name, err)
}
}
for _, sw := range team.Spec.DeadmanSwitches {
if seen[sw.Name] {
continue
}
w := desired[sw.Name]
if _, err := tc.CreateDeadmanSwitch(ctx, team.Status.TeamID, sw.Name, sw.Matcher, w.seconds, severityOrDefault(sw.Severity)); err != nil {
return nil, fmt.Errorf("POST deadman switch %q: %w", sw.Name, err)
}
}
return nil, nil
}
func severityOrDefault(s string) string {
if s == "" {
return "critical"
}
return s
}
+132 -118
View File
@@ -5,7 +5,6 @@ import (
"errors"
"fmt"
"net/http"
"time"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
@@ -15,33 +14,25 @@ import (
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
"sigs.k8s.io/controller-runtime/pkg/handler"
logf "sigs.k8s.io/controller-runtime/pkg/log"
"sigs.k8s.io/controller-runtime/pkg/reconcile"
"sigs.k8s.io/controller-runtime/pkg/recorder"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// teamFinalizerName cleans up the team-scoped credentials Secret this
// controller generates in the operator's own namespace on delete.
// teamFinalizerName deletes the team on the server before the CR goes.
const teamFinalizerName = "terdut.ryuvia.com/terdutteam"
// TerdutTeamReconciler reconciles a TerdutTeam object.
//
// Every child resolves its own teamRef/serverRef independently and never
// chains up through another controller (DESIGN.md §5) — this one talks
// directly to the TerdutServer it references and to terdut-server's API,
// never to TerdutServerReconciler.
// TerdutTeamReconciler reconciles a TerdutTeam: the team itself plus its
// escalation ladder and dead man's switches, all through the referenced
// TerdutServer's operator key.
type TerdutTeamReconciler struct {
client.Client
Scheme *runtime.Scheme
// OperatorNamespace is where every credentials Secret this controller
// reads (the referenced TerdutServer's) or writes (this team's own)
// lives (DESIGN.md §6) — never a TerdutTeam's or TerdutServer's own
// namespace.
OperatorNamespace string
Recorder recorder.EventRecorder
NewClient func(endpoint string) *tdclient.Client
}
@@ -50,10 +41,16 @@ type TerdutTeamReconciler struct {
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutteams/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutteams/finalizers,verbs=update
// +kubebuilder:rbac:groups=terdut.ryuvia.com,resources=terdutservers,verbs=get;list;watch
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch
// +kubebuilder:rbac:groups="",resources=namespaces,verbs=get;list;watch
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
// teamExternalID is the identity this CR gives its team on the server, so the
// team is found by who owns it and not by its (editable, global) display name.
func teamExternalID(team *terdutv1alpha1.TerdutTeam) string {
return team.Namespace + "/" + team.Name
}
func (r *TerdutTeamReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
log := logf.FromContext(ctx)
@@ -79,64 +76,56 @@ func (r *TerdutTeamReconciler) Reconcile(ctx context.Context, req ctrl.Request)
srv, resolveErr := r.resolveServerRef(ctx, &team)
if resolveErr != nil {
return r.setTeamNotReady(ctx, &team, resolveErr.reason, resolveErr.message, waitInterval)
return r.setTeamNotReady(ctx, &team, resolveErr.reason, resolveErr.message)
}
if srv.Status.CredentialsSecretRef == nil {
if !meta.IsStatusConditionTrue(srv.Status.Conditions, terdutv1alpha1.ConditionReady) {
return r.setTeamNotReady(ctx, &team, terdutv1alpha1.ReasonWaitingForServer,
fmt.Sprintf("TerdutServer %q is not Bootstrapped yet", srv.Name), waitInterval)
fmt.Sprintf("TerdutServer %q is not Ready yet", srv.Name))
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
instanceKey, err := readOperatorSecret(ctx, r.Client, r.OperatorNamespace, srv.Status.CredentialsSecretRef)
tc, err := r.serverClient(ctx, srv)
if err != nil {
return ctrl.Result{}, err
}
endpoint := serviceURL(srv)
team.Status.ServerEndpoint = endpoint
instanceClient := newClient(endpoint).WithToken(instanceKey)
if team.Status.TeamID == 0 {
if err := r.createOrAdoptTeam(ctx, &team, instanceClient); err != nil {
return ctrl.Result{}, err
}
}
if team.Status.CredentialsSecretRef == nil {
if err := r.mintTeamCredential(ctx, &team, instanceClient); err != nil {
return ctrl.Result{}, err
}
}
teamKey, err := readOperatorSecret(ctx, r.Client, r.OperatorNamespace, team.Status.CredentialsSecretRef)
// Idempotent on the external id: the same team comes back whether this is
// the first reconcile, a retry after a lost status write, or a repair after
// somebody deleted the team behind our back.
created, err := tc.CreateTeam(ctx, team.Spec.DisplayName, teamExternalID(&team))
if err != nil {
return ctrl.Result{}, err
if se, ok := errors.AsType[*tdclient.StatusError](err); ok && se.Code == http.StatusConflict {
return r.setTeamNotReady(ctx, &team, terdutv1alpha1.ReasonTeamNameTaken, nameTakenMessage(&team))
}
return ctrl.Result{}, fmt.Errorf("POST /api/teams: %w", err)
}
teamClient := newClient(endpoint).WithToken(teamKey)
team.Status.TeamID = created.ID
// Owner-gated on terdut-server, so this always runs with the
// team-scoped credential just minted above, never the instance-scoped
// one used to create the team (internal/api/teams.go's requireTeamOwner
// has no branch for an instance-scoped service account, confirmed
// against source). Applied unconditionally rather than diffed against a
// stored "last-applied" value: both calls are idempotent PUTs of the
// whole resource, the same "cheap because it's small" reasoning §5
// already applies to the escalation policy's whole-policy PUT.
if err := teamClient.RenameTeam(ctx, team.Status.TeamID, team.Spec.DisplayName); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d: %w", team.Status.TeamID, err)
if err := tc.RenameTeam(ctx, created.ID, team.Spec.DisplayName); err != nil {
if se, ok := errors.AsType[*tdclient.StatusError](err); ok && se.Code == http.StatusConflict {
return r.setTeamNotReady(ctx, &team, terdutv1alpha1.ReasonTeamNameTaken, nameTakenMessage(&team))
}
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d: %w", created.ID, err)
}
if err := teamClient.SetTeamOIDCGroups(ctx, team.Status.TeamID, team.Spec.OIDC.MemberGroup, team.Spec.OIDC.OwnerGroup); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/oidc-groups: %w", team.Status.TeamID, err)
if err := tc.SetTeamOIDCGroups(ctx, created.ID, team.Spec.OIDC.MemberGroup, team.Spec.OIDC.OwnerGroup); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/oidc-groups: %w", created.ID, err)
}
if condErr, err := r.reconcileEscalation(ctx, tc, &team); err != nil {
return ctrl.Result{}, err
} else if condErr != nil {
return r.setTeamNotReady(ctx, &team, condErr.reason, condErr.message)
}
if condErr, err := r.reconcileDeadmanSwitches(ctx, tc, &team); err != nil {
return ctrl.Result{}, err
} else if condErr != nil {
return r.setTeamNotReady(ctx, &team, condErr.reason, condErr.message)
}
meta.SetStatusCondition(&team.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
Status: metav1.ConditionTrue,
Reason: terdutv1alpha1.ReasonTeamAdopted,
Message: fmt.Sprintf("team %d ready, credentials in Secret %q", team.Status.TeamID, team.Status.CredentialsSecretRef.Name),
Message: fmt.Sprintf("team %d applied", team.Status.TeamID),
})
team.Status.ObservedGeneration = team.Generation
if err := r.Status().Update(ctx, &team); err != nil {
@@ -144,13 +133,31 @@ func (r *TerdutTeamReconciler) Reconcile(ctx context.Context, req ctrl.Request)
}
if r.Recorder != nil {
r.Recorder.Eventf(&team, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonTeamAdopted, terdutv1alpha1.ReasonTeamAdopted,
"team ready")
"team applied")
}
log.Info("TerdutTeam ready", "name", team.Name, "teamID", team.Status.TeamID)
return ctrl.Result{RequeueAfter: resyncInterval}, nil
}
func nameTakenMessage(team *terdutv1alpha1.TerdutTeam) string {
return fmt.Sprintf("the server already has a different team named %q; choose another spec.displayName, "+
"or remove that team", team.Spec.DisplayName)
}
// serverClient builds a client for srv, authenticated with its operator key.
func (r *TerdutTeamReconciler) serverClient(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (*tdclient.Client, error) {
key, err := readOperatorKey(ctx, r.Client, srv)
if err != nil {
return nil, err
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
return newClient(serviceURL(srv)).WithToken(key), nil
}
// teamError carries a condition reason/message, the same role databaseError
// plays for TerdutServer: an expected, requeue-and-retry outcome, not a
// reconcile failure.
@@ -165,6 +172,21 @@ func (e *teamError) Error() string { return e.message }
// consent check (DESIGN.md §4.6) when serverRef.namespace differs from this
// TerdutTeam's own.
func (r *TerdutTeamReconciler) resolveServerRef(ctx context.Context, team *terdutv1alpha1.TerdutTeam) (*terdutv1alpha1.TerdutServer, *teamError) {
srv, ns, getErr := r.getServerRef(ctx, team)
if getErr != nil {
return nil, getErr
}
if ns == team.Namespace {
return srv, nil
}
return r.checkConsent(ctx, team, srv, ns)
}
// getServerRef fetches the referenced TerdutServer and the namespace it was
// looked up in, without the consent check. Deleting a team uses it directly:
// consent gates what the operator will start acting on, not whether it may
// clean up after itself once the server owner has narrowed allowedTeams.
func (r *TerdutTeamReconciler) getServerRef(ctx context.Context, team *terdutv1alpha1.TerdutTeam) (*terdutv1alpha1.TerdutServer, string, *teamError) {
ns := team.Spec.ServerRef.Namespace
if ns == "" {
ns = team.Namespace
@@ -173,18 +195,19 @@ func (r *TerdutTeamReconciler) resolveServerRef(ctx context.Context, team *terdu
var srv terdutv1alpha1.TerdutServer
if err := r.Get(ctx, client.ObjectKey{Namespace: ns, Name: team.Spec.ServerRef.Name}, &srv); err != nil {
if apierrors.IsNotFound(err) {
return nil, &teamError{
return nil, ns, &teamError{
reason: terdutv1alpha1.ReasonServerRefNotFound,
message: fmt.Sprintf("TerdutServer %q not found in namespace %q", team.Spec.ServerRef.Name, ns),
}
}
return nil, &teamError{reason: terdutv1alpha1.ReasonServerRefNotFound, message: err.Error()}
}
if ns == team.Namespace {
return &srv, nil
return nil, ns, &teamError{reason: terdutv1alpha1.ReasonServerRefNotFound, message: err.Error()}
}
return &srv, ns, nil
}
// checkConsent applies DESIGN.md §4.6's allowedTeams gate to a cross-namespace
// reference.
func (r *TerdutTeamReconciler) checkConsent(ctx context.Context, team *terdutv1alpha1.TerdutTeam, srv *terdutv1alpha1.TerdutServer, ns string) (*terdutv1alpha1.TerdutServer, *teamError) {
var ownNamespace corev1.Namespace
if err := r.Get(ctx, client.ObjectKey{Name: team.Namespace}, &ownNamespace); err != nil {
return nil, &teamError{
@@ -203,11 +226,11 @@ func (r *TerdutTeamReconciler) resolveServerRef(ctx context.Context, team *terdu
ns, team.Spec.ServerRef.Name, team.Namespace),
}
}
return &srv, nil
return srv, nil
}
func (r *TerdutTeamReconciler) setTeamNotReady(
ctx context.Context, team *terdutv1alpha1.TerdutTeam, reason, message string, d time.Duration,
ctx context.Context, team *terdutv1alpha1.TerdutTeam, reason, message string,
) (ctrl.Result, error) {
meta.SetStatusCondition(&team.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady,
@@ -222,42 +245,18 @@ func (r *TerdutTeamReconciler) setTeamNotReady(
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeWarning, reason, reason, message)
}
return ctrl.Result{RequeueAfter: d}, nil
return ctrl.Result{RequeueAfter: waitInterval}, nil
}
// teamCredentialsSecretName/teamServiceAccountName follow the same
// <namespace>.<name>-suffix convention TerdutServer's secrets use
// (DESIGN.md §6 point 3), keyed on the TerdutTeam CR's own identity rather
// than its (mutable) displayName -- a service account's own name is
// globally unique across the whole install (internal/db/migrations/
// 014_service_accounts.sql's UNIQUE constraint, confirmed against source),
// so this has to be collision-safe the same way the credentials Secret
// names already are.
func teamCredentialsSecretName(team *terdutv1alpha1.TerdutTeam) string {
return fmt.Sprintf("%s.%s-team-credentials", team.Namespace, team.Name)
}
func teamServiceAccountName(team *terdutv1alpha1.TerdutTeam) string {
return fmt.Sprintf("terdut-team.%s.%s", team.Namespace, team.Name)
}
// reconcileTeamDelete cleans up the team-scoped credentials Secret and, if
// one was ever minted, deletes the team server-side first. If
// status.teamID or status.credentialsSecretRef was never set (the CR was
// deleted before reconciliation ever got that far), there is deliberately
// no attempt to clean up server-side: terdut-server's DELETE /api/teams/
// {teamID} is owner-gated (requireTeamOwner), and an instance-scoped
// credential -- the only one this controller would otherwise hold -- does
// not satisfy that check (confirmed against source, same finding as
// TEAM-LOOKUP.md's). A team created but never fully reconciled to Ready is
// left orphaned server-side for a human with real owner/admin access to
// clean up -- a known, documented limitation, not a silent gap.
// reconcileTeamDelete removes the team on the server, then the finalizer. The
// server refuses while the team has open incidents (409), which surfaces as a
// retried error: tidying up must not be how a live page disappears.
func (r *TerdutTeamReconciler) reconcileTeamDelete(ctx context.Context, team *terdutv1alpha1.TerdutTeam) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(team, teamFinalizerName) {
return ctrl.Result{}, nil
}
if team.Status.TeamID != 0 && team.Status.CredentialsSecretRef != nil {
if team.Status.TeamID != 0 {
if err := r.deleteTeamServerSide(ctx, team); err != nil {
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeWarning, "DeleteFailed", "DeleteFailed", err.Error())
@@ -266,40 +265,33 @@ func (r *TerdutTeamReconciler) reconcileTeamDelete(ctx context.Context, team *te
}
}
secret := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: teamCredentialsSecretName(team), Namespace: r.OperatorNamespace}}
if err := r.Delete(ctx, secret); err != nil && !apierrors.IsNotFound(err) {
return ctrl.Result{}, err
}
controllerutil.RemoveFinalizer(team, teamFinalizerName)
return ctrl.Result{}, r.Update(ctx, team)
}
func (r *TerdutTeamReconciler) deleteTeamServerSide(ctx context.Context, team *terdutv1alpha1.TerdutTeam) error {
srv, resolveErr := r.resolveServerRef(ctx, team)
if resolveErr != nil {
// The TerdutServer (or the namespace consent for it) is gone too --
// most likely the whole install is being torn down together.
// Nothing to delete against; proceed rather than block forever on
// a parent that no longer exists.
return nil
}
teamKey, err := readOperatorSecret(ctx, r.Client, r.OperatorNamespace, team.Status.CredentialsSecretRef)
if err != nil {
// Consent is deliberately not checked here: it gates what the operator
// starts acting on, not whether it may clean up after itself once the
// server owner has narrowed allowedTeams. Only a TerdutServer that is
// really gone lets the delete proceed (there is nothing to delete against);
// any other failure is retried rather than orphaning the team.
srv, ns, getErr := r.getServerRef(ctx, team)
if getErr != nil {
var probe terdutv1alpha1.TerdutServer
err := r.Get(ctx, client.ObjectKey{Namespace: ns, Name: team.Spec.ServerRef.Name}, &probe)
if apierrors.IsNotFound(err) {
return nil
}
return fmt.Errorf("reading TerdutServer %q/%q: %w", ns, team.Spec.ServerRef.Name, getErr)
}
tc, err := r.serverClient(ctx, srv)
if err != nil {
if apierrors.IsNotFound(err) {
return nil // the key Secret went with its TerdutServer
}
return err
}
newClient := r.NewClient
if newClient == nil {
newClient = tdclient.New
}
err = newClient(serviceURL(srv)).WithToken(teamKey).DeleteTeam(ctx, team.Status.TeamID)
if statusErr, ok := errors.AsType[*tdclient.StatusError](err); ok && statusErr.Code == http.StatusNotFound {
return nil
}
return err
return tc.DeleteTeam(ctx, team.Status.TeamID)
}
// SetupWithManager sets up the controller with the Manager.
@@ -312,6 +304,28 @@ func (r *TerdutTeamReconciler) SetupWithManager(mgr ctrl.Manager) error {
}
return ctrl.NewControllerManagedBy(mgr).
For(&terdutv1alpha1.TerdutTeam{}).
Watches(&terdutv1alpha1.TerdutServer{}, handler.EnqueueRequestsFromMapFunc(r.teamsForServer)).
Named("terdutteam").
Complete(r)
}
// teamsForServer re-reconciles every TerdutTeam that references a TerdutServer
// when it changes (becomes Ready, say), instead of polling for it.
func (r *TerdutTeamReconciler) teamsForServer(ctx context.Context, obj client.Object) []reconcile.Request {
var teams terdutv1alpha1.TerdutTeamList
if err := r.List(ctx, &teams); err != nil {
return nil
}
var out []reconcile.Request
for i := range teams.Items {
t := &teams.Items[i]
ns := t.Spec.ServerRef.Namespace
if ns == "" {
ns = t.Namespace
}
if t.Spec.ServerRef.Name == obj.GetName() && ns == obj.GetNamespace() {
out = append(out, reconcile.Request{NamespacedName: client.ObjectKeyFromObject(t)})
}
}
return out
}
+311 -77
View File
@@ -2,7 +2,6 @@ package controller
import (
"context"
"fmt"
"net/http/httptest"
. "github.com/onsi/ginkgo/v2"
@@ -17,6 +16,12 @@ import (
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// platformName is the display name most specs give their team.
const platformName = "platform"
// allNamespaces is allowedTeams.namespaces.from admitting every namespace.
const allNamespaces = "All"
var _ = Describe("TerdutTeam Controller", func() {
const operatorNamespace = "default"
@@ -33,13 +38,12 @@ var _ = Describe("TerdutTeam Controller", func() {
fake, fakeSrv = newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
srv = bootstrapReadyTerdutServer(ctx, uniqueName("ttserver"), fakeSrv.URL)
srv = readyTerdutServer(ctx, uniqueName("ttserver"))
reconciler = &TerdutTeamReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: operatorNamespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
Client: k8sClient,
Scheme: k8sClient.Scheme(),
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) },
}
teamName = uniqueName("team")
teamKey = types.NamespacedName{Name: teamName, Namespace: operatorNamespace}
@@ -52,13 +56,7 @@ var _ = Describe("TerdutTeam Controller", func() {
_ = k8sClient.Update(ctx, team)
_ = k8sClient.Delete(ctx, team)
}
_ = k8sClient.Delete(ctx, &corev1.Secret{ObjectMeta: metav1.ObjectMeta{
Name: fmt.Sprintf("%s.%s-team-credentials", operatorNamespace, teamName), Namespace: operatorNamespace,
}})
// The TerdutServer bootstrapReadyTerdutServer created in BeforeEach
// carries its own finalizer; clear it directly the same way, rather
// than relying on TerdutServerReconciler to ever run again here.
srvKey := types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}
if err := k8sClient.Get(ctx, srvKey, srv); err == nil {
srv.Finalizers = nil
@@ -66,7 +64,7 @@ var _ = Describe("TerdutTeam Controller", func() {
_ = k8sClient.Delete(ctx, srv)
}
_ = k8sClient.Delete(ctx, &corev1.Secret{ObjectMeta: metav1.ObjectMeta{
Name: fmt.Sprintf("%s.%s-instance-credentials", operatorNamespace, srv.Name), Namespace: operatorNamespace,
Name: srv.Name + "-operator-key", Namespace: operatorNamespace,
}})
})
@@ -95,39 +93,56 @@ var _ = Describe("TerdutTeam Controller", func() {
return terdutv1alpha1.TerdutServerRef{Name: srv.Name}
}
getTeam := func(ctx context.Context) *terdutv1alpha1.TerdutTeam {
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
return team
}
// reconcileToReady drives a freshly created team through the finalizer pass
// and one apply pass.
reconcileToReady := func(ctx context.Context) {
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // apply
}
Describe("the happy path", func() {
It("creates the team, mints its credential, and applies rename/oidc-groups", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // full create+mint+apply
It("creates the team under its CR identity with the operator key, and applies oidc groups", func(ctx SpecContext) {
// The fake only answers a request carrying the server's operator key.
var keySecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: srv.Name + "-operator-key", Namespace: operatorNamespace}, &keySecret)).To(Succeed())
fake.expectKey = string(keySecret.Data[operatorKeyDataKey])
team := &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: teamName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutTeamSpec{
ServerRef: sameNSRef(), DisplayName: platformName,
OIDC: terdutv1alpha1.TerdutTeamOIDC{MemberGroup: "members", OwnerGroup: "owners"},
},
}
Expect(k8sClient.Create(ctx, team)).To(Succeed())
reconcileToReady(ctx)
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionTrue))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonTeamAdopted))
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.TeamID).To(Equal(fake.teams["platform"]))
Expect(team.Status.CredentialsSecretRef).NotTo(BeNil())
var credsSecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: team.Status.CredentialsSecretRef.Name, Namespace: operatorNamespace,
}, &credsSecret)).To(Succeed())
Expect(credsSecret.Data[credentialsSecretDataKey]).NotTo(BeEmpty())
team = getTeam(ctx)
Expect(team.Status.TeamID).To(Equal(fake.teams[platformName]))
Expect(fake.teamExt).To(HaveKeyWithValue(operatorNamespace+"/"+teamName, team.Status.TeamID))
Expect(fake.teamOIDC[team.Status.TeamID]).To(Equal([2]string{"members", "owners"}))
})
})
Describe("waiting on the referenced TerdutServer", func() {
It("reports ServerRefNotFound when the TerdutServer doesn't exist", func(ctx SpecContext) {
createTeam(ctx, "orphan", terdutv1alpha1.TerdutServerRef{Name: testRefNotFoundName})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
reconcileToReady(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonServerRefNotFound))
})
It("reports WaitingForServer when the TerdutServer exists but isn't Bootstrapped yet", func(ctx SpecContext) {
It("reports WaitingForServer when the TerdutServer exists but isn't Ready yet", func(ctx SpecContext) {
unreadyName := uniqueName("ttserver-unready")
unready := &terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: unreadyName, Namespace: operatorNamespace},
@@ -141,8 +156,7 @@ var _ = Describe("TerdutTeam Controller", func() {
DeferCleanup(func() { _ = k8sClient.Delete(ctx, unready) })
createTeam(ctx, "waiting", terdutv1alpha1.TerdutServerRef{Name: unreadyName})
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx)
reconcileToReady(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonWaitingForServer))
})
@@ -181,7 +195,7 @@ var _ = Describe("TerdutTeam Controller", func() {
It("permits when the TerdutServer's allowedTeams.namespaces.from is All", func(ctx SpecContext) {
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}, srv)).To(Succeed())
srv.Spec.AllowedTeams.Namespaces.From = "All"
srv.Spec.AllowedTeams.Namespaces.From = allNamespaces
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
team := &terdutv1alpha1.TerdutTeam{
@@ -208,72 +222,292 @@ var _ = Describe("TerdutTeam Controller", func() {
})
})
Describe("adopt-on-409 recovery", func() {
It("adopts an already-created team instead of erroring", func(ctx SpecContext) {
Describe("team identity", func() {
It("refuses a name that belongs to a team this CR did not create", func(ctx SpecContext) {
fake.nextTeamID = 1
fake.teams["platform"] = 1
fake.teamNames[1] = "platform"
fake.teams[platformName] = 1
fake.teamNames[1] = platformName
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx) // finalizer
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionFalse))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonTeamNameTaken))
Expect(getTeam(ctx).Status.TeamID).To(BeZero())
Expect(fake.teamOIDC).To(BeEmpty(), "nothing is configured on a team that is not ours")
})
It("keeps the same team when spec.displayName changes", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
team := getTeam(ctx)
team.Spec.DisplayName = "platform-renamed"
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.TeamID).To(Equal(int64(1)))
Expect(getTeam(ctx).Status.TeamID).To(Equal(id))
Expect(fake.teamNames[id]).To(Equal("platform-renamed"))
Expect(fake.nextTeamID).To(Equal(id), "renamed in place, not recreated")
})
It("reports TeamNameTaken when a rename collides with another team", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
fake.nextTeamID++
fake.teams["other"] = fake.nextTeamID
fake.teamNames[fake.nextTeamID] = "other"
team := getTeam(ctx)
team.Spec.DisplayName = "other"
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonTeamNameTaken))
})
It("finds its team again when the status was lost", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
team := getTeam(ctx)
team.Status.TeamID = 0
Expect(k8sClient.Status().Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(getTeam(ctx).Status.TeamID).To(Equal(id))
Expect(fake.nextTeamID).To(Equal(id), "no second team was created")
})
It("recreates a team deleted behind its back", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
delete(fake.teams, platformName)
delete(fake.teamNames, id)
delete(fake.teamExt, operatorNamespace+"/"+teamName)
reconcileOnce(ctx)
Expect(getTeam(ctx).Status.TeamID).NotTo(Equal(id))
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
})
})
Describe("spec.escalation", func() {
ladder := func() *terdutv1alpha1.EscalationSpec {
return &terdutv1alpha1.EscalationSpec{
RepeatCount: 2, FallbackTopic: "oncall",
Levels: []terdutv1alpha1.EscalationLevel{
{Timeout: "5m", Targets: []terdutv1alpha1.EscalationTarget{{Kind: terdutv1alpha1.EscalationTargetOncall}}},
{Timeout: "15m", Targets: []terdutv1alpha1.EscalationTarget{{Kind: terdutv1alpha1.EscalationTargetUser, Username: "alice"}}},
},
}
}
createWithLadder := func(ctx context.Context, e *terdutv1alpha1.EscalationSpec) {
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: teamName, Namespace: operatorNamespace},
Spec: terdutv1alpha1.TerdutTeamSpec{ServerRef: sameNSRef(), DisplayName: platformName, Escalation: e},
})).To(Succeed())
}
It("PUTs the ladder with usernames for the server to resolve", func(ctx SpecContext) {
createWithLadder(ctx, ladder())
reconcileToReady(ctx)
got := fake.escalation[getTeam(ctx).Status.TeamID]
Expect(got.RepeatCount).To(Equal(int64(2)))
Expect(got.FallbackTopic).To(Equal("oncall"))
Expect(got.Levels).To(HaveLen(2))
Expect(got.Levels[0].TimeoutSeconds).To(Equal(int64(300)))
Expect(got.Levels[1].Targets[0]).To(Equal(tdclient.EscalationTargetRequest{Kind: "user", Username: "alice"}))
})
It("waits with UnknownUser while the server does not know a username", func(ctx SpecContext) {
fake.unknownUsers["alice"] = true
createWithLadder(ctx, ladder())
reconcileToReady(ctx)
cond := readyCondition(ctx)
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonUnknownUser))
Expect(cond.Message).To(ContainSubstring("alice"))
delete(fake.unknownUsers, "alice")
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
})
It("adopts an already-created team-scoped service account instead of erroring", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx) // finalizer
It("clears the ladder when spec.escalation is removed", func(ctx SpecContext) {
createWithLadder(ctx, ladder())
reconcileToReady(ctx)
// Pre-seed the service account the mint step is about to try to
// create, simulating an attempt that got this far before being
// interrupted.
fake.nextID = 1
saName := fmt.Sprintf("terdut-team.%s.%s", operatorNamespace, teamName)
fake.accounts[saName] = 1
fake.keyMints[1] = 1
team := getTeam(ctx)
team.Spec.Escalation = nil
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(fake.escalation[team.Status.TeamID].Levels).To(BeEmpty())
})
It("reports InvalidSpec for a duration that does not parse", func(ctx SpecContext) {
bad := ladder()
bad.Levels[0].Timeout = "5 minutes"
// The CRD pattern rejects it at admission when the API server enforces it;
// envtest does, so write it past the pattern by creating a valid one first.
bad.Levels[0].Timeout = "5m"
createWithLadder(ctx, bad)
reconcileToReady(ctx)
team := getTeam(ctx)
team.Spec.Escalation.Levels[0].Timeout = "0s"
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonInvalidSpec))
})
})
Describe("spec.deadmanSwitches", func() {
sw := func(name, matcher, timeout string) terdutv1alpha1.DeadmanSwitchSpec {
return terdutv1alpha1.DeadmanSwitchSpec{Name: name, Matcher: matcher, Timeout: timeout, Severity: "critical"}
}
setSwitches := func(ctx context.Context, switches ...terdutv1alpha1.DeadmanSwitchSpec) {
team := getTeam(ctx)
team.Spec.DeadmanSwitches = switches
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
}
names := func(teamID int64) []string {
out := make([]string, 0, len(fake.switches[teamID]))
for _, s := range fake.switches[teamID] {
out = append(out, s.Name)
}
return out
}
It("creates, updates and prunes switches by name", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
setSwitches(ctx, sw("watchdog", "alertname=Watchdog", "15m"), sw("edge", "alertname=Edge", "5m"))
Expect(names(id)).To(ConsistOf("watchdog", "edge"))
// Same name, new timeout: updated in place, same id.
var watchdogID int64
for _, s := range fake.switches[id] {
if s.Name == "watchdog" {
watchdogID = s.ID
}
}
setSwitches(ctx, sw("watchdog", "alertname=Watchdog", "30m"))
Expect(names(id)).To(ConsistOf("watchdog"))
Expect(fake.switches[id][watchdogID].TimeoutSeconds).To(Equal(int64(1800)))
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
})
It("removes a switch somebody added on the server", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
fake.nextSwitchID++
fake.switches[id] = map[int64]tdclient.DeadmanSwitch{
fake.nextSwitchID: {ID: fake.nextSwitchID, Name: "stray", Matcher: "alertname=X", TimeoutSeconds: 60, Severity: "critical"},
}
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
var credsSecret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: team.Status.CredentialsSecretRef.Name, Namespace: operatorNamespace,
}, &credsSecret)).To(Succeed())
// Minted fresh (mint2), not the pre-seeded account's original
// (never-issued-to-this-reconcile) key.
Expect(string(credsSecret.Data[credentialsSecretDataKey])).To(Equal("instance-key-1-mint2"))
Expect(fake.switches[id]).To(BeEmpty())
})
It("recreates a switch deleted on the server", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
id := getTeam(ctx).Status.TeamID
setSwitches(ctx, sw("watchdog", "alertname=Watchdog", "15m"))
fake.switches[id] = map[int64]tdclient.DeadmanSwitch{}
reconcileOnce(ctx)
Expect(names(id)).To(ConsistOf("watchdog"))
})
It("reports InvalidSpec for a timeout that does not parse", func(ctx SpecContext) {
createTeam(ctx, platformName, sameNSRef())
reconcileToReady(ctx)
setSwitches(ctx, sw("watchdog", "alertname=Watchdog", "0s"))
Expect(readyCondition(ctx).Reason).To(Equal(terdutv1alpha1.ReasonInvalidSpec))
})
})
Describe("deletion", func() {
It("deletes the team server-side and removes the credentials Secret", func(ctx SpecContext) {
It("deletes the team on the server and clears the finalizer", func(ctx SpecContext) {
createTeam(ctx, "to-delete", sameNSRef())
reconcileOnce(ctx)
reconcileOnce(ctx)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
reconcileToReady(ctx)
teamID := getTeam(ctx).Status.TeamID
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
teamID := team.Status.TeamID
credsName := team.Status.CredentialsSecretRef.Name
Expect(k8sClient.Delete(ctx, team)).To(Succeed())
Expect(k8sClient.Delete(ctx, getTeam(ctx))).To(Succeed())
reconcileOnce(ctx) // runs the finalizer
Expect(fake.teamDelete[teamID]).To(BeTrue())
Expect(k8sClient.Get(ctx, teamKey, &terdutv1alpha1.TerdutTeam{})).NotTo(Succeed(),
"the TerdutTeam should be gone once the finalizer clears")
})
err := k8sClient.Get(ctx, teamKey, team)
Expect(err).To(HaveOccurred(), "the TerdutTeam itself should be gone once the finalizer clears")
It("still deletes the team after the server's allowedTeams consent is revoked", func(ctx SpecContext) {
otherNS := uniqueName("ns")
Expect(k8sClient.Create(ctx, &corev1.Namespace{ObjectMeta: metav1.ObjectMeta{Name: otherNS}})).To(Succeed())
DeferCleanup(func() { _ = k8sClient.Delete(ctx, &corev1.Namespace{ObjectMeta: metav1.ObjectMeta{Name: otherNS}}) })
var leftover corev1.Secret
err = k8sClient.Get(ctx, types.NamespacedName{Name: credsName, Namespace: operatorNamespace}, &leftover)
Expect(err).To(HaveOccurred(), "the team-credentials Secret should have been cleaned up")
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}, srv)).To(Succeed())
srv.Spec.AllowedTeams.Namespaces.From = allNamespaces
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
crossKey := types.NamespacedName{Name: teamName, Namespace: otherNS}
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutTeam{
ObjectMeta: metav1.ObjectMeta{Name: teamName, Namespace: otherNS},
Spec: terdutv1alpha1.TerdutTeamSpec{
ServerRef: terdutv1alpha1.TerdutServerRef{Name: srv.Name, Namespace: operatorNamespace},
DisplayName: "revoked",
},
})).To(Succeed())
for range 2 {
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: crossKey})
Expect(err).NotTo(HaveOccurred())
}
cross := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, crossKey, cross)).To(Succeed())
teamID := cross.Status.TeamID
Expect(teamID).NotTo(BeZero())
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}, srv)).To(Succeed())
srv.Spec.AllowedTeams.Namespaces.From = "None"
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
Expect(k8sClient.Delete(ctx, cross)).To(Succeed())
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: crossKey})
Expect(err).NotTo(HaveOccurred())
Expect(fake.teamDelete[teamID]).To(BeTrue(), "revoking consent must not orphan the team")
})
It("lets the CR go when its TerdutServer is already gone", func(ctx SpecContext) {
createTeam(ctx, "to-delete", sameNSRef())
reconcileToReady(ctx)
serverKey := types.NamespacedName{Name: srv.Name, Namespace: operatorNamespace}
Expect(k8sClient.Get(ctx, serverKey, srv)).To(Succeed())
srv.Finalizers = nil
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
Expect(k8sClient.Delete(ctx, srv)).To(Succeed())
Expect(k8sClient.Delete(ctx, getTeam(ctx))).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, &terdutv1alpha1.TerdutTeam{})).NotTo(Succeed())
})
})
})
+18 -27
View File
@@ -8,6 +8,7 @@ import (
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
appsv1 "k8s.io/api/apps/v1"
"k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/types"
"sigs.k8s.io/controller-runtime/pkg/reconcile"
@@ -47,21 +48,15 @@ const (
testOperatorNamespace = "default"
)
// bootstrapReadyTerdutServer creates a TerdutServer with a bring-your-own
// DSN and drives it to Ready against fakeURL, the same three-pass sequence
// terdutserver_controller_test.go's own happy-path test exercises directly
// -- shared here so TerdutTeam's tests (which need a real, Ready
// TerdutServer to resolve against) don't duplicate it.
func bootstrapReadyTerdutServer(ctx context.Context, name, fakeURL string) *terdutv1alpha1.TerdutServer {
// readyTerdutServer creates a TerdutServer with a bring-your-own DSN and drives
// it to Ready the way the happy-path test does -- shared so the tests of
// everything that references a server (TerdutTeam, TerdutAlertSource) start
// from a real, Ready one.
func readyTerdutServer(ctx context.Context, name string) *terdutv1alpha1.TerdutServer {
GinkgoHelper()
namespace := testOperatorNamespace
reconciler := &TerdutServerReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: namespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeURL) },
}
reconciler := &TerdutServerReconciler{Client: k8sClient, Scheme: k8sClient.Scheme()}
objKey := types.NamespacedName{Name: name, Namespace: namespace}
srv := &terdutv1alpha1.TerdutServer{
@@ -74,9 +69,7 @@ func bootstrapReadyTerdutServer(ctx context.Context, name, fakeURL string) *terd
}
Expect(k8sClient.Create(ctx, srv)).To(Succeed())
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // finalizer
Expect(err).NotTo(HaveOccurred())
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // Deployment/Service
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // Deployment, Service, key
Expect(err).NotTo(HaveOccurred())
var deploy appsv1.Deployment
@@ -85,27 +78,25 @@ func bootstrapReadyTerdutServer(ctx context.Context, name, fakeURL string) *terd
deploy.Status.Replicas = 1
Expect(k8sClient.Status().Update(ctx, &deploy)).To(Succeed())
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // bootstrap
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // Ready
Expect(err).NotTo(HaveOccurred())
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(srv.Status.CredentialsSecretRef).NotTo(BeNil(), "test setup: TerdutServer %s/%s did not reach Bootstrapped", namespace, name)
Expect(meta.IsStatusConditionTrue(srv.Status.Conditions, terdutv1alpha1.ConditionReady)).To(BeTrue(),
"test setup: TerdutServer %s/%s did not reach Ready", namespace, name)
return srv
}
// readyTerdutTeam creates a TerdutTeam under srv (an already-Ready
// TerdutServer, e.g. from bootstrapReadyTerdutServer) and drives it to
// Ready against fakeURL -- shared by TerdutEscalationRule's and
// TerdutDeadmanSwitch's own tests, which both just need a resolvable
// teamRef (DESIGN.md §5), not TerdutTeam's own behavior.
// TerdutServer, from readyTerdutServer) and drives it to Ready against the fake
// at fakeURL.
func readyTerdutTeam(ctx context.Context, namespace, name string, srv *terdutv1alpha1.TerdutServer, fakeURL string) *terdutv1alpha1.TerdutTeam {
GinkgoHelper()
reconciler := &TerdutTeamReconciler{
Client: k8sClient,
Scheme: k8sClient.Scheme(),
OperatorNamespace: namespace,
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeURL) },
Client: k8sClient,
Scheme: k8sClient.Scheme(),
NewClient: func(string) *tdclient.Client { return tdclient.New(fakeURL) },
}
objKey := types.NamespacedName{Name: name, Namespace: namespace}
@@ -120,11 +111,11 @@ func readyTerdutTeam(ctx context.Context, namespace, name string, srv *terdutv1a
_, err := reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // finalizer
Expect(err).NotTo(HaveOccurred())
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // create+mint+apply
_, err = reconciler.Reconcile(ctx, reconcile.Request{NamespacedName: objKey}) // apply
Expect(err).NotTo(HaveOccurred())
Expect(k8sClient.Get(ctx, objKey, team)).To(Succeed())
Expect(team.Status.CredentialsSecretRef).NotTo(BeNil(), "test setup: TerdutTeam %s/%s did not reach Ready", namespace, name)
Expect(team.Status.TeamID).NotTo(BeZero(), "test setup: TerdutTeam %s/%s did not reach Ready", namespace, name)
return team
}
+28 -230
View File
@@ -10,6 +10,7 @@ package tdclient
import (
"context"
"encoding/json"
"errors"
"fmt"
"net/http"
"strings"
@@ -57,6 +58,16 @@ type StatusError struct {
Message string
}
// ignoreNotFound turns a 404 into success. Every DELETE here is idempotent: the
// thing already being gone is the outcome the caller wanted, and a finalizer
// that failed on it could never be removed.
func ignoreNotFound(err error) error {
if se, ok := errors.AsType[*StatusError](err); ok && se.Code == http.StatusNotFound {
return nil
}
return err
}
func (e *StatusError) Error() string {
if e.Message != "" {
return fmt.Sprintf("server returned %d: %s", e.Code, e.Message)
@@ -116,187 +127,23 @@ func (c *Client) do(req *http.Request, out any) error {
return nil
}
// Version calls GET /api/version — unauthenticated, per terdut-server's own
// router.go comment ("a client deciding whether it can talk to this server —
// terdut-tui, terdut-operator — needs to ask before it holds a credential
// for it"). Used here purely as a reachability probe: a bad endpoint fails
// here, clearly, rather than on whatever the controller tries first.
func (c *Client) Version(ctx context.Context) (string, error) {
req, err := c.newRequest(ctx, http.MethodGet, "/api/version", nil)
if err != nil {
return "", err
}
var v struct {
Version string `json:"version"`
}
if err := c.do(req, &v); err != nil {
return "", err
}
return v.Version, nil
}
// APIKey is the raw key a bootstrap or service-account-key mint hands back —
// the one moment its value exists outside the request that generated it.
// Mirrors terdut-server's models.APIKey/models.ServiceAccountKey shape
// (internal/models in that repo) for the fields this client actually reads.
type APIKey struct {
ID int64 `json:"id"`
Name string `json:"name"`
Key string `json:"key"`
CreatedAt string `json:"created_at"`
}
// BootstrapResult is /api/bootstrap's 201 response body.
type BootstrapResult struct {
User struct {
ID int64 `json:"id"`
Username string `json:"username"`
Email string `json:"email"`
} `json:"user"`
APIKey APIKey `json:"api_key"`
}
// Bootstrap calls POST /api/bootstrap — unauthenticated, single-shot per
// install (internal/api/users.go's handleBootstrap in terdut-server:
// gated on SELECT COUNT(*) FROM users). Returns the raw admin key directly;
// DESIGN.md §6 has the controller use it for exactly one further call
// (CreateServiceAccount) and discard it, never storing it as the lasting
// credential.
//
// A StatusError with Code 403 means this install already has a user —
// per this operator's design (DESIGN.md §1), that only happens if this
// exact TerdutServer's own controller already won this race on an earlier,
// interrupted reconcile; see the checkpoint-Secret handling in the
// controller, not a retry loop here.
func (c *Client) Bootstrap(ctx context.Context, username, email string) (*BootstrapResult, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/bootstrap", map[string]string{
"username": username,
"email": email,
})
if err != nil {
return nil, err
}
var result BootstrapResult
if err := c.do(req, &result); err != nil {
return nil, err
}
return &result, nil
}
// ServiceAccount mirrors terdut-server's models.ServiceAccount
// (internal/models/service_account.go), minus fields this client never
// reads.
type ServiceAccount struct {
ID int64 `json:"id"`
Name string `json:"name"`
Scope string `json:"scope"`
}
// CreateServiceAccountResult is POST /api/service-accounts' 201 response.
type CreateServiceAccountResult struct {
ServiceAccount ServiceAccount `json:"service_account"`
Key APIKey `json:"key"`
}
// CreateInstanceServiceAccount calls POST /api/service-accounts with
// scope "instance", authenticated with c's current token (the raw admin key
// from Bootstrap, for the operator's own first-ever call). A StatusError
// with Code 409 means a prior, interrupted attempt already created this
// name — DESIGN.md §6's adopt-rather-than-error rule: the caller should
// fall back to GetServiceAccountByName + CreateServiceAccountKey, not treat
// this as a hard failure.
func (c *Client) CreateInstanceServiceAccount(ctx context.Context, name string) (*CreateServiceAccountResult, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/service-accounts", map[string]string{
fieldName: name,
"scope": "instance",
})
if err != nil {
return nil, err
}
var result CreateServiceAccountResult
if err := c.do(req, &result); err != nil {
return nil, err
}
return &result, nil
}
// CreateTeamServiceAccount calls POST /api/service-accounts with scope
// "team" for teamID, authenticated with c's current token -- the
// TerdutServer's instance-scoped credential, per DESIGN.md §6 point 3: an
// instance-scoped caller may mint a team-scoped account against any team
// (internal/api/service_accounts.go's handleCreateServiceAccount,
// confirmed against source), which is what lets TerdutTeam's own
// controller do this without ever touching a human credential. Same
// adopt-on-409 contract as CreateInstanceServiceAccount.
func (c *Client) CreateTeamServiceAccount(ctx context.Context, name string, teamID int64) (*CreateServiceAccountResult, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/service-accounts", map[string]any{
fieldName: name,
"scope": "team",
"team_id": teamID,
})
if err != nil {
return nil, err
}
var result CreateServiceAccountResult
if err := c.do(req, &result); err != nil {
return nil, err
}
return &result, nil
}
// GetServiceAccountByName calls GET /api/service-accounts?name=... --
// authenticated (terdut-server's AuthMiddleware hard-rejects any
// unauthenticated request before this endpoint's own, more permissive
// internal check ever runs; DESIGN.md §6). Returns nil, nil if nothing
// matches, not an error -- the server's own distinction between "found
// nothing" and "the call failed".
func (c *Client) GetServiceAccountByName(ctx context.Context, name string) (*ServiceAccount, error) {
req, err := c.newRequest(ctx, http.MethodGet, "/api/service-accounts?name="+name, nil)
if err != nil {
return nil, err
}
var accounts []ServiceAccount
if err := c.do(req, &accounts); err != nil {
return nil, err
}
if len(accounts) == 0 {
return nil, nil
}
return &accounts[0], nil
}
// CreateServiceAccountKey calls POST /api/service-accounts/{id}/keys to
// mint an additional key on an existing account -- rotation (DESIGN.md §6
// point 6), and the adopt-on-409 recovery path in point 1.
func (c *Client) CreateServiceAccountKey(ctx context.Context, serviceAccountID int64, name string) (*APIKey, error) {
req, err := c.newRequest(ctx, http.MethodPost,
fmt.Sprintf("/api/service-accounts/%d/keys", serviceAccountID),
map[string]string{fieldName: name})
if err != nil {
return nil, err
}
var key APIKey
if err := c.do(req, &key); err != nil {
return nil, err
}
return &key, nil
}
// Team mirrors terdut-server's models.Team (internal/models/team.go), minus
// Role/Source, which are only ever populated for a human caller's own
// membership and never apply to a service account's view of a team.
type Team struct {
ID int64 `json:"id"`
Name string `json:"name"`
ID int64 `json:"id"`
Name string `json:"name"`
ExternalID *string `json:"external_id,omitempty"`
}
// CreateTeam calls POST /api/teams, authenticated with c's current token --
// the TerdutServer's instance-scoped credential (DESIGN.md §6 point 3). A
// StatusError with Code 409 means a prior, interrupted attempt already
// created this name — GetTeamByName (TEAM-LOOKUP.md) is the adopt-rather-
// than-error recovery, the same contract CreateInstanceServiceAccount has.
func (c *Client) CreateTeam(ctx context.Context, name string) (*Team, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/teams", map[string]string{fieldName: name})
// CreateTeam calls POST /api/teams with an external_id, which makes it
// idempotent: a team that already carries that id is returned (200) instead of
// created (201), so a client that crashed before recording the id finds its own
// team again. A StatusError with Code 409 means the name belongs to a different
// team. Needs an instance-scoped credential.
func (c *Client) CreateTeam(ctx context.Context, name, externalID string) (*Team, error) {
req, err := c.newRequest(ctx, http.MethodPost, "/api/teams",
map[string]string{fieldName: name, "external_id": externalID})
if err != nil {
return nil, err
}
@@ -307,24 +154,6 @@ func (c *Client) CreateTeam(ctx context.Context, name string) (*Team, error) {
return &team, nil
}
// GetTeamByName calls GET /api/teams?name=... (TEAM-LOOKUP.md) — open to
// any authenticated caller, not just the team's own members. Returns nil,
// nil on no match, not an error.
func (c *Client) GetTeamByName(ctx context.Context, name string) (*Team, error) {
req, err := c.newRequest(ctx, http.MethodGet, "/api/teams?name="+name, nil)
if err != nil {
return nil, err
}
var teams []Team
if err := c.do(req, &teams); err != nil {
return nil, err
}
if len(teams) == 0 {
return nil, nil
}
return &teams[0], nil
}
// RenameTeam calls PUT /api/teams/{teamID} — owner-gated server-side
// (requireTeamOwner), so c must hold this team's own team-scoped
// credential, not the instance-scoped one CreateTeam used.
@@ -346,7 +175,7 @@ func (c *Client) DeleteTeam(ctx context.Context, teamID int64) error {
if err != nil {
return err
}
return c.do(req, nil)
return ignoreNotFound(c.do(req, nil))
}
// SetTeamOIDCGroups calls PUT /api/teams/{teamID}/oidc-groups — owner-gated,
@@ -362,46 +191,15 @@ func (c *Client) SetTeamOIDCGroups(ctx context.Context, teamID int64, memberGrou
return c.do(req, nil)
}
// User mirrors terdut-server's models.User, minus fields this client never
// reads.
type User struct {
ID int64 `json:"id"`
Username string `json:"username"`
}
// GetUserByUsername calls GET /api/users and finds the one matching exactly
// -- confirmed open to any authenticated caller, not gated by team
// membership or admin (internal/api/router.go's own comment: "readable by
// anyone signed in"), so the team-scoped credential a TerdutEscalationRule's
// controller already holds is enough. There is no server-side filter, so
// this always fetches the whole list; terdut-server's own query has no
// pagination either (confirmed against source), so this matches what the
// server itself considers an acceptable cost. Returns nil, nil on no match.
func (c *Client) GetUserByUsername(ctx context.Context, username string) (*User, error) {
req, err := c.newRequest(ctx, http.MethodGet, "/api/users", nil)
if err != nil {
return nil, err
}
var users []User
if err := c.do(req, &users); err != nil {
return nil, err
}
for _, u := range users {
if u.Username == username {
return &u, nil
}
}
return nil, nil
}
// EscalationTargetRequest/EscalationLevelRequest/SetEscalationRequest mirror
// terdut-server's escalationTargetJSON/escalationLevelJSON/escalationJSON
// (internal/api/escalation.go) -- the PUT body, not the richer GET response
// (escalationView), which this client never needs to decode since the
// controller always computes its own desired state fresh from spec.
type EscalationTargetRequest struct {
Kind string `json:"kind"`
UserID *int64 `json:"user_id,omitempty"`
Kind string `json:"kind"`
// Username names the person for a "user" target; the server resolves it.
Username string `json:"username,omitempty"`
}
type EscalationLevelRequest struct {
@@ -498,7 +296,7 @@ func (c *Client) DeleteDeadmanSwitch(ctx context.Context, teamID, switchID int64
if err != nil {
return err
}
return c.do(req, nil)
return ignoreNotFound(c.do(req, nil))
}
// Integration mirrors terdut-server's models.Integration, minus
@@ -564,5 +362,5 @@ func (c *Client) DeleteIntegration(ctx context.Context, teamID, integrationID in
if err != nil {
return err
}
return c.do(req, nil)
return ignoreNotFound(c.do(req, nil))
}
+30 -69
View File
@@ -7,85 +7,46 @@ import (
"testing"
)
func TestVersion(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/api/version" {
t.Errorf("unexpected path %q", r.URL.Path)
}
// Unauthenticated per terdut-server's own router.go comment: no
// Authorization header should be required, and none is sent here.
w.Header().Set("Content-Type", "application/json")
w.Write([]byte(`{"version":"v0.20.0"}`)) //nolint:errcheck
}))
defer srv.Close()
got, err := New(srv.URL).Version(context.Background())
if err != nil {
t.Fatalf("Version() error = %v", err)
}
if got != "v0.20.0" {
t.Errorf("Version() = %q, want %q", got, "v0.20.0")
}
}
func TestVersionUnreachable(t *testing.T) {
// A closed server: connection refused, the same shape a bad
// spec.endpoint produces against a real cluster.
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {}))
srv.Close()
if _, err := New(srv.URL).Version(context.Background()); err == nil {
t.Fatal("Version() error = nil, want a connection error")
}
}
func TestVersionErrorStatus(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusInternalServerError)
w.Write([]byte(`{"error":"internal error"}`)) //nolint:errcheck
}))
defer srv.Close()
_, err := New(srv.URL).Version(context.Background())
if err == nil {
t.Fatal("Version() error = nil, want a StatusError")
}
var statusErr *StatusError
if !asStatusError(err, &statusErr) {
t.Fatalf("Version() error = %v (%T), want *StatusError", err, err)
}
if statusErr.Code != http.StatusInternalServerError {
t.Errorf("StatusError.Code = %d, want %d", statusErr.Code, http.StatusInternalServerError)
}
if statusErr.Message != "internal error" {
t.Errorf("StatusError.Message = %q, want %q", statusErr.Message, "internal error")
}
}
// asStatusError is errors.As without importing errors twice in a tiny test
// file — kept local since no other test here needs it.
func asStatusError(err error, target **StatusError) bool {
se, ok := err.(*StatusError)
if !ok {
return false
}
*target = se
return true
}
func TestWithTokenSetsAuthorizationHeader(t *testing.T) {
var gotAuth string
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
gotAuth = r.Header.Get("Authorization")
w.Write([]byte(`{"version":"v0.20.0"}`)) //nolint:errcheck
w.WriteHeader(http.StatusNoContent)
}))
defer srv.Close()
c := New(srv.URL).WithToken("tdsa_abc123")
if _, err := c.Version(context.Background()); err != nil {
t.Fatalf("Version() error = %v", err)
if err := c.DeleteTeam(context.Background(), 1); err != nil {
t.Fatalf("DeleteTeam() error = %v", err)
}
if want := "Bearer tdsa_abc123"; gotAuth != want {
t.Errorf("Authorization header = %q, want %q", gotAuth, want)
}
}
// Every delete is idempotent: the thing already being gone is success, any
// other failure is not.
func TestDeletesTreatNotFoundAsSuccess(t *testing.T) {
status := http.StatusNotFound
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(status)
}))
defer srv.Close()
c := New(srv.URL)
ctx := context.Background()
for name, call := range map[string]func() error{
"team": func() error { return c.DeleteTeam(ctx, 1) },
"switch": func() error { return c.DeleteDeadmanSwitch(ctx, 1, 2) },
"integration": func() error { return c.DeleteIntegration(ctx, 1, 2) },
} {
status = http.StatusNotFound
if err := call(); err != nil {
t.Errorf("%s: 404 should be success, got %v", name, err)
}
status = http.StatusConflict
if err := call(); err == nil {
t.Errorf("%s: 409 should be an error", name)
}
}
}
-103
View File
@@ -1,103 +0,0 @@
//go:build e2e
// +build e2e
package e2e
import (
"fmt"
"os"
"os/exec"
"testing"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"git.ryuvia.com/niklas/terdut-operator/test/utils"
)
var (
// managerImage is the manager image to be built and loaded for testing.
managerImage = "example.com/terdut-operator:v0.0.1"
// shouldCleanupCertManager tracks whether CertManager was installed by this suite.
shouldCleanupCertManager = false
)
// TestE2E runs the e2e test suite to validate the solution in an isolated environment.
// The default setup requires Kind and CertManager.
//
// To enable kubectl kuberc (use custom kubectl configurations), set: KUBECTL_KUBERC=true
// By default, kuberc is disabled to ensure consistent test behavior across different environments.
// To skip CertManager installation, set: CERT_MANAGER_INSTALL_SKIP=true
func TestE2E(t *testing.T) {
RegisterFailHandler(Fail)
_, _ = fmt.Fprintf(GinkgoWriter, "Starting terdut-operator e2e test suite\n")
RunSpecs(t, "e2e suite")
}
var _ = BeforeSuite(func() {
By("building the manager image")
cmd := exec.Command("make", "docker-build", fmt.Sprintf("IMG=%s", managerImage))
_, err := utils.Run(cmd)
ExpectWithOffset(1, err).NotTo(HaveOccurred(), "Failed to build the manager image")
// TODO(user): If you want to change the e2e test vendor from Kind,
// ensure the image is built and available, then remove the following block.
By("loading the manager image on Kind")
err = utils.LoadImageToKindClusterWithName(managerImage)
ExpectWithOffset(1, err).NotTo(HaveOccurred(), "Failed to load the manager image into Kind")
configureKubectlKubeRC()
setupCertManager()
})
var _ = AfterSuite(func() {
teardownCertManager()
})
// Disable kubectl kuberc by default for test isolation.
// This prevents local kubectl configurations from affecting test behavior.
// To enable kuberc, set: KUBECTL_KUBERC=true
func configureKubectlKubeRC() {
if os.Getenv("KUBECTL_KUBERC") != "true" {
By("disabling kubectl kuberc for test isolation")
err := os.Setenv("KUBECTL_KUBERC", "false")
ExpectWithOffset(1, err).NotTo(HaveOccurred(), "Failed to disable kubectl kuberc")
_, _ = fmt.Fprintf(GinkgoWriter,
"kubectl kuberc disabled for consistent test behavior (override with KUBECTL_KUBERC=true)\n")
} else {
_, _ = fmt.Fprintf(GinkgoWriter, "kubectl kuberc enabled (KUBECTL_KUBERC=true)\n")
}
}
// setupCertManager installs CertManager if needed for webhook tests.
// Skips installation if CERT_MANAGER_INSTALL_SKIP=true or if already present.
func setupCertManager() {
if os.Getenv("CERT_MANAGER_INSTALL_SKIP") == "true" {
_, _ = fmt.Fprintf(GinkgoWriter, "Skipping CertManager installation (CERT_MANAGER_INSTALL_SKIP=true)\n")
return
}
By("checking if CertManager is already installed")
if utils.IsCertManagerCRDsInstalled() {
_, _ = fmt.Fprintf(GinkgoWriter, "CertManager is already installed. Skipping installation.\n")
return
}
// Mark for cleanup before installation to handle interruptions and partial installs.
shouldCleanupCertManager = true
By("installing CertManager")
Expect(utils.InstallCertManager()).To(Succeed(), "Failed to install CertManager")
}
// teardownCertManager uninstalls CertManager if it was installed by setupCertManager.
// This ensures we only remove what we installed.
func teardownCertManager() {
if !shouldCleanupCertManager {
_, _ = fmt.Fprintf(GinkgoWriter, "Skipping CertManager cleanup (not installed by this suite)\n")
return
}
By("uninstalling CertManager")
utils.UninstallCertManager()
}
-323
View File
@@ -1,323 +0,0 @@
//go:build e2e
// +build e2e
package e2e
import (
"encoding/json"
"fmt"
"os"
"os/exec"
"path/filepath"
"time"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"git.ryuvia.com/niklas/terdut-operator/test/utils"
)
// namespace where the project is deployed in
const namespace = "terdut-operator-system"
// serviceAccountName created for the project
const serviceAccountName = "terdut-operator-controller-manager"
// metricsServiceName is the name of the metrics service of the project
const metricsServiceName = "terdut-operator-controller-manager-metrics-service"
// metricsRoleBindingName is the name of the RBAC that will be created to allow get the metrics data
const metricsRoleBindingName = "terdut-operator-metrics-binding"
var _ = Describe("Manager", Ordered, func() {
var controllerPodName string
// Before running the tests, set up the environment by creating the namespace,
// enforce the restricted security policy to the namespace, installing CRDs,
// and deploying the controller.
BeforeAll(func() {
By("creating manager namespace")
cmd := exec.Command("kubectl", "create", "ns", namespace)
_, err := utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to create namespace")
By("labeling the namespace to enforce the restricted security policy")
cmd = exec.Command("kubectl", "label", "--overwrite", "ns", namespace,
"pod-security.kubernetes.io/enforce=restricted")
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to label namespace with restricted policy")
By("installing CRDs")
cmd = exec.Command("make", "install")
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to install CRDs")
By("deploying the controller-manager")
cmd = exec.Command("make", "deploy", fmt.Sprintf("IMG=%s", managerImage))
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to deploy the controller-manager")
})
// After all tests have been executed, clean up by undeploying the controller, uninstalling CRDs,
// and deleting the namespace.
AfterAll(func() {
By("cleaning up the curl pod for metrics")
cmd := exec.Command("kubectl", "delete", "pod", "curl-metrics", "-n", namespace)
_, _ = utils.Run(cmd)
By("undeploying the controller-manager")
cmd = exec.Command("make", "undeploy")
_, _ = utils.Run(cmd)
By("uninstalling CRDs")
cmd = exec.Command("make", "uninstall")
_, _ = utils.Run(cmd)
By("removing manager namespace")
cmd = exec.Command("kubectl", "delete", "ns", namespace)
_, _ = utils.Run(cmd)
})
// After each test, check for failures and collect logs, events,
// and pod descriptions for debugging.
AfterEach(func() {
specReport := CurrentSpecReport()
if specReport.Failed() {
By("Fetching controller manager pod logs")
cmd := exec.Command("kubectl", "logs", controllerPodName, "-n", namespace)
controllerLogs, err := utils.Run(cmd)
if err == nil {
_, _ = fmt.Fprintf(GinkgoWriter, "Controller logs:\n %s", controllerLogs)
} else {
_, _ = fmt.Fprintf(GinkgoWriter, "Failed to get Controller logs: %s", err)
}
By("Fetching Kubernetes events")
cmd = exec.Command("kubectl", "get", "events", "-n", namespace, "--sort-by=.lastTimestamp")
eventsOutput, err := utils.Run(cmd)
if err == nil {
_, _ = fmt.Fprintf(GinkgoWriter, "Kubernetes events:\n%s", eventsOutput)
} else {
_, _ = fmt.Fprintf(GinkgoWriter, "Failed to get Kubernetes events: %s", err)
}
By("Fetching curl-metrics logs")
cmd = exec.Command("kubectl", "logs", "curl-metrics", "-n", namespace)
metricsOutput, err := utils.Run(cmd)
if err == nil {
_, _ = fmt.Fprintf(GinkgoWriter, "Metrics logs:\n %s", metricsOutput)
} else {
_, _ = fmt.Fprintf(GinkgoWriter, "Failed to get curl-metrics logs: %s", err)
}
By("Fetching controller manager pod description")
cmd = exec.Command("kubectl", "describe", "pod", controllerPodName, "-n", namespace)
podDescription, err := utils.Run(cmd)
if err == nil {
fmt.Println("Pod description:\n", podDescription)
} else {
fmt.Println("Failed to describe controller pod")
}
}
})
SetDefaultEventuallyTimeout(2 * time.Minute)
SetDefaultEventuallyPollingInterval(time.Second)
Context("Manager", func() {
It("should run successfully", func() {
By("validating that the controller-manager pod is running as expected")
verifyControllerUp := func(g Gomega) {
By("getting the name of the controller-manager pod")
cmd := exec.Command("kubectl", "get",
"pods", "-l", "control-plane=controller-manager",
"-o", "go-template={{ range .items }}"+
"{{ if not .metadata.deletionTimestamp }}"+
"{{ .metadata.name }}"+
"{{ \"\\n\" }}{{ end }}{{ end }}",
"-n", namespace,
)
podOutput, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred(), "Failed to retrieve controller-manager pod information")
podNames := utils.GetNonEmptyLines(podOutput)
g.Expect(podNames).To(HaveLen(1), "expected 1 controller pod running")
controllerPodName = podNames[0]
g.Expect(controllerPodName).To(ContainSubstring("controller-manager"))
By("validating the pod's status")
cmd = exec.Command("kubectl", "get",
"pods", controllerPodName, "-o", "jsonpath={.status.phase}",
"-n", namespace,
)
output, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred())
g.Expect(output).To(Equal("Running"), "Incorrect controller-manager pod status")
}
Eventually(verifyControllerUp).Should(Succeed())
})
It("should ensure the metrics endpoint is serving metrics", func() {
By("creating a ClusterRoleBinding for the service account to allow access to metrics")
cmd := exec.Command("kubectl", "create", "clusterrolebinding", metricsRoleBindingName,
"--clusterrole=terdut-operator-metrics-reader",
fmt.Sprintf("--serviceaccount=%s:%s", namespace, serviceAccountName),
)
_, err := utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to create ClusterRoleBinding")
By("validating that the metrics service is available")
cmd = exec.Command("kubectl", "get", "service", metricsServiceName, "-n", namespace)
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Metrics service should exist")
By("getting the service account token")
token, err := serviceAccountToken()
Expect(err).NotTo(HaveOccurred())
Expect(token).NotTo(BeEmpty())
By("ensuring the controller pod is ready")
verifyControllerPodReady := func(g Gomega) {
cmd := exec.Command("kubectl", "get", "pod", controllerPodName, "-n", namespace,
"-o", "jsonpath={.status.conditions[?(@.type=='Ready')].status}")
output, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred())
g.Expect(output).To(Equal("True"), "Controller pod not ready")
}
Eventually(verifyControllerPodReady, 3*time.Minute, time.Second).Should(Succeed())
By("verifying that the controller manager is serving the metrics server")
verifyMetricsServerStarted := func(g Gomega) {
cmd := exec.Command("kubectl", "logs", controllerPodName, "-n", namespace)
output, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred())
g.Expect(output).To(ContainSubstring("Serving metrics server"),
"Metrics server not yet started")
}
Eventually(verifyMetricsServerStarted, 3*time.Minute, time.Second).Should(Succeed())
// +kubebuilder:scaffold:e2e-metrics-webhooks-readiness
By("creating the curl-metrics pod to access the metrics endpoint")
cmd = exec.Command("kubectl", "run", "curl-metrics", "--restart=Never",
"--namespace", namespace,
"--image=curlimages/curl:latest",
"--overrides",
fmt.Sprintf(`{
"spec": {
"containers": [{
"name": "curl",
"image": "curlimages/curl:latest",
"command": ["/bin/sh", "-c"],
"args": [
"for i in $(seq 1 30); do curl -v -k -H 'Authorization: Bearer %s' https://%s.%s.svc.cluster.local:8443/metrics && exit 0 || sleep 2; done; exit 1"
],
"securityContext": {
"readOnlyRootFilesystem": true,
"allowPrivilegeEscalation": false,
"capabilities": {
"drop": ["ALL"]
},
"runAsNonRoot": true,
"runAsUser": 1000,
"seccompProfile": {
"type": "RuntimeDefault"
}
}
}],
"serviceAccountName": "%s"
}
}`, token, metricsServiceName, namespace, serviceAccountName))
_, err = utils.Run(cmd)
Expect(err).NotTo(HaveOccurred(), "Failed to create curl-metrics pod")
By("waiting for the curl-metrics pod to complete.")
verifyCurlUp := func(g Gomega) {
cmd := exec.Command("kubectl", "get", "pods", "curl-metrics",
"-o", "jsonpath={.status.phase}",
"-n", namespace)
output, err := utils.Run(cmd)
g.Expect(err).NotTo(HaveOccurred())
g.Expect(output).To(Equal("Succeeded"), "curl pod in wrong status")
}
Eventually(verifyCurlUp, 5*time.Minute).Should(Succeed())
By("getting the metrics by checking curl-metrics logs")
verifyMetricsAvailable := func(g Gomega) {
metricsOutput, err := getMetricsOutput()
g.Expect(err).NotTo(HaveOccurred(), "Failed to retrieve logs from curl pod")
g.Expect(metricsOutput).NotTo(BeEmpty())
g.Expect(metricsOutput).To(ContainSubstring("< HTTP/1.1 200 OK"))
}
Eventually(verifyMetricsAvailable, 2*time.Minute).Should(Succeed())
})
// +kubebuilder:scaffold:e2e-webhooks-checks
// TODO: Customize the e2e test suite with scenarios specific to your project.
// Consider applying sample/CR(s) and check their status and/or verifying
// the reconciliation by using the metrics, i.e.:
// metricsOutput, err := getMetricsOutput()
// Expect(err).NotTo(HaveOccurred(), "Failed to retrieve logs from curl pod")
// Expect(metricsOutput).To(ContainSubstring(
// fmt.Sprintf(`controller_runtime_reconcile_total{controller="%s",result="success"} 1`,
// strings.ToLower(<Kind>),
// ))
})
})
// serviceAccountToken returns a token for the specified service account in the given namespace.
// It uses the Kubernetes TokenRequest API to generate a token by directly sending a request
// and parsing the resulting token from the API response.
func serviceAccountToken() (string, error) {
const tokenRequestRawString = `{
"apiVersion": "authentication.k8s.io/v1",
"kind": "TokenRequest"
}`
By("creating temporary file to store the token request")
secretName := fmt.Sprintf("%s-token-request", serviceAccountName)
tokenRequestFile := filepath.Join("/tmp", secretName)
err := os.WriteFile(tokenRequestFile, []byte(tokenRequestRawString), os.FileMode(0o644))
if err != nil {
return "", err
}
var out string
verifyTokenCreation := func(g Gomega) {
By("executing kubectl command to create the token")
cmd := exec.Command("kubectl", "create", "--raw", fmt.Sprintf(
"/api/v1/namespaces/%s/serviceaccounts/%s/token",
namespace,
serviceAccountName,
), "-f", tokenRequestFile)
output, err := cmd.CombinedOutput()
g.Expect(err).NotTo(HaveOccurred())
By("parsing the JSON output to extract the token")
var token tokenRequest
err = json.Unmarshal(output, &token)
g.Expect(err).NotTo(HaveOccurred())
out = token.Status.Token
}
Eventually(verifyTokenCreation).Should(Succeed())
return out, err
}
// getMetricsOutput retrieves and returns the logs from the curl pod used to access the metrics endpoint.
func getMetricsOutput() (string, error) {
By("getting the curl-metrics logs")
cmd := exec.Command("kubectl", "logs", "curl-metrics", "-n", namespace)
return utils.Run(cmd)
}
// tokenRequest is a simplified representation of the Kubernetes TokenRequest API response,
// containing only the token field that we need to extract.
type tokenRequest struct {
Status struct {
Token string `json:"token"`
} `json:"status"`
}
-210
View File
@@ -1,210 +0,0 @@
package utils
import (
"bufio"
"bytes"
"fmt"
"os"
"os/exec"
"strings"
. "github.com/onsi/ginkgo/v2" // nolint:revive,staticcheck
)
const (
certmanagerVersion = "v1.21.1"
certmanagerURLTmpl = "https://github.com/cert-manager/cert-manager/releases/download/%s/cert-manager.yaml"
defaultKindBinary = "kind"
defaultKindCluster = "kind"
)
func warnError(err error) {
_, _ = fmt.Fprintf(GinkgoWriter, "warning: %v\n", err)
}
// Run executes the provided command within this context
func Run(cmd *exec.Cmd) (string, error) {
dir, _ := GetProjectDir()
cmd.Dir = dir
if err := os.Chdir(cmd.Dir); err != nil {
_, _ = fmt.Fprintf(GinkgoWriter, "chdir dir: %q\n", err)
}
cmd.Env = append(os.Environ(), "GO111MODULE=on")
command := strings.Join(cmd.Args, " ")
_, _ = fmt.Fprintf(GinkgoWriter, "running: %q\n", command)
output, err := cmd.CombinedOutput()
if err != nil {
return string(output), fmt.Errorf("%q failed with error %q: %w", command, string(output), err)
}
return string(output), nil
}
// UninstallCertManager uninstalls the cert manager
func UninstallCertManager() {
url := fmt.Sprintf(certmanagerURLTmpl, certmanagerVersion)
cmd := exec.Command("kubectl", "delete", "-f", url)
if _, err := Run(cmd); err != nil {
warnError(err)
}
// Delete leftover leases in kube-system (not cleaned by default)
kubeSystemLeases := []string{
"cert-manager-cainjector-leader-election",
"cert-manager-controller",
}
for _, lease := range kubeSystemLeases {
cmd = exec.Command("kubectl", "delete", "lease", lease,
"-n", "kube-system", "--ignore-not-found", "--force", "--grace-period=0")
if _, err := Run(cmd); err != nil {
warnError(err)
}
}
}
// InstallCertManager installs the cert manager bundle.
func InstallCertManager() error {
url := fmt.Sprintf(certmanagerURLTmpl, certmanagerVersion)
cmd := exec.Command("kubectl", "apply", "-f", url)
if _, err := Run(cmd); err != nil {
return err
}
// Wait for cert-manager-webhook to be ready, which can take time if cert-manager
// was re-installed after uninstalling on a cluster.
cmd = exec.Command("kubectl", "wait", "deployment.apps/cert-manager-webhook",
"--for", "condition=Available",
"--namespace", "cert-manager",
"--timeout", "5m",
)
_, err := Run(cmd)
return err
}
// IsCertManagerCRDsInstalled checks if any Cert Manager CRDs are installed
// by verifying the existence of key CRDs related to Cert Manager.
func IsCertManagerCRDsInstalled() bool {
// List of common Cert Manager CRDs
certManagerCRDs := []string{
"certificates.cert-manager.io",
"issuers.cert-manager.io",
"clusterissuers.cert-manager.io",
"certificaterequests.cert-manager.io",
"orders.acme.cert-manager.io",
"challenges.acme.cert-manager.io",
}
// Execute the kubectl command to get all CRDs
cmd := exec.Command("kubectl", "get", "crds")
output, err := Run(cmd)
if err != nil {
return false
}
// Check if any of the Cert Manager CRDs are present
crdList := GetNonEmptyLines(output)
for _, crd := range certManagerCRDs {
for _, line := range crdList {
if strings.Contains(line, crd) {
return true
}
}
}
return false
}
// LoadImageToKindClusterWithName loads a local docker image to the kind cluster
func LoadImageToKindClusterWithName(name string) error {
cluster := defaultKindCluster
if v, ok := os.LookupEnv("KIND_CLUSTER"); ok {
cluster = v
}
kindOptions := []string{"load", "docker-image", name, "--name", cluster}
kindBinary := defaultKindBinary
if v, ok := os.LookupEnv("KIND"); ok {
kindBinary = v
}
cmd := exec.Command(kindBinary, kindOptions...)
_, err := Run(cmd)
return err
}
// GetNonEmptyLines converts given command output string into individual objects
// according to line breakers, and ignores the empty elements in it.
func GetNonEmptyLines(output string) []string {
var res []string
elements := strings.SplitSeq(output, "\n")
for element := range elements {
if element != "" {
res = append(res, element)
}
}
return res
}
// GetProjectDir will return the directory where the project is
func GetProjectDir() (string, error) {
wd, err := os.Getwd()
if err != nil {
return wd, fmt.Errorf("failed to get current working directory: %w", err)
}
wd = strings.ReplaceAll(wd, "/test/e2e", "")
return wd, nil
}
// UncommentCode searches for target in the file and remove the comment prefix
// of the target content. The target content may span multiple lines.
func UncommentCode(filename, target, prefix string) error {
// false positive
// nolint:gosec
content, err := os.ReadFile(filename)
if err != nil {
return fmt.Errorf("failed to read file %q: %w", filename, err)
}
strContent := string(content)
idx := strings.Index(strContent, target)
if idx < 0 {
return fmt.Errorf("unable to find the code %q to be uncommented", target)
}
out := new(bytes.Buffer)
_, err = out.Write(content[:idx])
if err != nil {
return fmt.Errorf("failed to write to output: %w", err)
}
scanner := bufio.NewScanner(bytes.NewBufferString(target))
if !scanner.Scan() {
return nil
}
for {
if _, err = out.WriteString(strings.TrimPrefix(scanner.Text(), prefix)); err != nil {
return fmt.Errorf("failed to write to output: %w", err)
}
// Avoid writing a newline in case the previous line was the last in target.
if !scanner.Scan() {
break
}
if _, err = out.WriteString("\n"); err != nil {
return fmt.Errorf("failed to write to output: %w", err)
}
}
if _, err = out.Write(content[idx+len(target):]); err != nil {
return fmt.Errorf("failed to write to output: %w", err)
}
// false positive
// nolint:gosec
if err = os.WriteFile(filename, out.Bytes(), 0644); err != nil {
return fmt.Errorf("failed to write file %q: %w", filename, err)
}
return nil
}