b0a431f2a4294b30f1a967c8960f36f6240ceebe
10 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
50ce5bcec0 |
examples/demo: run terdut-server v0.43.0, with alerts from two clusters
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the web UI since: the queue and incident layouts, the rota and escalation pages, the theme toggle, and the cluster chip, filter and page titles (v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the floor, which is what the replicas setting actually depends on. fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus external label plus `cluster` in Alertmanager's group_by (terdut-server's README, "Several clusters, one team"). It goes on the alert's labels and groupLabels, and into the group key and the fingerprint, so the same alert in two clusters is two incidents and not one. Unset, the payload is exactly what it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in both, so the queue has a chip and a filter to show. run-demo.sh also failed on its second run, though it says it is safe to re-run: it expected HTTP 409 when alice already exists, but a spent invite is answered with 403 "invite link is not usable" before the username is ever checked. It now tries to log alice in first and skips the signup if that works. Checked on the kind cluster: the server rolled to v0.43.0, every CR became Ready and Adopted (server, both teams, both escalation rules, both dead man's switches, both alert sources), and /api/incidents/clusters, /api/incidents?cluster=prod-us and the incident titles came back as expected. No operator code changed, so this needs no operator release. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
822c80dda6 |
examples/demo: split the Ready wait so escalation rules wait on alice too
wait_for_ready waited for every demo object at once, including terdutescalationrule-platform, which names alice as a level-1 target -- but alice does not exist yet at that point in main(): she is created by redeem_platform_invite, which ran after wait_for_ready. terdut-server resolves every named username at reconcile time, not just when an escalation actually fires, so that CR could never reach Ready before alice did, and main() had no step in between to create her. Split into wait_for_objects (the shared loop, now taking its object list as arguments) plus two callers: wait_for_teams_ready, covering just the server and the two teams redeem_platform_invite/join_payments_team need, run before alice exists; wait_for_remaining_ready, covering the escalation rules, dead man's switches and alert sources, run after. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
0ee7ede648 |
examples/demo: match v0.4.0's new spec.replicas default and bump terdut-server
replicas: 1 and tag: v0.34.0 were both correct when written, but the CRD's own default moved to 2 in v0.4.0 (same release this demo is meant to show off), and v0.34.0 predates v0.36.0's advisory locks that make a second replica safe instead of racing the first. Left as-is, the demo would have been the one place in this repo demonstrating the exact unsafe combination the CRD's own doc comment warns against: more than one replica against an image that doesn't guard the sweeper/notifier/migration-runner singletons. replicas is now stated explicitly as 2 rather than dropped to pick up the default silently, matching every other field in this file's own habit of spelling out what it depends on. tag moves to v0.36.0 specifically -- the first version where the lock landed -- with the comment keeping v0.34.0's original reasoning (the service-account race fix) alongside the new one, since v0.36.0 still carries that fix forward. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
4aa4f17c42 |
examples/demo: bump terdut-server to v0.34.0 (the service-account fix)
Required for this demo to actually exercise the fix for niklas/terdut-operator#3 -- v0.33.2 still has the authorization gap this demo hit live (callerMayManageServiceAccount had no branch letting an instance-scoped account adopt a team-scoped account's key). |
||
|
|
a0ea13955e |
TerdutTeam: mint and surface a real invite link (spec.invite)
The actual fix for the human-onboarding gap niklas/terdut-server#23 found -- not a terdut-server change at all. A team-scoped credential is already owner-equivalent for POST/GET/DELETE /api/teams/{teamID}/invites (requireTeamOwner's synthetic-membership mechanism, ratified not accidental per that repo's SERVICE-ACCOUNTS.md), and invite redemption bypasses signup_mode entirely -- this TerdutTeam controller just never grew a feature to use either fact. New spec.invite{enabled, role (member|owner, default member), maxUses (1-100, default 1)} and status.inviteSecretRef. The Secret lives in the TerdutTeam's OWN namespace, not the operator's: unlike status.credentialsSecretRef (a durable, high-privilege credential, kept operator-side per DESIGN.md §6), an invite is bounded and limited-use, meant for this namespace's own human operators to read and hand out -- same precedent as TerdutAlertSource's status.webhookURLSecretRef, same- namespace and OwnerReference'd so deleting the TerdutTeam garbage-collects it automatically. internal/controller/terdutteam_invite.go: mints on first spec.invite.enabled, refreshes a day ahead of terdut-server's fixed 7-day TTL (reading the Secret's own stored expiresAt, no extra server round-trip per reconcile), revokes server-side and deletes the Secret when flipped back to false. A lost invite Secret is silently re-minted rather than treated as unrecoverable the way TerdutAlertSource's webhook key is -- nothing external holds a durable dependency on one specific invite link staying stable, it's read once by one human and handed out. New tdclient.Invite/CreateInvite/RevokeInvite. New envtest coverage: mint into the team's own namespace, refresh-before-expiry, revoke-on-disable (internal/controller/terdutteam_controller_test.go's new "spec.invite" Describe block), plus the fake server growing invite support (terdutserver_controller_test.go) -- its handleTeamSubPath dispatcher was split further (deadman switches into their own handleDeadmanSubPath, matching the existing handleIntegrationSubPath precedent) to stay under golangci-lint's gocyclo threshold with the new route added. examples/demo updated to prove this end to end: 02-team-platform.yaml turns on spec.invite; run-demo.sh's bootstrap_login/join_demo_teams (the psql signup_mode flip + a direct team_members INSERT) are replaced by redeem_platform_invite (reads status.inviteSecretRef, a real POST /api/signup with the invite token) and join_payments_team (POST /api/teams/{teamID}/members using Payments' own credential and alice's user id resolved via GET /api/users, deliberately not given its own spec.invite, so the demo shows both onboarding paths this feature unlocks) -- zero kubectl exec/psql calls remain anywhere in the script. README.md's "First login" section rewritten to match; it no longer documents the admin-token curl call that 403s against current terdut-server (niklas/terdut-server#23). Depends on niklas/terdut-server#24 (the callerMayManageServiceAccount fix for terdut-operator#3) being released before this is deployed for real -- not required to build or test this change itself, since the envtest fake never modeled that authorization gap to begin with. |
||
|
|
2a08a8cd8e |
examples/demo: add run-demo.sh, an automated kind-cluster demo
One script, two modes (run-demo.sh / run-demo.sh --teardown), that takes a fresh empty kind cluster all the way to a working demo: creates the cluster if needed, helm-installs this chart, applies every CRD kind in this directory, waits for all nine objects to go Ready, then does what the README's own first-login section cannot (see niklas/terdut-server#23 and niklas/terdut-operator#3 -- no service-account credential this operator holds can ever call /api/admin/settings or POST /api/users) by reaching into the demo's own throwaway Postgres directly: flips signup_mode to open, signs alice up for real over the ordinary signup endpoint, and joins her to both Platform and Payments (open signup always creates its own new team, never joins an existing one by name, so without this she'd have a working login that can't see a single incident this demo fires -- /api/incidents and /api/alerts are both scoped to the caller's own team memberships). Finishes by port-forwarding the service and firing fire-alerts.sh at both teams, so a fresh run already has visible incidents waiting in the web UI. Verified end to end against a real kind cluster, including a second, genuinely-fresh run that hit niklas/terdut-operator#3 live (terdutteam- platform wedged in the 403 retry loop that issue describes) -- confirmed the script itself fails cleanly on that (clear FAILED message, correct exit code, no orphaned port-forward) rather than hanging or leaving a mess, which is the most this script can do about a bug in the operator it's driving. |
||
|
|
6a699d4341 |
Let TerdutServer customize its pod, and never manage its own ingress
spec.pod (api/v1alpha1/terdutserver_types.go): annotations, nodeSelector,
tolerations, affinity, topologySpreadConstraints, resources, pod and
container securityContext, serviceAccountName, extraEnv/extraEnvFrom,
extraVolumes/extraVolumeMounts, imagePullSecrets, and an optional
disruptionBudget. All direct corev1 passthrough -- no wrapper types buy
anything for any of these, matching how CloudNativePG and the Zalando
postgres-operator both expose the same knobs, and matching this repo's
own SweeperSpec precedent ("wrap only when a round-trip through a
different type buys something"). affinity is pure user-supplied
passthrough, not a toggle-plus-generated-default the way a multi-replica
cluster operator's pod anti-affinity usually is: this operator never
auto-generates one, since spec.replicas above 1 isn't a supported
topology (the sweeper/notifier singleton constraint). Considered and
declined for this round: priorityClassName, pod labels beyond
annotations, and a HorizontalPodAutoscaler -- the last of those would
directly contradict the singleton constraint above.
disruptionBudget is the one field here that isn't a plain PodTemplateSpec
knob: when set, the controller now reconciles a PodDisruptionBudget
selecting the TerdutServer's own pods (new terdutserver_pdb.go); clearing
it deletes any it previously created. New RBAC marker on
poddisruptionbudgets to match.
Driven by a public-release pass: looking past this project's own use case
at what a mature, general-purpose operator CRD exposes here (researched
against Zalando postgres-operator and CloudNativePG specifically), not
just the fields this install happened to need.
Separately, and found while answering a question about exposing
TerdutServer through Istio instead of Gateway API: spec.networking's own
doc comment quietly promised a Gateway API HTTPRoute this operator would
build eventually ("a near-term follow-up, not deferred"). That promise is
wrong for a public release -- an operator managing someone's ingress
mechanism for them is a worse default than not touching it at all, and a
surprise HTTPRoute appearing once that follow-up eventually landed would
have been exactly backwards for an Istio (or plain-Ingress, or
intentionally-unexposed) install. Made the non-goal explicit and
permanent instead (DESIGN.md §1), removed the dead `gatewayListener`
field it was the only consumer of (zero runtime call sites anywhere --
setting it already had no effect, so this is a schema cleanup, not a
behavior change), and corrected ROADMAP.md's framing. hostname/servicePort
stay: both are live (TERDUT_PUBLIC_URL, container/Service port), this
operator just never acts on hostname for exposure. Added
examples/networking (Gateway API HTTPRoute, Istio VirtualService) showing
how to expose the plain ClusterIP Service the operator already creates --
outside the operator itself, as illustrations, not as something
examples/demo applies automatically.
No new terdut-server version requirement: both changes are CRD/controller-
only, nothing about the API this operator's bootstrap flow depends on
changed.
|
||
|
|
d9315322fc |
examples/demo: fix two real bugs this exact demo just hit live
1. Renamed every object this demo creates (TerdutServer, Postgres
Secret/Deployment/Service) from terdut-demo[-postgres] to
terdut-operator-demo[-postgres]. The user applied this kit into the
already-live "terdut-demo" namespace -- the real operator exercise
from earlier in this repo's own history -- and this demo's own
TerdutServer/Postgres objects shared that exact name. The TerdutServer
apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
while the live object already had postgresClusterRef violates "exactly
one of" and the API server refused it), and the real Postgres Service
was never touched (confirmed live: still Zalando's own spilo selector,
endpoint still the real StatefulSet pod) -- but the Postgres Secret and
Deployment, having no such protection, were created as brand new,
extra, crash-looping objects sitting right next to the real ones.
Prefixing every name this demo creates means a repeat of this exact
mistake no longer collides with anything, documented directly in
README.md now.
2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
(added responding to a PodSecurity "restricted" warning) took
CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
entrypoint needs to chown/chmod the data directory before it drops
privileges itself -- confirmed in a real crashed pod's logs: `chmod:
/var/run/postgresql: Operation not permitted`. kubectl apply
--dry-run=server, which is as far as this got verified before, only
checks admission policy; it was never actually booted. Removed the
capability drop and verified for real this time: applied just
00-postgres.yaml alone into a disposable namespace, waited for the pod
to go Ready, read its logs ("database system is ready to accept
connections"), then deleted that namespace.
|
||
|
|
ae97d28444 |
Add wait-for-postgres init container to the generated Deployment
This operator's own Deployment template crash-looped a few times against a from-scratch postgres-operator cluster still doing initdb and Patroni leader election -- exactly the gap examples/demo's own README just documented for it. terdut-server's ping-retry budget on startup (internal/db/db.go in that repo) is sized for a much shorter, different race (NetworkPolicy propagation, a few seconds), not genuine first-time cluster creation, so it exhausted and the process exited before ever binding its HTTP port -- a startupProbe cannot fix that, since the crash happens before there is anything to probe. Same root cause and same fix as charts/terdut-server's own deployment.yaml template as of that repo's v0.33.2. waitForPostgresContainer reuses dbEnv unchanged: both of resolveDatabaseEnv's two paths (DSN, postgresClusterRef) put TERDUT_DB_DSN first, so it's already exactly what pg_isready needs, and pg_isready needs no credentials, so dbEnv's optional PGPASSWORD riding along too is harmless rather than load-bearing. Covered by the existing envtest suite (asserts on Containers[0], the main container, unaffected by adding InitContainers) -- `make test` passes unchanged, 71.7% coverage on internal/controller. Updates examples/demo's own README, which no longer needs to warn about this. |
||
|
|
ca2cd2c645 |
Add examples/demo: one of every CRD, plus a script to fire alerts at it
A self-contained demo kit: a TerdutServer against a throwaway, bare Postgres (bring-your-own DSN -- simplest path to stand up from nothing, ROADMAP.md Stage 1's own note), two TerdutTeams, and each team's own TerdutEscalationRule/TerdutDeadmanSwitch/TerdutAlertSource, so every CRD this operator manages is exercised together rather than in isolation the way config/samples' one-of-each already does. fire-alerts.sh sends terdut-server's own amPayload/amAlert shape (read from internal/api/alertmanager.go in that repo, not guessed from its docs) at whichever TerdutAlertSource's generated webhook Secret it reads the key out of -- high-cpu/disk-full/pod-crash scenarios to open and resolve incidents, and a heartbeat scenario matching each team's dead man's switch matcher, so stopping it demonstrates the switch noticing silence on its own. Verified server-side (kubectl apply --dry-run=server -k examples/demo) against this operator's own dev cluster, which already has these CRDs installed: every object validates. The one warning that cluster's "restricted" PodSecurity raises (postgres:17-alpine's entrypoint needs to start as root before it drops privileges itself) is noted inline in 00-postgres.yaml rather than worked around -- not a real production pattern, and this Postgres exists only to be thrown away with the rest of the demo namespace. README.md walks through: applying, watching status, why a few early CrashLoopBackOff restarts on terdut-demo itself are expected (this operator's Deployment template has no wait-for-postgres init container yet, unlike charts/terdut-server's chart as of v0.33.2), reaching the web UI (port-forward -- spec.networking.hostname is accepted but nothing creates an HTTPRoute for it yet), turning on open signup with the operator's own generated admin token since the bootstrap-created account has no password, firing alerts, and tearing down. |