16 Commits

Author SHA1 Message Date
Niklas Ye 88172ade29 Set the chart's placeholder version to 0.3.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 1m53s
Release / test (push) Successful in 1m45s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 5m27s
Release / scan-image (push) Successful in 2s
2026-10-02 22:19:18 +02:00
niklas 478ae6284a Merge pull request 'TerdutTeam: mint and surface a real invite link (spec.invite)' (#5) from terdutteam-invite-minting into main
CI / test (push) Has been cancelled
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
2026-10-02 20:17:21 +00:00
niklas fcc80b32d5 Merge pull request 'examples/demo: add run-demo.sh, an automated kind-cluster demo' (#4) from examples-demo/run-demo-script into main
CI / chart (push) Successful in 2s
CI / security (push) Successful in 1m31s
CI / test (push) Successful in 3m30s
2026-10-02 20:13:36 +00:00
Niklas Ye 4aa4f17c42 examples/demo: bump terdut-server to v0.34.0 (the service-account fix)
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m18s
CI / test (pull_request) Successful in 3m40s
Required for this demo to actually exercise the fix for
niklas/terdut-operator#3 -- v0.33.2 still has the authorization gap this
demo hit live (callerMayManageServiceAccount had no branch letting an
instance-scoped account adopt a team-scoped account's key).
2026-10-02 22:13:26 +02:00
Niklas Ye a0ea13955e TerdutTeam: mint and surface a real invite link (spec.invite)
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 52s
CI / test (pull_request) Successful in 2m43s
The actual fix for the human-onboarding gap niklas/terdut-server#23 found --
not a terdut-server change at all. A team-scoped credential is already
owner-equivalent for POST/GET/DELETE /api/teams/{teamID}/invites
(requireTeamOwner's synthetic-membership mechanism, ratified not
accidental per that repo's SERVICE-ACCOUNTS.md), and invite redemption
bypasses signup_mode entirely -- this TerdutTeam controller just never
grew a feature to use either fact.

New spec.invite{enabled, role (member|owner, default member), maxUses
(1-100, default 1)} and status.inviteSecretRef. The Secret lives in the
TerdutTeam's OWN namespace, not the operator's: unlike
status.credentialsSecretRef (a durable, high-privilege credential, kept
operator-side per DESIGN.md §6), an invite is bounded and limited-use,
meant for this namespace's own human operators to read and hand out --
same precedent as TerdutAlertSource's status.webhookURLSecretRef, same-
namespace and OwnerReference'd so deleting the TerdutTeam garbage-collects
it automatically.

internal/controller/terdutteam_invite.go: mints on first spec.invite.enabled,
refreshes a day ahead of terdut-server's fixed 7-day TTL (reading the
Secret's own stored expiresAt, no extra server round-trip per reconcile),
revokes server-side and deletes the Secret when flipped back to false. A
lost invite Secret is silently re-minted rather than treated as
unrecoverable the way TerdutAlertSource's webhook key is -- nothing
external holds a durable dependency on one specific invite link staying
stable, it's read once by one human and handed out.

New tdclient.Invite/CreateInvite/RevokeInvite. New envtest coverage: mint
into the team's own namespace, refresh-before-expiry, revoke-on-disable
(internal/controller/terdutteam_controller_test.go's new "spec.invite"
Describe block), plus the fake server growing invite support
(terdutserver_controller_test.go) -- its handleTeamSubPath dispatcher was
split further (deadman switches into their own handleDeadmanSubPath,
matching the existing handleIntegrationSubPath precedent) to stay under
golangci-lint's gocyclo threshold with the new route added.

examples/demo updated to prove this end to end: 02-team-platform.yaml
turns on spec.invite; run-demo.sh's bootstrap_login/join_demo_teams (the
psql signup_mode flip + a direct team_members INSERT) are replaced by
redeem_platform_invite (reads status.inviteSecretRef, a real POST
/api/signup with the invite token) and join_payments_team (POST
/api/teams/{teamID}/members using Payments' own credential and alice's
user id resolved via GET /api/users, deliberately not given its own
spec.invite, so the demo shows both onboarding paths this feature
unlocks) -- zero kubectl exec/psql calls remain anywhere in the script.
README.md's "First login" section rewritten to match; it no longer
documents the admin-token curl call that 403s against current
terdut-server (niklas/terdut-server#23).

Depends on niklas/terdut-server#24 (the callerMayManageServiceAccount fix
for terdut-operator#3) being released before this is deployed for real --
not required to build or test this change itself, since the envtest fake
never modeled that authorization gap to begin with.
2026-10-02 22:04:56 +02:00
Niklas Ye 2a08a8cd8e examples/demo: add run-demo.sh, an automated kind-cluster demo
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 1m5s
CI / test (pull_request) Successful in 2m49s
One script, two modes (run-demo.sh / run-demo.sh --teardown), that takes a
fresh empty kind cluster all the way to a working demo: creates the
cluster if needed, helm-installs this chart, applies every CRD kind in
this directory, waits for all nine objects to go Ready, then does what
the README's own first-login section cannot (see niklas/terdut-server#23
and niklas/terdut-operator#3 -- no service-account credential this
operator holds can ever call /api/admin/settings or POST /api/users) by
reaching into the demo's own throwaway Postgres directly: flips
signup_mode to open, signs alice up for real over the ordinary signup
endpoint, and joins her to both Platform and Payments (open signup always
creates its own new team, never joins an existing one by name, so
without this she'd have a working login that can't see a single incident
this demo fires -- /api/incidents and /api/alerts are both scoped to the
caller's own team memberships). Finishes by port-forwarding the service
and firing fire-alerts.sh at both teams, so a fresh run already has
visible incidents waiting in the web UI.

Verified end to end against a real kind cluster, including a second,
genuinely-fresh run that hit niklas/terdut-operator#3 live (terdutteam-
platform wedged in the 403 retry loop that issue describes) -- confirmed
the script itself fails cleanly on that (clear FAILED message, correct
exit code, no orphaned port-forward) rather than hanging or leaving a
mess, which is the most this script can do about a bug in the operator
it's driving.
2026-10-02 21:23:44 +02:00
Niklas Ye 375b5ed2f7 Set the chart's placeholder version to 0.2.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 57s
CI / test (push) Successful in 2m13s
Release / test (push) Successful in 1m41s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m10s
Release / scan-image (push) Successful in 3s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree
heading for v0.2.0 that still says 0.1.2 tells its reader something
false. Same as 738c210 and 0e3118d before it.
2026-10-02 18:51:00 +02:00
Niklas Ye 6a699d4341 Let TerdutServer customize its pod, and never manage its own ingress
spec.pod (api/v1alpha1/terdutserver_types.go): annotations, nodeSelector,
tolerations, affinity, topologySpreadConstraints, resources, pod and
container securityContext, serviceAccountName, extraEnv/extraEnvFrom,
extraVolumes/extraVolumeMounts, imagePullSecrets, and an optional
disruptionBudget. All direct corev1 passthrough -- no wrapper types buy
anything for any of these, matching how CloudNativePG and the Zalando
postgres-operator both expose the same knobs, and matching this repo's
own SweeperSpec precedent ("wrap only when a round-trip through a
different type buys something"). affinity is pure user-supplied
passthrough, not a toggle-plus-generated-default the way a multi-replica
cluster operator's pod anti-affinity usually is: this operator never
auto-generates one, since spec.replicas above 1 isn't a supported
topology (the sweeper/notifier singleton constraint). Considered and
declined for this round: priorityClassName, pod labels beyond
annotations, and a HorizontalPodAutoscaler -- the last of those would
directly contradict the singleton constraint above.

disruptionBudget is the one field here that isn't a plain PodTemplateSpec
knob: when set, the controller now reconciles a PodDisruptionBudget
selecting the TerdutServer's own pods (new terdutserver_pdb.go); clearing
it deletes any it previously created. New RBAC marker on
poddisruptionbudgets to match.

Driven by a public-release pass: looking past this project's own use case
at what a mature, general-purpose operator CRD exposes here (researched
against Zalando postgres-operator and CloudNativePG specifically), not
just the fields this install happened to need.

Separately, and found while answering a question about exposing
TerdutServer through Istio instead of Gateway API: spec.networking's own
doc comment quietly promised a Gateway API HTTPRoute this operator would
build eventually ("a near-term follow-up, not deferred"). That promise is
wrong for a public release -- an operator managing someone's ingress
mechanism for them is a worse default than not touching it at all, and a
surprise HTTPRoute appearing once that follow-up eventually landed would
have been exactly backwards for an Istio (or plain-Ingress, or
intentionally-unexposed) install. Made the non-goal explicit and
permanent instead (DESIGN.md §1), removed the dead `gatewayListener`
field it was the only consumer of (zero runtime call sites anywhere --
setting it already had no effect, so this is a schema cleanup, not a
behavior change), and corrected ROADMAP.md's framing. hostname/servicePort
stay: both are live (TERDUT_PUBLIC_URL, container/Service port), this
operator just never acts on hostname for exposure. Added
examples/networking (Gateway API HTTPRoute, Istio VirtualService) showing
how to expose the plain ClusterIP Service the operator already creates --
outside the operator itself, as illustrations, not as something
examples/demo applies automatically.

No new terdut-server version requirement: both changes are CRD/controller-
only, nothing about the API this operator's bootstrap flow depends on
changed.
2026-10-02 18:50:38 +02:00
Niklas Ye d9315322fc examples/demo: fix two real bugs this exact demo just hit live
CI / chart (push) Successful in 1s
CI / security (push) Successful in 3m24s
CI / test (push) Successful in 10m45s
1. Renamed every object this demo creates (TerdutServer, Postgres
   Secret/Deployment/Service) from terdut-demo[-postgres] to
   terdut-operator-demo[-postgres]. The user applied this kit into the
   already-live "terdut-demo" namespace -- the real operator exercise
   from earlier in this repo's own history -- and this demo's own
   TerdutServer/Postgres objects shared that exact name. The TerdutServer
   apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
   while the live object already had postgresClusterRef violates "exactly
   one of" and the API server refused it), and the real Postgres Service
   was never touched (confirmed live: still Zalando's own spilo selector,
   endpoint still the real StatefulSet pod) -- but the Postgres Secret and
   Deployment, having no such protection, were created as brand new,
   extra, crash-looping objects sitting right next to the real ones.
   Prefixing every name this demo creates means a repeat of this exact
   mistake no longer collides with anything, documented directly in
   README.md now.

2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
   (added responding to a PodSecurity "restricted" warning) took
   CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
   entrypoint needs to chown/chmod the data directory before it drops
   privileges itself -- confirmed in a real crashed pod's logs: `chmod:
   /var/run/postgresql: Operation not permitted`. kubectl apply
   --dry-run=server, which is as far as this got verified before, only
   checks admission policy; it was never actually booted. Removed the
   capability drop and verified for real this time: applied just
   00-postgres.yaml alone into a disposable namespace, waited for the pod
   to go Ready, read its logs ("database system is ready to accept
   connections"), then deleted that namespace.
2026-10-02 13:25:47 +02:00
Niklas Ye 5b45cf72e1 Bump go.opentelemetry.io/otel to v1.45.0: v1.44.0 carries GO-2026-6505
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m22s
CI / test (push) Successful in 5m0s
Release / test (push) Successful in 5m26s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 7m22s
Release / scan-image (push) Successful in 34s
Exporter config logging may leak endpoint URLs in info logs
(otlptrace/otlptracegrpc/sdk, transitively through grpc's own otel
instrumentation -- all indirect in go.mod, nothing imports these by
name). govulncheck flagged it reachable through real call chains
(tdclient.Client.DeleteIntegration, cmd/main.go's own init), caught by
ci.yaml's security job while cutting v0.1.2 (run 897) -- same pattern as
terdut-server's own da48814 for grpc's CVE-2026-84445.

go.opentelemetry.io/otel, /metric, /sdk, /sdk/metric, /trace,
/exporters/otlp/otlptrace, /exporters/otlp/otlptrace/otlptracegrpc all
moved 1.44.0 -> 1.45.0 together, plus go-logr/logr's own patch bump and
proto/otlp + genproto that go mod tidy pulled along with them. Verified:
go build, full test suite (71.7% coverage unchanged), golangci-lint,
helm-lint, and govulncheck itself now reporting zero reachable
vulnerabilities.
2026-10-02 12:59:40 +02:00
Niklas Ye 738c210505 Set the chart's placeholder version to 0.1.2
CI / chart (push) Successful in 1s
CI / security (push) Failing after 1m4s
CI / test (push) Successful in 2m11s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree heading
for v0.1.2 that still says 0.1.1 tells its reader something false. Same
as 0e3118d before it.
2026-10-02 12:56:16 +02:00
Niklas Ye 46ba0e8d5c CLAUDE.md: the wrapper-chart one-time step is done, not still pending
CI / chart (push) Successful in 1s
CI / test (push) Has been cancelled
CI / security (push) Has been cancelled
Stale since this repo's actual first release (v0.1.1, 2026-10-01) already
did it -- Ryuvia/charts/terdut-operator already exists and already pins
v0.1.1. Caught while about to repeat the same wrong assumption for this
release.
2026-10-02 12:55:23 +02:00
Niklas Ye ae97d28444 Add wait-for-postgres init container to the generated Deployment
CI / chart (push) Successful in 1s
CI / security (push) Failing after 57s
CI / test (push) Successful in 2m0s
This operator's own Deployment template crash-looped a few times against
a from-scratch postgres-operator cluster still doing initdb and Patroni
leader election -- exactly the gap examples/demo's own README just
documented for it. terdut-server's ping-retry budget on startup
(internal/db/db.go in that repo) is sized for a much shorter, different
race (NetworkPolicy propagation, a few seconds), not genuine first-time
cluster creation, so it exhausted and the process exited before ever
binding its HTTP port -- a startupProbe cannot fix that, since the crash
happens before there is anything to probe. Same root cause and same fix
as charts/terdut-server's own deployment.yaml template as of that repo's
v0.33.2.

waitForPostgresContainer reuses dbEnv unchanged: both of
resolveDatabaseEnv's two paths (DSN, postgresClusterRef) put
TERDUT_DB_DSN first, so it's already exactly what pg_isready needs, and
pg_isready needs no credentials, so dbEnv's optional PGPASSWORD riding
along too is harmless rather than load-bearing.

Covered by the existing envtest suite (asserts on Containers[0], the main
container, unaffected by adding InitContainers) -- `make test` passes
unchanged, 71.7% coverage on internal/controller. Updates examples/demo's
own README, which no longer needs to warn about this.
2026-10-02 10:11:55 +02:00
Niklas Ye ca2cd2c645 Add examples/demo: one of every CRD, plus a script to fire alerts at it
CI / chart (push) Successful in 1s
CI / security (push) Failing after 1m7s
CI / test (push) Successful in 2m37s
A self-contained demo kit: a TerdutServer against a throwaway, bare
Postgres (bring-your-own DSN -- simplest path to stand up from nothing,
ROADMAP.md Stage 1's own note), two TerdutTeams, and each team's own
TerdutEscalationRule/TerdutDeadmanSwitch/TerdutAlertSource, so every CRD
this operator manages is exercised together rather than in isolation the
way config/samples' one-of-each already does.

fire-alerts.sh sends terdut-server's own amPayload/amAlert shape (read
from internal/api/alertmanager.go in that repo, not guessed from its
docs) at whichever TerdutAlertSource's generated webhook Secret it reads
the key out of -- high-cpu/disk-full/pod-crash scenarios to open and
resolve incidents, and a heartbeat scenario matching each team's dead
man's switch matcher, so stopping it demonstrates the switch noticing
silence on its own.

Verified server-side (kubectl apply --dry-run=server -k examples/demo)
against this operator's own dev cluster, which already has these CRDs
installed: every object validates. The one warning that cluster's
"restricted" PodSecurity raises (postgres:17-alpine's entrypoint needs to
start as root before it drops privileges itself) is noted inline in
00-postgres.yaml rather than worked around -- not a real production
pattern, and this Postgres exists only to be thrown away with the rest of
the demo namespace.

README.md walks through: applying, watching status, why a few early
CrashLoopBackOff restarts on terdut-demo itself are expected (this
operator's Deployment template has no wait-for-postgres init container
yet, unlike charts/terdut-server's chart as of v0.33.2), reaching the web
UI (port-forward -- spec.networking.hostname is accepted but nothing
creates an HTTPRoute for it yet), turning on open signup with the
operator's own generated admin token since the bootstrap-created account
has no password, firing alerts, and tearing down.
2026-10-02 10:03:49 +02:00
Niklas Ye a75b23c4ad DESIGN.md: record the missing OIDC trustEmail field, found exercising a real second install
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 2m14s
Found while standing up terdut-demo (Ryuvia/charts#275), a second real
TerdutServer against the same Authentik provider as production: OIDCSpec
has no trustEmail override, so a demo install copying production's OIDC
config otherwise verbatim silently runs with the wrong default for it.
Not fixed here -- recorded in §13 as a real, found gap, not a decision,
same as the mid-life teamRef note already there.
2026-10-01 19:39:20 +02:00
Niklas Ye 4ab04d29a8 DESIGN.md: document operator mode and what it deliberately doesn't lock
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m3s
CI / test (push) Successful in 2m22s
No operator-mode section existed here before -- terdut-server's own
README.md documents the feature, but this repo's design doc never
mentioned it. Added as §6 point 7, confirmed against source
(internal/api/middleware.go's OperatorModeBlock, router.go's opMode
wrapper): it blocks human writes to exactly the resources this
operator's CRDs manage (team identity, OIDC-group binding, escalation,
dead man's switches, integrations), and nothing else -- team membership,
invites, and the on-call schedule/rota stay human-editable regardless,
confirmed from the router rather than assumed from the README's prose
alone.
2026-10-01 19:16:53 +02:00
44 changed files with 10207 additions and 132 deletions
+6 -8
View File
@@ -36,11 +36,9 @@ charts --force` after `config/` changes, then re-review — `--force` does not t
importantly the optional `terdutServer` block, DESIGN.md §10). It installs the operator + importantly the optional `terdutServer` block, DESIGN.md §10). It installs the operator +
CRDs + RBAC, and optionally one `TerdutServer` CR (`terdutServer.enabled`, off by default). CRDs + RBAC, and optionally one `TerdutServer` CR (`terdutServer.enabled`, off by default).
**One manual step the release skill's own automation does not cover**: `release-preflight` That one-time manual step — hand-creating the initial `terdut-operator/` wrapper entry
expects an existing `terdut-operator/` entry under `Ryuvia/charts` to bump on release under `Ryuvia/charts`, since `chart-bump` only ever bumps an existing one — is done. It
(steps 8-10 of the skill). There is no such entry yet — this repo's first-ever release happened during this repo's actual first release (v0.1.1, 2026-10-01; v0.1.0 published but
can publish its own image and chart (the `test`/`image`/`chart`/`scan-image` jobs), but never deployed anywhere, after its own `scan-image` found a CVE in the grpc version it had
the wrapper-chart bump and PR will fail until someone creates that initial wrapper entry just bumped to). Every release since bumps that wrapper entry like any other onboarded
in `Ryuvia/charts` by hand, the same one-time step every other onboarded repo already had repo's.
done for it before its own first release. That's a deliberate decision to deploy this
operator for real, not something to do as a side effect of finishing this stage.
+96 -8
View File
@@ -50,6 +50,14 @@ it just created, never a server something else already bootstrapped first.
that needs a live look at another object (e.g. "does this teamRef exist") that needs a live look at another object (e.g. "does this teamRef exist")
is a status condition, not an admission rejection — keeps v1 to a is a status condition, not an admission rejection — keeps v1 to a
controller-only deployment with no cert-manager/webhook dependency. controller-only deployment with no cert-manager/webhook dependency.
- **Never manages external exposure/ingress for `TerdutServer`, in any
form — a permanent non-goal, not a staged one.** The operator creates a
plain `ClusterIP` Service (§4.1) and stops there: no `Ingress`, no
Gateway API `HTTPRoute`, no Istio `VirtualService`, nothing. Some
installs won't expose `TerdutServer` outside the cluster at all; others
will use whichever of those mechanisms already fits their cluster. That
choice belongs to whoever deploys it, not to this operator — see
`examples/networking` for worked (but not operator-managed) examples.
## 2. The README's open questions, resolved ## 2. The README's open questions, resolved
@@ -149,7 +157,6 @@ spec:
networking: networking:
hostname: terdut.example.com hostname: terdut.example.com
servicePort: 8080 servicePort: 8080
gatewayListener: "" # same semantics as chart's networking.listener
database: database:
dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef
passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password} passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password}
@@ -188,6 +195,20 @@ spec:
# selector: # required, and only meaningful, when from: Selector # selector: # required, and only meaningful, when from: Selector
# matchLabels: # matchLabels:
# terdut.ryuvia.com/allowed: "true" # terdut.ryuvia.com/allowed: "true"
# Pod-level customization of the Deployment -- all optional, direct corev1
# passthrough throughout (see PodSpec's own doc comment). A representative
# subset:
pod:
resources:
requests: {cpu: 100m, memory: 128Mi}
limits: {memory: 256Mi}
tolerations:
- key: dedicated
operator: Equal
value: terdut
effect: NoSchedule
disruptionBudget:
minAvailable: 1 # mutually exclusive with maxUnavailable
status: status:
conditions: [...] # Ready, DatabaseReady, Bootstrapped conditions: [...] # Ready, DatabaseReady, Bootstrapped
observedGeneration: 3 observedGeneration: 3
@@ -214,6 +235,25 @@ per-name allowlist (no "and only these teams") — namespace-level consent is
the right granularity here, same reasoning as `ListenerSet`: the namespace the right granularity here, same reasoning as `ListenerSet`: the namespace
is the tenancy boundary, not the object. is the tenancy boundary, not the object.
`spec.pod` is pod-level customization of the Deployment, all optional and
directly reusing corev1 types wherever corev1 already models the knob
exactly (`tolerations`, `affinity`, `topologySpreadConstraints`,
`resources`, `securityContext`/`containerSecurityContext`, `extraEnv`/
`extraEnvFrom`, `extraVolumes`/`extraVolumeMounts`, `imagePullSecrets`) —
no custom wrapper buys anything for any of these, matching how
CloudNativePG and the Zalando postgres-operator both expose the same
knobs. `affinity` is pure user-supplied passthrough, not a
toggle-plus-generated-default the way a multi-replica-aware operator's
pod anti-affinity typically is: this operator never auto-generates
affinity of its own, since `replicas` above 1 isn't a supported topology
(the sweeper/notifier singleton constraint, §4.1's own illustrative YAML
comment). `spec.pod.disruptionBudget` is the one field here that isn't a
straight PodTemplateSpec knob — when set, the controller reconciles a
`PodDisruptionBudget` selecting this `TerdutServer`'s pods; clearing it
deletes any it previously created (§7). `minAvailable`/`maxUnavailable`
are mutually exclusive, `+kubebuilder:validation:XValidation`-guarded the
same way as `spec.database`'s own `dsn`/`postgresClusterRef` rule.
### 4.2 `TerdutTeam` ### 4.2 `TerdutTeam`
```yaml ```yaml
@@ -566,13 +606,40 @@ when nothing ever crosses into a tenant namespace in the first place.
Secret is updated in place. No DB-level workaround, no re-triggering a Secret is updated in place. No DB-level workaround, no re-triggering a
single-shot endpoint that can't fire twice (which is what made rotation single-shot endpoint that can't fire twice (which is what made rotation
unworkable under the old `/api/bootstrap`-only design). unworkable under the old `/api/bootstrap`-only design).
7. **Operator mode** (`TERDUT_OPERATOR_MODE`, terdut-server's own
deploy-time flag, off by default) is the complementary half of this
trust model: it makes terdut-server itself refuse a *human* write (a
session or a user's own API key) on a route, while a service account's
— this operator's — still goes through. Confirmed against source
(`internal/api/middleware.go`'s `OperatorModeBlock`,
`internal/api/router.go`'s `opMode` wrapper): the blocked set is exactly
team create/rename/delete, a team's OIDC-group binding, its escalation
policy, its dead man's switches, and its integrations — precisely the
resources `TerdutTeam`, `TerdutEscalationRule`, `TerdutDeadmanSwitch` and
`TerdutAlertSource` manage, and nothing more. **Deliberately not
blocked, confirmed against the same router**: team membership and
invites (the router's own comment: "membership is deliberately never
gitops-managed"), and the on-call schedule/rota
(`/api/teams/{teamID}/schedule`, `/api/schedule/current`) — neither
route carries the `opMode` wrapper at all. A human can still add or
remove a team member, or assign who's on call, on a server running in
operator mode; only the CRD-shaped resources above are locked to
GitOps. This isn't a gap to close — it's the same boundary §4.2 already
draws for membership, confirmed to hold on the server side too, not
just stated as an intent here.
## 7. Ownership, status, garbage collection ## 7. Ownership, status, garbage collection
- Every generated object that lives in the *same* namespace as the CR that - Every generated object that lives in the *same* namespace as the CR that
caused it (Deployment, Service, webhook Secret) carries a standard caused it (Deployment, Service, webhook Secret, and `TerdutServer`'s own
`metav1.OwnerReference` — GC handles these, no finalizer needed. The two PodDisruptionBudget) carries a standard `metav1.OwnerReference` — GC
credential Secrets from §6 are the one exception: they live in the handles these, no finalizer needed. PodDisruptionBudget is the one
member of that list that's conditionally created/deleted rather than
always present: it exists only while `spec.pod.disruptionBudget` is set,
and the controller deletes it itself the moment that field is cleared
(it doesn't wait on GC for that case, only for the `TerdutServer` being
deleted outright). The two credential Secrets from §6 are the one
exception to OwnerReference-based cleanup generally: they live in the
operator's own namespace regardless of where their owning CR lives, so operator's own namespace regardless of where their owning CR lives, so
`OwnerReference` doesn't apply (cross-namespace) and cleanup instead runs `OwnerReference` doesn't apply (cross-namespace) and cleanup instead runs
through that CR's finalizer directly, alongside the server-side DELETE through that CR's finalizer directly, alongside the server-side DELETE
@@ -619,10 +686,10 @@ documented and tested operationally:
- The operator's own ServiceAccount needs, per namespace it's granted: - The operator's own ServiceAccount needs, per namespace it's granted:
`get/list/watch/create/update/patch/delete` on `Deployments`, `Services` `get/list/watch/create/update/patch/delete` on `Deployments`, `Services`
it owns, and `get/list/watch` on `postgresql.acid.zalan.do` (optional, and `PodDisruptionBudgets` it owns, and `get/list/watch` on
degrade gracefully if absent per §8), plus cluster-wide `get/list` on `postgresql.acid.zalan.do` (optional, degrade gracefully if absent per
`Namespace` (labels only, for `allowedTeams: {from: Selector}` evaluation §8), plus cluster-wide `get/list` on `Namespace` (labels only, for
— §4.1, §4.6). `allowedTeams: {from: Selector}` evaluation — §4.1, §4.6).
- **Two different `Secret` scopes, not one — corrected from an earlier draft - **Two different `Secret` scopes, not one — corrected from an earlier draft
of this section.** That earlier draft said `Secret` access was "scoped to of this section.** That earlier draft said `Secret` access was "scoped to
the operator's own namespace only... nowhere else," reasoning that with the operator's own namespace only... nowhere else," reasoning that with
@@ -746,8 +813,29 @@ what it was, a separate install, until someone deletes it.
integration/policy/switch to a different team in place regardless, so integration/policy/switch to a different team in place regardless, so
retargeting one onto a live child isn't a supported operation in v1 — retargeting one onto a live child isn't a supported operation in v1 —
delete and recreate the CR instead. delete and recreate the CR instead.
- `TerdutServerSpec.OIDC` has no `trustEmail` field (nor `usernameClaim`,
`emailClaim`, `groupsClaim` — the "rather than being added here
speculatively" fields its own doc comment already names), unlike
`charts/terdut-server`'s own chart, which sets `oidc.trustEmail: true`
for the production install specifically because Authentik reports
`email_verified: false` and without it a user's first SSO sign-in
creates a second, empty account instead of linking to their existing
one (`terdut-server/README.md`'s own account of this). Found by
actually trying to stand up a second real `TerdutServer` against the
same Authentik provider (`terdut-demo`, `Ryuvia/charts#275`), not by
inspection: that install's OIDC config is otherwise a straight copy of
production's and runs with terdut-server's own default (`trustEmail:
false`) regardless, since the CRD has nowhere to put the override.
Worth closing if a second real OIDC install becomes routine rather than
a one-off exercise.
- Gitops-managed team *membership* (see §4.2). - Gitops-managed team *membership* (see §4.2).
- Automatic Deployment restart on upstream Postgres credential rotation. - Automatic Deployment restart on upstream Postgres credential rotation.
- `spec.pod.priorityClassName`, pod-label passthrough beyond
`spec.pod.annotations`, and a HorizontalPodAutoscaler for `TerdutServer`
— all considered alongside §4.1's `spec.pod` and explicitly left out of
that round: an HPA in particular would actively contradict
`spec.replicas`'s own stance that this operator doesn't support more
than one replica (the sweeper/notifier singleton constraint).
- Admission webhooks / CEL-only validation limits (e.g. verifying a - Admission webhooks / CEL-only validation limits (e.g. verifying a
`teamRef` exists at admission time rather than surfacing it as a status `teamRef` exists at admission time rather than surfacing it as a status
condition after the fact). condition after the fact).
+8
View File
@@ -34,3 +34,11 @@ DESIGN.md §4.1, §4.6.
- `teamRef` (DESIGN.md §4.5) - `teamRef` (DESIGN.md §4.5)
- URL/key are generated by the server at creation and surfaced only via a - URL/key are generated by the server at creation and surfaced only via a
generated Secret, never set explicitly generated Secret, never set explicitly
## Demo
[`examples/demo`](./examples/demo) wires one of every CRD above together
— two teams, each with an escalation rule, a dead man's switch and an
alert source — plus a script that fires synthetic Alertmanager webhooks
at it, so you can watch real incidents open, escalate and resolve without
a real Alertmanager anywhere in the picture.
+10 -8
View File
@@ -72,15 +72,17 @@ New commits build forward over the old ones; no git history rewrite.
individual keys, so there's nothing to undo there regardless. individual keys, so there's nothing to undo there regardless.
- RBAC: read-only watch on `postgresql.acid.zalan.do`, degrading gracefully - RBAC: read-only watch on `postgresql.acid.zalan.do`, degrading gracefully
if that CRD isn't installed (§8, §9). if that CRD isn't installed (§8, §9).
- Shipped, scoped down from §8's full ambition in two ways, both called out - Shipped, scoped down from §8's full ambition in one way, called out in
in code rather than silently dropped: no live watch on the Zalando- code rather than silently dropped: no live watch on the Zalando-
generated credentials Secret for rotation (relies on the periodic resync generated credentials Secret for rotation (relies on the periodic resync
to notice eventually, higher latency than a watch); no Gateway API to notice eventually, higher latency than a watch). A near-term
`HTTPRoute` creation from `spec.networking.hostname`/`gatewayListener` follow-up, not deferred to a later stage.
(needs the Gateway API types as a new dependency, and nothing about - External exposure (a Gateway API `HTTPRoute` from
proving a `TerdutServer` boots and bootstraps a real server depends on `spec.networking.hostname`) was originally sketched here too, as a
external ingress existing). Both are near-term follow-ups, not deferred second near-term follow-up alongside the one above. It's since become an
to a later stage. explicit, permanent non-goal instead (DESIGN.md §1): the operator will
never manage ingress/exposure for `TerdutServer` in any form. See
`examples/networking` for how to do that yourself.
- `envtest` covering Deployment/Service reconciliation and both database - `envtest` covering Deployment/Service reconciliation and both database
paths — the Zalando path needs that CRD's schema vendored into the test paths — the Zalando path needs that CRD's schema vendored into the test
environment (there's no real `postgres-operator` controller in `envtest`, environment (there's no real `postgres-operator` controller in `envtest`,
+125 -20
View File
@@ -1,8 +1,10 @@
package v1alpha1 package v1alpha1
import ( import (
corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/util/intstr"
) )
// SecretKeyRef names one data key inside a Secret. Every use of this type in // SecretKeyRef names one data key inside a Secret. Every use of this type in
@@ -29,33 +31,28 @@ type ImageSpec struct {
Tag string `json:"tag"` Tag string `json:"tag"`
} }
// NetworkingSpec is how this TerdutServer is reached from outside the // NetworkingSpec configures terdut-server itself and the plain ClusterIP
// cluster. // Service the operator creates in front of it. It does not expose
// // TerdutServer outside the cluster in any way, and never will (DESIGN.md
// hostname/gatewayListener describe the intended Gateway API HTTPRoute // §1 -- a permanent non-goal, not a staged one): exposing it is entirely up
// (matching charts/terdut-server's own templates/httpproxy.yaml, despite its // to whoever deploys it -- a Gateway API HTTPRoute, a plain Ingress, an
// name — that chart carries a Gateway API HTTPRoute, not a Contour // Istio VirtualService, or nothing at all if it should stay cluster-
// HTTPProxy), but creating that HTTPRoute isn't implemented yet: it needs // internal. See examples/networking for worked examples against the
// the Gateway API types as a new dependency, and nothing about proving a // Service this creates.
// TerdutServer boots and bootstraps a real server depends on external
// ingress existing. Tracked as a near-term follow-up, not deferred to a
// later ROADMAP.md stage the way Deployment/database/bootstrap once were.
type NetworkingSpec struct { type NetworkingSpec struct {
// hostname the HTTPRoute will carry once it exists. // hostname is terdut-server's own public URL (TERDUT_PUBLIC_URL) --
// used for absolute links terdut-server generates itself
// (notifications, OIDC redirect URIs), not read by this operator for
// anything ingress-related. Set it to whatever hostname your own
// exposure mechanism, if any, actually serves this on.
// +optional // +optional
Hostname string `json:"hostname,omitempty"` Hostname string `json:"hostname,omitempty"`
// servicePort is both the Service's port and the HTTPRoute's backend // servicePort is both the container's port and the ClusterIP Service's
// port once it exists. Defaults to 8080, matching the chart's own // port. Defaults to 8080, matching the chart's own service.port default.
// service.port default.
// +kubebuilder:default=8080 // +kubebuilder:default=8080
// +optional // +optional
ServicePort int32 `json:"servicePort,omitempty"` ServicePort int32 `json:"servicePort,omitempty"`
// gatewayListener is the HTTPRoute's sectionName once it exists. Empty
// attaches to every matching listener, including plaintext HTTP.
// +optional
GatewayListener string `json:"gatewayListener,omitempty"`
} }
// PostgresClusterRef names a Zalando postgres-operator `postgresql` CR // PostgresClusterRef names a Zalando postgres-operator `postgresql` CR
@@ -190,6 +187,109 @@ type AllowedTeams struct {
Namespaces AllowedTeamsNamespaces `json:"namespaces,omitempty"` Namespaces AllowedTeamsNamespaces `json:"namespaces,omitempty"`
} }
// PodDisruptionBudgetSpec configures an optional PodDisruptionBudget for
// this TerdutServer's Deployment. Exactly one of minAvailable or
// maxUnavailable may be set, matching policyv1.PodDisruptionBudgetSpec's own
// upstream rule (both wrap intstr.IntOrString unchanged here -- this is pure
// passthrough, not reshaped) and mirroring DatabaseSpec's own
// dsn/postgresClusterRef mutual-exclusion pattern. Clearing this field
// deletes any PodDisruptionBudget the controller previously created for this
// TerdutServer (DESIGN.md §7).
// +kubebuilder:validation:XValidation:rule="(has(self.minAvailable) ? 1 : 0) + (has(self.maxUnavailable) ? 1 : 0) == 1",message="exactly one of minAvailable or maxUnavailable must be set"
type PodDisruptionBudgetSpec struct {
// minAvailable -- mutually exclusive with maxUnavailable.
// +optional
MinAvailable *intstr.IntOrString `json:"minAvailable,omitempty"`
// maxUnavailable -- mutually exclusive with minAvailable.
// +optional
MaxUnavailable *intstr.IntOrString `json:"maxUnavailable,omitempty"`
}
// PodSpec is pod-level customization of the Deployment this TerdutServer
// creates. Fields here directly reuse corev1 types wherever corev1 already
// models the knob exactly, rather than wrapping (unlike SecretKeyRef's own
// "wrap only when a round-trip through a different type buys something"
// standard would suggest at first glance -- none of these do: Tolerations,
// Affinity, TopologySpreadConstraints, Resources, SecurityContext, EnvVar,
// EnvFromSource, Volume, VolumeMount and LocalObjectReference are all passed
// straight through to the pod template with no added semantics, matching how
// CloudNativePG and the Zalando postgres-operator both expose the same
// knobs).
type PodSpec struct {
// annotations are merged onto the pod template's own metadata.
// Operator-managed labels (labelsFor) are never touched by this field.
// +optional
Annotations map[string]string `json:"annotations,omitempty"`
// +optional
NodeSelector map[string]string `json:"nodeSelector,omitempty"`
// +optional
Tolerations []corev1.Toleration `json:"tolerations,omitempty"`
// affinity covers node affinity, pod affinity and pod anti-affinity in
// one field -- unlike a multi-replica-aware operator, this one never
// generates a default anti-affinity itself (replicas above 1 isn't a
// supported topology, see TerdutServerSpec.Replicas's own doc comment),
// so this is pure user-supplied passthrough, not a toggle-plus-generated-
// default.
// +optional
Affinity *corev1.Affinity `json:"affinity,omitempty"`
// +optional
TopologySpreadConstraints []corev1.TopologySpreadConstraint `json:"topologySpreadConstraints,omitempty"`
// resources applied to the main terdut-server container. Unset today --
// this field closes a pre-existing gap, not a behavior change for
// anyone not setting it.
// +optional
Resources corev1.ResourceRequirements `json:"resources,omitempty"`
// securityContext is pod-level.
// +optional
SecurityContext *corev1.PodSecurityContext `json:"securityContext,omitempty"`
// containerSecurityContext applies to the main terdut-server container
// only -- not wait-for-postgres, which runs a stock postgres image this
// operator doesn't control the entrypoint of. No implicit defaults are
// merged underneath it.
// +optional
ContainerSecurityContext *corev1.SecurityContext `json:"containerSecurityContext,omitempty"`
// serviceAccountName. Defaults to "default", same as any pod that
// doesn't set it.
// +optional
ServiceAccountName string `json:"serviceAccountName,omitempty"`
// extraEnv is appended after the fixed env vars buildEnv produces.
// +optional
ExtraEnv []corev1.EnvVar `json:"extraEnv,omitempty"`
// +optional
ExtraEnvFrom []corev1.EnvFromSource `json:"extraEnvFrom,omitempty"`
// extraVolumes are added to the pod spec; pair with extraVolumeMounts to
// actually mount one on the main container.
// +optional
ExtraVolumes []corev1.Volume `json:"extraVolumes,omitempty"`
// extraVolumeMounts are added to the main terdut-server container only
// -- not wait-for-postgres.
// +optional
ExtraVolumeMounts []corev1.VolumeMount `json:"extraVolumeMounts,omitempty"`
// +optional
ImagePullSecrets []corev1.LocalObjectReference `json:"imagePullSecrets,omitempty"`
// disruptionBudget, when set, causes the controller to reconcile a
// policyv1.PodDisruptionBudget selecting this TerdutServer's pods.
// Removing this field deletes any PodDisruptionBudget the controller
// previously created.
// +optional
DisruptionBudget *PodDisruptionBudgetSpec `json:"disruptionBudget,omitempty"`
}
// TerdutServerSpec defines the desired state of TerdutServer. // TerdutServerSpec defines the desired state of TerdutServer.
// //
// The operator creates and owns every TerdutServer it manages (DESIGN.md // The operator creates and owns every TerdutServer it manages (DESIGN.md
@@ -237,6 +337,11 @@ type TerdutServerSpec struct {
// this CRD's schema doesn't need a breaking change to grow it later. // this CRD's schema doesn't need a breaking change to grow it later.
// +optional // +optional
AllowedTeams AllowedTeams `json:"allowedTeams,omitempty"` AllowedTeams AllowedTeams `json:"allowedTeams,omitempty"`
// pod is pod-level customization of the Deployment this TerdutServer
// creates (DESIGN.md §4.1).
// +optional
Pod PodSpec `json:"pod,omitempty"`
} }
// Condition types this controller sets on TerdutServer. // Condition types this controller sets on TerdutServer.
+63
View File
@@ -27,6 +27,42 @@ type TerdutTeamOIDC struct {
OwnerGroup string `json:"ownerGroup,omitempty"` OwnerGroup string `json:"ownerGroup,omitempty"`
} }
// TerdutTeamInvite requests a standing invite link into this team, minted
// with the team's own team-scoped credential — requireTeamOwner already
// treats that credential as owner-equivalent for every /invites route
// (ratified, not a gap, as of terdut-server's SERVICE-ACCOUNTS.md). This is
// the real answer to "how does a human ever get a first login on a
// password-only, operator-managed install" (terdut-server#23): no signup_mode
// flip, no admin token, just a link redeemed the same way anyone else's
// invite would be.
type TerdutTeamInvite struct {
// enabled mints (and keeps refreshed ahead of terdut-server's own fixed
// 7-day TTL) an invite link while true. Flipping it back to false
// revokes the current one server-side rather than leaving it to expire
// on its own.
// +optional
Enabled bool `json:"enabled,omitempty"`
// role is what the invite grants: member or owner. Defaults to member —
// owner by default would make every invite link a standing
// administrative credential for the team, a much bigger blast radius
// than "let a human see the queue".
// +optional
// +kubebuilder:validation:Enum=member;owner
// +kubebuilder:default=member
Role string `json:"role,omitempty"`
// maxUses bounds how many times this link may be redeemed before it
// stops working, mirroring terdut-server's own 1-100 range
// (POST /api/teams/{teamID}/invites). Defaults to 1: a link meant for
// one specific person, not a standing door.
// +optional
// +kubebuilder:validation:Minimum=1
// +kubebuilder:validation:Maximum=100
// +kubebuilder:default=1
MaxUses int64 `json:"maxUses,omitempty"`
}
// TerdutTeamSpec defines the desired state of TerdutTeam. // TerdutTeamSpec defines the desired state of TerdutTeam.
type TerdutTeamSpec struct { type TerdutTeamSpec struct {
// serverRef names the TerdutServer this team belongs to. // serverRef names the TerdutServer this team belongs to.
@@ -43,6 +79,9 @@ type TerdutTeamSpec struct {
// +optional // +optional
OIDC TerdutTeamOIDC `json:"oidc,omitempty"` OIDC TerdutTeamOIDC `json:"oidc,omitempty"`
// +optional
Invite TerdutTeamInvite `json:"invite,omitempty"`
} }
// Condition reasons this controller sets. // Condition reasons this controller sets.
@@ -62,6 +101,19 @@ const (
ReasonTeamAdopted = "Adopted" ReasonTeamAdopted = "Adopted"
) )
// Condition reasons for spec.invite reconciliation (TerdutTeamInvite). Not
// surfaced on the Ready condition itself — an invite is a convenience, not
// a dependency anything else in this team's own readiness waits on — but
// recorded as Events and readable via `kubectl describe`.
const (
// ReasonInviteMinted: spec.invite.enabled is true and status.inviteSecretRef
// is populated and live.
ReasonInviteMinted = "InviteMinted"
// ReasonInviteRevoked: spec.invite.enabled flipped back to false and the
// server-side invite was revoked (or there was nothing to revoke).
ReasonInviteRevoked = "InviteRevoked"
)
// TerdutTeamStatus defines the observed state of TerdutTeam. // TerdutTeamStatus defines the observed state of TerdutTeam.
type TerdutTeamStatus struct { type TerdutTeamStatus struct {
// +listType=map // +listType=map
@@ -87,6 +139,17 @@ type TerdutTeamStatus struct {
// +optional // +optional
ServerEndpoint string `json:"serverEndpoint,omitempty"` ServerEndpoint string `json:"serverEndpoint,omitempty"`
// inviteSecretRef is this team's current invite link, if spec.invite.enabled.
// Unlike credentialsSecretRef, this lives in the TerdutTeam's OWN
// namespace, not the operator's: an invite is bounded, limited-use, and
// meant for this namespace's own human operators to read and hand out,
// not a durable high-privilege credential — same shape as
// TerdutAlertSource's status.webhookURLSecretRef, not TerdutServer's
// cross-namespace credentialsSecretRef. Nil whenever spec.invite.enabled
// is false or unset.
// +optional
InviteSecretRef *LocalSecretRef `json:"inviteSecretRef,omitempty"`
// +optional // +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"` ObservedGeneration int64 `json:"observedGeneration,omitempty"`
} }
+146
View File
@@ -5,8 +5,10 @@
package v1alpha1 package v1alpha1
import ( import (
corev1 "k8s.io/api/core/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/util/intstr"
) )
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
@@ -210,6 +212,128 @@ func (in *OIDCSpec) DeepCopy() *OIDCSpec {
return out return out
} }
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *PodDisruptionBudgetSpec) DeepCopyInto(out *PodDisruptionBudgetSpec) {
*out = *in
if in.MinAvailable != nil {
in, out := &in.MinAvailable, &out.MinAvailable
*out = new(intstr.IntOrString)
**out = **in
}
if in.MaxUnavailable != nil {
in, out := &in.MaxUnavailable, &out.MaxUnavailable
*out = new(intstr.IntOrString)
**out = **in
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PodDisruptionBudgetSpec.
func (in *PodDisruptionBudgetSpec) DeepCopy() *PodDisruptionBudgetSpec {
if in == nil {
return nil
}
out := new(PodDisruptionBudgetSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *PodSpec) DeepCopyInto(out *PodSpec) {
*out = *in
if in.Annotations != nil {
in, out := &in.Annotations, &out.Annotations
*out = make(map[string]string, len(*in))
for key, val := range *in {
(*out)[key] = val
}
}
if in.NodeSelector != nil {
in, out := &in.NodeSelector, &out.NodeSelector
*out = make(map[string]string, len(*in))
for key, val := range *in {
(*out)[key] = val
}
}
if in.Tolerations != nil {
in, out := &in.Tolerations, &out.Tolerations
*out = make([]corev1.Toleration, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.Affinity != nil {
in, out := &in.Affinity, &out.Affinity
*out = new(corev1.Affinity)
(*in).DeepCopyInto(*out)
}
if in.TopologySpreadConstraints != nil {
in, out := &in.TopologySpreadConstraints, &out.TopologySpreadConstraints
*out = make([]corev1.TopologySpreadConstraint, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
in.Resources.DeepCopyInto(&out.Resources)
if in.SecurityContext != nil {
in, out := &in.SecurityContext, &out.SecurityContext
*out = new(corev1.PodSecurityContext)
(*in).DeepCopyInto(*out)
}
if in.ContainerSecurityContext != nil {
in, out := &in.ContainerSecurityContext, &out.ContainerSecurityContext
*out = new(corev1.SecurityContext)
(*in).DeepCopyInto(*out)
}
if in.ExtraEnv != nil {
in, out := &in.ExtraEnv, &out.ExtraEnv
*out = make([]corev1.EnvVar, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ExtraEnvFrom != nil {
in, out := &in.ExtraEnvFrom, &out.ExtraEnvFrom
*out = make([]corev1.EnvFromSource, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ExtraVolumes != nil {
in, out := &in.ExtraVolumes, &out.ExtraVolumes
*out = make([]corev1.Volume, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ExtraVolumeMounts != nil {
in, out := &in.ExtraVolumeMounts, &out.ExtraVolumeMounts
*out = make([]corev1.VolumeMount, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ImagePullSecrets != nil {
in, out := &in.ImagePullSecrets, &out.ImagePullSecrets
*out = make([]corev1.LocalObjectReference, len(*in))
copy(*out, *in)
}
if in.DisruptionBudget != nil {
in, out := &in.DisruptionBudget, &out.DisruptionBudget
*out = new(PodDisruptionBudgetSpec)
(*in).DeepCopyInto(*out)
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PodSpec.
func (in *PodSpec) DeepCopy() *PodSpec {
if in == nil {
return nil
}
out := new(PodSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *PostgresClusterRef) DeepCopyInto(out *PostgresClusterRef) { func (in *PostgresClusterRef) DeepCopyInto(out *PostgresClusterRef) {
*out = *in *out = *in
@@ -643,6 +767,7 @@ func (in *TerdutServerSpec) DeepCopyInto(out *TerdutServerSpec) {
in.Notify.DeepCopyInto(&out.Notify) in.Notify.DeepCopyInto(&out.Notify)
in.OIDC.DeepCopyInto(&out.OIDC) in.OIDC.DeepCopyInto(&out.OIDC)
in.AllowedTeams.DeepCopyInto(&out.AllowedTeams) in.AllowedTeams.DeepCopyInto(&out.AllowedTeams)
in.Pod.DeepCopyInto(&out.Pod)
} }
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutServerSpec. // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutServerSpec.
@@ -709,6 +834,21 @@ func (in *TerdutTeam) DeepCopyObject() runtime.Object {
return nil return nil
} }
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutTeamInvite) DeepCopyInto(out *TerdutTeamInvite) {
*out = *in
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamInvite.
func (in *TerdutTeamInvite) DeepCopy() *TerdutTeamInvite {
if in == nil {
return nil
}
out := new(TerdutTeamInvite)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutTeamList) DeepCopyInto(out *TerdutTeamList) { func (in *TerdutTeamList) DeepCopyInto(out *TerdutTeamList) {
*out = *in *out = *in
@@ -776,6 +916,7 @@ func (in *TerdutTeamSpec) DeepCopyInto(out *TerdutTeamSpec) {
*out = *in *out = *in
out.ServerRef = in.ServerRef out.ServerRef = in.ServerRef
out.OIDC = in.OIDC out.OIDC = in.OIDC
out.Invite = in.Invite
} }
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamSpec. // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamSpec.
@@ -803,6 +944,11 @@ func (in *TerdutTeamStatus) DeepCopyInto(out *TerdutTeamStatus) {
*out = new(SecretKeyRef) *out = new(SecretKeyRef)
**out = **in **out = **in
} }
if in.InviteSecretRef != nil {
in, out := &in.InviteSecretRef, &out.InviteSecretRef
*out = new(LocalSecretRef)
**out = **in
}
} }
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamStatus. // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamStatus.
+2 -2
View File
@@ -6,8 +6,8 @@ type: application
# These fields decide nothing: `make helm-package` passes --version and # These fields decide nothing: `make helm-package` passes --version and
# --app-version from the release tag (same reasoning as terdut-server's own # --app-version from the release tag (same reasoning as terdut-server's own
# chart). They're for whoever reads the tree before a tag exists. # chart). They're for whoever reads the tree before a tag exists.
version: 0.1.1 version: 0.3.0
appVersion: "v0.1.1" appVersion: "v0.3.0"
keywords: keywords:
- kubernetes - kubernetes
File diff suppressed because it is too large Load Diff
@@ -63,6 +63,47 @@ spec:
create rule, via TEAM-LOOKUP.md). create rule, via TEAM-LOOKUP.md).
minLength: 1 minLength: 1
type: string type: string
invite:
description: |-
TerdutTeamInvite requests a standing invite link into this team, minted
with the team's own team-scoped credential — requireTeamOwner already
treats that credential as owner-equivalent for every /invites route
(ratified, not a gap, as of terdut-server's SERVICE-ACCOUNTS.md). This is
the real answer to "how does a human ever get a first login on a
password-only, operator-managed install" (terdut-server#23): no signup_mode
flip, no admin token, just a link redeemed the same way anyone else's
invite would be.
properties:
enabled:
description: |-
enabled mints (and keeps refreshed ahead of terdut-server's own fixed
7-day TTL) an invite link while true. Flipping it back to false
revokes the current one server-side rather than leaving it to expire
on its own.
type: boolean
maxUses:
default: 1
description: |-
maxUses bounds how many times this link may be redeemed before it
stops working, mirroring terdut-server's own 1-100 range
(POST /api/teams/{teamID}/invites). Defaults to 1: a link meant for
one specific person, not a standing door.
format: int64
maximum: 100
minimum: 1
type: integer
role:
default: member
description: |-
role is what the invite grants: member or owner. Defaults to member —
owner by default would make every invite link a standing
administrative credential for the team, a much bigger blast radius
than "let a human see the queue".
enum:
- member
- owner
type: string
type: object
oidc: oidc:
description: |- description: |-
TerdutTeamOIDC binds which identity-provider groups grant membership and TerdutTeamOIDC binds which identity-provider groups grant membership and
@@ -171,6 +212,24 @@ spec:
- key - key
- name - name
type: object type: object
inviteSecretRef:
description: |-
inviteSecretRef is this team's current invite link, if spec.invite.enabled.
Unlike credentialsSecretRef, this lives in the TerdutTeam's OWN
namespace, not the operator's: an invite is bounded, limited-use, and
meant for this namespace's own human operators to read and hand out,
not a durable high-privilege credential — same shape as
TerdutAlertSource's status.webhookURLSecretRef, not TerdutServer's
cross-namespace credentialsSecretRef. Nil whenever spec.invite.enabled
is false or unset.
properties:
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- name
type: object
observedGeneration: observedGeneration:
format: int64 format: int64
type: integer type: integer
@@ -58,6 +58,18 @@ rules:
verbs: verbs:
- create - create
- patch - patch
- apiGroups:
- policy
resources:
- poddisruptionbudgets
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups: - apiGroups:
- terdut.ryuvia.com - terdut.ryuvia.com
resources: resources:
@@ -51,4 +51,8 @@ spec:
allowedTeams: allowedTeams:
{{- toYaml . | nindent 4 }} {{- toYaml . | nindent 4 }}
{{- end }} {{- end }}
{{- with .Values.terdutServer.pod }}
pod:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- end }} {{- end }}
+22 -3
View File
@@ -236,9 +236,12 @@ terdutServer:
replicas: 1 replicas: 1
networking: networking:
## Required when terdutServer.enabled -- the hostname a future ## Required when terdutServer.enabled -- terdut-server's own public
## HTTPRoute will carry (see NetworkingSpec's own doc comment: creating ## URL, used for absolute links it generates itself (notifications,
## that HTTPRoute isn't implemented yet). ## OIDC redirect URIs). This chart/operator never creates any
## ingress/HTTPRoute for it -- see examples/networking in the repo for
## how to expose the Service this chart's TerdutServer CR causes to
## be created, if you want to expose it at all.
# hostname: "" # hostname: ""
servicePort: 8080 servicePort: 8080
@@ -270,3 +273,19 @@ terdutServer:
# allowedTeams: {} # allowedTeams: {}
## Pod-level customization of the Deployment this TerdutServer creates --
## see api/v1alpha1/terdutserver_types.go's PodSpec for the full shape
## (nodeSelector, tolerations, affinity, topologySpreadConstraints,
## securityContext, containerSecurityContext, serviceAccountName,
## extraEnv/extraEnvFrom, extraVolumes/extraVolumeMounts,
## imagePullSecrets, disruptionBudget). All optional; resources is the
## one most installs will want to set, since the container otherwise
## runs with no requests/limits at all:
# pod:
# resources:
# requests:
# cpu: 100m
# memory: 128Mi
# limits:
# memory: 256Mi
File diff suppressed because it is too large Load Diff
@@ -60,6 +60,47 @@ spec:
create rule, via TEAM-LOOKUP.md). create rule, via TEAM-LOOKUP.md).
minLength: 1 minLength: 1
type: string type: string
invite:
description: |-
TerdutTeamInvite requests a standing invite link into this team, minted
with the team's own team-scoped credential — requireTeamOwner already
treats that credential as owner-equivalent for every /invites route
(ratified, not a gap, as of terdut-server's SERVICE-ACCOUNTS.md). This is
the real answer to "how does a human ever get a first login on a
password-only, operator-managed install" (terdut-server#23): no signup_mode
flip, no admin token, just a link redeemed the same way anyone else's
invite would be.
properties:
enabled:
description: |-
enabled mints (and keeps refreshed ahead of terdut-server's own fixed
7-day TTL) an invite link while true. Flipping it back to false
revokes the current one server-side rather than leaving it to expire
on its own.
type: boolean
maxUses:
default: 1
description: |-
maxUses bounds how many times this link may be redeemed before it
stops working, mirroring terdut-server's own 1-100 range
(POST /api/teams/{teamID}/invites). Defaults to 1: a link meant for
one specific person, not a standing door.
format: int64
maximum: 100
minimum: 1
type: integer
role:
default: member
description: |-
role is what the invite grants: member or owner. Defaults to member —
owner by default would make every invite link a standing
administrative credential for the team, a much bigger blast radius
than "let a human see the queue".
enum:
- member
- owner
type: string
type: object
oidc: oidc:
description: |- description: |-
TerdutTeamOIDC binds which identity-provider groups grant membership and TerdutTeamOIDC binds which identity-provider groups grant membership and
@@ -168,6 +209,24 @@ spec:
- key - key
- name - name
type: object type: object
inviteSecretRef:
description: |-
inviteSecretRef is this team's current invite link, if spec.invite.enabled.
Unlike credentialsSecretRef, this lives in the TerdutTeam's OWN
namespace, not the operator's: an invite is bounded, limited-use, and
meant for this namespace's own human operators to read and hand out,
not a durable high-privilege credential — same shape as
TerdutAlertSource's status.webhookURLSecretRef, not TerdutServer's
cross-namespace credentialsSecretRef. Nil whenever spec.invite.enabled
is false or unset.
properties:
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- name
type: object
observedGeneration: observedGeneration:
format: int64 format: int64
type: integer type: integer
+12
View File
@@ -52,6 +52,18 @@ rules:
verbs: verbs:
- create - create
- patch - patch
- apiGroups:
- policy
resources:
- poddisruptionbudgets
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups: - apiGroups:
- terdut.ryuvia.com - terdut.ryuvia.com
resources: resources:
@@ -31,3 +31,12 @@ spec:
timeout: 15m timeout: 15m
severity: critical severity: critical
passwordLogin: true passwordLogin: true
# Pod-level customization, all optional -- see PodSpec in
# api/v1alpha1/terdutserver_types.go for the full shape (tolerations,
# affinity, topologySpreadConstraints, securityContext,
# serviceAccountName, extraEnv/extraVolumes, imagePullSecrets,
# disruptionBudget, ...). Example:
# pod:
# resources:
# requests: {cpu: 100m, memory: 128Mi}
# limits: {memory: 256Mi}
+96
View File
@@ -0,0 +1,96 @@
# Demo-only Postgres: a bare Deployment+Service+Secret, not the Zalando
# postgres-operator path (DatabaseSpec.postgresClusterRef, DESIGN.md §8).
# Bring-your-own DSN is the simpler of the two paths to stand up from
# nothing (ROADMAP.md Stage 1's own note), which is all this needs to be.
#
# emptyDir, one replica, a password sitting in a plaintext Secret below --
# none of that is how you'd run Postgres for real. It exists only so
# 01-server.yaml has something to talk to. Throw the whole demo namespace
# away when you're done; nothing here is meant to survive that.
apiVersion: v1
kind: Secret
metadata:
name: terdut-operator-demo-postgres
type: Opaque
stringData:
password: demo-not-a-real-password
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: terdut-operator-demo-postgres
labels:
app: terdut-operator-demo-postgres
spec:
replicas: 1
# Recreate, not RollingUpdate: emptyDir means a new pod starts with an
# empty database anyway, and two Postgres pods would never agree on one
# emptyDir each.
strategy:
type: Recreate
selector:
matchLabels:
app: terdut-operator-demo-postgres
template:
metadata:
labels:
app: terdut-operator-demo-postgres
spec:
containers:
- name: postgres
image: postgres:17-alpine
# Partial, deliberately: the official image's entrypoint needs to
# start as root to chown/chmod the data directory before it drops
# privileges itself (gosu, to the postgres user) -- forcing
# runAsNonRoot would just refuse to start the container, and
# dropping all capabilities (an earlier version of this file did)
# takes CAP_CHOWN/CAP_FOWNER away from that same root user, which
# is a different way of breaking the identical startup step:
# confirmed the hard way, as `chmod: /var/run/postgresql:
# Operation not permitted` in a real pod's logs, not caught by
# `kubectl apply --dry-run=server` -- that only checks admission
# policy, never whether the container actually boots. A
# "restricted" PodSecurity namespace warns on the remaining gap
# (no runAsNonRoot) rather than blocking, which is an acceptable
# tradeoff for Postgres that exists only to be thrown away with
# the rest of this demo.
securityContext:
allowPrivilegeEscalation: false
seccompProfile:
type: RuntimeDefault
ports:
- name: postgres
containerPort: 5432
env:
- name: POSTGRES_USER
value: terdut
- name: POSTGRES_DB
value: terdut
- name: POSTGRES_PASSWORD
valueFrom:
secretKeyRef:
name: terdut-operator-demo-postgres
key: password
volumeMounts:
- name: data
mountPath: /var/lib/postgresql/data
subPath: pgdata
readinessProbe:
exec:
command: ["pg_isready", "-U", "terdut"]
initialDelaySeconds: 5
volumes:
- name: data
emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
name: terdut-operator-demo-postgres
spec:
selector:
app: terdut-operator-demo-postgres
ports:
- name: postgres
port: 5432
targetPort: postgres
+37
View File
@@ -0,0 +1,37 @@
# The one TerdutServer this whole demo runs against. Everything else in
# this directory (teams, escalation rules, dead man's switches, alert
# sources) references it by name.
#
# The operator never creates any ingress/HTTPRoute for this TerdutServer --
# that's a permanent non-goal (DESIGN.md §1, NetworkingSpec's own doc
# comment), not a missing feature. This demo reaches it only by
# port-forwarding its Service, same name as this object (see README.md);
# see ../networking for worked examples of exposing it yourself instead.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer
metadata:
name: terdut-operator-demo
spec:
image:
repository: git.ryuvia.com/niklas/terdut-server
# v0.34.0: fixes callerMayManageServiceAccount so an instance-scoped
# service account can adopt/rotate a key on a team-scoped account it
# didn't just create in the same call -- without this, terdutteam-*
# can wedge permanently on exactly the crash-window race this demo
# hit live (niklas/terdut-operator#3).
tag: v0.34.0
replicas: 1
networking:
hostname: terdut-operator-demo.example
servicePort: 8080
database:
dsn: "postgres://terdut@terdut-operator-demo-postgres:5432/terdut?sslmode=disable"
passwordSecretRef:
name: terdut-operator-demo-postgres
key: password
sweeper:
staleAfter: 6h
archiveAfter: 168h
# No oidc block: password login only, so there's nothing external to
# register a redirect URI with before this demo can sign in.
passwordLogin: true
+23
View File
@@ -0,0 +1,23 @@
# Two teams (this one and 03-team-payments.yaml) so the demo shows
# per-team isolation -- separate incident lists, separate escalation
# ladders, separate alert sources -- rather than one team standing in for
# everything.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
name: terdutteam-platform
spec:
# serverRef.namespace omitted: both this and terdut-operator-demo (01-server.yaml)
# live in whatever namespace you apply this directory into, which is the
# common case and needs no allowedTeams consent on the TerdutServer side
# (DESIGN.md §4.1, §4.6).
serverRef:
name: terdut-operator-demo
displayName: Platform
# No oidc block: this demo is password-login only (01-server.yaml).
# A real invite link, minted via this team's own credential, is how
# run-demo.sh's alice actually gets in -- no signup_mode flip, no admin
# token (see niklas/terdut-server#23's resolution). Role/maxUses left at
# their defaults (member, 1): one link for one person.
invite:
enabled: true
+14
View File
@@ -0,0 +1,14 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
name: terdutteam-payments
spec:
serverRef:
name: terdut-operator-demo
displayName: Payments
# No spec.invite here, deliberately: run-demo.sh joins alice to this team
# through POST /api/teams/{teamID}/members instead (this team's own
# credential, same owner-equivalent reach spec.invite relies on, plus her
# user id resolved via GET /api/users), once she already has an account
# from Platform's invite -- showing both onboarding paths this feature
# unlocks, not just the one.
+24
View File
@@ -0,0 +1,24 @@
# One per team is the rule (DESIGN.md §4.3) -- a second TerdutEscalationRule
# naming the same teamRef would just clobber this one on the next reconcile.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutEscalationRule
metadata:
name: terdutescalationrule-platform
spec:
teamRef:
name: terdutteam-platform
repeatCount: 2
fallbackTopic: platform-fallback
levels:
# username is required iff kind is "user", rejected otherwise -- CEL
# validation at apply time (api/v1alpha1/terdutescalationrule_types.go).
# alice won't exist on a fresh demo install -- see README.md for
# creating a real user if you want this level to mean something, or
# just watch it fall through to oncall after 5m.
- timeout: 5m
targets:
- kind: user
username: alice
- timeout: 10m
targets:
- kind: oncall
+13
View File
@@ -0,0 +1,13 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutEscalationRule
metadata:
name: terdutescalationrule-payments
spec:
teamRef:
name: terdutteam-payments
repeatCount: 1
fallbackTopic: payments-fallback
levels:
- timeout: 5m
targets:
- kind: oncall
+17
View File
@@ -0,0 +1,17 @@
# A per-team dead man's switch (DESIGN.md §4.4) -- a different thing from
# 01-server.yaml's spec.deadman, which this demo leaves unset so this CRD
# is what you're actually seeing reconcile. fire-alerts.sh's "heartbeat"
# scenario sends a matching alert; stop sending it and terdut-server
# itself opens an incident once `timeout` passes with no heartbeat.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutDeadmanSwitch
metadata:
name: terdutdeadmanswitch-platform
spec:
teamRef:
name: terdutteam-platform
# name omitted -- terdut-server derives one from the matcher's own
# canonical form (DESIGN.md §4.4).
matcher: "alertname=PlatformWatchdog"
timeout: 15m
severity: critical
+10
View File
@@ -0,0 +1,10 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutDeadmanSwitch
metadata:
name: terdutdeadmanswitch-payments
spec:
teamRef:
name: terdutteam-payments
matcher: "alertname=PaymentsWatchdog"
timeout: 15m
severity: critical
@@ -0,0 +1,18 @@
# The webhook URL/key fire-alerts.sh sends to, for the Platform team.
# terdut-server shows the key exactly once, at creation, and never again
# (DESIGN.md §4.5) -- this object's status.webhookURLSecretRef names the
# generated Secret holding it (keys "url" and "key"), which is what
# fire-alerts.sh reads. See README.md before applying this: it's the one
# object in this directory whose Secret you can't just re-read if you
# miss it.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutAlertSource
metadata:
name: terdutalertsource-platform
spec:
teamRef:
name: terdutteam-platform
# kind defaults to "alertmanager" -- the only value terdut-server
# supports today.
kind: alertmanager
name: platform-demo-alertmanager
@@ -0,0 +1,9 @@
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutAlertSource
metadata:
name: terdutalertsource-payments
spec:
teamRef:
name: terdutteam-payments
kind: alertmanager
name: payments-demo-alertmanager
+159
View File
@@ -0,0 +1,159 @@
# Demo
Every CRD this operator reconciles, wired into one working install: one
`TerdutServer`, two `TerdutTeam`s (Platform and Payments) each with their
own `TerdutEscalationRule`, `TerdutDeadmanSwitch` and `TerdutAlertSource`,
plus a script that fires synthetic Alertmanager webhooks at it so you can
watch real incidents appear, escalate and resolve.
This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
Postgres with `emptyDir` storage and a password committed in this
directory. Throw the whole namespace away when you're done.
**Want this fully automated instead of walking through it by hand?**
`./run-demo.sh` does everything below itself, against a fresh (or
already-set-up) `kind` cluster — creates the cluster, installs the
operator, applies every CR here, signs `alice` in for real, and fires a
few alerts. `./run-demo.sh --help` for the knobs, `./run-demo.sh
--teardown` to tear it back down. The rest of this file is the manual
walkthrough it automates.
## Prerequisites
- The operator and its CRDs installed and running (`make install
deploy IMG=...`, or `charts/terdut-operator` — see this repo's own
README.md/DESIGN.md), pointed at a cluster you're fine creating
throwaway resources in. A `kind` cluster is the easy choice.
- `kubectl`, `jq`, `curl` on your path.
**Apply this into a namespace of its own.** Every object name in this
directory is prefixed `terdut-operator-demo` specifically so applying it
by mistake into some other namespace that already has unrelated objects
doesn't collide with them -- but that only helps if this directory's own
objects don't collide with *each other* across two applies. Applying it
twice into two different namespaces is fine; applying it a second time
into a namespace that already has something else named `terdut-demo` (a
real install from following `terdut-operator`'s own repo along, say) is
exactly the mistake this prefix exists to avoid, and it only works if you
don't override these names yourself.
## Apply it
```sh
kubectl create namespace terdut-operator-demo
kubectl apply -n terdut-operator-demo -k .
```
Listed and applied in dependency order (server → team → everything that
`teamRef`s it), but you don't have to preserve that order yourself:
every controller here re-queues and waits rather than failing when a ref
isn't resolvable yet (`kubectl describe` shows `Reason: TeamRefNotFound` /
`WaitingForTeam` while that settles).
Watch it converge:
```sh
kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitches,terdutalertsources \
-n terdut-operator-demo
```
The operator's own Deployment template now carries a `wait-for-postgres`
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
so `terdut-operator-demo`'s pod should come up clean even against this brand-new
Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
Once `terdut-operator-demo`'s own `Ready` condition is `True`, everything downstream
of it should settle within a reconcile interval or two.
## See the web UI
The operator never creates any external exposure for a `TerdutServer` --
that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this
demo just reaches it the simplest way there is:
```sh
kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
```
and open http://localhost:8080. See `../networking` for worked examples of
exposing it for real (Gateway API, Istio, or a plain `Ingress`) instead.
### First login
The operator's own bootstrap (DESIGN.md §6) creates the first user through
`/api/bootstrap` and immediately mints itself a service-account token from
it, then discards the bootstrap user's own key — nobody ever signs in as
that account, and `signup_mode` stays `invite_only` by default. **Don't try
to flip it via the operator's own token**: that token is a service account,
and `/api/admin/settings` is deliberately human-only on terdut-server
(`niklas/terdut-server#23` has the full reasoning — widening that gate was
the wrong fix).
The real path in: `02-team-platform.yaml` turns on `spec.invite`, so
Platform's own `TerdutTeam` mints a real invite link with its own
already-working team-scoped credential (the same reach that lets it manage
its own escalation policy, dead man's switches and integrations — owner-
equivalent, confirmed in terdut-server's `SERVICE-ACCOUNTS.md`). Invite
redemption bypasses `signup_mode` entirely, so this needs no admin
credential at all:
```sh
secretname=$(kubectl -n terdut-operator-demo get terdutteam terdutteam-platform \
-o jsonpath='{.status.inviteSecretRef.name}')
url=$(kubectl -n terdut-operator-demo get secret "$secretname" -o jsonpath='{.data.url}' | base64 -d)
echo "$url" # open this, or POST /api/signup with {"invite": "<the token after invite=>", ...}
```
`04-escalation-platform.yaml` names a user `alice` at its first escalation
level — sign up as `alice` if you want that level to mean something rather
than falling through to on-call after 5 minutes. `run-demo.sh` does exactly
this automatically (and also joins `alice` to Payments, which deliberately
has no `spec.invite` of its own — see that file's comment for the second
onboarding path this demonstrates).
## Fire some alerts
In another terminal, with the port-forward above still running:
```sh
export NAMESPACE=terdut-operator-demo
./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform disk-full
./fire-alerts.sh payments pod-crash
# Watch it open an incident, escalate per 04/05-escalation-*.yaml's
# levels, and show up on the Platform/Payments team's own incident list.
./fire-alerts.sh platform high-cpu resolve
```
`fire-alerts.sh -h` (or any bad argument) prints the full scenario list.
Each `(team, scenario)` pair is one stable fingerprint, so firing the same
one twice updates the same alert (a real re-fire) and `resolve` closes
exactly that one.
### Dead man's switches
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
alert on a 15-minute timeout:
```sh
./fire-alerts.sh platform heartbeat
```
Keep sending that (e.g. a `watch -n 60`) and nothing happens — that's the
point. Stop sending it and, 15 minutes after the last one, terdut-server
opens a `critical` incident on its own, with no webhook involved: proof
the switch is watching for silence, not for a signal.
## Tear down
```sh
kubectl delete namespace terdut-operator-demo
```
The operator's own finalizers clean up everything cross-namespace
(credentials Secrets in the operator's namespace, server-side team/rule/
integration rows) before this namespace's objects actually disappear —
give it a few seconds past the `kubectl delete` returning.
+138
View File
@@ -0,0 +1,138 @@
#!/usr/bin/env bash
# Sends a synthetic Alertmanager v4 webhook payload at the "platform" or
# "payments" demo team's TerdutAlertSource, so terdut-server opens (or
# resolves) an incident exactly the way it would for a real Alertmanager.
#
# The payload shape here is amPayload/amAlert, read straight out of
# terdut-server's own internal/api/alertmanager.go rather than guessed from
# its docs -- version/status/groupKey/groupLabels, and alerts[] carrying
# status/labels/annotations/startsAt/endsAt/generatorURL/fingerprint.
#
# Why this reads the webhook key out of a kubectl Secret instead of using
# the "url" key already in it: that URL is built from spec.networking.hostname
# (TERDUT_PUBLIC_URL), and nothing in this demo stands up real ingress for
# it (01-server.yaml's own comment) -- so it resolves nowhere. The key
# alone, against whatever you've actually port-forwarded BASE_URL to below,
# is the one part of that URL still usable here.
#
# Usage:
# ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
#
# Prerequisites: kubectl context pointed at the demo namespace, jq, curl,
# and (in another terminal) a running:
# kubectl port-forward svc/terdut-operator-demo 8080:8080
set -euo pipefail
NAMESPACE="${NAMESPACE:-}"
BASE_URL="${BASE_URL:-http://localhost:8080}"
usage() {
cat >&2 <<'EOF'
usage: fire-alerts.sh <platform|payments> <scenario> [resolve]
scenarios:
high-cpu warning -- CPU usage above 90% for 10 minutes
disk-full critical -- disk usage above 95%
pod-crash error -- a pod crash-looping
heartbeat critical -- the team's dead man's switch heartbeat
(matches the matcher in 06/07-deadman-*.yaml -- send this
repeatedly to keep the switch alive, or stop sending it and
watch terdut-server open an incident on its own once
`timeout` passes with no heartbeat. "resolve" is not a valid
third argument for this scenario: a heartbeat is only ever
firing.)
env vars:
NAMESPACE kubectl -n for reading the webhook Secret (required)
BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080)
EOF
exit 1
}
[ $# -ge 2 ] || usage
team="$1" scenario="$2" verb="${3:-fire}"
[ -n "$NAMESPACE" ] || { echo "fire-alerts.sh: set NAMESPACE" >&2; exit 1; }
case "$team" in
platform|payments) ;;
*) usage ;;
esac
case "$scenario" in
high-cpu) alertname=TerdutDemoHighCPU severity=warning summary="CPU usage above 90% for 10 minutes" ;;
disk-full) alertname=TerdutDemoDiskFull severity=critical summary="Disk usage above 95% on /data" ;;
pod-crash) alertname=TerdutDemoPodCrashLooping severity=error summary="Pod web-7f8b9 is crash-looping (5 restarts in 10m)" ;;
heartbeat)
# Must match 06-deadman-platform.yaml / 07-deadman-payments.yaml's own
# matcher exactly -- that's what makes this a heartbeat rather than a
# third ordinary alert.
case "$team" in
platform) alertname=PlatformWatchdog ;;
payments) alertname=PaymentsWatchdog ;;
esac
severity=critical summary="demo heartbeat"
[ "$verb" = fire ] || { echo "fire-alerts.sh: heartbeat is only ever fired, never resolved -- just stop sending it" >&2; exit 1; }
;;
*) usage ;;
esac
case "$verb" in
fire) status=firing ;;
resolve) status=resolved ;;
*) usage ;;
esac
secret_name="terdutalertsource-${team}-terdut-webhook"
key="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.key}' | base64 -d)"
[ -n "$key" ] || { echo "fire-alerts.sh: empty key read from Secret $secret_name -- has 08/09-alertsource-*.yaml reconciled yet?" >&2; exit 1; }
# Stable per (team, scenario) so a resolve targets the same alert a fire
# created: terdut-server correlates on (team_id, fingerprint), not on
# anything else in the payload. Real Alertmanager computes this from the
# alert's label set; a fixed string plays the same role here.
fingerprint="demo-${team}-${scenario}"
now="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
if [ "$status" = firing ]; then
ends_at="0001-01-01T00:00:00Z" # Alertmanager's own "not resolved" zero value
else
ends_at="$now"
fi
payload="$(jq -n \
--arg status "$status" \
--arg groupKey "demo:${team}:${scenario}" \
--arg alertname "$alertname" \
--arg team "$team" \
--arg severity "$severity" \
--arg summary "$summary" \
--arg startsAt "$now" \
--arg endsAt "$ends_at" \
--arg fingerprint "$fingerprint" \
'{
version: "4",
status: $status,
groupKey: $groupKey,
groupLabels: { alertname: $alertname, team: $team },
alerts: [{
status: $status,
labels: { alertname: $alertname, severity: $severity, team: $team, instance: "demo" },
annotations: { summary: $summary },
startsAt: $startsAt,
endsAt: $endsAt,
generatorURL: "https://example.com/demo",
fingerprint: $fingerprint
}]
}')"
url="${BASE_URL}/api/integrations/${key}/alertmanager"
echo "POST $url (team=$team scenario=$scenario status=$status)" >&2
code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \
-X POST "$url" -H 'Content-Type: application/json' -d "$payload")"
echo "-> HTTP $code" >&2
cat /tmp/fire-alerts-response.json >&2
echo >&2
if [ "$code" != "200" ]; then
exit 1
fi
+17
View File
@@ -0,0 +1,17 @@
## kubectl apply -n <your-demo-namespace> -k examples/demo
##
## Listed in apply order even though kustomize itself doesn't need that --
## a human reading this file top-to-bottom should see the same dependency
## order the controllers themselves require (server before team, team
## before everything that teamRefs it).
resources:
- 00-postgres.yaml
- 01-server.yaml
- 02-team-platform.yaml
- 03-team-payments.yaml
- 04-escalation-platform.yaml
- 05-escalation-payments.yaml
- 06-deadman-platform.yaml
- 07-deadman-payments.yaml
- 08-alertsource-platform.yaml
- 09-alertsource-payments.yaml
+333
View File
@@ -0,0 +1,333 @@
#!/usr/bin/env bash
# Stands up the complete terdut demo (terdut-operator + every CRD kind this
# repo ships + a working local login + a few synthetic incidents) on a kind
# cluster, fully automated. Password login only -- this demo kit's own
# 01-server.yaml carries no oidc: block at all, so there is nothing to
# disable; OIDC is simply absent.
#
# Usage:
# ./run-demo.sh deploy the whole demo (idempotent: safe to
# re-run against a cluster that already has it)
# ./run-demo.sh --teardown delete the demo namespace (and, if
# TEARDOWN_CLUSTER=true, the kind cluster too)
# ./run-demo.sh --help
#
# All of the defaults below are overridable as environment variables.
set -euo pipefail
CLUSTER_NAME="${CLUSTER_NAME:-terdut-demo}"
NAMESPACE="${NAMESPACE:-terdut-operator-demo}"
OPERATOR_NAMESPACE="${OPERATOR_NAMESPACE:-terdut-operator-system}"
RELEASE_NAME="${RELEASE_NAME:-terdut-operator}"
ALICE_USERNAME="${ALICE_USERNAME:-alice}"
ALICE_EMAIL="${ALICE_EMAIL:-alice@terdut-demo.local}"
DEMO_PASSWORD="${DEMO_PASSWORD:-terdut-demo-1234}"
BASE_URL="${BASE_URL:-http://localhost:8080}"
LOCAL_PORT="${LOCAL_PORT:-8080}"
HELM_TIMEOUT="${HELM_TIMEOUT:-180s}"
WAIT_TIMEOUT="${WAIT_TIMEOUT:-180s}"
TEARDOWN_CLUSTER="${TEARDOWN_CLUSTER:-false}"
PF_PIDFILE="${PF_PIDFILE:-/tmp/terdut-demo-port-forward.pid}"
PF_LOGFILE="${PF_LOGFILE:-/tmp/terdut-demo-port-forward.log}"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
CHART_DIR="$(cd "$SCRIPT_DIR/../.." && pwd)/charts/terdut-operator"
# ---------------------------------------------------------------------------
# helpers
# ---------------------------------------------------------------------------
log() { printf '[run-demo] %s\n' "$*" >&2; }
die() { printf '[run-demo] FAILED: %s\n' "$*" >&2; exit 1; }
usage() {
cat <<EOF
Usage: $0 [--teardown|--help]
Deploys (or tears down) the complete terdut demo on a kind cluster.
See the top of this file for every overridable environment variable.
EOF
}
# Only kills a port-forward THIS invocation started, so a successful run
# never has its background job reaped by its own exit trap.
STARTED_PF_PID=""
cleanup_on_failure() {
local rc=$?
if [ "$rc" -ne 0 ] && [ -n "$STARTED_PF_PID" ]; then
log "run failed -- stopping the port-forward it started (pid $STARTED_PF_PID)"
kill "$STARTED_PF_PID" 2>/dev/null || true
fi
exit "$rc"
}
trap cleanup_on_failure EXIT
# ---------------------------------------------------------------------------
# steps
# ---------------------------------------------------------------------------
preflight() {
local missing=()
for bin in kubectl kind helm jq curl; do
command -v "$bin" >/dev/null 2>&1 || missing+=("$bin")
done
if [ "${#missing[@]}" -gt 0 ]; then
die "missing required tools on PATH: ${missing[*]}"
fi
}
ensure_kind_cluster() {
case "$(kind get clusters 2>/dev/null)" in
*"$CLUSTER_NAME"*)
log "kind cluster '$CLUSTER_NAME' already exists, skipping creation" ;;
*)
log "creating kind cluster '$CLUSTER_NAME'"
kind create cluster --name "$CLUSTER_NAME" ;;
esac
kubectl config use-context "kind-${CLUSTER_NAME}" >/dev/null
}
install_operator() {
log "installing terdut-operator into namespace $OPERATOR_NAMESPACE"
helm upgrade --install "$RELEASE_NAME" "$CHART_DIR" \
--namespace "$OPERATOR_NAMESPACE" --create-namespace \
--wait --timeout "$HELM_TIMEOUT" \
|| die "helm install of terdut-operator did not become ready"
}
apply_demo() {
log "creating namespace $NAMESPACE"
kubectl create namespace "$NAMESPACE" --dry-run=client -o yaml | kubectl apply -f - >/dev/null
log "applying demo CRs (00-09) into $NAMESPACE"
kubectl apply -n "$NAMESPACE" -k "$SCRIPT_DIR" >/dev/null
}
wait_for_ready() {
local objects=(
"terdutserver/terdut-operator-demo"
"terdutteam/terdutteam-platform"
"terdutteam/terdutteam-payments"
"terdutescalationrule/terdutescalationrule-platform"
"terdutescalationrule/terdutescalationrule-payments"
"terdutdeadmanswitch/terdutdeadmanswitch-platform"
"terdutdeadmanswitch/terdutdeadmanswitch-payments"
"terdutalertsource/terdutalertsource-platform"
"terdutalertsource/terdutalertsource-payments"
)
local obj
for obj in "${objects[@]}"; do
log "waiting for $obj to become Ready"
kubectl wait --for=condition=Ready --timeout "$WAIT_TIMEOUT" -n "$NAMESPACE" "$obj" >/dev/null \
|| die "timed out waiting for $obj -- try: kubectl describe -n $NAMESPACE $obj"
done
}
start_port_forward() {
# A stale pidfile from an earlier run would otherwise collide with us on
# $LOCAL_PORT -- if that pid is still alive, stop it first.
if [ -f "$PF_PIDFILE" ]; then
local old_pid
old_pid="$(cat "$PF_PIDFILE" 2>/dev/null || true)"
if [ -n "$old_pid" ] && kill -0 "$old_pid" 2>/dev/null; then
log "stopping stale port-forward from a previous run (pid $old_pid)"
kill "$old_pid" 2>/dev/null || true
sleep 1
fi
rm -f "$PF_PIDFILE"
fi
log "starting port-forward svc/terdut-operator-demo ${LOCAL_PORT}:8080"
kubectl -n "$NAMESPACE" port-forward svc/terdut-operator-demo "${LOCAL_PORT}:8080" \
>"$PF_LOGFILE" 2>&1 &
STARTED_PF_PID=$!
echo "$STARTED_PF_PID" > "$PF_PIDFILE"
local tries=0
until curl -sf -o /dev/null "http://localhost:${LOCAL_PORT}/healthz"; do
tries=$((tries + 1))
if [ "$tries" -ge 30 ]; then
die "port-forward never became ready -- see $PF_LOGFILE"
fi
sleep 1
done
log "port-forward ready (pid $STARTED_PF_PID, log $PF_LOGFILE)"
}
# Redeems Platform's own invite link -- minted by its TerdutTeam
# (02-team-platform.yaml's spec.invite.enabled, reconciled through that
# team's own already-working team-scoped credential, which requireTeamOwner
# already treats as owner-equivalent for /invites -- ratified, not a
# workaround, in terdut-server's SERVICE-ACCOUNTS.md) and redeemed through
# the ordinary signup endpoint. Invite redemption bypasses signup_mode
# entirely (terdut-server's internal/api/signup.go), so this needs no admin
# credential, no signup_mode flip, and no direct Postgres access at all --
# unlike an earlier version of this script, before terdut-operator grew
# this feature (see niklas/terdut-server#23).
redeem_platform_invite() {
log "reading Platform's invite link"
local secret_name invite_url invite_token tries=0
until secret_name="$(kubectl -n "$NAMESPACE" get terdutteam terdutteam-platform \
-o jsonpath='{.status.inviteSecretRef.name}' 2>/dev/null)" && [ -n "$secret_name" ]; do
tries=$((tries + 1))
[ "$tries" -lt 30 ] || die "terdutteam-platform never reported status.inviteSecretRef -- check spec.invite.enabled and kubectl describe it"
sleep 1
done
invite_url="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.url}' | base64 -d)"
invite_token="${invite_url##*invite=}"
[ -n "$invite_token" ] || die "could not parse an invite token out of $invite_url"
log "signing up ${ALICE_USERNAME} via Platform's invite"
local body resp_file code
body="$(jq -n \
--arg u "$ALICE_USERNAME" --arg e "$ALICE_EMAIL" \
--arg p "$DEMO_PASSWORD" --arg i "$invite_token" \
'{username: $u, email: $e, password: $p, invite: $i}')"
resp_file="$(mktemp)"
code="$(curl -sS -o "$resp_file" -w '%{http_code}' \
-X POST "${BASE_URL}/api/signup" -H 'Content-Type: application/json' -d "$body")"
case "$code" in
201) log "created local account ${ALICE_USERNAME} (joined Platform)" ;;
409) log "account ${ALICE_USERNAME} already exists, skipping (re-run detected)" ;;
*) die "signup failed (HTTP $code): $(cat "$resp_file")" ;;
esac
rm -f "$resp_file"
}
# redeem_platform_invite's signup already used up alice's one signup -- a
# second POST /api/signup would just 409 on the taken username, it doesn't
# join an existing account to another team. Payments is joined through the
# ordinary team-scoped member endpoint instead, using that team's own
# already-working team-scoped credential (owner-equivalent, same reach the
# invite route above relies on) and alice's user id resolved via
# GET /api/users -- the same lookup tdclient.GetUserByUsername does
# operator-side. The endpoint upserts on (team_id, user_id), so this is
# naturally idempotent across re-runs with no separate conflict handling
# needed.
join_payments_team() {
log "adding ${ALICE_USERNAME} to Payments"
local payments_id payments_secret payments_key alice_id resp_file code
payments_id="$(kubectl -n "$NAMESPACE" get terdutteam terdutteam-payments -o jsonpath='{.status.teamID}')"
[ -n "$payments_id" ] || die "could not read status.teamID off terdutteam-payments"
payments_secret="$(kubectl -n "$NAMESPACE" get terdutteam terdutteam-payments \
-o jsonpath='{.status.credentialsSecretRef.name}')"
[ -n "$payments_secret" ] || die "terdutteam-payments has no status.credentialsSecretRef yet"
payments_key="$(kubectl -n "$OPERATOR_NAMESPACE" get secret "$payments_secret" -o jsonpath='{.data.token}' | base64 -d)"
alice_id="$(curl -sS -H "Authorization: Bearer $payments_key" "${BASE_URL}/api/users" \
| jq -r --arg u "$ALICE_USERNAME" '.[] | select(.username == $u) | .id')"
[ -n "$alice_id" ] || die "could not resolve ${ALICE_USERNAME}'s user id via GET /api/users"
resp_file="$(mktemp)"
code="$(curl -sS -o "$resp_file" -w '%{http_code}' \
-X POST "${BASE_URL}/api/teams/${payments_id}/members" \
-H "Authorization: Bearer $payments_key" -H 'Content-Type: application/json' \
-d "$(jq -n --argjson id "$alice_id" '{user_id: $id, role: "member"}')")"
[ "$code" = "204" ] || die "adding ${ALICE_USERNAME} to Payments failed (HTTP $code): $(cat "$resp_file")"
log "added ${ALICE_USERNAME} to Payments"
rm -f "$resp_file"
}
fire_demo_alerts() {
log "firing representative demo alerts"
NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$SCRIPT_DIR/fire-alerts.sh" platform high-cpu
NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$SCRIPT_DIR/fire-alerts.sh" platform disk-full
NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$SCRIPT_DIR/fire-alerts.sh" payments pod-crash
}
print_summary() {
cat <<EOF
terdut demo is up.
Web UI: http://localhost:${LOCAL_PORT}
Login: ${ALICE_USERNAME} / ${DEMO_PASSWORD}
Cluster: kind-${CLUSTER_NAME}
Namespace: ${NAMESPACE}
Port-forward pid: ${STARTED_PF_PID:-$(cat "$PF_PIDFILE" 2>/dev/null || echo unknown)} (log: ${PF_LOGFILE})
stop it with: kill \$(cat ${PF_PIDFILE})
Fire more alerts:
export NAMESPACE=${NAMESPACE} BASE_URL=${BASE_URL}
./fire-alerts.sh platform high-cpu
./fire-alerts.sh platform high-cpu resolve
./fire-alerts.sh payments heartbeat # send repeatedly (e.g. every minute)
# to keep a dead man's switch alive;
# stop sending it and, 15 minutes
# later, terdut-server opens a
# critical incident on its own.
Tear down:
$0 --teardown
# add TEARDOWN_CLUSTER=true to also delete the kind cluster itself
EOF
}
teardown() {
if [ -f "$PF_PIDFILE" ]; then
local pid
pid="$(cat "$PF_PIDFILE" 2>/dev/null || true)"
if [ -n "$pid" ] && kill -0 "$pid" 2>/dev/null; then
log "stopping port-forward (pid $pid)"
kill "$pid" 2>/dev/null || true
fi
rm -f "$PF_PIDFILE"
fi
log "deleting namespace $NAMESPACE"
kubectl delete namespace "$NAMESPACE" --ignore-not-found --wait=true --timeout "$WAIT_TIMEOUT"
if [ "$TEARDOWN_CLUSTER" = "true" ]; then
log "deleting kind cluster $CLUSTER_NAME"
kind delete cluster --name "$CLUSTER_NAME"
else
log "leaving kind cluster '$CLUSTER_NAME' and the operator install in place" \
"(set TEARDOWN_CLUSTER=true to also delete the cluster)"
fi
}
# ---------------------------------------------------------------------------
# main
# ---------------------------------------------------------------------------
main() {
case "${1:-}" in
--teardown)
preflight
teardown
trap - EXIT
exit 0
;;
--help|-h)
usage
trap - EXIT
exit 0
;;
"") ;;
*)
usage
die "unknown argument: $1"
;;
esac
preflight
ensure_kind_cluster
install_operator
apply_demo
wait_for_ready
start_port_forward
redeem_platform_invite
join_payments_team
fire_demo_alerts
print_summary
# Success: leave the port-forward running, don't let the EXIT trap kill it.
trap - EXIT
}
main "$@"
+31
View File
@@ -0,0 +1,31 @@
# Exposing a TerdutServer
This operator never manages external exposure/ingress for `TerdutServer`,
in any form — a permanent non-goal (`DESIGN.md` §1), not a missing
feature. Some installs won't expose it outside the cluster at all (see
`examples/demo`, which just port-forwards); others will put it behind
whatever their cluster already uses. That choice is entirely yours, not
the operator's.
The only contract the operator gives you to build on: a plain `ClusterIP`
Service, named after the `TerdutServer` CR (same name, same namespace),
with a port named `http` (`spec.networking.servicePort`, default `8080`).
Everything here targets exactly that Service — none of it is applied by
`examples/demo`'s `kustomization.yaml`, and none of it depends on anything
the operator creates beyond that one Service.
Pick whichever matches your cluster:
- **`httproute.yaml`** — a [Gateway API](https://gateway-api.sigs.k8s.io/)
`HTTPRoute`, attached to a `Gateway` you already have.
- **`istio-virtualservice.yaml`** — an Istio `VirtualService`, attached to
a `Gateway` (Istio's own CRD, not Gateway API's) you already have.
- **A plain `Ingress`** needs no example here — it's the same idea with
one fewer layer of indirection: an `Ingress` with a single rule whose
`backend.service.name`/`port.name` point at the `TerdutServer`'s Service
and `http` port.
Remember to set `spec.networking.hostname` on the `TerdutServer` itself to
whatever hostname you expose it on — that's not read by the operator for
any of this, but terdut-server uses it for its own absolute links
(notifications, OIDC redirect URIs).
+21
View File
@@ -0,0 +1,21 @@
# Gateway API HTTPRoute exposing a TerdutServer through a Gateway you
# already have (not something this operator creates or watches -- see
# ../networking/README.md). Replace terdut-operator-demo and the Gateway
# reference/hostname with your own; terdut-operator-demo matches
# ../demo/01-server.yaml, if you're layering this onto that demo.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: terdut-operator-demo
namespace: terdut-operator-demo
spec:
parentRefs:
- name: my-gateway # an existing Gateway in this namespace (or
# namespace: gateway-ns # a different one, if the Gateway allows it)
sectionName: https # the listener to attach to, if it's picky
hostnames:
- terdut-operator-demo.example # matches networking.hostname on the CR
rules:
- backendRefs:
- name: terdut-operator-demo # the Service the operator created --
port: 8080 # same name as the TerdutServer CR
@@ -0,0 +1,23 @@
# Istio VirtualService exposing a TerdutServer through an Istio Gateway
# you already have (Istio's own Gateway CRD, not Gateway API's -- not
# something this operator creates or watches, see ../networking/README.md).
# Replace terdut-operator-demo and the gateway reference/hostname with your
# own; terdut-operator-demo matches ../demo/01-server.yaml, if you're
# layering this onto that demo.
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: terdut-operator-demo
namespace: terdut-operator-demo
spec:
hosts:
- terdut-operator-demo.example # matches networking.hostname on the CR
gateways:
- my-gateway-namespace/my-gateway # an existing istio Gateway
http:
- route:
- destination:
host: terdut-operator-demo.terdut-operator-demo.svc.cluster.local
port:
number: 8080 # the Service's port -- same name as the
# TerdutServer CR, default servicePort
+10 -10
View File
@@ -25,7 +25,7 @@ require (
github.com/felixge/httpsnoop v1.0.4 // indirect github.com/felixge/httpsnoop v1.0.4 // indirect
github.com/fsnotify/fsnotify v1.9.0 // indirect github.com/fsnotify/fsnotify v1.9.0 // indirect
github.com/fxamacker/cbor/v2 v2.9.1 // indirect github.com/fxamacker/cbor/v2 v2.9.1 // indirect
github.com/go-logr/logr v1.4.3 // indirect github.com/go-logr/logr v1.4.4 // indirect
github.com/go-logr/stdr v1.2.2 // indirect github.com/go-logr/stdr v1.2.2 // indirect
github.com/go-logr/zapr v1.3.0 // indirect github.com/go-logr/zapr v1.3.0 // indirect
github.com/go-openapi/jsonpointer v1.0.0 // indirect github.com/go-openapi/jsonpointer v1.0.0 // indirect
@@ -64,13 +64,13 @@ require (
github.com/x448/float16 v0.8.4 // indirect github.com/x448/float16 v0.8.4 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 // indirect
go.opentelemetry.io/otel v1.44.0 // indirect go.opentelemetry.io/otel v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 // indirect go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 // indirect
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 // indirect go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 // indirect
go.opentelemetry.io/otel/metric v1.44.0 // indirect go.opentelemetry.io/otel/metric v1.45.0 // indirect
go.opentelemetry.io/otel/sdk v1.44.0 // indirect go.opentelemetry.io/otel/sdk v1.45.0 // indirect
go.opentelemetry.io/otel/trace v1.44.0 // indirect go.opentelemetry.io/otel/trace v1.45.0 // indirect
go.opentelemetry.io/proto/otlp v1.10.0 // indirect go.opentelemetry.io/proto/otlp v1.11.0 // indirect
go.uber.org/multierr v1.11.0 // indirect go.uber.org/multierr v1.11.0 // indirect
go.uber.org/zap v1.27.1 // indirect go.uber.org/zap v1.27.1 // indirect
go.yaml.in/yaml/v2 v2.4.4 // indirect go.yaml.in/yaml/v2 v2.4.4 // indirect
@@ -86,8 +86,8 @@ require (
golang.org/x/time v0.15.0 // indirect golang.org/x/time v0.15.0 // indirect
golang.org/x/tools v0.48.0 // indirect golang.org/x/tools v0.48.0 // indirect
gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect gomodules.xyz/jsonpatch/v2 v2.4.0 // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa // indirect google.golang.org/genproto/googleapis/api v0.0.0-20260803160001-6ac0973c030d // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa // indirect google.golang.org/genproto/googleapis/rpc v0.0.0-20260803160001-6ac0973c030d // indirect
google.golang.org/grpc v1.83.2 // indirect google.golang.org/grpc v1.83.2 // indirect
google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af // indirect google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af // indirect
gopkg.in/evanphx/json-patch.v4 v4.13.0 // indirect gopkg.in/evanphx/json-patch.v4 v4.13.0 // indirect
+22 -22
View File
@@ -36,8 +36,8 @@ github.com/gkampitakis/go-diff v1.3.2/go.mod h1:LLgOrpqleQe26cte8s36HTWcTmMEur6O
github.com/gkampitakis/go-snaps v0.5.15 h1:amyJrvM1D33cPHwVrjo9jQxX8g/7E2wYdZ+01KS3zGE= github.com/gkampitakis/go-snaps v0.5.15 h1:amyJrvM1D33cPHwVrjo9jQxX8g/7E2wYdZ+01KS3zGE=
github.com/gkampitakis/go-snaps v0.5.15/go.mod h1:HNpx/9GoKisdhw9AFOBT1N7DBs9DiHo/hGheFGBZ+mc= github.com/gkampitakis/go-snaps v0.5.15/go.mod h1:HNpx/9GoKisdhw9AFOBT1N7DBs9DiHo/hGheFGBZ+mc=
github.com/go-logr/logr v1.2.2/go.mod h1:jdQByPbusPIv2/zmleS9BjJVeZ6kBagPoEUsqbVz/1A= github.com/go-logr/logr v1.2.2/go.mod h1:jdQByPbusPIv2/zmleS9BjJVeZ6kBagPoEUsqbVz/1A=
github.com/go-logr/logr v1.4.3 h1:CjnDlHq8ikf6E492q6eKboGOC0T8CDaOvkHCIg8idEI= github.com/go-logr/logr v1.4.4 h1:tG4xh9yMsRCAiodLVTxyrkzSZ9+o0L1Kg/+cPVcbP/8=
github.com/go-logr/logr v1.4.3/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY= github.com/go-logr/logr v1.4.4/go.mod h1:9T104GzyrTigFIr8wt5mBrctHMim0Nb2HLGrmQ40KvY=
github.com/go-logr/stdr v1.2.2 h1:hSWxHoqTgW2S2qGc0LTAI563KZ5YKYRhT3MFKZMbjag= github.com/go-logr/stdr v1.2.2 h1:hSWxHoqTgW2S2qGc0LTAI563KZ5YKYRhT3MFKZMbjag=
github.com/go-logr/stdr v1.2.2/go.mod h1:mMo/vtBO5dYbehREoey6XUKy/eSumjCCveDpRre4VKE= github.com/go-logr/stdr v1.2.2/go.mod h1:mMo/vtBO5dYbehREoey6XUKy/eSumjCCveDpRre4VKE=
github.com/go-logr/zapr v1.3.0 h1:XGdV8XW8zdwFiwOA2Dryh1gj2KRQyOOoNmBy4EplIcQ= github.com/go-logr/zapr v1.3.0 h1:XGdV8XW8zdwFiwOA2Dryh1gj2KRQyOOoNmBy4EplIcQ=
@@ -168,22 +168,22 @@ go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ
go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y= go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo= go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0 h1:8tvICD4vSTOOsNrsI4Ljf6C+6UKvpTEH5XY3JMoyPoo=
go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI= go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.69.0/go.mod h1:z9+yiacE0IHRqM4qFfkbt/JYlmYXgss8GY/jXoNuPJI=
go.opentelemetry.io/otel v1.44.0 h1:JjwHmHpA4iZ3wBxluu2fbbE7j4kqlE8jXyAyPXH7HqU= go.opentelemetry.io/otel v1.45.0 h1:pdrWmLHofpubmArBv1LgFSv1Z0Ie/ppdZzu+kUN5EeU=
go.opentelemetry.io/otel v1.44.0/go.mod h1:BMgjTHL9WPRlRjL2oZCBTL4whCGtXch2H4BhOPIAyYc= go.opentelemetry.io/otel v1.45.0/go.mod h1:XZxIqPapzEYnhNSScF5DIqXhm/rYi0FzCe2XddAwZfQ=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0 h1:4YsVu3B8+3qtWYYrsUYgn0OG78pN0rnNPRGX4SbokQI= go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0 h1:QRefszxJmfPdjXUUm3j6iDzY03mTPXMjqErFqQ67vUg=
go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.44.0/go.mod h1:+wnlSn0mD1ADVMe3v9Z/WIaiz6q6gL2J/ejaAmdmv80= go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.45.0/go.mod h1:Tiz03lTBVBrm7eWZBOidzEaYaJa8tjwGUGv6d8mlTyk=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0 h1:qazEJlUOQzhCpzQpFETGby7EdqjI1wsd0W+6Gg1SCTU= go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0 h1:fG5MCxGz8+2VtrN/WgqSpJFctVz24gpxj8CxkKmc8Ww=
go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.44.0/go.mod h1:fOD2Yefuxixkx3ahVNf0O/PERb6r4OlbxfATVnYvzCo= go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.45.0/go.mod h1:BmAYTn+3ysbRe+IU2msxmf5Rx3g6DHvex+tWI3LdhYI=
go.opentelemetry.io/otel/metric v1.44.0 h1:1w0gILTcHdr3YI+ixLyjemwrVnsMURbTZFrSYCdDdmc= go.opentelemetry.io/otel/metric v1.45.0 h1:7Eg1uH7CJ5cXv9is6tnBe1FI6rj1nwUdbFypRm3br/M=
go.opentelemetry.io/otel/metric v1.44.0/go.mod h1:8O7hanEPBNgEMmybD3s2VBKcgWOCsA6tzHBPODAiquo= go.opentelemetry.io/otel/metric v1.45.0/go.mod h1:HAPbm1nd3p1PmFH7v2dR+6BjXxw+Lq4a2+pndMAm08s=
go.opentelemetry.io/otel/sdk v1.44.0 h1:nHYwb9lK+fJPU/dnT6s7W7Z8itMWyqrnVfbheVYrZ58= go.opentelemetry.io/otel/sdk v1.45.0 h1:4VVSMgQ83dUgW2aoX5f6JgLvHwIvzcuLnF9lUdCSpCw=
go.opentelemetry.io/otel/sdk v1.44.0/go.mod h1:Osuydd3Se74nqjAKxid74N5eC+jfEqfTegHRnq58oK0= go.opentelemetry.io/otel/sdk v1.45.0/go.mod h1:Sr40LgXV7DsKMMJMKOhUWOgMWTfAaqvm2kF0g7ilwuA=
go.opentelemetry.io/otel/sdk/metric v1.44.0 h1:3LlKgI+VjbVsjNRFZJZAJ30WjXC5VkNRks6si09iEfI= go.opentelemetry.io/otel/sdk/metric v1.45.0 h1:oVFszMfyj1Am6s24Vtc7wBb8BKLcwepJjNEYILuiE3o=
go.opentelemetry.io/otel/sdk/metric v1.44.0/go.mod h1:5B5pMARnXxKhltooO4xUuCBorl65a4EpnTalObqOigA= go.opentelemetry.io/otel/sdk/metric v1.45.0/go.mod h1:vUWUxDZvu1WVRj8JA8S0AdhsPrZoDpA2DdZauIh4mDA=
go.opentelemetry.io/otel/trace v1.44.0 h1:jxF5CsGYCe74MCRx2X4g7WsY/VBKRqqpNvXlX/6gtIk= go.opentelemetry.io/otel/trace v1.45.0 h1:l/mP6Uv7oNO7/TblbhpbgMidxhq1uO/rPsikOyVhxag=
go.opentelemetry.io/otel/trace v1.44.0/go.mod h1:oLl1jrMQAVo6v3GAggN+1VH9VIz9iUSvW53sW1Q8PIE= go.opentelemetry.io/otel/trace v1.45.0/go.mod h1:qoJJA2xNMnxRrdISU/kLtfUH2wNeQbiv+jhs/CxI8bc=
go.opentelemetry.io/proto/otlp v1.10.0 h1:IQRWgT5srOCYfiWnpqUYz9CVmbO8bFmKcwYxpuCSL2g= go.opentelemetry.io/proto/otlp v1.11.0 h1:5rrYs0Ykyj50sdU/JU0x8etU+LubXWb+gED6TbEdMIk=
go.opentelemetry.io/proto/otlp v1.10.0/go.mod h1:/CV4QoCR/S9yaPj8utp3lvQPoqMtxXdzn7ozvvozVqk= go.opentelemetry.io/proto/otlp v1.11.0/go.mod h1:SmVizdCOAm3XBtG1g1NnOdhW6jtddT72hLMhv8VwA8E=
go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto= go.uber.org/goleak v1.3.0 h1:2K3zAYmnTNqV73imy9J1T3WC+gmCePx2hEGkimedGto=
go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE= go.uber.org/goleak v1.3.0/go.mod h1:CoHD4mav9JJNrW/WLlf7HGZPjdw8EucARQHekz1X6bE=
go.uber.org/multierr v1.11.0 h1:blXXJkSxSSfBVBlC76pxqeO+LN3aDfLQo+309xJstO0= go.uber.org/multierr v1.11.0 h1:blXXJkSxSSfBVBlC76pxqeO+LN3aDfLQo+309xJstO0=
@@ -218,10 +218,10 @@ gomodules.xyz/jsonpatch/v2 v2.4.0 h1:Ci3iUJyx9UeRx7CeFN8ARgGbkESwJK+KB9lLcWxY/Zw
gomodules.xyz/jsonpatch/v2 v2.4.0/go.mod h1:AH3dM2RI6uoBZxn3LVrfvJ3E0/9dG4cSrbuBJT4moAY= gomodules.xyz/jsonpatch/v2 v2.4.0/go.mod h1:AH3dM2RI6uoBZxn3LVrfvJ3E0/9dG4cSrbuBJT4moAY=
gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4= gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4=
gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E= gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E=
google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa h1:Kjn0N0tCrDgiAFW+lGO4JZ3ck44CehvJQMAwj9QF0G8= google.golang.org/genproto/googleapis/api v0.0.0-20260803160001-6ac0973c030d h1:FarXi840EJWSHYTN3ERkADbPWjl307+FGrA22KAVjjc=
google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa/go.mod h1:q4lMZS6kskjT5HvCPrnnypcDPVJqT/f4nfxmkE7gryY= google.golang.org/genproto/googleapis/api v0.0.0-20260803160001-6ac0973c030d/go.mod h1:K/+WGbmBY7aNW1HDw1fJnKYo10i0DkAX6pows00dLig=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa h1:mZHHdPZl0dbGHCflZgAq/Q468DWVFcU2whhB2KAo8fk= google.golang.org/genproto/googleapis/rpc v0.0.0-20260803160001-6ac0973c030d h1:IL4hdHzcUv2l/gcg98/Rj3FbtE6axwqslOW8SW0C+S0=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8= google.golang.org/genproto/googleapis/rpc v0.0.0-20260803160001-6ac0973c030d/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8=
google.golang.org/grpc v1.83.2 h1:EManeRomTObA0BU7I8vXgg/78uE5MJ9M8B39EX2WscU= google.golang.org/grpc v1.83.2 h1:EManeRomTObA0BU7I8vXgg/78uE5MJ9M8B39EX2WscU=
google.golang.org/grpc v1.83.2/go.mod h1:YPI1hK3kDked6iHvgX3tR0y+nX/qpMFKhPgFsokw1S8= google.golang.org/grpc v1.83.2/go.mod h1:YPI1hK3kDked6iHvgX3tR0y+nX/qpMFKhPgFsokw1S8=
google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af h1:+5/Sw3GsDNlEmu7TfklWKPdQ0Ykja5VEmq2i817+jbI= google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af h1:+5/Sw3GsDNlEmu7TfklWKPdQ0Ykja5VEmq2i817+jbI=
@@ -8,6 +8,7 @@ import (
appsv1 "k8s.io/api/apps/v1" appsv1 "k8s.io/api/apps/v1"
corev1 "k8s.io/api/core/v1" corev1 "k8s.io/api/core/v1"
policyv1 "k8s.io/api/policy/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors" apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta" "k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
@@ -97,6 +98,7 @@ type TerdutServerReconciler struct {
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups="",resources=services,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups="",resources=services,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=policy,resources=poddisruptionbudgets,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=acid.zalan.do,resources=postgresqls,verbs=get;list;watch // +kubebuilder:rbac:groups=acid.zalan.do,resources=postgresqls,verbs=get;list;watch
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch // +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
@@ -137,6 +139,9 @@ func (r *TerdutServerReconciler) Reconcile(ctx context.Context, req ctrl.Request
if err := r.reconcileService(ctx, &srv); err != nil { if err := r.reconcileService(ctx, &srv); err != nil {
return ctrl.Result{}, err return ctrl.Result{}, err
} }
if err := r.reconcilePodDisruptionBudget(ctx, &srv); err != nil {
return ctrl.Result{}, err
}
meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{ meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionDatabaseReady, Type: terdutv1alpha1.ConditionDatabaseReady,
@@ -246,6 +251,7 @@ func (r *TerdutServerReconciler) SetupWithManager(mgr ctrl.Manager) error {
For(&terdutv1alpha1.TerdutServer{}). For(&terdutv1alpha1.TerdutServer{}).
Owns(&appsv1.Deployment{}). Owns(&appsv1.Deployment{}).
Owns(&corev1.Service{}). Owns(&corev1.Service{}).
Owns(&policyv1.PodDisruptionBudget{}).
Named("terdutserver"). Named("terdutserver").
Complete(r) Complete(r)
} }
@@ -9,15 +9,20 @@ import (
"strconv" "strconv"
"strings" "strings"
"sync" "sync"
"time"
. "github.com/onsi/ginkgo/v2" . "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega" . "github.com/onsi/gomega"
appsv1 "k8s.io/api/apps/v1" appsv1 "k8s.io/api/apps/v1"
corev1 "k8s.io/api/core/v1" corev1 "k8s.io/api/core/v1"
policyv1 "k8s.io/api/policy/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta" "k8s.io/apimachinery/pkg/api/meta"
"k8s.io/apimachinery/pkg/api/resource"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
"k8s.io/apimachinery/pkg/types" "k8s.io/apimachinery/pkg/types"
"k8s.io/apimachinery/pkg/util/intstr"
ctrl "sigs.k8s.io/controller-runtime" ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/reconcile" "sigs.k8s.io/controller-runtime/pkg/reconcile"
@@ -31,6 +36,11 @@ import (
const ( const (
fakeVersionString = "test" fakeVersionString = "test"
errJSONKey = "error" errJSONKey = "error"
// deadmanSwitchesPath is the literal path (not Printf'd like the others
// below) shared by the exact-match collection route and the dispatcher
// that routes into it -- goconst flags three occurrences of the same
// string, so this is that string, once.
deadmanSwitchesPath = "/deadman/switches"
) )
// fakeTerdutServer reproduces the exact stateful semantics of // fakeTerdutServer reproduces the exact stateful semantics of
@@ -79,6 +89,14 @@ type fakeTerdutServer struct {
nextIntegrationID int64 nextIntegrationID int64
integrations map[int64]map[int64]tdclient.Integration // teamID -> integrationID -> integration integrations map[int64]map[int64]tdclient.Integration // teamID -> integrationID -> integration
integrationDelete map[int64]bool // integrationID -> true once DELETEd, for 404-on-redelete integrationDelete map[int64]bool // integrationID -> true once DELETEd, for 404-on-redelete
// invites/nextInviteID/inviteDelete back TerdutTeam's own invite-minting
// feature -- no unique constraint on an invite server-side either (every
// POST mints a brand new row, confirmed against source), same keyed-by-id
// shape as switches/integrations.
nextInviteID int64
invites map[int64]map[int64]tdclient.Invite // teamID -> inviteID -> invite
inviteDelete map[int64]bool // inviteID -> true once DELETEd, for 404-on-redelete
} }
func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) { func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) {
@@ -96,6 +114,9 @@ func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) {
integrations: map[int64]map[int64]tdclient.Integration{}, integrations: map[int64]map[int64]tdclient.Integration{},
integrationDelete: map[int64]bool{}, integrationDelete: map[int64]bool{},
invites: map[int64]map[int64]tdclient.Invite{},
inviteDelete: map[int64]bool{},
} }
return f, httptest.NewServer(f) return f, httptest.NewServer(f)
} }
@@ -261,7 +282,26 @@ func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Requ
f.escalation[id] = req f.escalation[id] = req
w.WriteHeader(http.StatusNoContent) w.WriteHeader(http.StatusNoContent)
case rest == "/deadman/switches" && r.Method == http.MethodGet: case rest == deadmanSwitchesPath || strings.HasPrefix(rest, deadmanSwitchesPath+"/"):
f.handleDeadmanSubPath(w, r, id, rest)
case rest == "/integrations" || strings.HasPrefix(rest, "/integrations/"):
f.handleIntegrationSubPath(w, r, id, rest)
case rest == "/invites" || strings.HasPrefix(rest, "/invites/"):
f.handleInviteSubPath(w, r, id, rest)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleDeadmanSubPath answers GET/POST /api/teams/{id}/deadman/switches and
// PUT/DELETE .../deadman/switches/{switchID} -- split out of
// handleTeamSubPath for the same gocyclo reason as handleIntegrationSubPath.
func (f *fakeTerdutServer) handleDeadmanSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == deadmanSwitchesPath && r.Method == http.MethodGet:
existing := f.switches[id] existing := f.switches[id]
out := make([]tdclient.DeadmanSwitch, 0, len(existing)) out := make([]tdclient.DeadmanSwitch, 0, len(existing))
for _, s := range existing { for _, s := range existing {
@@ -269,7 +309,7 @@ func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Requ
} }
writeJSON(w, http.StatusOK, out) writeJSON(w, http.StatusOK, out)
case rest == "/deadman/switches" && r.Method == http.MethodPost: case rest == deadmanSwitchesPath && r.Method == http.MethodPost:
var req deadmanSwitchFakeRequest var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req) _ = json.NewDecoder(r.Body).Decode(&req)
f.nextSwitchID++ f.nextSwitchID++
@@ -324,8 +364,48 @@ func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Requ
f.switchDelete[switchID] = true f.switchDelete[switchID] = true
w.WriteHeader(http.StatusNoContent) w.WriteHeader(http.StatusNoContent)
case rest == "/integrations" || strings.HasPrefix(rest, "/integrations/"): default:
f.handleIntegrationSubPath(w, r, id, rest) w.WriteHeader(http.StatusNotFound)
}
}
// handleInviteSubPath answers POST /api/teams/{id}/invites and
// DELETE .../invites/{inviteID} -- split out for the same gocyclo reason as
// handleIntegrationSubPath.
func (f *fakeTerdutServer) handleInviteSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "/invites" && r.Method == http.MethodPost:
var req struct {
Role string `json:"role"`
MaxUses int64 `json:"max_uses"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
f.nextInviteID++
inviteID := f.nextInviteID
inv := tdclient.Invite{
ID: inviteID, TeamID: id, Role: req.Role, MaxUses: req.MaxUses,
ExpiresAt: time.Now().Add(7 * 24 * time.Hour),
URL: fmt.Sprintf("https://terdut.example.invalid/signup?invite=invite-token-%d", inviteID),
}
if f.invites[id] == nil {
f.invites[id] = map[int64]tdclient.Invite{}
}
f.invites[id][inviteID] = inv
writeJSON(w, http.StatusCreated, inv)
case strings.HasPrefix(rest, "/invites/") && r.Method == http.MethodDelete:
inviteID, ok := parseTrailingID(rest, "/invites/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.invites[id][inviteID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.invites[id], inviteID)
f.inviteDelete[inviteID] = true
w.WriteHeader(http.StatusNoContent)
default: default:
w.WriteHeader(http.StatusNotFound) w.WriteHeader(http.StatusNotFound)
@@ -705,6 +785,95 @@ var _ = Describe("TerdutServer Controller", func() {
}) })
}) })
Describe("spec.pod", func() {
It("wires pod-level customization onto the right spot on the Deployment", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
spec := dsnSpec()
qty := resource.MustParse("250m")
spec.Pod = terdutv1alpha1.PodSpec{
Resources: corev1.ResourceRequirements{Requests: corev1.ResourceList{corev1.ResourceCPU: qty}},
Tolerations: []corev1.Toleration{{Key: "dedicated", Operator: corev1.TolerationOpEqual, Value: "terdut", Effect: corev1.TaintEffectNoSchedule}},
ExtraEnv: []corev1.EnvVar{{Name: "EXTRA_FLAG", Value: "on"}},
ServiceAccountName: "terdut-server-custom",
ExtraVolumes: []corev1.Volume{{Name: "extra-ca", VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}},
ExtraVolumeMounts: []corev1.VolumeMount{{Name: "extra-ca", MountPath: "/etc/extra-ca"}},
}
createServer(ctx, spec)
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service/PDB
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
podSpec := deploy.Spec.Template.Spec
Expect(podSpec.Tolerations).To(ConsistOf(spec.Pod.Tolerations))
Expect(podSpec.ServiceAccountName).To(Equal("terdut-server-custom"))
Expect(podSpec.Volumes).To(ConsistOf(spec.Pod.ExtraVolumes))
main := podSpec.Containers[0]
Expect(main.Name).To(Equal("terdut-server"))
Expect(main.Resources).To(Equal(spec.Pod.Resources))
Expect(main.VolumeMounts).To(ConsistOf(spec.Pod.ExtraVolumeMounts))
Expect(main.Env).To(ContainElement(corev1.EnvVar{Name: "EXTRA_FLAG", Value: "on"}))
initContainer := podSpec.InitContainers[0]
Expect(initContainer.Name).To(Equal("wait-for-postgres"))
Expect(initContainer.VolumeMounts).To(BeEmpty(), "extraVolumeMounts must not leak onto wait-for-postgres")
})
})
Describe("spec.pod.disruptionBudget", func() {
It("creates an owned PodDisruptionBudget when set, and deletes it once cleared", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
spec := dsnSpec()
minAvail := intstr.FromInt32(1)
spec.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{MinAvailable: &minAvail}
createServer(ctx, spec)
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service/PDB
var pdb policyv1.PodDisruptionBudget
Expect(k8sClient.Get(ctx, objKey, &pdb)).To(Succeed())
Expect(pdb.Spec.Selector.MatchLabels).To(Equal(labelsFor(&terdutv1alpha1.TerdutServer{ObjectMeta: metav1.ObjectMeta{Name: name}})))
Expect(pdb.Spec.MinAvailable).To(Equal(&minAvail))
Expect(pdb.OwnerReferences).To(ContainElement(HaveField("Name", name)))
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
srv.Spec.Pod.DisruptionBudget = nil
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
reconcileOnce(ctx)
err := k8sClient.Get(ctx, objKey, &policyv1.PodDisruptionBudget{})
Expect(apierrors.IsNotFound(err)).To(BeTrue(), "PodDisruptionBudget should be deleted once spec.pod.disruptionBudget is cleared")
})
It("rejects both minAvailable and maxUnavailable set together, and neither set", func(ctx SpecContext) {
bothSet := dsnSpec()
minAvail, maxUnavail := intstr.FromInt32(1), intstr.FromInt32(1)
bothSet.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{MinAvailable: &minAvail, MaxUnavailable: &maxUnavail}
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: name + "-both", Namespace: operatorNamespace},
Spec: bothSet,
})).To(HaveOccurred())
neitherSet := dsnSpec()
neitherSet.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{}
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: name + "-neither", Namespace: operatorNamespace},
Spec: neitherSet,
})).To(HaveOccurred())
})
})
Describe("deletion", func() { Describe("deletion", func() {
It("removes the credentials and checkpoint Secrets and the finalizer", func(ctx SpecContext) { It("removes the credentials and checkpoint Secrets and the finalizer", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer() fake, fakeSrv := newFakeTerdutServer()
+60 -7
View File
@@ -48,10 +48,26 @@ func (r *TerdutServerReconciler) reconcileDeployment(
// overlapping during a rollout would both page for the same // overlapping during a rollout would both page for the same
// incident (matches the chart's own deployment.yaml comment). // incident (matches the chart's own deployment.yaml comment).
deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RecreateDeploymentStrategyType} deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RecreateDeploymentStrategyType}
pod := srv.Spec.Pod
deploy.Spec.Template = corev1.PodTemplateSpec{ deploy.Spec.Template = corev1.PodTemplateSpec{
ObjectMeta: metav1.ObjectMeta{Labels: labels}, // pod.Annotations is assigned directly, not merged -- nothing
// else sets pod-template annotations today. If a future change
// needs the controller to set one of its own (e.g. a
// Prometheus-scrape annotation), this needs to become a real
// map merge with a stated precedence rather than silently
// clobbering one side.
ObjectMeta: metav1.ObjectMeta{Labels: labels, Annotations: pod.Annotations},
Spec: corev1.PodSpec{ Spec: corev1.PodSpec{
EnableServiceLinks: new(false), EnableServiceLinks: new(false),
NodeSelector: pod.NodeSelector,
Tolerations: pod.Tolerations,
Affinity: pod.Affinity,
TopologySpreadConstraints: pod.TopologySpreadConstraints,
SecurityContext: pod.SecurityContext,
ServiceAccountName: pod.ServiceAccountName,
ImagePullSecrets: pod.ImagePullSecrets,
InitContainers: []corev1.Container{waitForPostgresContainer(dbEnv)},
Volumes: pod.ExtraVolumes,
Containers: []corev1.Container{{ Containers: []corev1.Container{{
Name: "terdut-server", Name: "terdut-server",
Image: fmt.Sprintf("%s:%s", srv.Spec.Image.Repository, srv.Spec.Image.Tag), Image: fmt.Sprintf("%s:%s", srv.Spec.Image.Repository, srv.Spec.Image.Tag),
@@ -60,9 +76,13 @@ func (r *TerdutServerReconciler) reconcileDeployment(
ContainerPort: servicePort(srv), ContainerPort: servicePort(srv),
Protocol: corev1.ProtocolTCP, Protocol: corev1.ProtocolTCP,
}}, }},
Env: buildEnv(srv, dbEnv), Env: buildEnv(srv, dbEnv),
LivenessProbe: healthzProbe(), EnvFrom: pod.ExtraEnvFrom,
ReadinessProbe: healthzProbe(), VolumeMounts: pod.ExtraVolumeMounts,
Resources: pod.Resources,
SecurityContext: pod.ContainerSecurityContext,
LivenessProbe: healthzProbe(),
ReadinessProbe: healthzProbe(),
}}, }},
}, },
} }
@@ -104,6 +124,36 @@ func servicePort(srv *terdutv1alpha1.TerdutServer) int32 {
return srv.Spec.Networking.ServicePort return srv.Spec.Networking.ServicePort
} }
// waitForPostgresContainer blocks the main container from starting until
// Postgres accepts connections, matching charts/terdut-server's own
// deployment.yaml template as of v0.33.2 (that repo's CLAUDE.md/release
// notes) -- that chart grew this the moment this exact Deployment, created
// by this controller, crash-looped a few times against a from-scratch
// postgres-operator cluster still doing initdb and Patroni leader election:
// terdut-server's own ping-retry budget on startup (internal/db/db.go) is
// sized for a much shorter, different race (NetworkPolicy propagation, a
// few seconds), not for genuine first-time cluster creation, so it
// exhausted and the process exited before ever binding its HTTP port -- a
// startupProbe cannot help there, since the crash happens before there is
// anything to probe.
//
// Reuses dbEnv unchanged: both of resolveDatabaseEnv's paths put
// TERDUT_DB_DSN first (terdutserver_database.go), so it's already exactly
// what pg_isready needs, and pg_isready needs no credentials -- it reports
// PQPING_OK on anything that amounts to a Postgres backend answering,
// including an auth challenge -- so including dbEnv's optional PGPASSWORD
// here too is harmless rather than load-bearing.
func waitForPostgresContainer(dbEnv []corev1.EnvVar) corev1.Container {
return corev1.Container{
Name: "wait-for-postgres",
Image: "postgres:17-alpine",
Env: dbEnv,
Command: []string{"sh", "-c",
`until pg_isready -d "$TERDUT_DB_DSN"; do echo "wait-for-postgres: not ready yet, retrying in 2s"; sleep 2; done`,
},
}
}
func healthzProbe() *corev1.Probe { func healthzProbe() *corev1.Probe {
return &corev1.Probe{ return &corev1.Probe{
ProbeHandler: corev1.ProbeHandler{ ProbeHandler: corev1.ProbeHandler{
@@ -120,7 +170,10 @@ func healthzProbe() *corev1.Probe {
// field-for-field (confirmed against that source, not reconstructed from // field-for-field (confirmed against that source, not reconstructed from
// DESIGN.md's illustrative YAML alone) — dbEnv (TERDUT_DB_DSN, optionally // DESIGN.md's illustrative YAML alone) — dbEnv (TERDUT_DB_DSN, optionally
// PGPASSWORD) comes from resolveDatabaseEnv, since which of §8's two paths // PGPASSWORD) comes from resolveDatabaseEnv, since which of §8's two paths
// produced it doesn't matter past this point. // produced it doesn't matter past this point. spec.pod.extraEnv is appended
// last, after every fixed var -- this is the one place that owns "what env
// this container gets," so the escape hatch lives here rather than being
// appended separately in reconcileDeployment.
func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.EnvVar { func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.EnvVar {
env := []corev1.EnvVar{{Name: "TERDUT_ADDR", Value: fmt.Sprintf(":%d", servicePort(srv))}} env := []corev1.EnvVar{{Name: "TERDUT_ADDR", Value: fmt.Sprintf(":%d", servicePort(srv))}}
env = append(env, dbEnv...) env = append(env, dbEnv...)
@@ -199,7 +252,7 @@ func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.
} }
} }
return env return append(env, srv.Spec.Pod.ExtraEnv...)
} }
func secretEnvSource(ref *terdutv1alpha1.SecretKeyRef) *corev1.EnvVarSource { func secretEnvSource(ref *terdutv1alpha1.SecretKeyRef) *corev1.EnvVarSource {
+38
View File
@@ -0,0 +1,38 @@
package controller
import (
"context"
policyv1 "k8s.io/api/policy/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
)
// reconcilePodDisruptionBudget creates/updates the PodDisruptionBudget
// spec.pod.disruptionBudget asks for, or deletes a previously-created one
// when the field has been cleared -- the one conditionally-created child
// object in this controller (Deployment/Service are unconditional). Owned
// by srv, same plain-OwnerReference shape as Deployment/Service (DESIGN.md
// §7): same namespace, GC handles it, no finalizer needed.
func (r *TerdutServerReconciler) reconcilePodDisruptionBudget(ctx context.Context, srv *terdutv1alpha1.TerdutServer) error {
pdb := &policyv1.PodDisruptionBudget{ObjectMeta: metav1.ObjectMeta{Name: srv.Name, Namespace: srv.Namespace}}
spec := srv.Spec.Pod.DisruptionBudget
if spec == nil {
if err := r.Delete(ctx, pdb); err != nil && !apierrors.IsNotFound(err) {
return err
}
return nil
}
_, err := controllerutil.CreateOrUpdate(ctx, r.Client, pdb, func() error {
pdb.Spec.Selector = &metav1.LabelSelector{MatchLabels: labelsFor(srv)}
pdb.Spec.MinAvailable = spec.MinAvailable
pdb.Spec.MaxUnavailable = spec.MaxUnavailable
return controllerutil.SetControllerReference(srv, pdb, r.Scheme)
})
return err
}
@@ -131,6 +131,9 @@ func (r *TerdutTeamReconciler) Reconcile(ctx context.Context, req ctrl.Request)
if err := teamClient.SetTeamOIDCGroups(ctx, team.Status.TeamID, team.Spec.OIDC.MemberGroup, team.Spec.OIDC.OwnerGroup); err != nil { if err := teamClient.SetTeamOIDCGroups(ctx, team.Status.TeamID, team.Spec.OIDC.MemberGroup, team.Spec.OIDC.OwnerGroup); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/oidc-groups: %w", team.Status.TeamID, err) return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/oidc-groups: %w", team.Status.TeamID, err)
} }
if err := r.reconcileInvite(ctx, &team, teamClient); err != nil {
return ctrl.Result{}, err
}
meta.SetStatusCondition(&team.Status.Conditions, metav1.Condition{ meta.SetStatusCondition(&team.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady, Type: terdutv1alpha1.ConditionReady,
@@ -4,6 +4,7 @@ import (
"context" "context"
"fmt" "fmt"
"net/http/httptest" "net/http/httptest"
"time"
. "github.com/onsi/ginkgo/v2" . "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega" . "github.com/onsi/gomega"
@@ -251,6 +252,88 @@ var _ = Describe("TerdutTeam Controller", func() {
}) })
}) })
Describe("spec.invite", func() {
It("mints a link into the TerdutTeam's own namespace, not the operator's", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // create+mint+apply
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
team.Spec.Invite.Enabled = true
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.InviteSecretRef).NotTo(BeNil())
var secret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: team.Status.InviteSecretRef.Name, Namespace: team.Namespace,
}, &secret)).To(Succeed())
Expect(string(secret.Data[inviteSecretURLKey])).To(ContainSubstring("invite="))
Expect(string(secret.Data[inviteSecretInviteIDKey])).To(Equal("1"))
inv := fake.invites[team.Status.TeamID][1]
Expect(inv.Role).To(Equal("member"), "default role")
Expect(inv.MaxUses).To(Equal(int64(1)), "default max uses")
})
It("refreshes a link that's within a day of terdut-server's 7-day TTL", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx)
reconcileOnce(ctx)
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
team.Spec.Invite.Enabled = true
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
firstSecretName := team.Status.InviteSecretRef.Name
var secret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: firstSecretName, Namespace: team.Namespace}, &secret)).To(Succeed())
// Simulate the stored link being within the refresh window of
// expiry, the way it genuinely would be six days from now,
// without the test waiting six days.
secret.Data[inviteSecretExpiresAtKey] = []byte(time.Now().Add(12 * time.Hour).Format(time.RFC3339)) // inside inviteRefreshWindow
Expect(k8sClient.Update(ctx, &secret)).To(Succeed())
reconcileOnce(ctx)
Expect(fake.inviteDelete[1]).To(BeTrue(), "the stale invite should have been revoked")
Expect(fake.invites[team.Status.TeamID]).To(HaveKey(int64(2)), "a replacement should have been minted")
})
It("revokes the invite when spec.invite.enabled flips back to false", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx)
reconcileOnce(ctx)
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
team.Spec.Invite.Enabled = true
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
secretName := team.Status.InviteSecretRef.Name
team.Spec.Invite.Enabled = false
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(fake.inviteDelete[1]).To(BeTrue())
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.InviteSecretRef).To(BeNil())
var secret corev1.Secret
err := k8sClient.Get(ctx, types.NamespacedName{Name: secretName, Namespace: team.Namespace}, &secret)
Expect(err).To(HaveOccurred(), "the invite Secret should have been deleted")
})
})
Describe("deletion", func() { Describe("deletion", func() {
It("deletes the team server-side and removes the credentials Secret", func(ctx SpecContext) { It("deletes the team server-side and removes the credentials Secret", func(ctx SpecContext) {
createTeam(ctx, "to-delete", sameNSRef()) createTeam(ctx, "to-delete", sameNSRef())
+149
View File
@@ -0,0 +1,149 @@
package controller
import (
"context"
"fmt"
"strconv"
"time"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// Data keys inside the generated invite Secret, following the same naming
// shape as TerdutAlertSource's webhookSecret*Key constants.
const (
inviteSecretURLKey = "url"
inviteSecretInviteIDKey = "inviteID"
inviteSecretExpiresAtKey = "expiresAt"
)
// inviteRefreshWindow is how far ahead of expiry this controller mints a
// replacement link, so a human reading status.inviteSecretRef never finds a
// dead link mid-use. terdut-server's invite TTL is a fixed, unconfigurable
// 7 days (internal/api/signup.go's inviteTTL) -- refreshing a full day
// ahead of that leaves comfortable margin against this controller's own
// 5-minute resync interval ever being delayed.
const inviteRefreshWindow = 24 * time.Hour
func inviteSecretName(team *terdutv1alpha1.TerdutTeam) string {
return team.Name + "-terdut-invite"
}
// reconcileInvite applies spec.invite against teamClient -- this team's own
// team-scoped credential, already owner-equivalent for every /invites route
// (terdut-server's SERVICE-ACCOUNTS.md, ratified not accidental). Mints,
// refreshes ahead of expiry, or revokes, entirely independent of this
// team's own Ready condition: an invite is a convenience for onboarding a
// human, never something anything else in this reconcile waits on.
func (r *TerdutTeamReconciler) reconcileInvite(ctx context.Context, team *terdutv1alpha1.TerdutTeam, teamClient *tdclient.Client) error {
if !team.Spec.Invite.Enabled {
return r.revokeInvite(ctx, team, teamClient)
}
secretName := inviteSecretName(team)
var secret corev1.Secret
err := r.Get(ctx, client.ObjectKey{Namespace: team.Namespace, Name: secretName}, &secret)
switch {
case err == nil:
expiresAt, parseErr := time.Parse(time.RFC3339, string(secret.Data[inviteSecretExpiresAtKey]))
if parseErr == nil && time.Until(expiresAt) > inviteRefreshWindow {
team.Status.InviteSecretRef = &terdutv1alpha1.LocalSecretRef{Name: secretName}
return nil // still fresh, nothing to do this reconcile
}
// Expired, about to expire, or unreadable: mint a replacement.
// Revoke the old row by id first (best-effort) so a leaked old link
// stops working immediately rather than lingering unrevoked until
// its own TTL -- failure here is not fatal, since the replacement
// below is what actually matters.
if oldID, idErr := strconv.ParseInt(string(secret.Data[inviteSecretInviteIDKey]), 10, 64); idErr == nil {
_ = teamClient.RevokeInvite(ctx, team.Status.TeamID, oldID)
}
return r.mintInvite(ctx, team, teamClient, secretName)
case apierrors.IsNotFound(err):
// Low stakes, unlike TerdutAlertSource's webhook URL: nothing
// external holds a durable dependency on one specific invite link
// staying stable the way an Alertmanager config depends on a
// webhook URL -- it's read once by one human and handed out. So
// this silently re-mints rather than failing closed the way
// TerdutAlertSource's ReasonWebhookSecretLost does for its Secret.
return r.mintInvite(ctx, team, teamClient, secretName)
default:
return err
}
}
func (r *TerdutTeamReconciler) mintInvite(
ctx context.Context, team *terdutv1alpha1.TerdutTeam, teamClient *tdclient.Client, secretName string,
) error {
role := team.Spec.Invite.Role
if role == "" {
role = "member"
}
maxUses := team.Spec.Invite.MaxUses
if maxUses == 0 {
maxUses = 1
}
inv, err := teamClient.CreateInvite(ctx, team.Status.TeamID, role, maxUses)
if err != nil {
return fmt.Errorf("POST /api/teams/%d/invites: %w", team.Status.TeamID, err)
}
secret := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: secretName, Namespace: team.Namespace}}
if _, err := controllerutil.CreateOrUpdate(ctx, r.Client, secret, func() error {
secret.Data = map[string][]byte{
inviteSecretURLKey: []byte(inv.URL),
inviteSecretInviteIDKey: []byte(strconv.FormatInt(inv.ID, 10)),
inviteSecretExpiresAtKey: []byte(inv.ExpiresAt.Format(time.RFC3339)),
}
return controllerutil.SetControllerReference(team, secret, r.Scheme)
}); err != nil {
return err
}
team.Status.InviteSecretRef = &terdutv1alpha1.LocalSecretRef{Name: secretName}
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonInviteMinted, terdutv1alpha1.ReasonInviteMinted,
"invite link minted into Secret %q", secretName)
}
return nil
}
// revokeInvite tears down spec.invite's Secret and server-side row when
// spec.invite.enabled is false (or was never set). Same-namespace and
// OwnerReference'd, so deleting the TerdutTeam itself already garbage-
// collects this Secret -- this path exists for the narrower case of
// flipping enabled back to false on an otherwise-live TerdutTeam.
func (r *TerdutTeamReconciler) revokeInvite(ctx context.Context, team *terdutv1alpha1.TerdutTeam, teamClient *tdclient.Client) error {
if team.Status.InviteSecretRef == nil {
return nil
}
name := team.Status.InviteSecretRef.Name
var secret corev1.Secret
err := r.Get(ctx, client.ObjectKey{Namespace: team.Namespace, Name: name}, &secret)
switch {
case err == nil:
if id, idErr := strconv.ParseInt(string(secret.Data[inviteSecretInviteIDKey]), 10, 64); idErr == nil {
if err := teamClient.RevokeInvite(ctx, team.Status.TeamID, id); err != nil {
return fmt.Errorf("DELETE /api/teams/%d/invites/%d: %w", team.Status.TeamID, id, err)
}
}
if err := r.Delete(ctx, &secret); err != nil && !apierrors.IsNotFound(err) {
return err
}
case !apierrors.IsNotFound(err):
return err
}
team.Status.InviteSecretRef = nil
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonInviteRevoked, terdutv1alpha1.ReasonInviteRevoked,
"invite link revoked")
}
return nil
}
+47
View File
@@ -566,3 +566,50 @@ func (c *Client) DeleteIntegration(ctx context.Context, teamID, integrationID in
} }
return c.do(req, nil) return c.do(req, nil)
} }
// Invite is a standing link into a team (POST /api/teams/{teamID}/invites'
// own response shape). URL carries the raw token exactly once, at creation
// -- terdut-server never shows it again (same one-time-shown shape as an
// integration's webhook key) -- so a caller that needs it later has to have
// kept this response, not re-fetched it.
type Invite struct {
ID int64 `json:"id"`
TeamID int64 `json:"team_id"`
Role string `json:"role"`
ExpiresAt time.Time `json:"expires_at"`
MaxUses int64 `json:"max_uses"`
URL string `json:"url,omitempty"`
}
// CreateInvite calls POST /api/teams/{teamID}/invites -- owner-gated
// (requireTeamOwner), so c must hold this team's own team-scoped
// credential, which already satisfies that check via its synthetic owner
// membership (terdut-server's SERVICE-ACCOUNTS.md). No conflict handling
// needed: unlike a team or a service account, an invite has no unique name
// to collide on -- every call mints a brand new row.
func (c *Client) CreateInvite(ctx context.Context, teamID int64, role string, maxUses int64) (*Invite, error) {
req, err := c.newRequest(ctx, http.MethodPost, fmt.Sprintf("/api/teams/%d/invites", teamID),
map[string]any{"role": role, "max_uses": maxUses})
if err != nil {
return nil, err
}
var inv Invite
if err := c.do(req, &inv); err != nil {
return nil, err
}
return &inv, nil
}
// RevokeInvite calls DELETE /api/teams/{teamID}/invites/{inviteID} -- same
// credential requirement as CreateInvite. A 404 (already revoked, or never
// existed) is the caller's to treat as success if it wants to, the same way
// DeleteTeam's own 404 handling works -- this method itself just reports
// whatever terdut-server said.
func (c *Client) RevokeInvite(ctx context.Context, teamID, inviteID int64) error {
req, err := c.newRequest(ctx, http.MethodDelete,
fmt.Sprintf("/api/teams/%d/invites/%d", teamID, inviteID), nil)
if err != nil {
return err
}
return c.do(req, nil)
}