17 Commits

Author SHA1 Message Date
Niklas Ye b0a431f2a4 Set the chart's placeholder version to 0.5.0
CI / chart (push) Successful in 2s
CI / security (push) Successful in 53s
Release / chart (push) Successful in 3s
CI / test (push) Successful in 2m19s
Release / test (push) Successful in 1m41s
Release / image (push) Successful in 6m30s
Release / scan-image (push) Successful in 4s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with d502485 (0.4.0) and 88172ad (0.3.0) before it, because a
tree heading for v0.5.0 that still says 0.4.0 tells its reader
something false.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 21:10:27 +02:00
Niklas Ye 62664c93ff Keep the instance credential across a TerdutServer delete, and adopt it on recreate
Deleting a TerdutServer removed the credential Secrets but never touched the
database, so a recreated one found a server that was already bootstrapped and
no key for it: /api/bootstrap answered 403 and the operator stopped at
BootstrapStateLost, whose message and DESIGN.md both said "delete and
recreate". That is how the terdut-demo install on the cluster got stuck on
2026-10-03: Helm's cleanupOnFail deleted its TerdutServer after a failed
upgrade, the recreate found the bootstrapped database, and it sat at Ready:
False for five days until the database was reset by hand. Recreating cannot
fix it, because the finalizer clears Secrets and the database is not its to
reset, so "a fresh create starts clean" was only ever true when the database
went with it.

spec.credentials.deletionPolicy is Retain by default: the finalizer keeps the
instance credential Secret (Delete removes it, as before). The bootstrap
checkpoint is always removed. Before calling /api/bootstrap, reconcile now
looks for the retained Secret and asks the server for the operator's own
service account with its token. Accepted: adopt it and skip bootstrap.
Rejected with 401/403: the Secret outlived a database reset, so ignore it and
bootstrap like a first install, which replaces it. Any other error retries.
terdut-server's own tests already call that endpoint with an instance-scoped
key, so the permission is not new.

BootstrapStateLost is still the answer when the server is bootstrapped and no
credential it accepts survives, but its message now names the Secret to
restore and says that recreating does not clear the database. DESIGN.md §6
says the same, and the chart passes the setting through as
terdutServer.credentials.deletionPolicy.

A retained Secret of a TerdutServer that is gone for good is an orphan to
delete by hand. It is inert: nothing adopts it unless the server accepts the
token.

Checked on the kind demo with a locally built image against the real
terdut-server v0.43.0: deleting the TerdutServer kept the Secret, recreating it
reached Ready with the same credential (identical hash) and both TerdutTeams
came back Ready with their original ids. The controller specs cover adoption,
a rejected token after a reset, the bootstrapped-and-rejected failure, and
both deletion policies.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 21:09:56 +02:00
niklas 6572f63157 Merge pull request 'examples/demo: run terdut-server v0.43.0, with alerts from two clusters' (#6) from demo-matches-v0.4.0-crd into main
CI / chart (push) Successful in 2s
CI / security (push) Successful in 51s
CI / test (push) Successful in 2m26s
Reviewed-on: #6
2026-10-08 17:10:06 +00:00
Niklas Ye 50ce5bcec0 examples/demo: run terdut-server v0.43.0, with alerts from two clusters
CI / test (pull_request) Successful in 6m35s
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m6s
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the
web UI since: the queue and incident layouts, the rota and escalation
pages, the theme toggle, and the cluster chip, filter and page titles
(v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the
floor, which is what the replicas setting actually depends on.

fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus
external label plus `cluster` in Alertmanager's group_by (terdut-server's
README, "Several clusters, one team"). It goes on the alert's labels and
groupLabels, and into the group key and the fingerprint, so the same alert in
two clusters is two incidents and not one. Unset, the payload is exactly what
it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in
both, so the queue has a chip and a filter to show.

run-demo.sh also failed on its second run, though it says it is safe to
re-run: it expected HTTP 409 when alice already exists, but a spent invite
is answered with 403 "invite link is not usable" before the username is ever
checked. It now tries to log alice in first and skips the signup if that works.

Checked on the kind cluster: the server rolled to v0.43.0, every CR became
Ready and Adopted (server, both teams, both escalation rules, both dead man's
switches, both alert sources), and /api/incidents/clusters,
/api/incidents?cluster=prod-us and the incident titles came back as expected.
No operator code changed, so this needs no operator release.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-08 18:55:02 +02:00
Niklas Ye 822c80dda6 examples/demo: split the Ready wait so escalation rules wait on alice too
wait_for_ready waited for every demo object at once, including
terdutescalationrule-platform, which names alice as a level-1 target --
but alice does not exist yet at that point in main(): she is created by
redeem_platform_invite, which ran after wait_for_ready. terdut-server
resolves every named username at reconcile time, not just when an
escalation actually fires, so that CR could never reach Ready before
alice did, and main() had no step in between to create her.

Split into wait_for_objects (the shared loop, now taking its object list
as arguments) plus two callers: wait_for_teams_ready, covering just the
server and the two teams redeem_platform_invite/join_payments_team
need, run before alice exists; wait_for_remaining_ready, covering the
escalation rules, dead man's switches and alert sources, run after.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 18:15:42 +02:00
Niklas Ye 0ee7ede648 examples/demo: match v0.4.0's new spec.replicas default and bump terdut-server
replicas: 1 and tag: v0.34.0 were both correct when written, but the CRD's
own default moved to 2 in v0.4.0 (same release this demo is meant to show
off), and v0.34.0 predates v0.36.0's advisory locks that make a second
replica safe instead of racing the first. Left as-is, the demo would have
been the one place in this repo demonstrating the exact unsafe combination
the CRD's own doc comment warns against: more than one replica against an
image that doesn't guard the sweeper/notifier/migration-runner singletons.

replicas is now stated explicitly as 2 rather than dropped to pick up the
default silently, matching every other field in this file's own habit of
spelling out what it depends on. tag moves to v0.36.0 specifically -- the
first version where the lock landed -- with the comment keeping v0.34.0's
original reasoning (the service-account race fix) alongside the new one,
since v0.36.0 still carries that fix forward.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 18:13:09 +02:00
Niklas Ye d50248531c Set the chart's placeholder version to 0.4.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 2m10s
Release / test (push) Successful in 7m54s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m25s
Release / scan-image (push) Successful in 36s
2026-10-03 16:22:15 +02:00
Niklas Ye 4007f54279 Default TerdutServer.spec.replicas to 2 and switch to RollingUpdate
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Has been cancelled
Mirrors charts/terdut-server's own deployment.yaml change: v0.36.0 put
the sweeper, the notifier and the migration runner each behind a
Postgres advisory lock, and gave incident creation its own conflict
resolution, so the Recreate strategy and replicas-stays-at-1 guidance
this controller carried (explicitly tracking that chart's comment)
are no longer load-bearing.

spec.replicas' +kubebuilder:default moves 1 -> 2 (config/crd/bases and
the chart's CRD template regenerated via controller-gen and
kubebuilder's helm plugin respectively, then hand-verified identical
to the generator's own output rather than trusting a bulk regen --
the plugin's --output-dir charts writes a fresh charts/chart scaffold
rather than updating charts/terdut-operator in place, so only the
diff was taken, not the whole tree). terdutserver_deployment.go's
same-value fallback (reachable only for a TerdutServer stored before
this default existed) moves with it, and its Strategy changes from
Recreate to RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicas: 2, already
zero-downtime.

DESIGN.md's three places asserting multi-replica isn't a supported
topology (the illustrative spec.replicas YAML, spec.pod.affinity's
rationale, and the HPA deferred-feature note) are corrected to match;
the HPA note now gives its own standing reason (no scaling metric or
bounds decided yet) rather than a contradiction that no longer holds.

The chart's optional terdutServer.replicas sample value moves 1 -> 2
alongside it. image.tag must be v0.36.0 or newer for any of this to
hold -- stated in both the CRD field's doc comment and the chart
value's comment, not enforced in code, same stance the chart takes on
every other version-coupled assumption.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:35:18 +02:00
Niklas Ye 88172ade29 Set the chart's placeholder version to 0.3.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
CI / test (push) Successful in 1m53s
Release / test (push) Successful in 1m45s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 5m27s
Release / scan-image (push) Successful in 2s
2026-10-02 22:19:18 +02:00
niklas 478ae6284a Merge pull request 'TerdutTeam: mint and surface a real invite link (spec.invite)' (#5) from terdutteam-invite-minting into main
CI / test (push) Has been cancelled
CI / chart (push) Successful in 1s
CI / security (push) Successful in 59s
2026-10-02 20:17:21 +00:00
niklas fcc80b32d5 Merge pull request 'examples/demo: add run-demo.sh, an automated kind-cluster demo' (#4) from examples-demo/run-demo-script into main
CI / chart (push) Successful in 2s
CI / security (push) Successful in 1m31s
CI / test (push) Successful in 3m30s
2026-10-02 20:13:36 +00:00
Niklas Ye 4aa4f17c42 examples/demo: bump terdut-server to v0.34.0 (the service-account fix)
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Successful in 1m18s
CI / test (pull_request) Successful in 3m40s
Required for this demo to actually exercise the fix for
niklas/terdut-operator#3 -- v0.33.2 still has the authorization gap this
demo hit live (callerMayManageServiceAccount had no branch letting an
instance-scoped account adopt a team-scoped account's key).
2026-10-02 22:13:26 +02:00
Niklas Ye a0ea13955e TerdutTeam: mint and surface a real invite link (spec.invite)
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 52s
CI / test (pull_request) Successful in 2m43s
The actual fix for the human-onboarding gap niklas/terdut-server#23 found --
not a terdut-server change at all. A team-scoped credential is already
owner-equivalent for POST/GET/DELETE /api/teams/{teamID}/invites
(requireTeamOwner's synthetic-membership mechanism, ratified not
accidental per that repo's SERVICE-ACCOUNTS.md), and invite redemption
bypasses signup_mode entirely -- this TerdutTeam controller just never
grew a feature to use either fact.

New spec.invite{enabled, role (member|owner, default member), maxUses
(1-100, default 1)} and status.inviteSecretRef. The Secret lives in the
TerdutTeam's OWN namespace, not the operator's: unlike
status.credentialsSecretRef (a durable, high-privilege credential, kept
operator-side per DESIGN.md §6), an invite is bounded and limited-use,
meant for this namespace's own human operators to read and hand out --
same precedent as TerdutAlertSource's status.webhookURLSecretRef, same-
namespace and OwnerReference'd so deleting the TerdutTeam garbage-collects
it automatically.

internal/controller/terdutteam_invite.go: mints on first spec.invite.enabled,
refreshes a day ahead of terdut-server's fixed 7-day TTL (reading the
Secret's own stored expiresAt, no extra server round-trip per reconcile),
revokes server-side and deletes the Secret when flipped back to false. A
lost invite Secret is silently re-minted rather than treated as
unrecoverable the way TerdutAlertSource's webhook key is -- nothing
external holds a durable dependency on one specific invite link staying
stable, it's read once by one human and handed out.

New tdclient.Invite/CreateInvite/RevokeInvite. New envtest coverage: mint
into the team's own namespace, refresh-before-expiry, revoke-on-disable
(internal/controller/terdutteam_controller_test.go's new "spec.invite"
Describe block), plus the fake server growing invite support
(terdutserver_controller_test.go) -- its handleTeamSubPath dispatcher was
split further (deadman switches into their own handleDeadmanSubPath,
matching the existing handleIntegrationSubPath precedent) to stay under
golangci-lint's gocyclo threshold with the new route added.

examples/demo updated to prove this end to end: 02-team-platform.yaml
turns on spec.invite; run-demo.sh's bootstrap_login/join_demo_teams (the
psql signup_mode flip + a direct team_members INSERT) are replaced by
redeem_platform_invite (reads status.inviteSecretRef, a real POST
/api/signup with the invite token) and join_payments_team (POST
/api/teams/{teamID}/members using Payments' own credential and alice's
user id resolved via GET /api/users, deliberately not given its own
spec.invite, so the demo shows both onboarding paths this feature
unlocks) -- zero kubectl exec/psql calls remain anywhere in the script.
README.md's "First login" section rewritten to match; it no longer
documents the admin-token curl call that 403s against current
terdut-server (niklas/terdut-server#23).

Depends on niklas/terdut-server#24 (the callerMayManageServiceAccount fix
for terdut-operator#3) being released before this is deployed for real --
not required to build or test this change itself, since the envtest fake
never modeled that authorization gap to begin with.
2026-10-02 22:04:56 +02:00
Niklas Ye 2a08a8cd8e examples/demo: add run-demo.sh, an automated kind-cluster demo
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 1m5s
CI / test (pull_request) Successful in 2m49s
One script, two modes (run-demo.sh / run-demo.sh --teardown), that takes a
fresh empty kind cluster all the way to a working demo: creates the
cluster if needed, helm-installs this chart, applies every CRD kind in
this directory, waits for all nine objects to go Ready, then does what
the README's own first-login section cannot (see niklas/terdut-server#23
and niklas/terdut-operator#3 -- no service-account credential this
operator holds can ever call /api/admin/settings or POST /api/users) by
reaching into the demo's own throwaway Postgres directly: flips
signup_mode to open, signs alice up for real over the ordinary signup
endpoint, and joins her to both Platform and Payments (open signup always
creates its own new team, never joins an existing one by name, so
without this she'd have a working login that can't see a single incident
this demo fires -- /api/incidents and /api/alerts are both scoped to the
caller's own team memberships). Finishes by port-forwarding the service
and firing fire-alerts.sh at both teams, so a fresh run already has
visible incidents waiting in the web UI.

Verified end to end against a real kind cluster, including a second,
genuinely-fresh run that hit niklas/terdut-operator#3 live (terdutteam-
platform wedged in the 403 retry loop that issue describes) -- confirmed
the script itself fails cleanly on that (clear FAILED message, correct
exit code, no orphaned port-forward) rather than hanging or leaving a
mess, which is the most this script can do about a bug in the operator
it's driving.
2026-10-02 21:23:44 +02:00
Niklas Ye 375b5ed2f7 Set the chart's placeholder version to 0.2.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 57s
CI / test (push) Successful in 2m13s
Release / test (push) Successful in 1m41s
Release / chart (push) Successful in 2s
Release / image (push) Successful in 6m10s
Release / scan-image (push) Successful in 3s
make helm-package passes --version and --app-version from the tag, so
these fields decide nothing about what is published -- but a tree
heading for v0.2.0 that still says 0.1.2 tells its reader something
false. Same as 738c210 and 0e3118d before it.
2026-10-02 18:51:00 +02:00
Niklas Ye 6a699d4341 Let TerdutServer customize its pod, and never manage its own ingress
spec.pod (api/v1alpha1/terdutserver_types.go): annotations, nodeSelector,
tolerations, affinity, topologySpreadConstraints, resources, pod and
container securityContext, serviceAccountName, extraEnv/extraEnvFrom,
extraVolumes/extraVolumeMounts, imagePullSecrets, and an optional
disruptionBudget. All direct corev1 passthrough -- no wrapper types buy
anything for any of these, matching how CloudNativePG and the Zalando
postgres-operator both expose the same knobs, and matching this repo's
own SweeperSpec precedent ("wrap only when a round-trip through a
different type buys something"). affinity is pure user-supplied
passthrough, not a toggle-plus-generated-default the way a multi-replica
cluster operator's pod anti-affinity usually is: this operator never
auto-generates one, since spec.replicas above 1 isn't a supported
topology (the sweeper/notifier singleton constraint). Considered and
declined for this round: priorityClassName, pod labels beyond
annotations, and a HorizontalPodAutoscaler -- the last of those would
directly contradict the singleton constraint above.

disruptionBudget is the one field here that isn't a plain PodTemplateSpec
knob: when set, the controller now reconciles a PodDisruptionBudget
selecting the TerdutServer's own pods (new terdutserver_pdb.go); clearing
it deletes any it previously created. New RBAC marker on
poddisruptionbudgets to match.

Driven by a public-release pass: looking past this project's own use case
at what a mature, general-purpose operator CRD exposes here (researched
against Zalando postgres-operator and CloudNativePG specifically), not
just the fields this install happened to need.

Separately, and found while answering a question about exposing
TerdutServer through Istio instead of Gateway API: spec.networking's own
doc comment quietly promised a Gateway API HTTPRoute this operator would
build eventually ("a near-term follow-up, not deferred"). That promise is
wrong for a public release -- an operator managing someone's ingress
mechanism for them is a worse default than not touching it at all, and a
surprise HTTPRoute appearing once that follow-up eventually landed would
have been exactly backwards for an Istio (or plain-Ingress, or
intentionally-unexposed) install. Made the non-goal explicit and
permanent instead (DESIGN.md §1), removed the dead `gatewayListener`
field it was the only consumer of (zero runtime call sites anywhere --
setting it already had no effect, so this is a schema cleanup, not a
behavior change), and corrected ROADMAP.md's framing. hostname/servicePort
stay: both are live (TERDUT_PUBLIC_URL, container/Service port), this
operator just never acts on hostname for exposure. Added
examples/networking (Gateway API HTTPRoute, Istio VirtualService) showing
how to expose the plain ClusterIP Service the operator already creates --
outside the operator itself, as illustrations, not as something
examples/demo applies automatically.

No new terdut-server version requirement: both changes are CRD/controller-
only, nothing about the API this operator's bootstrap flow depends on
changed.
2026-10-02 18:50:38 +02:00
Niklas Ye d9315322fc examples/demo: fix two real bugs this exact demo just hit live
CI / chart (push) Successful in 1s
CI / security (push) Successful in 3m24s
CI / test (push) Successful in 10m45s
1. Renamed every object this demo creates (TerdutServer, Postgres
   Secret/Deployment/Service) from terdut-demo[-postgres] to
   terdut-operator-demo[-postgres]. The user applied this kit into the
   already-live "terdut-demo" namespace -- the real operator exercise
   from earlier in this repo's own history -- and this demo's own
   TerdutServer/Postgres objects shared that exact name. The TerdutServer
   apply was rejected outright (DatabaseSpec's own CEL rule: adding dsn
   while the live object already had postgresClusterRef violates "exactly
   one of" and the API server refused it), and the real Postgres Service
   was never touched (confirmed live: still Zalando's own spilo selector,
   endpoint still the real StatefulSet pod) -- but the Postgres Secret and
   Deployment, having no such protection, were created as brand new,
   extra, crash-looping objects sitting right next to the real ones.
   Prefixing every name this demo creates means a repeat of this exact
   mistake no longer collides with anything, documented directly in
   README.md now.

2. The actual crash itself, independent of (1): capabilities.drop: ["ALL"]
   (added responding to a PodSecurity "restricted" warning) took
   CAP_CHOWN/CAP_FOWNER away from the root user postgres:17-alpine's own
   entrypoint needs to chown/chmod the data directory before it drops
   privileges itself -- confirmed in a real crashed pod's logs: `chmod:
   /var/run/postgresql: Operation not permitted`. kubectl apply
   --dry-run=server, which is as far as this got verified before, only
   checks admission policy; it was never actually booted. Removed the
   capability drop and verified for real this time: applied just
   00-postgres.yaml alone into a disposable namespace, waited for the pod
   to go Ready, read its logs ("database system is ready to accept
   connections"), then deleted that namespace.
2026-10-02 13:25:47 +02:00
34 changed files with 10107 additions and 228 deletions
+101 -22
View File
@@ -50,6 +50,14 @@ it just created, never a server something else already bootstrapped first.
that needs a live look at another object (e.g. "does this teamRef exist") that needs a live look at another object (e.g. "does this teamRef exist")
is a status condition, not an admission rejection — keeps v1 to a is a status condition, not an admission rejection — keeps v1 to a
controller-only deployment with no cert-manager/webhook dependency. controller-only deployment with no cert-manager/webhook dependency.
- **Never manages external exposure/ingress for `TerdutServer`, in any
form — a permanent non-goal, not a staged one.** The operator creates a
plain `ClusterIP` Service (§4.1) and stops there: no `Ingress`, no
Gateway API `HTTPRoute`, no Istio `VirtualService`, nothing. Some
installs won't expose `TerdutServer` outside the cluster at all; others
will use whichever of those mechanisms already fits their cluster. That
choice belongs to whoever deploys it, not to this operator — see
`examples/networking` for worked (but not operator-managed) examples.
## 2. The README's open questions, resolved ## 2. The README's open questions, resolved
@@ -145,11 +153,10 @@ spec:
image: image:
repository: git.ryuvia.com/niklas/terdut-server repository: git.ryuvia.com/niklas/terdut-server
tag: v0.9.3 tag: v0.9.3
replicas: 1 # terdut-server is not horizontally-scale-tested; keep the field, default 1 replicas: 2 # default since terdut-server v0.36.0's advisory locks; see TerdutServerSpec.Replicas
networking: networking:
hostname: terdut.example.com hostname: terdut.example.com
servicePort: 8080 servicePort: 8080
gatewayListener: "" # same semantics as chart's networking.listener
database: database:
dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef dsn: "postgres://terdut@terdut-postgres:5432/terdut?sslmode=require" # mutually exclusive with postgresClusterRef
passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password} passwordSecretRef: {name: terdut.terdut-postgres.credentials.postgresql.acid.zalan.do, key: password}
@@ -162,6 +169,8 @@ spec:
matchers: "alertname=Watchdog" matchers: "alertname=Watchdog"
timeout: 15m timeout: 15m
severity: critical severity: critical
credentials:
deletionPolicy: Retain # Retain (default) | Delete -- what deleting this CR does to the instance credential Secret; see §6 point 2
notify: notify:
ntfyURL: "http://ntfy.ntfy.svc.cluster.local" ntfyURL: "http://ntfy.ntfy.svc.cluster.local"
fallbackTopic: "" fallbackTopic: ""
@@ -188,6 +197,20 @@ spec:
# selector: # required, and only meaningful, when from: Selector # selector: # required, and only meaningful, when from: Selector
# matchLabels: # matchLabels:
# terdut.ryuvia.com/allowed: "true" # terdut.ryuvia.com/allowed: "true"
# Pod-level customization of the Deployment -- all optional, direct corev1
# passthrough throughout (see PodSpec's own doc comment). A representative
# subset:
pod:
resources:
requests: {cpu: 100m, memory: 128Mi}
limits: {memory: 256Mi}
tolerations:
- key: dedicated
operator: Equal
value: terdut
effect: NoSchedule
disruptionBudget:
minAvailable: 1 # mutually exclusive with maxUnavailable
status: status:
conditions: [...] # Ready, DatabaseReady, Bootstrapped conditions: [...] # Ready, DatabaseReady, Bootstrapped
observedGeneration: 3 observedGeneration: 3
@@ -214,6 +237,25 @@ per-name allowlist (no "and only these teams") — namespace-level consent is
the right granularity here, same reasoning as `ListenerSet`: the namespace the right granularity here, same reasoning as `ListenerSet`: the namespace
is the tenancy boundary, not the object. is the tenancy boundary, not the object.
`spec.pod` is pod-level customization of the Deployment, all optional and
directly reusing corev1 types wherever corev1 already models the knob
exactly (`tolerations`, `affinity`, `topologySpreadConstraints`,
`resources`, `securityContext`/`containerSecurityContext`, `extraEnv`/
`extraEnvFrom`, `extraVolumes`/`extraVolumeMounts`, `imagePullSecrets`) —
no custom wrapper buys anything for any of these, matching how
CloudNativePG and the Zalando postgres-operator both expose the same
knobs. `affinity` is pure user-supplied passthrough, not a
toggle-plus-generated-default the way a multi-replica-aware operator's
pod anti-affinity typically is: even though `replicas` now defaults to 2
(terdut-server v0.36.0's advisory locks made that safe, §4.1's own
illustrative YAML comment), this operator still never auto-generates
affinity of its own. `spec.pod.disruptionBudget` is the one field here that isn't a
straight PodTemplateSpec knob — when set, the controller reconciles a
`PodDisruptionBudget` selecting this `TerdutServer`'s pods; clearing it
deletes any it previously created (§7). `minAvailable`/`maxUnavailable`
are mutually exclusive, `+kubebuilder:validation:XValidation`-guarded the
same way as `spec.database`'s own `dsn`/`postgresClusterRef` rule.
### 4.2 `TerdutTeam` ### 4.2 `TerdutTeam`
```yaml ```yaml
@@ -489,6 +531,18 @@ when nothing ever crosses into a tenant namespace in the first place.
the same rigor §5's general idempotent-create rule already applies the same rigor §5's general idempotent-create rule already applies
elsewhere: elsewhere:
- `status.credentialsSecretRef` already set: done, nothing to do. - `status.credentialsSecretRef` already set: done, nothing to do.
- Otherwise, look for the instance credential Secret itself
(`<namespace>.<name>-instance-credentials`, point 2 below): a
`TerdutServer` deleted and recreated under the same name leaves it
behind by default (`spec.credentials.deletionPolicy: Retain`), and the
server it logs in to has not changed, because deleting a
`TerdutServer` never touches its database. If it exists, ask the server
for the operator's own service account with that token
(`GET /api/service-accounts?name=terdut-operator`). Accepted: adopt it,
set `status.credentialsSecretRef`, and skip everything below. Rejected
with a `401`/`403`: it is stale (a database reset since), so ignore it
and carry on; the steps below replace it. Any other failure is a retry,
not a guess.
- Otherwise, check for an intermediate - Otherwise, check for an intermediate
`<namespace>.<name>-bootstrap-admin` Secret in `<namespace>.<name>-bootstrap-admin` Secret in
the operator's own namespace first. If it exists, its key is a still- the operator's own namespace first. If it exists, its key is a still-
@@ -498,15 +552,21 @@ when nothing ever crosses into a tenant namespace in the first place.
on `201`, immediately checkpoint its response's raw admin key on `201`, immediately checkpoint its response's raw admin key
(`{"user": ..., "api_key": {"key": "<raw>", ...}}`) into that Secret (`{"user": ..., "api_key": {"key": "<raw>", ...}}`) into that Secret
before doing anything else with it. A `403` with neither before doing anything else with it. A `403` with neither
`status.credentialsSecretRef` nor this checkpoint Secret present is `status.credentialsSecretRef`, nor an instance credential the server
the one genuinely pathological case left (the checkpoint deleted out accepts, nor this checkpoint Secret present is the one genuinely
from under a reconcile already past this point) — handled the same pathological case left: the server's database is already bootstrapped
way the design already handles unrecoverable server-issued material and no credential for it survives here (the checkpoint deleted out from
elsewhere (§5's webhook-Secret-loss rule): fail closed, under a reconcile already past this point, `deletionPolicy: Delete`, or
`Ready: False, reason: BootstrapStateLost`, with the same recovery as the Secret removed by hand). Fail closed, `Ready: False, reason:
that case, delete and recreate the `TerdutServer` (its finalizer tears BootstrapStateLost`, with a message that names the instance credential
down the Deployment/database-backing and server-side rows; a fresh Secret to restore. Deleting and recreating the `TerdutServer` does **not**
create starts clean) — not a workaround peculiar to this one path. recover from it: the finalizer removes Secrets only, and the database
(the one thing that still says "already bootstrapped") is not the
operator's to reset. Restore the Secret from a copy, or reset the
server's database and then recreate the `TerdutServer`, which then
bootstraps like a first install. (This section used to say a fresh
create "starts clean"; that was only ever true when the database went
with it.)
- With an admin key in hand (fresh or checkpointed): `POST - With an admin key in hand (fresh or checkpointed): `POST
/api/service-accounts {name: "terdut-operator", scope: "instance"}`. /api/service-accounts {name: "terdut-operator", scope: "instance"}`.
A `409` here means a prior attempt got this far before being A `409` here means a prior attempt got this far before being
@@ -523,10 +583,15 @@ when nothing ever crosses into a tenant namespace in the first place.
under a fixed data key, `token`), referenced back from under a fixed data key, `token`), referenced back from
`TerdutServer.status.credentialsSecretRef: {name, key}` (§4.1). No `TerdutServer.status.credentialsSecretRef: {name, key}` (§4.1). No
`OwnerReference` (those can't cross namespaces, and this Secret doesn't `OwnerReference` (those can't cross namespaces, and this Secret doesn't
share a namespace with the `TerdutServer` that caused it); the `TerdutServer`'s share a namespace with the `TerdutServer` that caused it). What the
finalizer deletes this Secret directly as part of its own teardown, `TerdutServer`'s finalizer does with it is `spec.credentials.deletionPolicy`:
the same way it already has to clean up the server-side resources it `Retain` (the default) leaves it for a recreated `TerdutServer` to adopt
created (§5's general finalizer rule extends naturally to this Secret). (point 1), `Delete` removes it directly as part of the teardown. There are
no server-side resources to undo either way: the bootstrap user and
service account have no delete verb in terdut-server's API. The bootstrap
checkpoint Secret is always removed. A retained Secret of a `TerdutServer`
that is gone for good is an orphan to delete by hand, and it is inert:
adoption asks the server to accept the token first.
3. When a `TerdutTeam` first becomes `Ready` (its `serverRef` resolved, 3. When a `TerdutTeam` first becomes `Ready` (its `serverRef` resolved,
`allowedTeams` satisfied if cross-namespace), its controller uses the `allowedTeams` satisfied if cross-namespace), its controller uses the
`TerdutServer`'s instance-scoped credential (read from the operator's own `TerdutServer`'s instance-scoped credential (read from the operator's own
@@ -591,9 +656,15 @@ when nothing ever crosses into a tenant namespace in the first place.
## 7. Ownership, status, garbage collection ## 7. Ownership, status, garbage collection
- Every generated object that lives in the *same* namespace as the CR that - Every generated object that lives in the *same* namespace as the CR that
caused it (Deployment, Service, webhook Secret) carries a standard caused it (Deployment, Service, webhook Secret, and `TerdutServer`'s own
`metav1.OwnerReference` — GC handles these, no finalizer needed. The two PodDisruptionBudget) carries a standard `metav1.OwnerReference` — GC
credential Secrets from §6 are the one exception: they live in the handles these, no finalizer needed. PodDisruptionBudget is the one
member of that list that's conditionally created/deleted rather than
always present: it exists only while `spec.pod.disruptionBudget` is set,
and the controller deletes it itself the moment that field is cleared
(it doesn't wait on GC for that case, only for the `TerdutServer` being
deleted outright). The two credential Secrets from §6 are the one
exception to OwnerReference-based cleanup generally: they live in the
operator's own namespace regardless of where their owning CR lives, so operator's own namespace regardless of where their owning CR lives, so
`OwnerReference` doesn't apply (cross-namespace) and cleanup instead runs `OwnerReference` doesn't apply (cross-namespace) and cleanup instead runs
through that CR's finalizer directly, alongside the server-side DELETE through that CR's finalizer directly, alongside the server-side DELETE
@@ -640,10 +711,10 @@ documented and tested operationally:
- The operator's own ServiceAccount needs, per namespace it's granted: - The operator's own ServiceAccount needs, per namespace it's granted:
`get/list/watch/create/update/patch/delete` on `Deployments`, `Services` `get/list/watch/create/update/patch/delete` on `Deployments`, `Services`
it owns, and `get/list/watch` on `postgresql.acid.zalan.do` (optional, and `PodDisruptionBudgets` it owns, and `get/list/watch` on
degrade gracefully if absent per §8), plus cluster-wide `get/list` on `postgresql.acid.zalan.do` (optional, degrade gracefully if absent per
`Namespace` (labels only, for `allowedTeams: {from: Selector}` evaluation §8), plus cluster-wide `get/list` on `Namespace` (labels only, for
— §4.1, §4.6). `allowedTeams: {from: Selector}` evaluation — §4.1, §4.6).
- **Two different `Secret` scopes, not one — corrected from an earlier draft - **Two different `Secret` scopes, not one — corrected from an earlier draft
of this section.** That earlier draft said `Secret` access was "scoped to of this section.** That earlier draft said `Secret` access was "scoped to
the operator's own namespace only... nowhere else," reasoning that with the operator's own namespace only... nowhere else," reasoning that with
@@ -784,6 +855,14 @@ what it was, a separate install, until someone deletes it.
a one-off exercise. a one-off exercise.
- Gitops-managed team *membership* (see §4.2). - Gitops-managed team *membership* (see §4.2).
- Automatic Deployment restart on upstream Postgres credential rotation. - Automatic Deployment restart on upstream Postgres credential rotation.
- `spec.pod.priorityClassName`, pod-label passthrough beyond
`spec.pod.annotations`, and a HorizontalPodAutoscaler for `TerdutServer`
— all considered alongside §4.1's `spec.pod` and left out of that round.
An HPA no longer contradicts anything now that `spec.replicas` defaults
to 2 (terdut-server v0.36.0's advisory locks), but it is still a
separate, not-yet-made decision: a fixed replica count has no scaling
metric, min/max bounds, or cooldown behaviour to get right, and nobody
has asked for it yet.
- Admission webhooks / CEL-only validation limits (e.g. verifying a - Admission webhooks / CEL-only validation limits (e.g. verifying a
`teamRef` exists at admission time rather than surfacing it as a status `teamRef` exists at admission time rather than surfacing it as a status
condition after the fact). condition after the fact).
+10 -8
View File
@@ -72,15 +72,17 @@ New commits build forward over the old ones; no git history rewrite.
individual keys, so there's nothing to undo there regardless. individual keys, so there's nothing to undo there regardless.
- RBAC: read-only watch on `postgresql.acid.zalan.do`, degrading gracefully - RBAC: read-only watch on `postgresql.acid.zalan.do`, degrading gracefully
if that CRD isn't installed (§8, §9). if that CRD isn't installed (§8, §9).
- Shipped, scoped down from §8's full ambition in two ways, both called out - Shipped, scoped down from §8's full ambition in one way, called out in
in code rather than silently dropped: no live watch on the Zalando- code rather than silently dropped: no live watch on the Zalando-
generated credentials Secret for rotation (relies on the periodic resync generated credentials Secret for rotation (relies on the periodic resync
to notice eventually, higher latency than a watch); no Gateway API to notice eventually, higher latency than a watch). A near-term
`HTTPRoute` creation from `spec.networking.hostname`/`gatewayListener` follow-up, not deferred to a later stage.
(needs the Gateway API types as a new dependency, and nothing about - External exposure (a Gateway API `HTTPRoute` from
proving a `TerdutServer` boots and bootstraps a real server depends on `spec.networking.hostname`) was originally sketched here too, as a
external ingress existing). Both are near-term follow-ups, not deferred second near-term follow-up alongside the one above. It's since become an
to a later stage. explicit, permanent non-goal instead (DESIGN.md §1): the operator will
never manage ingress/exposure for `TerdutServer` in any form. See
`examples/networking` for how to do that yourself.
- `envtest` covering Deployment/Service reconciliation and both database - `envtest` covering Deployment/Service reconciliation and both database
paths — the Zalando path needs that CRD's schema vendored into the test paths — the Zalando path needs that CRD's schema vendored into the test
environment (there's no real `postgres-operator` controller in `envtest`, environment (there's no real `postgres-operator` controller in `envtest`,
+175 -27
View File
@@ -1,8 +1,10 @@
package v1alpha1 package v1alpha1
import ( import (
corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/util/intstr"
) )
// SecretKeyRef names one data key inside a Secret. Every use of this type in // SecretKeyRef names one data key inside a Secret. Every use of this type in
@@ -21,6 +23,37 @@ type SecretKeyRef struct {
Key string `json:"key"` Key string `json:"key"`
} }
// CredentialsDeletionPolicy is what deleting a TerdutServer does to the
// instance credential Secret the operator generated for it.
// +kubebuilder:validation:Enum=Retain;Delete
type CredentialsDeletionPolicy string
const (
// CredentialsRetain keeps the Secret when the TerdutServer is deleted, so
// a TerdutServer recreated with the same name and namespace against the
// same database adopts it again instead of finding a server it cannot
// log in to. Deleting a TerdutServer never touches its database, so the
// operator's service account is still there to be reused. The default.
CredentialsRetain CredentialsDeletionPolicy = "Retain"
// CredentialsDelete removes the Secret with the TerdutServer. Choose it
// when the database goes too, or when the credential must not outlive the
// object.
CredentialsDelete CredentialsDeletionPolicy = "Delete"
)
// CredentialsSpec configures the lifecycle of the generated instance
// credential.
type CredentialsSpec struct {
// deletionPolicy: whether the instance credential Secret is kept
// (Retain, the default) or removed (Delete) when this TerdutServer is
// deleted. A kept Secret is only ever adopted after the server accepts
// its token, so one left over from a database that has since been reset
// is ignored and replaced.
// +kubebuilder:default=Retain
// +optional
DeletionPolicy CredentialsDeletionPolicy `json:"deletionPolicy,omitempty"`
}
// ImageSpec is the terdut-server image to run. // ImageSpec is the terdut-server image to run.
type ImageSpec struct { type ImageSpec struct {
// +kubebuilder:validation:MinLength=1 // +kubebuilder:validation:MinLength=1
@@ -29,33 +62,28 @@ type ImageSpec struct {
Tag string `json:"tag"` Tag string `json:"tag"`
} }
// NetworkingSpec is how this TerdutServer is reached from outside the // NetworkingSpec configures terdut-server itself and the plain ClusterIP
// cluster. // Service the operator creates in front of it. It does not expose
// // TerdutServer outside the cluster in any way, and never will (DESIGN.md
// hostname/gatewayListener describe the intended Gateway API HTTPRoute // §1 -- a permanent non-goal, not a staged one): exposing it is entirely up
// (matching charts/terdut-server's own templates/httpproxy.yaml, despite its // to whoever deploys it -- a Gateway API HTTPRoute, a plain Ingress, an
// name — that chart carries a Gateway API HTTPRoute, not a Contour // Istio VirtualService, or nothing at all if it should stay cluster-
// HTTPProxy), but creating that HTTPRoute isn't implemented yet: it needs // internal. See examples/networking for worked examples against the
// the Gateway API types as a new dependency, and nothing about proving a // Service this creates.
// TerdutServer boots and bootstraps a real server depends on external
// ingress existing. Tracked as a near-term follow-up, not deferred to a
// later ROADMAP.md stage the way Deployment/database/bootstrap once were.
type NetworkingSpec struct { type NetworkingSpec struct {
// hostname the HTTPRoute will carry once it exists. // hostname is terdut-server's own public URL (TERDUT_PUBLIC_URL) --
// used for absolute links terdut-server generates itself
// (notifications, OIDC redirect URIs), not read by this operator for
// anything ingress-related. Set it to whatever hostname your own
// exposure mechanism, if any, actually serves this on.
// +optional // +optional
Hostname string `json:"hostname,omitempty"` Hostname string `json:"hostname,omitempty"`
// servicePort is both the Service's port and the HTTPRoute's backend // servicePort is both the container's port and the ClusterIP Service's
// port once it exists. Defaults to 8080, matching the chart's own // port. Defaults to 8080, matching the chart's own service.port default.
// service.port default.
// +kubebuilder:default=8080 // +kubebuilder:default=8080
// +optional // +optional
ServicePort int32 `json:"servicePort,omitempty"` ServicePort int32 `json:"servicePort,omitempty"`
// gatewayListener is the HTTPRoute's sectionName once it exists. Empty
// attaches to every matching listener, including plaintext HTTP.
// +optional
GatewayListener string `json:"gatewayListener,omitempty"`
} }
// PostgresClusterRef names a Zalando postgres-operator `postgresql` CR // PostgresClusterRef names a Zalando postgres-operator `postgresql` CR
@@ -190,6 +218,109 @@ type AllowedTeams struct {
Namespaces AllowedTeamsNamespaces `json:"namespaces,omitempty"` Namespaces AllowedTeamsNamespaces `json:"namespaces,omitempty"`
} }
// PodDisruptionBudgetSpec configures an optional PodDisruptionBudget for
// this TerdutServer's Deployment. Exactly one of minAvailable or
// maxUnavailable may be set, matching policyv1.PodDisruptionBudgetSpec's own
// upstream rule (both wrap intstr.IntOrString unchanged here -- this is pure
// passthrough, not reshaped) and mirroring DatabaseSpec's own
// dsn/postgresClusterRef mutual-exclusion pattern. Clearing this field
// deletes any PodDisruptionBudget the controller previously created for this
// TerdutServer (DESIGN.md §7).
// +kubebuilder:validation:XValidation:rule="(has(self.minAvailable) ? 1 : 0) + (has(self.maxUnavailable) ? 1 : 0) == 1",message="exactly one of minAvailable or maxUnavailable must be set"
type PodDisruptionBudgetSpec struct {
// minAvailable -- mutually exclusive with maxUnavailable.
// +optional
MinAvailable *intstr.IntOrString `json:"minAvailable,omitempty"`
// maxUnavailable -- mutually exclusive with minAvailable.
// +optional
MaxUnavailable *intstr.IntOrString `json:"maxUnavailable,omitempty"`
}
// PodSpec is pod-level customization of the Deployment this TerdutServer
// creates. Fields here directly reuse corev1 types wherever corev1 already
// models the knob exactly, rather than wrapping (unlike SecretKeyRef's own
// "wrap only when a round-trip through a different type buys something"
// standard would suggest at first glance -- none of these do: Tolerations,
// Affinity, TopologySpreadConstraints, Resources, SecurityContext, EnvVar,
// EnvFromSource, Volume, VolumeMount and LocalObjectReference are all passed
// straight through to the pod template with no added semantics, matching how
// CloudNativePG and the Zalando postgres-operator both expose the same
// knobs).
type PodSpec struct {
// annotations are merged onto the pod template's own metadata.
// Operator-managed labels (labelsFor) are never touched by this field.
// +optional
Annotations map[string]string `json:"annotations,omitempty"`
// +optional
NodeSelector map[string]string `json:"nodeSelector,omitempty"`
// +optional
Tolerations []corev1.Toleration `json:"tolerations,omitempty"`
// affinity covers node affinity, pod affinity and pod anti-affinity in
// one field -- even though replicas now defaults to 2 (see
// TerdutServerSpec.Replicas's own doc comment), this operator still
// never generates a default anti-affinity of its own the way a
// multi-replica-aware operator typically would, so this stays pure
// user-supplied passthrough, not a toggle-plus-generated-default.
// +optional
Affinity *corev1.Affinity `json:"affinity,omitempty"`
// +optional
TopologySpreadConstraints []corev1.TopologySpreadConstraint `json:"topologySpreadConstraints,omitempty"`
// resources applied to the main terdut-server container. Unset today --
// this field closes a pre-existing gap, not a behavior change for
// anyone not setting it.
// +optional
Resources corev1.ResourceRequirements `json:"resources,omitempty"`
// securityContext is pod-level.
// +optional
SecurityContext *corev1.PodSecurityContext `json:"securityContext,omitempty"`
// containerSecurityContext applies to the main terdut-server container
// only -- not wait-for-postgres, which runs a stock postgres image this
// operator doesn't control the entrypoint of. No implicit defaults are
// merged underneath it.
// +optional
ContainerSecurityContext *corev1.SecurityContext `json:"containerSecurityContext,omitempty"`
// serviceAccountName. Defaults to "default", same as any pod that
// doesn't set it.
// +optional
ServiceAccountName string `json:"serviceAccountName,omitempty"`
// extraEnv is appended after the fixed env vars buildEnv produces.
// +optional
ExtraEnv []corev1.EnvVar `json:"extraEnv,omitempty"`
// +optional
ExtraEnvFrom []corev1.EnvFromSource `json:"extraEnvFrom,omitempty"`
// extraVolumes are added to the pod spec; pair with extraVolumeMounts to
// actually mount one on the main container.
// +optional
ExtraVolumes []corev1.Volume `json:"extraVolumes,omitempty"`
// extraVolumeMounts are added to the main terdut-server container only
// -- not wait-for-postgres.
// +optional
ExtraVolumeMounts []corev1.VolumeMount `json:"extraVolumeMounts,omitempty"`
// +optional
ImagePullSecrets []corev1.LocalObjectReference `json:"imagePullSecrets,omitempty"`
// disruptionBudget, when set, causes the controller to reconcile a
// policyv1.PodDisruptionBudget selecting this TerdutServer's pods.
// Removing this field deletes any PodDisruptionBudget the controller
// previously created.
// +optional
DisruptionBudget *PodDisruptionBudgetSpec `json:"disruptionBudget,omitempty"`
}
// TerdutServerSpec defines the desired state of TerdutServer. // TerdutServerSpec defines the desired state of TerdutServer.
// //
// The operator creates and owns every TerdutServer it manages (DESIGN.md // The operator creates and owns every TerdutServer it manages (DESIGN.md
@@ -201,10 +332,14 @@ type TerdutServerSpec struct {
// +required // +required
Image ImageSpec `json:"image"` Image ImageSpec `json:"image"`
// replicas. terdut-server is not horizontally-scale-tested; keep this // replicas. Defaults to 2: terdut-server v0.36.0 put the sweeper, the
// at its default of 1 unless you've verified otherwise -- the sweeper // notifier and the migration runner each behind a Postgres advisory
// and the notifier are unsynchronised singletons. // lock, and gave incident creation its own conflict resolution, so
// +kubebuilder:default=1 // more than one replica no longer double-pages, races a migration, or
// drops a webhook payload. image.tag must be v0.36.0 or newer for
// that to hold -- an older terdut-server has none of these guards,
// and this field does not check the tag for you.
// +kubebuilder:default=2
// +optional // +optional
Replicas int32 `json:"replicas,omitempty"` Replicas int32 `json:"replicas,omitempty"`
@@ -220,6 +355,11 @@ type TerdutServerSpec struct {
// +optional // +optional
Deadman DeadmanSpec `json:"deadman,omitempty"` Deadman DeadmanSpec `json:"deadman,omitempty"`
// credentials: what happens to the instance credential this operator
// generates for the server.
// +optional
Credentials CredentialsSpec `json:"credentials,omitempty"`
// +optional // +optional
Notify NotifySpec `json:"notify,omitempty"` Notify NotifySpec `json:"notify,omitempty"`
@@ -237,6 +377,11 @@ type TerdutServerSpec struct {
// this CRD's schema doesn't need a breaking change to grow it later. // this CRD's schema doesn't need a breaking change to grow it later.
// +optional // +optional
AllowedTeams AllowedTeams `json:"allowedTeams,omitempty"` AllowedTeams AllowedTeams `json:"allowedTeams,omitempty"`
// pod is pod-level customization of the Deployment this TerdutServer
// creates (DESIGN.md §4.1).
// +optional
Pod PodSpec `json:"pod,omitempty"`
} }
// Condition types this controller sets on TerdutServer. // Condition types this controller sets on TerdutServer.
@@ -273,9 +418,12 @@ const (
// ReasonBootstrapStateLost: a checkpointed admin credential // ReasonBootstrapStateLost: a checkpointed admin credential
// (DESIGN.md §6) was lost after being used but before the lasting // (DESIGN.md §6) was lost after being used but before the lasting
// credential it was for could be persisted -- the one genuinely // credential it was for could be persisted -- the one genuinely
// pathological case in the self-registration flow. Fail-closed, same // pathological case in the self-registration flow -- or the server's
// recovery as DESIGN.md §5's webhook-Secret-loss rule: delete and // database is already bootstrapped and no credential for it survives
// recreate this TerdutServer. // (spec.credentials.deletionPolicy: Delete, or the Secret removed by
// hand). Fail-closed: the operator cannot mint a credential, and
// deleting and recreating the TerdutServer does not clear the database.
// Restore the Secret, or reset the server's database.
ReasonBootstrapStateLost = "BootstrapStateLost" ReasonBootstrapStateLost = "BootstrapStateLost"
// ReasonAdopted: the happy path. A working credential is in hand, the // ReasonAdopted: the happy path. A working credential is in hand, the
// Deployment has a ready replica, and the database (if postgresClusterRef) // Deployment has a ready replica, and the database (if postgresClusterRef)
+63
View File
@@ -27,6 +27,42 @@ type TerdutTeamOIDC struct {
OwnerGroup string `json:"ownerGroup,omitempty"` OwnerGroup string `json:"ownerGroup,omitempty"`
} }
// TerdutTeamInvite requests a standing invite link into this team, minted
// with the team's own team-scoped credential — requireTeamOwner already
// treats that credential as owner-equivalent for every /invites route
// (ratified, not a gap, as of terdut-server's SERVICE-ACCOUNTS.md). This is
// the real answer to "how does a human ever get a first login on a
// password-only, operator-managed install" (terdut-server#23): no signup_mode
// flip, no admin token, just a link redeemed the same way anyone else's
// invite would be.
type TerdutTeamInvite struct {
// enabled mints (and keeps refreshed ahead of terdut-server's own fixed
// 7-day TTL) an invite link while true. Flipping it back to false
// revokes the current one server-side rather than leaving it to expire
// on its own.
// +optional
Enabled bool `json:"enabled,omitempty"`
// role is what the invite grants: member or owner. Defaults to member —
// owner by default would make every invite link a standing
// administrative credential for the team, a much bigger blast radius
// than "let a human see the queue".
// +optional
// +kubebuilder:validation:Enum=member;owner
// +kubebuilder:default=member
Role string `json:"role,omitempty"`
// maxUses bounds how many times this link may be redeemed before it
// stops working, mirroring terdut-server's own 1-100 range
// (POST /api/teams/{teamID}/invites). Defaults to 1: a link meant for
// one specific person, not a standing door.
// +optional
// +kubebuilder:validation:Minimum=1
// +kubebuilder:validation:Maximum=100
// +kubebuilder:default=1
MaxUses int64 `json:"maxUses,omitempty"`
}
// TerdutTeamSpec defines the desired state of TerdutTeam. // TerdutTeamSpec defines the desired state of TerdutTeam.
type TerdutTeamSpec struct { type TerdutTeamSpec struct {
// serverRef names the TerdutServer this team belongs to. // serverRef names the TerdutServer this team belongs to.
@@ -43,6 +79,9 @@ type TerdutTeamSpec struct {
// +optional // +optional
OIDC TerdutTeamOIDC `json:"oidc,omitempty"` OIDC TerdutTeamOIDC `json:"oidc,omitempty"`
// +optional
Invite TerdutTeamInvite `json:"invite,omitempty"`
} }
// Condition reasons this controller sets. // Condition reasons this controller sets.
@@ -62,6 +101,19 @@ const (
ReasonTeamAdopted = "Adopted" ReasonTeamAdopted = "Adopted"
) )
// Condition reasons for spec.invite reconciliation (TerdutTeamInvite). Not
// surfaced on the Ready condition itself — an invite is a convenience, not
// a dependency anything else in this team's own readiness waits on — but
// recorded as Events and readable via `kubectl describe`.
const (
// ReasonInviteMinted: spec.invite.enabled is true and status.inviteSecretRef
// is populated and live.
ReasonInviteMinted = "InviteMinted"
// ReasonInviteRevoked: spec.invite.enabled flipped back to false and the
// server-side invite was revoked (or there was nothing to revoke).
ReasonInviteRevoked = "InviteRevoked"
)
// TerdutTeamStatus defines the observed state of TerdutTeam. // TerdutTeamStatus defines the observed state of TerdutTeam.
type TerdutTeamStatus struct { type TerdutTeamStatus struct {
// +listType=map // +listType=map
@@ -87,6 +139,17 @@ type TerdutTeamStatus struct {
// +optional // +optional
ServerEndpoint string `json:"serverEndpoint,omitempty"` ServerEndpoint string `json:"serverEndpoint,omitempty"`
// inviteSecretRef is this team's current invite link, if spec.invite.enabled.
// Unlike credentialsSecretRef, this lives in the TerdutTeam's OWN
// namespace, not the operator's: an invite is bounded, limited-use, and
// meant for this namespace's own human operators to read and hand out,
// not a durable high-privilege credential — same shape as
// TerdutAlertSource's status.webhookURLSecretRef, not TerdutServer's
// cross-namespace credentialsSecretRef. Nil whenever spec.invite.enabled
// is false or unset.
// +optional
InviteSecretRef *LocalSecretRef `json:"inviteSecretRef,omitempty"`
// +optional // +optional
ObservedGeneration int64 `json:"observedGeneration,omitempty"` ObservedGeneration int64 `json:"observedGeneration,omitempty"`
} }
+162
View File
@@ -5,8 +5,10 @@
package v1alpha1 package v1alpha1
import ( import (
corev1 "k8s.io/api/core/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/util/intstr"
) )
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
@@ -45,6 +47,21 @@ func (in *AllowedTeamsNamespaces) DeepCopy() *AllowedTeamsNamespaces {
return out return out
} }
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *CredentialsSpec) DeepCopyInto(out *CredentialsSpec) {
*out = *in
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new CredentialsSpec.
func (in *CredentialsSpec) DeepCopy() *CredentialsSpec {
if in == nil {
return nil
}
out := new(CredentialsSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *DatabaseSpec) DeepCopyInto(out *DatabaseSpec) { func (in *DatabaseSpec) DeepCopyInto(out *DatabaseSpec) {
*out = *in *out = *in
@@ -210,6 +227,128 @@ func (in *OIDCSpec) DeepCopy() *OIDCSpec {
return out return out
} }
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *PodDisruptionBudgetSpec) DeepCopyInto(out *PodDisruptionBudgetSpec) {
*out = *in
if in.MinAvailable != nil {
in, out := &in.MinAvailable, &out.MinAvailable
*out = new(intstr.IntOrString)
**out = **in
}
if in.MaxUnavailable != nil {
in, out := &in.MaxUnavailable, &out.MaxUnavailable
*out = new(intstr.IntOrString)
**out = **in
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PodDisruptionBudgetSpec.
func (in *PodDisruptionBudgetSpec) DeepCopy() *PodDisruptionBudgetSpec {
if in == nil {
return nil
}
out := new(PodDisruptionBudgetSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *PodSpec) DeepCopyInto(out *PodSpec) {
*out = *in
if in.Annotations != nil {
in, out := &in.Annotations, &out.Annotations
*out = make(map[string]string, len(*in))
for key, val := range *in {
(*out)[key] = val
}
}
if in.NodeSelector != nil {
in, out := &in.NodeSelector, &out.NodeSelector
*out = make(map[string]string, len(*in))
for key, val := range *in {
(*out)[key] = val
}
}
if in.Tolerations != nil {
in, out := &in.Tolerations, &out.Tolerations
*out = make([]corev1.Toleration, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.Affinity != nil {
in, out := &in.Affinity, &out.Affinity
*out = new(corev1.Affinity)
(*in).DeepCopyInto(*out)
}
if in.TopologySpreadConstraints != nil {
in, out := &in.TopologySpreadConstraints, &out.TopologySpreadConstraints
*out = make([]corev1.TopologySpreadConstraint, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
in.Resources.DeepCopyInto(&out.Resources)
if in.SecurityContext != nil {
in, out := &in.SecurityContext, &out.SecurityContext
*out = new(corev1.PodSecurityContext)
(*in).DeepCopyInto(*out)
}
if in.ContainerSecurityContext != nil {
in, out := &in.ContainerSecurityContext, &out.ContainerSecurityContext
*out = new(corev1.SecurityContext)
(*in).DeepCopyInto(*out)
}
if in.ExtraEnv != nil {
in, out := &in.ExtraEnv, &out.ExtraEnv
*out = make([]corev1.EnvVar, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ExtraEnvFrom != nil {
in, out := &in.ExtraEnvFrom, &out.ExtraEnvFrom
*out = make([]corev1.EnvFromSource, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ExtraVolumes != nil {
in, out := &in.ExtraVolumes, &out.ExtraVolumes
*out = make([]corev1.Volume, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ExtraVolumeMounts != nil {
in, out := &in.ExtraVolumeMounts, &out.ExtraVolumeMounts
*out = make([]corev1.VolumeMount, len(*in))
for i := range *in {
(*in)[i].DeepCopyInto(&(*out)[i])
}
}
if in.ImagePullSecrets != nil {
in, out := &in.ImagePullSecrets, &out.ImagePullSecrets
*out = make([]corev1.LocalObjectReference, len(*in))
copy(*out, *in)
}
if in.DisruptionBudget != nil {
in, out := &in.DisruptionBudget, &out.DisruptionBudget
*out = new(PodDisruptionBudgetSpec)
(*in).DeepCopyInto(*out)
}
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PodSpec.
func (in *PodSpec) DeepCopy() *PodSpec {
if in == nil {
return nil
}
out := new(PodSpec)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *PostgresClusterRef) DeepCopyInto(out *PostgresClusterRef) { func (in *PostgresClusterRef) DeepCopyInto(out *PostgresClusterRef) {
*out = *in *out = *in
@@ -640,9 +779,11 @@ func (in *TerdutServerSpec) DeepCopyInto(out *TerdutServerSpec) {
in.Database.DeepCopyInto(&out.Database) in.Database.DeepCopyInto(&out.Database)
out.Sweeper = in.Sweeper out.Sweeper = in.Sweeper
out.Deadman = in.Deadman out.Deadman = in.Deadman
out.Credentials = in.Credentials
in.Notify.DeepCopyInto(&out.Notify) in.Notify.DeepCopyInto(&out.Notify)
in.OIDC.DeepCopyInto(&out.OIDC) in.OIDC.DeepCopyInto(&out.OIDC)
in.AllowedTeams.DeepCopyInto(&out.AllowedTeams) in.AllowedTeams.DeepCopyInto(&out.AllowedTeams)
in.Pod.DeepCopyInto(&out.Pod)
} }
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutServerSpec. // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutServerSpec.
@@ -709,6 +850,21 @@ func (in *TerdutTeam) DeepCopyObject() runtime.Object {
return nil return nil
} }
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutTeamInvite) DeepCopyInto(out *TerdutTeamInvite) {
*out = *in
}
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamInvite.
func (in *TerdutTeamInvite) DeepCopy() *TerdutTeamInvite {
if in == nil {
return nil
}
out := new(TerdutTeamInvite)
in.DeepCopyInto(out)
return out
}
// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil.
func (in *TerdutTeamList) DeepCopyInto(out *TerdutTeamList) { func (in *TerdutTeamList) DeepCopyInto(out *TerdutTeamList) {
*out = *in *out = *in
@@ -776,6 +932,7 @@ func (in *TerdutTeamSpec) DeepCopyInto(out *TerdutTeamSpec) {
*out = *in *out = *in
out.ServerRef = in.ServerRef out.ServerRef = in.ServerRef
out.OIDC = in.OIDC out.OIDC = in.OIDC
out.Invite = in.Invite
} }
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamSpec. // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamSpec.
@@ -803,6 +960,11 @@ func (in *TerdutTeamStatus) DeepCopyInto(out *TerdutTeamStatus) {
*out = new(SecretKeyRef) *out = new(SecretKeyRef)
**out = **in **out = **in
} }
if in.InviteSecretRef != nil {
in, out := &in.InviteSecretRef, &out.InviteSecretRef
*out = new(LocalSecretRef)
**out = **in
}
} }
// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamStatus. // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TerdutTeamStatus.
+2 -2
View File
@@ -6,8 +6,8 @@ type: application
# These fields decide nothing: `make helm-package` passes --version and # These fields decide nothing: `make helm-package` passes --version and
# --app-version from the release tag (same reasoning as terdut-server's own # --app-version from the release tag (same reasoning as terdut-server's own
# chart). They're for whoever reads the tree before a tag exists. # chart). They're for whoever reads the tree before a tag exists.
version: 0.1.2 version: 0.5.0
appVersion: "v0.1.2" appVersion: "v0.5.0"
keywords: keywords:
- kubernetes - kubernetes
File diff suppressed because it is too large Load Diff
@@ -63,6 +63,47 @@ spec:
create rule, via TEAM-LOOKUP.md). create rule, via TEAM-LOOKUP.md).
minLength: 1 minLength: 1
type: string type: string
invite:
description: |-
TerdutTeamInvite requests a standing invite link into this team, minted
with the team's own team-scoped credential — requireTeamOwner already
treats that credential as owner-equivalent for every /invites route
(ratified, not a gap, as of terdut-server's SERVICE-ACCOUNTS.md). This is
the real answer to "how does a human ever get a first login on a
password-only, operator-managed install" (terdut-server#23): no signup_mode
flip, no admin token, just a link redeemed the same way anyone else's
invite would be.
properties:
enabled:
description: |-
enabled mints (and keeps refreshed ahead of terdut-server's own fixed
7-day TTL) an invite link while true. Flipping it back to false
revokes the current one server-side rather than leaving it to expire
on its own.
type: boolean
maxUses:
default: 1
description: |-
maxUses bounds how many times this link may be redeemed before it
stops working, mirroring terdut-server's own 1-100 range
(POST /api/teams/{teamID}/invites). Defaults to 1: a link meant for
one specific person, not a standing door.
format: int64
maximum: 100
minimum: 1
type: integer
role:
default: member
description: |-
role is what the invite grants: member or owner. Defaults to member —
owner by default would make every invite link a standing
administrative credential for the team, a much bigger blast radius
than "let a human see the queue".
enum:
- member
- owner
type: string
type: object
oidc: oidc:
description: |- description: |-
TerdutTeamOIDC binds which identity-provider groups grant membership and TerdutTeamOIDC binds which identity-provider groups grant membership and
@@ -171,6 +212,24 @@ spec:
- key - key
- name - name
type: object type: object
inviteSecretRef:
description: |-
inviteSecretRef is this team's current invite link, if spec.invite.enabled.
Unlike credentialsSecretRef, this lives in the TerdutTeam's OWN
namespace, not the operator's: an invite is bounded, limited-use, and
meant for this namespace's own human operators to read and hand out,
not a durable high-privilege credential — same shape as
TerdutAlertSource's status.webhookURLSecretRef, not TerdutServer's
cross-namespace credentialsSecretRef. Nil whenever spec.invite.enabled
is false or unset.
properties:
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- name
type: object
observedGeneration: observedGeneration:
format: int64 format: int64
type: integer type: integer
@@ -58,6 +58,18 @@ rules:
verbs: verbs:
- create - create
- patch - patch
- apiGroups:
- policy
resources:
- poddisruptionbudgets
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups: - apiGroups:
- terdut.ryuvia.com - terdut.ryuvia.com
resources: resources:
@@ -38,6 +38,10 @@ spec:
deadman: deadman:
{{- toYaml . | nindent 4 }} {{- toYaml . | nindent 4 }}
{{- end }} {{- end }}
{{- with .Values.terdutServer.credentials }}
credentials:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- with .Values.terdutServer.notify }} {{- with .Values.terdutServer.notify }}
notify: notify:
{{- toYaml . | nindent 4 }} {{- toYaml . | nindent 4 }}
@@ -51,4 +55,8 @@ spec:
allowedTeams: allowedTeams:
{{- toYaml . | nindent 4 }} {{- toYaml . | nindent 4 }}
{{- end }} {{- end }}
{{- with .Values.terdutServer.pod }}
pod:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- end }} {{- end }}
+34 -4
View File
@@ -233,12 +233,19 @@ terdutServer:
## Required when terdutServer.enabled. ## Required when terdutServer.enabled.
# tag: "" # tag: ""
replicas: 1 ## Safe above 1 since terdut-server v0.36.0 (image.tag above must be that or
## newer): the sweeper, notifier and migration runner are each behind a
## Postgres advisory lock, and incident creation resolves its own insert
## conflict, matching this CRD's own spec.replicas default.
replicas: 2
networking: networking:
## Required when terdutServer.enabled -- the hostname a future ## Required when terdutServer.enabled -- terdut-server's own public
## HTTPRoute will carry (see NetworkingSpec's own doc comment: creating ## URL, used for absolute links it generates itself (notifications,
## that HTTPRoute isn't implemented yet). ## OIDC redirect URIs). This chart/operator never creates any
## ingress/HTTPRoute for it -- see examples/networking in the repo for
## how to expose the Service this chart's TerdutServer CR causes to
## be created, if you want to expose it at all.
# hostname: "" # hostname: ""
servicePort: 8080 servicePort: 8080
@@ -263,6 +270,13 @@ terdutServer:
# matchers: "alertname=Watchdog" # matchers: "alertname=Watchdog"
# timeout: 15m # timeout: 15m
# severity: critical # severity: critical
# What happens to the instance credential the operator generates for this
# server when the TerdutServer is deleted. Retain (the default) keeps it, so
# a TerdutServer recreated with the same name against the same database
# adopts it again; Delete removes it with the TerdutServer. A kept Secret
# is only adopted if the server accepts its token.
# credentials:
# deletionPolicy: Retain
# notify: {} # notify: {}
# oidc: {} # oidc: {}
@@ -270,3 +284,19 @@ terdutServer:
# allowedTeams: {} # allowedTeams: {}
## Pod-level customization of the Deployment this TerdutServer creates --
## see api/v1alpha1/terdutserver_types.go's PodSpec for the full shape
## (nodeSelector, tolerations, affinity, topologySpreadConstraints,
## securityContext, containerSecurityContext, serviceAccountName,
## extraEnv/extraEnvFrom, extraVolumes/extraVolumeMounts,
## imagePullSecrets, disruptionBudget). All optional; resources is the
## one most installs will want to set, since the container otherwise
## runs with no requests/limits at all:
# pod:
# resources:
# requests:
# cpu: 100m
# memory: 128Mi
# limits:
# memory: 256Mi
File diff suppressed because it is too large Load Diff
@@ -60,6 +60,47 @@ spec:
create rule, via TEAM-LOOKUP.md). create rule, via TEAM-LOOKUP.md).
minLength: 1 minLength: 1
type: string type: string
invite:
description: |-
TerdutTeamInvite requests a standing invite link into this team, minted
with the team's own team-scoped credential — requireTeamOwner already
treats that credential as owner-equivalent for every /invites route
(ratified, not a gap, as of terdut-server's SERVICE-ACCOUNTS.md). This is
the real answer to "how does a human ever get a first login on a
password-only, operator-managed install" (terdut-server#23): no signup_mode
flip, no admin token, just a link redeemed the same way anyone else's
invite would be.
properties:
enabled:
description: |-
enabled mints (and keeps refreshed ahead of terdut-server's own fixed
7-day TTL) an invite link while true. Flipping it back to false
revokes the current one server-side rather than leaving it to expire
on its own.
type: boolean
maxUses:
default: 1
description: |-
maxUses bounds how many times this link may be redeemed before it
stops working, mirroring terdut-server's own 1-100 range
(POST /api/teams/{teamID}/invites). Defaults to 1: a link meant for
one specific person, not a standing door.
format: int64
maximum: 100
minimum: 1
type: integer
role:
default: member
description: |-
role is what the invite grants: member or owner. Defaults to member —
owner by default would make every invite link a standing
administrative credential for the team, a much bigger blast radius
than "let a human see the queue".
enum:
- member
- owner
type: string
type: object
oidc: oidc:
description: |- description: |-
TerdutTeamOIDC binds which identity-provider groups grant membership and TerdutTeamOIDC binds which identity-provider groups grant membership and
@@ -168,6 +209,24 @@ spec:
- key - key
- name - name
type: object type: object
inviteSecretRef:
description: |-
inviteSecretRef is this team's current invite link, if spec.invite.enabled.
Unlike credentialsSecretRef, this lives in the TerdutTeam's OWN
namespace, not the operator's: an invite is bounded, limited-use, and
meant for this namespace's own human operators to read and hand out,
not a durable high-privilege credential — same shape as
TerdutAlertSource's status.webhookURLSecretRef, not TerdutServer's
cross-namespace credentialsSecretRef. Nil whenever spec.invite.enabled
is false or unset.
properties:
name:
description: name is the Secret's name.
minLength: 1
type: string
required:
- name
type: object
observedGeneration: observedGeneration:
format: int64 format: int64
type: integer type: integer
+12
View File
@@ -52,6 +52,18 @@ rules:
verbs: verbs:
- create - create
- patch - patch
- apiGroups:
- policy
resources:
- poddisruptionbudgets
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups: - apiGroups:
- terdut.ryuvia.com - terdut.ryuvia.com
resources: resources:
@@ -31,3 +31,12 @@ spec:
timeout: 15m timeout: 15m
severity: critical severity: critical
passwordLogin: true passwordLogin: true
# Pod-level customization, all optional -- see PodSpec in
# api/v1alpha1/terdutserver_types.go for the full shape (tolerations,
# affinity, topologySpreadConstraints, securityContext,
# serviceAccountName, extraEnv/extraVolumes, imagePullSecrets,
# disruptionBudget, ...). Example:
# pod:
# resources:
# requests: {cpu: 100m, memory: 128Mi}
# limits: {memory: 256Mi}
+21 -16
View File
@@ -10,7 +10,7 @@
apiVersion: v1 apiVersion: v1
kind: Secret kind: Secret
metadata: metadata:
name: terdut-demo-postgres name: terdut-operator-demo-postgres
type: Opaque type: Opaque
stringData: stringData:
password: demo-not-a-real-password password: demo-not-a-real-password
@@ -18,9 +18,9 @@ stringData:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
name: terdut-demo-postgres name: terdut-operator-demo-postgres
labels: labels:
app: terdut-demo-postgres app: terdut-operator-demo-postgres
spec: spec:
replicas: 1 replicas: 1
# Recreate, not RollingUpdate: emptyDir means a new pod starts with an # Recreate, not RollingUpdate: emptyDir means a new pod starts with an
@@ -30,27 +30,32 @@ spec:
type: Recreate type: Recreate
selector: selector:
matchLabels: matchLabels:
app: terdut-demo-postgres app: terdut-operator-demo-postgres
template: template:
metadata: metadata:
labels: labels:
app: terdut-demo-postgres app: terdut-operator-demo-postgres
spec: spec:
containers: containers:
- name: postgres - name: postgres
image: postgres:17-alpine image: postgres:17-alpine
# Partial, deliberately: the official image's entrypoint needs to # Partial, deliberately: the official image's entrypoint needs to
# start as root to chown the data directory before it drops # start as root to chown/chmod the data directory before it drops
# privileges itself (gosu, to the postgres user) -- forcing # privileges itself (gosu, to the postgres user) -- forcing
# runAsNonRoot here would just refuse to start the container. A # runAsNonRoot would just refuse to start the container, and
# "restricted" PodSecurity namespace warns on that gap rather # dropping all capabilities (an earlier version of this file did)
# than blocking (confirmed server-side against this operator's # takes CAP_CHOWN/CAP_FOWNER away from that same root user, which
# own dev cluster), which is an acceptable tradeoff for Postgres # is a different way of breaking the identical startup step:
# that exists only to be thrown away with the rest of this demo. # confirmed the hard way, as `chmod: /var/run/postgresql:
# Operation not permitted` in a real pod's logs, not caught by
# `kubectl apply --dry-run=server` -- that only checks admission
# policy, never whether the container actually boots. A
# "restricted" PodSecurity namespace warns on the remaining gap
# (no runAsNonRoot) rather than blocking, which is an acceptable
# tradeoff for Postgres that exists only to be thrown away with
# the rest of this demo.
securityContext: securityContext:
allowPrivilegeEscalation: false allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
seccompProfile: seccompProfile:
type: RuntimeDefault type: RuntimeDefault
ports: ports:
@@ -64,7 +69,7 @@ spec:
- name: POSTGRES_PASSWORD - name: POSTGRES_PASSWORD
valueFrom: valueFrom:
secretKeyRef: secretKeyRef:
name: terdut-demo-postgres name: terdut-operator-demo-postgres
key: password key: password
volumeMounts: volumeMounts:
- name: data - name: data
@@ -81,10 +86,10 @@ spec:
apiVersion: v1 apiVersion: v1
kind: Service kind: Service
metadata: metadata:
name: terdut-demo-postgres name: terdut-operator-demo-postgres
spec: spec:
selector: selector:
app: terdut-demo-postgres app: terdut-operator-demo-postgres
ports: ports:
- name: postgres - name: postgres
port: 5432 port: 5432
+25 -11
View File
@@ -2,27 +2,41 @@
# this directory (teams, escalation rules, dead man's switches, alert # this directory (teams, escalation rules, dead man's switches, alert
# sources) references it by name. # sources) references it by name.
# #
# networking.hostname is accepted but not yet acted on: creating the # The operator never creates any ingress/HTTPRoute for this TerdutServer --
# HTTPRoute for it isn't implemented yet (api/v1alpha1/terdutserver_types.go, # that's a permanent non-goal (DESIGN.md §1, NetworkingSpec's own doc
# NetworkingSpec's own doc comment) -- this TerdutServer is reachable from # comment), not a missing feature. This demo reaches it only by
# outside the cluster only by port-forwarding its Service, same name as # port-forwarding its Service, same name as this object (see README.md);
# this object (see README.md). # see ../networking for worked examples of exposing it yourself instead.
apiVersion: terdut.ryuvia.com/v1alpha1 apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutServer kind: TerdutServer
metadata: metadata:
name: terdut-demo name: terdut-operator-demo
spec: spec:
image: image:
repository: git.ryuvia.com/niklas/terdut-server repository: git.ryuvia.com/niklas/terdut-server
tag: v0.33.2 # v0.36.0 is the floor now that replicas below is 2 (this demo pins
replicas: 1 # the current release, v0.43.0, so it shows the current web UI too): that
# release put the sweeper, the notifier and the migration runner each
# behind a Postgres advisory lock, and gave incident creation its own
# conflict resolution, which is what makes a second replica safe
# instead of racing the first. (Still carries v0.34.0's fix too --
# callerMayManageServiceAccount, so an instance-scoped service account
# can adopt/rotate a key on a team-scoped account it didn't just create
# in the same call -- without which terdutteam-* can wedge permanently
# on the crash-window race this demo hit live, niklas/terdut-operator#3.)
tag: v0.43.0
# Matches this CRD's own spec.replicas default (v0.4.0) -- stated
# explicitly, like every other field in this file, rather than left to
# the default. RollingUpdate follows automatically; this operator does
# not expose Strategy as a spec field.
replicas: 2
networking: networking:
hostname: terdut-demo.example hostname: terdut-operator-demo.example
servicePort: 8080 servicePort: 8080
database: database:
dsn: "postgres://terdut@terdut-demo-postgres:5432/terdut?sslmode=disable" dsn: "postgres://terdut@terdut-operator-demo-postgres:5432/terdut?sslmode=disable"
passwordSecretRef: passwordSecretRef:
name: terdut-demo-postgres name: terdut-operator-demo-postgres
key: password key: password
sweeper: sweeper:
staleAfter: 6h staleAfter: 6h
+8 -2
View File
@@ -7,11 +7,17 @@ kind: TerdutTeam
metadata: metadata:
name: terdutteam-platform name: terdutteam-platform
spec: spec:
# serverRef.namespace omitted: both this and terdut-demo (01-server.yaml) # serverRef.namespace omitted: both this and terdut-operator-demo (01-server.yaml)
# live in whatever namespace you apply this directory into, which is the # live in whatever namespace you apply this directory into, which is the
# common case and needs no allowedTeams consent on the TerdutServer side # common case and needs no allowedTeams consent on the TerdutServer side
# (DESIGN.md §4.1, §4.6). # (DESIGN.md §4.1, §4.6).
serverRef: serverRef:
name: terdut-demo name: terdut-operator-demo
displayName: Platform displayName: Platform
# No oidc block: this demo is password-login only (01-server.yaml). # No oidc block: this demo is password-login only (01-server.yaml).
# A real invite link, minted via this team's own credential, is how
# run-demo.sh's alice actually gets in -- no signup_mode flip, no admin
# token (see niklas/terdut-server#23's resolution). Role/maxUses left at
# their defaults (member, 1): one link for one person.
invite:
enabled: true
+7 -1
View File
@@ -4,5 +4,11 @@ metadata:
name: terdutteam-payments name: terdutteam-payments
spec: spec:
serverRef: serverRef:
name: terdut-demo name: terdut-operator-demo
displayName: Payments displayName: Payments
# No spec.invite here, deliberately: run-demo.sh joins alice to this team
# through POST /api/teams/{teamID}/members instead (this team's own
# credential, same owner-equivalent reach spec.invite relies on, plus her
# user id resolved via GET /api/users), once she already has an account
# from Platform's invite -- showing both onboarding paths this feature
# unlocks, not just the one.
+65 -27
View File
@@ -10,6 +10,14 @@ This is a demo kit, not a reference deployment: `00-postgres.yaml` runs
Postgres with `emptyDir` storage and a password committed in this Postgres with `emptyDir` storage and a password committed in this
directory. Throw the whole namespace away when you're done. directory. Throw the whole namespace away when you're done.
**Want this fully automated instead of walking through it by hand?**
`./run-demo.sh` does everything below itself, against a fresh (or
already-set-up) `kind` cluster — creates the cluster, installs the
operator, applies every CR here, signs `alice` in for real, and fires a
few alerts. `./run-demo.sh --help` for the knobs, `./run-demo.sh
--teardown` to tear it back down. The rest of this file is the manual
walkthrough it automates.
## Prerequisites ## Prerequisites
- The operator and its CRDs installed and running (`make install - The operator and its CRDs installed and running (`make install
@@ -18,6 +26,17 @@ directory. Throw the whole namespace away when you're done.
throwaway resources in. A `kind` cluster is the easy choice. throwaway resources in. A `kind` cluster is the easy choice.
- `kubectl`, `jq`, `curl` on your path. - `kubectl`, `jq`, `curl` on your path.
**Apply this into a namespace of its own.** Every object name in this
directory is prefixed `terdut-operator-demo` specifically so applying it
by mistake into some other namespace that already has unrelated objects
doesn't collide with them -- but that only helps if this directory's own
objects don't collide with *each other* across two applies. Applying it
twice into two different namespaces is fine; applying it a second time
into a namespace that already has something else named `terdut-demo` (a
real install from following `terdut-operator`'s own repo along, say) is
exactly the mistake this prefix exists to avoid, and it only works if you
don't override these names yourself.
## Apply it ## Apply it
```sh ```sh
@@ -40,52 +59,57 @@ kubectl get terdutservers,terdutteams,terdutescalationrules,terdutdeadmanswitche
The operator's own Deployment template now carries a `wait-for-postgres` The operator's own Deployment template now carries a `wait-for-postgres`
init container (same fix as `charts/terdut-server`'s chart as of v0.33.2), init container (same fix as `charts/terdut-server`'s chart as of v0.33.2),
so `terdut-demo`'s pod should come up clean even against this brand-new so `terdut-operator-demo`'s pod should come up clean even against this brand-new
Postgres doing its very first boot — no `CrashLoopBackOff` expected here. Postgres doing its very first boot — no `CrashLoopBackOff` expected here.
Once `terdut-demo`'s own `Ready` condition is `True`, everything downstream Once `terdut-operator-demo`'s own `Ready` condition is `True`, everything downstream
of it should settle within a reconcile interval or two. of it should settle within a reconcile interval or two.
## See the web UI ## See the web UI
The operator doesn't create any external exposure yet The operator never creates any external exposure for a `TerdutServer` --
(`NetworkingSpec`'s own doc comment in `api/v1alpha1/terdutserver_types.go` that's a permanent non-goal (DESIGN.md §1), not a missing feature, so this
— `spec.networking.hostname` is accepted but nothing acts on it), so: demo just reaches it the simplest way there is:
```sh ```sh
kubectl -n terdut-operator-demo port-forward svc/terdut-demo 8080:8080 kubectl -n terdut-operator-demo port-forward svc/terdut-operator-demo 8080:8080
``` ```
and open http://localhost:8080. and open http://localhost:8080. See `../networking` for worked examples of
exposing it for real (Gateway API, Istio, or a plain `Ingress`) instead.
### First login ### First login
The operator's own bootstrap (DESIGN.md §6) creates the first user through The operator's own bootstrap (DESIGN.md §6) creates the first user through
`/api/bootstrap` and immediately mints itself a service-account token from `/api/bootstrap` and immediately mints itself a service-account token from
it — that account has no password, so there's nothing to sign in with yet. it, then discards the bootstrap user's own key — nobody ever signs in as
`signup_mode` also defaults to `invite_only`, so open signup needs turning that account, and `signup_mode` stays `invite_only` by default. **Don't try
on first, using the admin token the operator generated for itself: to flip it via the operator's own token**: that token is a service account,
and `/api/admin/settings` is deliberately human-only on terdut-server
(`niklas/terdut-server#23` has the full reasoning — widening that gate was
the wrong fix).
The real path in: `02-team-platform.yaml` turns on `spec.invite`, so
Platform's own `TerdutTeam` mints a real invite link with its own
already-working team-scoped credential (the same reach that lets it manage
its own escalation policy, dead man's switches and integrations — owner-
equivalent, confirmed in terdut-server's `SERVICE-ACCOUNTS.md`). Invite
redemption bypasses `signup_mode` entirely, so this needs no admin
credential at all:
```sh ```sh
# Which namespace the operator itself runs in: secretname=$(kubectl -n terdut-operator-demo get terdutteam terdutteam-platform \
kubectl get deploy -A -l control-plane=controller-manager -o jsonpath='{.status.inviteSecretRef.name}')
url=$(kubectl -n terdut-operator-demo get secret "$secretname" -o jsonpath='{.data.url}' | base64 -d)
# The Secret holding the operator's own admin token for this TerdutServer echo "$url" # open this, or POST /api/signup with {"invite": "<the token after invite=>", ...}
# (cross-namespace from terdut-operator-demo, per DESIGN.md §7):
secretname=$(kubectl -n terdut-operator-demo get terdutserver terdut-demo \
-o jsonpath='{.status.credentialsSecretRef.name}')
token=$(kubectl -n <operator-namespace-from-above> get secret "$secretname" \
-o jsonpath='{.data.token}' | base64 -d)
curl -X PUT http://localhost:8080/api/admin/settings \
-H "Authorization: Bearer $token" -H 'Content-Type: application/json' \
-d '{"signup_mode":"open"}'
``` ```
Then sign up through the UI as a normal human account. `04-escalation-platform.yaml` `04-escalation-platform.yaml` names a user `alice` at its first escalation
names a user `alice` at its first escalation level — sign up as `alice` if level — sign up as `alice` if you want that level to mean something rather
you want that level to mean something rather than falling through to than falling through to on-call after 5 minutes. `run-demo.sh` does exactly
on-call after 5 minutes. this automatically (and also joins `alice` to Payments, which deliberately
has no `spec.invite` of its own — see that file's comment for the second
onboarding path this demonstrates).
## Fire some alerts ## Fire some alerts
@@ -109,6 +133,20 @@ Each `(team, scenario)` pair is one stable fingerprint, so firing the same
one twice updates the same alert (a real re-fire) and `resolve` closes one twice updates the same alert (a real re-fire) and `resolve` closes
exactly that one. exactly that one.
Set `CLUSTER` to send the alert as if it came from one of several clusters:
```sh
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-us ./fire-alerts.sh platform high-cpu # a second incident, not a join
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
```
It stands in for a Prometheus external label plus `cluster` in Alertmanager's
`group_by` (terdut-server's README, "Several clusters, one team"): the web UI
then shows the cluster chip on each incident and a cluster filter in the
queue. `CLUSTER` is part of the fingerprint, so resolve with the same value you
fired with. `./run-demo.sh` fires its alerts across `prod-eu` and `prod-us`.
### Dead man's switches ### Dead man's switches
`06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat `06-deadman-platform.yaml` / `07-deadman-payments.yaml` expect a heartbeat
+19 -7
View File
@@ -16,14 +16,23 @@
# is the one part of that URL still usable here. # is the one part of that URL still usable here.
# #
# Usage: # Usage:
# ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve] # [CLUSTER=prod-eu] ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
#
# CLUSTER stands in for a Prometheus externalLabel plus `cluster` in
# Alertmanager's group_by (terdut-server's README, "Several clusters, one
# team"): it is put on the alert's labels and on groupLabels, so the incident
# carries it and the web UI shows the cluster chip and the queue's cluster
# filter. It is also part of the group key and the fingerprint, which is what
# keeps the same alert in two clusters from joining one incident. Unset, the
# alert is sent exactly as before.
# #
# Prerequisites: kubectl context pointed at the demo namespace, jq, curl, # Prerequisites: kubectl context pointed at the demo namespace, jq, curl,
# and (in another terminal) a running: # and (in another terminal) a running:
# kubectl port-forward svc/terdut-demo 8080:8080 # kubectl port-forward svc/terdut-operator-demo 8080:8080
set -euo pipefail set -euo pipefail
NAMESPACE="${NAMESPACE:-}" NAMESPACE="${NAMESPACE:-}"
CLUSTER="${CLUSTER:-}"
BASE_URL="${BASE_URL:-http://localhost:8080}" BASE_URL="${BASE_URL:-http://localhost:8080}"
usage() { usage() {
@@ -45,6 +54,8 @@ scenarios:
env vars: env vars:
NAMESPACE kubectl -n for reading the webhook Secret (required) NAMESPACE kubectl -n for reading the webhook Secret (required)
BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080) BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080)
CLUSTER optional cluster name, e.g. prod-eu: sent as a `cluster` label and
group label, so the UI shows where the incident came from
EOF EOF
exit 1 exit 1
} }
@@ -90,7 +101,7 @@ key="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.key
# created: terdut-server correlates on (team_id, fingerprint), not on # created: terdut-server correlates on (team_id, fingerprint), not on
# anything else in the payload. Real Alertmanager computes this from the # anything else in the payload. Real Alertmanager computes this from the
# alert's label set; a fixed string plays the same role here. # alert's label set; a fixed string plays the same role here.
fingerprint="demo-${team}-${scenario}" fingerprint="demo-${team}-${scenario}${CLUSTER:+-$CLUSTER}"
now="$(date -u +%Y-%m-%dT%H:%M:%SZ)" now="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
if [ "$status" = firing ]; then if [ "$status" = firing ]; then
@@ -101,7 +112,8 @@ fi
payload="$(jq -n \ payload="$(jq -n \
--arg status "$status" \ --arg status "$status" \
--arg groupKey "demo:${team}:${scenario}" \ --arg groupKey "demo:${team}:${scenario}${CLUSTER:+:$CLUSTER}" \
--arg cluster "$CLUSTER" \
--arg alertname "$alertname" \ --arg alertname "$alertname" \
--arg team "$team" \ --arg team "$team" \
--arg severity "$severity" \ --arg severity "$severity" \
@@ -113,10 +125,10 @@ payload="$(jq -n \
version: "4", version: "4",
status: $status, status: $status,
groupKey: $groupKey, groupKey: $groupKey,
groupLabels: { alertname: $alertname, team: $team }, groupLabels: ({ alertname: $alertname, team: $team } + (if $cluster != "" then { cluster: $cluster } else {} end)),
alerts: [{ alerts: [{
status: $status, status: $status,
labels: { alertname: $alertname, severity: $severity, team: $team, instance: "demo" }, labels: ({ alertname: $alertname, severity: $severity, team: $team, instance: "demo" } + (if $cluster != "" then { cluster: $cluster } else {} end)),
annotations: { summary: $summary }, annotations: { summary: $summary },
startsAt: $startsAt, startsAt: $startsAt,
endsAt: $endsAt, endsAt: $endsAt,
@@ -126,7 +138,7 @@ payload="$(jq -n \
}')" }')"
url="${BASE_URL}/api/integrations/${key}/alertmanager" url="${BASE_URL}/api/integrations/${key}/alertmanager"
echo "POST $url (team=$team scenario=$scenario status=$status)" >&2 echo "POST $url (team=$team scenario=$scenario status=$status${CLUSTER:+ cluster=$CLUSTER})" >&2
code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \ code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \
-X POST "$url" -H 'Content-Type: application/json' -d "$payload")" -X POST "$url" -H 'Content-Type: application/json' -d "$payload")"
echo "-> HTTP $code" >&2 echo "-> HTTP $code" >&2
+365
View File
@@ -0,0 +1,365 @@
#!/usr/bin/env bash
# Stands up the complete terdut demo (terdut-operator + every CRD kind this
# repo ships + a working local login + a few synthetic incidents) on a kind
# cluster, fully automated. Password login only -- this demo kit's own
# 01-server.yaml carries no oidc: block at all, so there is nothing to
# disable; OIDC is simply absent.
#
# Usage:
# ./run-demo.sh deploy the whole demo (idempotent: safe to
# re-run against a cluster that already has it)
# ./run-demo.sh --teardown delete the demo namespace (and, if
# TEARDOWN_CLUSTER=true, the kind cluster too)
# ./run-demo.sh --help
#
# All of the defaults below are overridable as environment variables.
set -euo pipefail
CLUSTER_NAME="${CLUSTER_NAME:-terdut-demo}"
NAMESPACE="${NAMESPACE:-terdut-operator-demo}"
OPERATOR_NAMESPACE="${OPERATOR_NAMESPACE:-terdut-operator-system}"
RELEASE_NAME="${RELEASE_NAME:-terdut-operator}"
ALICE_USERNAME="${ALICE_USERNAME:-alice}"
ALICE_EMAIL="${ALICE_EMAIL:-alice@terdut-demo.local}"
DEMO_PASSWORD="${DEMO_PASSWORD:-terdut-demo-1234}"
BASE_URL="${BASE_URL:-http://localhost:8080}"
LOCAL_PORT="${LOCAL_PORT:-8080}"
HELM_TIMEOUT="${HELM_TIMEOUT:-180s}"
WAIT_TIMEOUT="${WAIT_TIMEOUT:-180s}"
TEARDOWN_CLUSTER="${TEARDOWN_CLUSTER:-false}"
PF_PIDFILE="${PF_PIDFILE:-/tmp/terdut-demo-port-forward.pid}"
PF_LOGFILE="${PF_LOGFILE:-/tmp/terdut-demo-port-forward.log}"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
CHART_DIR="$(cd "$SCRIPT_DIR/../.." && pwd)/charts/terdut-operator"
# ---------------------------------------------------------------------------
# helpers
# ---------------------------------------------------------------------------
log() { printf '[run-demo] %s\n' "$*" >&2; }
die() { printf '[run-demo] FAILED: %s\n' "$*" >&2; exit 1; }
usage() {
cat <<EOF
Usage: $0 [--teardown|--help]
Deploys (or tears down) the complete terdut demo on a kind cluster.
See the top of this file for every overridable environment variable.
EOF
}
# Only kills a port-forward THIS invocation started, so a successful run
# never has its background job reaped by its own exit trap.
STARTED_PF_PID=""
cleanup_on_failure() {
local rc=$?
if [ "$rc" -ne 0 ] && [ -n "$STARTED_PF_PID" ]; then
log "run failed -- stopping the port-forward it started (pid $STARTED_PF_PID)"
kill "$STARTED_PF_PID" 2>/dev/null || true
fi
exit "$rc"
}
trap cleanup_on_failure EXIT
# ---------------------------------------------------------------------------
# steps
# ---------------------------------------------------------------------------
preflight() {
local missing=()
for bin in kubectl kind helm jq curl; do
command -v "$bin" >/dev/null 2>&1 || missing+=("$bin")
done
if [ "${#missing[@]}" -gt 0 ]; then
die "missing required tools on PATH: ${missing[*]}"
fi
}
ensure_kind_cluster() {
case "$(kind get clusters 2>/dev/null)" in
*"$CLUSTER_NAME"*)
log "kind cluster '$CLUSTER_NAME' already exists, skipping creation" ;;
*)
log "creating kind cluster '$CLUSTER_NAME'"
kind create cluster --name "$CLUSTER_NAME" ;;
esac
kubectl config use-context "kind-${CLUSTER_NAME}" >/dev/null
}
install_operator() {
log "installing terdut-operator into namespace $OPERATOR_NAMESPACE"
helm upgrade --install "$RELEASE_NAME" "$CHART_DIR" \
--namespace "$OPERATOR_NAMESPACE" --create-namespace \
--wait --timeout "$HELM_TIMEOUT" \
|| die "helm install of terdut-operator did not become ready"
}
apply_demo() {
log "creating namespace $NAMESPACE"
kubectl create namespace "$NAMESPACE" --dry-run=client -o yaml | kubectl apply -f - >/dev/null
log "applying demo CRs (00-09) into $NAMESPACE"
kubectl apply -n "$NAMESPACE" -k "$SCRIPT_DIR" >/dev/null
}
wait_for_objects() {
local obj
for obj in "$@"; do
log "waiting for $obj to become Ready"
kubectl wait --for=condition=Ready --timeout "$WAIT_TIMEOUT" -n "$NAMESPACE" "$obj" >/dev/null \
|| die "timed out waiting for $obj -- try: kubectl describe -n $NAMESPACE $obj"
done
}
# Just the server and the two teams -- everything redeem_platform_invite and
# join_payments_team need. Deliberately NOT the escalation rules here: this
# demo kit's own terdutescalationrule-platform names alice as a level-1
# target, and that CR cannot reach Ready until alice actually exists
# (terdut-server resolves every named username at reconcile time, not just
# at escalation time) -- a real dependency this script has to satisfy by
# creating her first, not something kubectl wait can be told to ignore.
wait_for_teams_ready() {
wait_for_objects \
"terdutserver/terdut-operator-demo" \
"terdutteam/terdutteam-platform" \
"terdutteam/terdutteam-payments"
}
# Everything that was waiting on alice (or just on the teams above, now
# already satisfied) to exist.
wait_for_remaining_ready() {
wait_for_objects \
"terdutescalationrule/terdutescalationrule-platform" \
"terdutescalationrule/terdutescalationrule-payments" \
"terdutdeadmanswitch/terdutdeadmanswitch-platform" \
"terdutdeadmanswitch/terdutdeadmanswitch-payments" \
"terdutalertsource/terdutalertsource-platform" \
"terdutalertsource/terdutalertsource-payments"
}
start_port_forward() {
# A stale pidfile from an earlier run would otherwise collide with us on
# $LOCAL_PORT -- if that pid is still alive, stop it first.
if [ -f "$PF_PIDFILE" ]; then
local old_pid
old_pid="$(cat "$PF_PIDFILE" 2>/dev/null || true)"
if [ -n "$old_pid" ] && kill -0 "$old_pid" 2>/dev/null; then
log "stopping stale port-forward from a previous run (pid $old_pid)"
kill "$old_pid" 2>/dev/null || true
sleep 1
fi
rm -f "$PF_PIDFILE"
fi
log "starting port-forward svc/terdut-operator-demo ${LOCAL_PORT}:8080"
kubectl -n "$NAMESPACE" port-forward svc/terdut-operator-demo "${LOCAL_PORT}:8080" \
>"$PF_LOGFILE" 2>&1 &
STARTED_PF_PID=$!
echo "$STARTED_PF_PID" > "$PF_PIDFILE"
local tries=0
until curl -sf -o /dev/null "http://localhost:${LOCAL_PORT}/healthz"; do
tries=$((tries + 1))
if [ "$tries" -ge 30 ]; then
die "port-forward never became ready -- see $PF_LOGFILE"
fi
sleep 1
done
log "port-forward ready (pid $STARTED_PF_PID, log $PF_LOGFILE)"
}
# Redeems Platform's own invite link -- minted by its TerdutTeam
# (02-team-platform.yaml's spec.invite.enabled, reconciled through that
# team's own already-working team-scoped credential, which requireTeamOwner
# already treats as owner-equivalent for /invites -- ratified, not a
# workaround, in terdut-server's SERVICE-ACCOUNTS.md) and redeemed through
# the ordinary signup endpoint. Invite redemption bypasses signup_mode
# entirely (terdut-server's internal/api/signup.go), so this needs no admin
# credential, no signup_mode flip, and no direct Postgres access at all --
# unlike an earlier version of this script, before terdut-operator grew
# this feature (see niklas/terdut-server#23).
redeem_platform_invite() {
log "reading Platform's invite link"
local secret_name invite_url invite_token tries=0
until secret_name="$(kubectl -n "$NAMESPACE" get terdutteam terdutteam-platform \
-o jsonpath='{.status.inviteSecretRef.name}' 2>/dev/null)" && [ -n "$secret_name" ]; do
tries=$((tries + 1))
[ "$tries" -lt 30 ] || die "terdutteam-platform never reported status.inviteSecretRef -- check spec.invite.enabled and kubectl describe it"
sleep 1
done
invite_url="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.url}' | base64 -d)"
invite_token="${invite_url##*invite=}"
[ -n "$invite_token" ] || die "could not parse an invite token out of $invite_url"
# A re-run: the invite was spent by the first run, and the server answers a
# spent invite with 403 before it ever looks at the username, so the 409
# handled below never arrives. If alice can already sign in, she exists.
local login_code
login_code="$(curl -sS -o /dev/null -w '%{http_code}' \
-X POST "${BASE_URL}/api/login" -H 'Content-Type: application/json' \
-d "$(jq -n --arg u "$ALICE_USERNAME" --arg p "$DEMO_PASSWORD" '{username: $u, password: $p}')")"
if [ "$login_code" = "200" ]; then
log "account ${ALICE_USERNAME} already exists and can sign in, skipping signup (re-run detected)"
return 0
fi
log "signing up ${ALICE_USERNAME} via Platform's invite"
local body resp_file code
body="$(jq -n \
--arg u "$ALICE_USERNAME" --arg e "$ALICE_EMAIL" \
--arg p "$DEMO_PASSWORD" --arg i "$invite_token" \
'{username: $u, email: $e, password: $p, invite: $i}')"
resp_file="$(mktemp)"
code="$(curl -sS -o "$resp_file" -w '%{http_code}' \
-X POST "${BASE_URL}/api/signup" -H 'Content-Type: application/json' -d "$body")"
case "$code" in
201) log "created local account ${ALICE_USERNAME} (joined Platform)" ;;
409) log "account ${ALICE_USERNAME} already exists, skipping (re-run detected)" ;;
*) die "signup failed (HTTP $code): $(cat "$resp_file")" ;;
esac
rm -f "$resp_file"
}
# redeem_platform_invite's signup already used up alice's one signup -- a
# second POST /api/signup would just 409 on the taken username, it doesn't
# join an existing account to another team. Payments is joined through the
# ordinary team-scoped member endpoint instead, using that team's own
# already-working team-scoped credential (owner-equivalent, same reach the
# invite route above relies on) and alice's user id resolved via
# GET /api/users -- the same lookup tdclient.GetUserByUsername does
# operator-side. The endpoint upserts on (team_id, user_id), so this is
# naturally idempotent across re-runs with no separate conflict handling
# needed.
join_payments_team() {
log "adding ${ALICE_USERNAME} to Payments"
local payments_id payments_secret payments_key alice_id resp_file code
payments_id="$(kubectl -n "$NAMESPACE" get terdutteam terdutteam-payments -o jsonpath='{.status.teamID}')"
[ -n "$payments_id" ] || die "could not read status.teamID off terdutteam-payments"
payments_secret="$(kubectl -n "$NAMESPACE" get terdutteam terdutteam-payments \
-o jsonpath='{.status.credentialsSecretRef.name}')"
[ -n "$payments_secret" ] || die "terdutteam-payments has no status.credentialsSecretRef yet"
payments_key="$(kubectl -n "$OPERATOR_NAMESPACE" get secret "$payments_secret" -o jsonpath='{.data.token}' | base64 -d)"
alice_id="$(curl -sS -H "Authorization: Bearer $payments_key" "${BASE_URL}/api/users" \
| jq -r --arg u "$ALICE_USERNAME" '.[] | select(.username == $u) | .id')"
[ -n "$alice_id" ] || die "could not resolve ${ALICE_USERNAME}'s user id via GET /api/users"
resp_file="$(mktemp)"
code="$(curl -sS -o "$resp_file" -w '%{http_code}' \
-X POST "${BASE_URL}/api/teams/${payments_id}/members" \
-H "Authorization: Bearer $payments_key" -H 'Content-Type: application/json' \
-d "$(jq -n --argjson id "$alice_id" '{user_id: $id, role: "member"}')")"
[ "$code" = "204" ] || die "adding ${ALICE_USERNAME} to Payments failed (HTTP $code): $(cat "$resp_file")"
log "added ${ALICE_USERNAME} to Payments"
rm -f "$resp_file"
}
fire_demo_alerts() {
log "firing representative demo alerts"
# Two clusters, so the queue shows the cluster chip and offers its filter.
# high-cpu fires in both: the same alert in two clusters is two incidents.
local fire="$SCRIPT_DIR/fire-alerts.sh"
CLUSTER=prod-eu NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform high-cpu
CLUSTER=prod-us NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform high-cpu
CLUSTER=prod-us NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" platform disk-full
CLUSTER=prod-eu NAMESPACE="$NAMESPACE" BASE_URL="$BASE_URL" "$fire" payments pod-crash
}
print_summary() {
cat <<EOF
terdut demo is up.
Web UI: http://localhost:${LOCAL_PORT}
Login: ${ALICE_USERNAME} / ${DEMO_PASSWORD}
Cluster: kind-${CLUSTER_NAME}
Namespace: ${NAMESPACE}
Port-forward pid: ${STARTED_PF_PID:-$(cat "$PF_PIDFILE" 2>/dev/null || echo unknown)} (log: ${PF_LOGFILE})
stop it with: kill \$(cat ${PF_PIDFILE})
Fire more alerts:
export NAMESPACE=${NAMESPACE} BASE_URL=${BASE_URL}
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu
CLUSTER=prod-eu ./fire-alerts.sh platform high-cpu resolve
./fire-alerts.sh payments heartbeat # send repeatedly (e.g. every minute)
# to keep a dead man's switch alive;
# stop sending it and, 15 minutes
# later, terdut-server opens a
# critical incident on its own.
Tear down:
$0 --teardown
# add TEARDOWN_CLUSTER=true to also delete the kind cluster itself
EOF
}
teardown() {
if [ -f "$PF_PIDFILE" ]; then
local pid
pid="$(cat "$PF_PIDFILE" 2>/dev/null || true)"
if [ -n "$pid" ] && kill -0 "$pid" 2>/dev/null; then
log "stopping port-forward (pid $pid)"
kill "$pid" 2>/dev/null || true
fi
rm -f "$PF_PIDFILE"
fi
log "deleting namespace $NAMESPACE"
kubectl delete namespace "$NAMESPACE" --ignore-not-found --wait=true --timeout "$WAIT_TIMEOUT"
if [ "$TEARDOWN_CLUSTER" = "true" ]; then
log "deleting kind cluster $CLUSTER_NAME"
kind delete cluster --name "$CLUSTER_NAME"
else
log "leaving kind cluster '$CLUSTER_NAME' and the operator install in place" \
"(set TEARDOWN_CLUSTER=true to also delete the cluster)"
fi
}
# ---------------------------------------------------------------------------
# main
# ---------------------------------------------------------------------------
main() {
case "${1:-}" in
--teardown)
preflight
teardown
trap - EXIT
exit 0
;;
--help|-h)
usage
trap - EXIT
exit 0
;;
"") ;;
*)
usage
die "unknown argument: $1"
;;
esac
preflight
ensure_kind_cluster
install_operator
apply_demo
wait_for_teams_ready
start_port_forward
redeem_platform_invite
join_payments_team
wait_for_remaining_ready
fire_demo_alerts
print_summary
# Success: leave the port-forward running, don't let the EXIT trap kill it.
trap - EXIT
}
main "$@"
+31
View File
@@ -0,0 +1,31 @@
# Exposing a TerdutServer
This operator never manages external exposure/ingress for `TerdutServer`,
in any form — a permanent non-goal (`DESIGN.md` §1), not a missing
feature. Some installs won't expose it outside the cluster at all (see
`examples/demo`, which just port-forwards); others will put it behind
whatever their cluster already uses. That choice is entirely yours, not
the operator's.
The only contract the operator gives you to build on: a plain `ClusterIP`
Service, named after the `TerdutServer` CR (same name, same namespace),
with a port named `http` (`spec.networking.servicePort`, default `8080`).
Everything here targets exactly that Service — none of it is applied by
`examples/demo`'s `kustomization.yaml`, and none of it depends on anything
the operator creates beyond that one Service.
Pick whichever matches your cluster:
- **`httproute.yaml`** — a [Gateway API](https://gateway-api.sigs.k8s.io/)
`HTTPRoute`, attached to a `Gateway` you already have.
- **`istio-virtualservice.yaml`** — an Istio `VirtualService`, attached to
a `Gateway` (Istio's own CRD, not Gateway API's) you already have.
- **A plain `Ingress`** needs no example here — it's the same idea with
one fewer layer of indirection: an `Ingress` with a single rule whose
`backend.service.name`/`port.name` point at the `TerdutServer`'s Service
and `http` port.
Remember to set `spec.networking.hostname` on the `TerdutServer` itself to
whatever hostname you expose it on — that's not read by the operator for
any of this, but terdut-server uses it for its own absolute links
(notifications, OIDC redirect URIs).
+21
View File
@@ -0,0 +1,21 @@
# Gateway API HTTPRoute exposing a TerdutServer through a Gateway you
# already have (not something this operator creates or watches -- see
# ../networking/README.md). Replace terdut-operator-demo and the Gateway
# reference/hostname with your own; terdut-operator-demo matches
# ../demo/01-server.yaml, if you're layering this onto that demo.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: terdut-operator-demo
namespace: terdut-operator-demo
spec:
parentRefs:
- name: my-gateway # an existing Gateway in this namespace (or
# namespace: gateway-ns # a different one, if the Gateway allows it)
sectionName: https # the listener to attach to, if it's picky
hostnames:
- terdut-operator-demo.example # matches networking.hostname on the CR
rules:
- backendRefs:
- name: terdut-operator-demo # the Service the operator created --
port: 8080 # same name as the TerdutServer CR
@@ -0,0 +1,23 @@
# Istio VirtualService exposing a TerdutServer through an Istio Gateway
# you already have (Istio's own Gateway CRD, not Gateway API's -- not
# something this operator creates or watches, see ../networking/README.md).
# Replace terdut-operator-demo and the gateway reference/hostname with your
# own; terdut-operator-demo matches ../demo/01-server.yaml, if you're
# layering this onto that demo.
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: terdut-operator-demo
namespace: terdut-operator-demo
spec:
hosts:
- terdut-operator-demo.example # matches networking.hostname on the CR
gateways:
- my-gateway-namespace/my-gateway # an existing istio Gateway
http:
- route:
- destination:
host: terdut-operator-demo.terdut-operator-demo.svc.cluster.local
port:
number: 8080 # the Service's port -- same name as the
# TerdutServer CR, default servicePort
+75 -10
View File
@@ -16,18 +16,24 @@ import (
) )
// bootstrapStateLostError is DESIGN.md §6's one genuinely pathological // bootstrapStateLostError is DESIGN.md §6's one genuinely pathological
// case: a checkpointed admin credential was used and then lost before the // case: the server's database is already bootstrapped, and no credential for
// lasting credential it was for could be persisted. Distinct from a plain // it survives here -- a checkpointed admin key was used and then lost before
// error so Reconcile can route it to a Ready: False condition (the // the lasting credential could be persisted, or the instance credential
// documented recovery is delete-and-recreate, not an automatic retry) rather // Secret was deleted (spec.credentials.deletionPolicy: Delete, or by hand).
// than treating it as a transient reconcile failure. // Distinct from a plain error so Reconcile can route it to a Ready: False
type bootstrapStateLostError struct{ detail string } // condition rather than treating it as a transient reconcile failure: no
// retry can fix it.
type bootstrapStateLostError struct {
detail string
secretName string // the instance credential Secret that would have fixed it
}
func (e *bootstrapStateLostError) Error() string { func (e *bootstrapStateLostError) Error() string {
return fmt.Sprintf( return fmt.Sprintf(
"server reports already bootstrapped, but neither status.credentialsSecretRef nor a "+ "server reports already bootstrapped, but there is no credential for it here: %s. "+
"checkpointed admin credential exist here: %s. This TerdutServer cannot recover a "+ "The operator cannot mint one on its own, and deleting and recreating this TerdutServer does "+
"credential on its own; delete and recreate it", e.detail) "not clear the database. Restore Secret %q in the operator's namespace if you have a copy, "+
"or reset the server's database and recreate this TerdutServer", e.detail, e.secretName)
} }
// reconcileBootstrap implements DESIGN.md §6 point 1's self-registration // reconcileBootstrap implements DESIGN.md §6 point 1's self-registration
@@ -36,6 +42,17 @@ func (e *bootstrapStateLostError) Error() string {
// srv.Status.CredentialsSecretRef is nil and the Deployment has a ready // srv.Status.CredentialsSecretRef is nil and the Deployment has a ready
// replica. // replica.
func (r *TerdutServerReconciler) reconcileBootstrap(ctx context.Context, srv *terdutv1alpha1.TerdutServer) error { func (r *TerdutServerReconciler) reconcileBootstrap(ctx context.Context, srv *terdutv1alpha1.TerdutServer) error {
// A TerdutServer recreated against a database that is already
// bootstrapped: its instance credential may still be here (the default
// deletion policy keeps it), and then there is nothing to bootstrap.
adopted, err := r.adoptRetainedCredentials(ctx, srv)
if err != nil {
return err
}
if adopted {
return nil
}
adminKey, err := r.getOrCreateCheckpointedAdminKey(ctx, srv) adminKey, err := r.getOrCreateCheckpointedAdminKey(ctx, srv)
if err != nil { if err != nil {
return err return err
@@ -64,6 +81,51 @@ func (r *TerdutServerReconciler) reconcileBootstrap(ctx context.Context, srv *te
return nil return nil
} }
// adoptRetainedCredentials looks for the instance credential Secret an
// earlier TerdutServer of the same name and namespace left behind
// (spec.credentials.deletionPolicy: Retain), and adopts it if the server
// still accepts its token. Reports whether it did.
//
// The token is tried before it is trusted: a Secret that outlived a database
// reset holds a key the server has never heard of, and adopting that would
// make every later call 401. A rejected token (401/403) is not an error here;
// it just means there is nothing to adopt, and bootstrap proceeds as it would
// for a first install -- which succeeds against a freshly reset database and
// replaces the Secret. Anything else (the server unreachable, a 5xx) is
// returned for a retry rather than guessed at.
func (r *TerdutServerReconciler) adoptRetainedCredentials(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (bool, error) {
name := credentialsSecretName(srv)
var sec corev1.Secret
if err := r.Get(ctx, client.ObjectKey{Namespace: r.OperatorNamespace, Name: name}, &sec); err != nil {
if apierrors.IsNotFound(err) {
return false, nil
}
return false, err
}
token := string(sec.Data[credentialsSecretDataKey])
if token == "" {
return false, nil
}
// The operator's own service account is the one thing this credential
// is for, so asking the server for it by name checks the token and that
// it belongs to that account in the same call.
sa, err := r.NewClient(serviceURL(srv)).WithToken(token).GetServiceAccountByName(ctx, serviceAccountName)
if err != nil {
if statusErr, ok := errors.AsType[*tdclient.StatusError](err); ok &&
(statusErr.Code == http.StatusUnauthorized || statusErr.Code == http.StatusForbidden) {
return false, nil
}
return false, fmt.Errorf("checking the retained credential in Secret %q: %w", name, err)
}
if sa == nil {
return false, nil
}
srv.Status.CredentialsSecretRef = &terdutv1alpha1.SecretKeyRef{Name: name, Key: credentialsSecretDataKey}
return true, nil
}
// getOrCreateCheckpointedAdminKey returns a usable admin key: from the // getOrCreateCheckpointedAdminKey returns a usable admin key: from the
// checkpoint Secret if an earlier, interrupted attempt already got one, or // checkpoint Secret if an earlier, interrupted attempt already got one, or
// freshly from /api/bootstrap, immediately checkpointed before it's used // freshly from /api/bootstrap, immediately checkpointed before it's used
@@ -89,7 +151,10 @@ func (r *TerdutServerReconciler) getOrCreateCheckpointedAdminKey(ctx context.Con
// status.credentialsSecretRef) means a prior reconcile already // status.credentialsSecretRef) means a prior reconcile already
// won this exact race and its checkpoint was lost afterward -- // won this exact race and its checkpoint was lost afterward --
// the one case §6 doesn't try to paper over. // the one case §6 doesn't try to paper over.
return "", &bootstrapStateLostError{detail: "/api/bootstrap returned 403"} return "", &bootstrapStateLostError{
detail: "/api/bootstrap returned 403",
secretName: credentialsSecretName(srv),
}
} }
return "", fmt.Errorf("POST /api/bootstrap: %w", err) return "", fmt.Errorf("POST /api/bootstrap: %w", err)
} }
+29 -10
View File
@@ -8,6 +8,7 @@ import (
appsv1 "k8s.io/api/apps/v1" appsv1 "k8s.io/api/apps/v1"
corev1 "k8s.io/api/core/v1" corev1 "k8s.io/api/core/v1"
policyv1 "k8s.io/api/policy/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors" apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta" "k8s.io/apimachinery/pkg/api/meta"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
@@ -38,10 +39,10 @@ const resyncInterval = 5 * time.Minute
// than "someone edited something out of band." // than "someone edited something out of band."
const waitInterval = 15 * time.Second const waitInterval = 15 * time.Second
// finalizerName cleans up the credentials Secret(s) this controller // finalizerName cleans up the Secret(s) this controller generates in the
// generates in the operator's own namespace on delete — the Deployment and // operator's own namespace on delete (which of them, spec.credentials.
// Service are owned (OwnerReference, DESIGN.md §7) and need no finalizer of // deletionPolicy decides) — the Deployment and Service are owned
// their own. // (OwnerReference, DESIGN.md §7) and need no finalizer of their own.
const finalizerName = "terdut.ryuvia.com/terdutserver" const finalizerName = "terdut.ryuvia.com/terdutserver"
// serviceAccountName is the name the operator registers itself under // serviceAccountName is the name the operator registers itself under
@@ -97,6 +98,7 @@ type TerdutServerReconciler struct {
// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups="",resources=services,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups="",resources=services,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=policy,resources=poddisruptionbudgets,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=acid.zalan.do,resources=postgresqls,verbs=get;list;watch // +kubebuilder:rbac:groups=acid.zalan.do,resources=postgresqls,verbs=get;list;watch
// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch // +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch
@@ -137,6 +139,9 @@ func (r *TerdutServerReconciler) Reconcile(ctx context.Context, req ctrl.Request
if err := r.reconcileService(ctx, &srv); err != nil { if err := r.reconcileService(ctx, &srv); err != nil {
return ctrl.Result{}, err return ctrl.Result{}, err
} }
if err := r.reconcilePodDisruptionBudget(ctx, &srv); err != nil {
return ctrl.Result{}, err
}
meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{ meta.SetStatusCondition(&srv.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionDatabaseReady, Type: terdutv1alpha1.ConditionDatabaseReady,
@@ -210,18 +215,31 @@ func (r *TerdutServerReconciler) setNotReady(
return ctrl.Result{RequeueAfter: d}, nil return ctrl.Result{RequeueAfter: d}, nil
} }
// reconcileDelete cleans up the credentials Secret(s) this controller // reconcileDelete cleans up the Secrets this controller generated in the
// generated in the operator's own namespace. The Deployment and Service are // operator's own namespace. The Deployment and Service are owned
// owned (OwnerReference, DESIGN.md §7) and need no attention here — normal // (OwnerReference, DESIGN.md §7) and need no attention here — normal GC
// GC handles them. There is no server-side "delete this install" call to // handles them. There is no server-side "delete this install" call to
// make: bootstrap created a user and a service account, and terdut-server's // make: bootstrap created a user and a service account, and terdut-server's
// API has no way to delete either (only to revoke individual keys), so // API has no way to delete either (only to revoke individual keys), so
// there is nothing meaningful to undo there either. // there is nothing meaningful to undo there either, and the database is
// never touched.
//
// That is why the instance credential is kept by default
// (spec.credentials.deletionPolicy: Retain). Everything it logs in to
// outlives the TerdutServer, so a recreated one finds a bootstrapped server
// it has no key for -- unless the key is still here to be adopted (see
// adoptRetainedCredentials). The bootstrap checkpoint is always removed: it
// is a short-lived admin key, and the instance credential is all that is
// needed afterwards.
func (r *TerdutServerReconciler) reconcileDelete(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (ctrl.Result, error) { func (r *TerdutServerReconciler) reconcileDelete(ctx context.Context, srv *terdutv1alpha1.TerdutServer) (ctrl.Result, error) {
if !controllerutil.ContainsFinalizer(srv, finalizerName) { if !controllerutil.ContainsFinalizer(srv, finalizerName) {
return ctrl.Result{}, nil return ctrl.Result{}, nil
} }
for _, name := range []string{checkpointSecretName(srv), credentialsSecretName(srv)} { names := []string{checkpointSecretName(srv)}
if srv.Spec.Credentials.DeletionPolicy == terdutv1alpha1.CredentialsDelete {
names = append(names, credentialsSecretName(srv))
}
for _, name := range names {
sec := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: r.OperatorNamespace}} sec := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: r.OperatorNamespace}}
if err := r.Delete(ctx, sec); err != nil && !apierrors.IsNotFound(err) { if err := r.Delete(ctx, sec); err != nil && !apierrors.IsNotFound(err) {
return ctrl.Result{}, err return ctrl.Result{}, err
@@ -246,6 +264,7 @@ func (r *TerdutServerReconciler) SetupWithManager(mgr ctrl.Manager) error {
For(&terdutv1alpha1.TerdutServer{}). For(&terdutv1alpha1.TerdutServer{}).
Owns(&appsv1.Deployment{}). Owns(&appsv1.Deployment{}).
Owns(&corev1.Service{}). Owns(&corev1.Service{}).
Owns(&policyv1.PodDisruptionBudget{}).
Named("terdutserver"). Named("terdutserver").
Complete(r) Complete(r)
} }
@@ -9,15 +9,20 @@ import (
"strconv" "strconv"
"strings" "strings"
"sync" "sync"
"time"
. "github.com/onsi/ginkgo/v2" . "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega" . "github.com/onsi/gomega"
appsv1 "k8s.io/api/apps/v1" appsv1 "k8s.io/api/apps/v1"
corev1 "k8s.io/api/core/v1" corev1 "k8s.io/api/core/v1"
policyv1 "k8s.io/api/policy/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
"k8s.io/apimachinery/pkg/api/meta" "k8s.io/apimachinery/pkg/api/meta"
"k8s.io/apimachinery/pkg/api/resource"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
"k8s.io/apimachinery/pkg/types" "k8s.io/apimachinery/pkg/types"
"k8s.io/apimachinery/pkg/util/intstr"
ctrl "sigs.k8s.io/controller-runtime" ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/reconcile" "sigs.k8s.io/controller-runtime/pkg/reconcile"
@@ -31,6 +36,11 @@ import (
const ( const (
fakeVersionString = "test" fakeVersionString = "test"
errJSONKey = "error" errJSONKey = "error"
// deadmanSwitchesPath is the literal path (not Printf'd like the others
// below) shared by the exact-match collection route and the dispatcher
// that routes into it -- goconst flags three occurrences of the same
// string, so this is that string, once.
deadmanSwitchesPath = "/deadman/switches"
) )
// fakeTerdutServer reproduces the exact stateful semantics of // fakeTerdutServer reproduces the exact stateful semantics of
@@ -44,9 +54,13 @@ type fakeTerdutServer struct {
mu sync.Mutex mu sync.Mutex
bootstrapped bool bootstrapped bool
bootstrap403 bool // force every /api/bootstrap call to 403, even the first bootstrap403 bool // force every /api/bootstrap call to 403, even the first
nextID int64 // rejectedTokens: bearer tokens the fake answers 401 on GET
accounts map[string]int64 // name -> id // /api/service-accounts, the way a server that never issued the key
keyMints map[int64]int // id -> number of keys minted so far // (a database reset since) would.
rejectedTokens map[string]bool
nextID int64
accounts map[string]int64 // name -> id
keyMints map[int64]int // id -> number of keys minted so far
nextTeamID int64 nextTeamID int64
teams map[string]int64 // name -> id teams map[string]int64 // name -> id
@@ -79,6 +93,14 @@ type fakeTerdutServer struct {
nextIntegrationID int64 nextIntegrationID int64
integrations map[int64]map[int64]tdclient.Integration // teamID -> integrationID -> integration integrations map[int64]map[int64]tdclient.Integration // teamID -> integrationID -> integration
integrationDelete map[int64]bool // integrationID -> true once DELETEd, for 404-on-redelete integrationDelete map[int64]bool // integrationID -> true once DELETEd, for 404-on-redelete
// invites/nextInviteID/inviteDelete back TerdutTeam's own invite-minting
// feature -- no unique constraint on an invite server-side either (every
// POST mints a brand new row, confirmed against source), same keyed-by-id
// shape as switches/integrations.
nextInviteID int64
invites map[int64]map[int64]tdclient.Invite // teamID -> inviteID -> invite
inviteDelete map[int64]bool // inviteID -> true once DELETEd, for 404-on-redelete
} }
func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) { func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) {
@@ -96,6 +118,9 @@ func newFakeTerdutServer() (*fakeTerdutServer, *httptest.Server) {
integrations: map[int64]map[int64]tdclient.Integration{}, integrations: map[int64]map[int64]tdclient.Integration{},
integrationDelete: map[int64]bool{}, integrationDelete: map[int64]bool{},
invites: map[int64]map[int64]tdclient.Invite{},
inviteDelete: map[int64]bool{},
} }
return f, httptest.NewServer(f) return f, httptest.NewServer(f)
} }
@@ -149,6 +174,10 @@ func (f *fakeTerdutServer) ServeHTTP(w http.ResponseWriter, r *http.Request) {
}) })
case r.URL.Path == "/api/service-accounts" && r.Method == http.MethodGet: case r.URL.Path == "/api/service-accounts" && r.Method == http.MethodGet:
if f.rejectedTokens[strings.TrimPrefix(r.Header.Get("Authorization"), "Bearer ")] {
writeJSON(w, http.StatusUnauthorized, map[string]string{errJSONKey: "invalid or expired API key"})
return
}
name := r.URL.Query().Get("name") name := r.URL.Query().Get("name")
id, exists := f.accounts[name] id, exists := f.accounts[name]
if !exists { if !exists {
@@ -261,7 +290,26 @@ func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Requ
f.escalation[id] = req f.escalation[id] = req
w.WriteHeader(http.StatusNoContent) w.WriteHeader(http.StatusNoContent)
case rest == "/deadman/switches" && r.Method == http.MethodGet: case rest == deadmanSwitchesPath || strings.HasPrefix(rest, deadmanSwitchesPath+"/"):
f.handleDeadmanSubPath(w, r, id, rest)
case rest == "/integrations" || strings.HasPrefix(rest, "/integrations/"):
f.handleIntegrationSubPath(w, r, id, rest)
case rest == "/invites" || strings.HasPrefix(rest, "/invites/"):
f.handleInviteSubPath(w, r, id, rest)
default:
w.WriteHeader(http.StatusNotFound)
}
}
// handleDeadmanSubPath answers GET/POST /api/teams/{id}/deadman/switches and
// PUT/DELETE .../deadman/switches/{switchID} -- split out of
// handleTeamSubPath for the same gocyclo reason as handleIntegrationSubPath.
func (f *fakeTerdutServer) handleDeadmanSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == deadmanSwitchesPath && r.Method == http.MethodGet:
existing := f.switches[id] existing := f.switches[id]
out := make([]tdclient.DeadmanSwitch, 0, len(existing)) out := make([]tdclient.DeadmanSwitch, 0, len(existing))
for _, s := range existing { for _, s := range existing {
@@ -269,7 +317,7 @@ func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Requ
} }
writeJSON(w, http.StatusOK, out) writeJSON(w, http.StatusOK, out)
case rest == "/deadman/switches" && r.Method == http.MethodPost: case rest == deadmanSwitchesPath && r.Method == http.MethodPost:
var req deadmanSwitchFakeRequest var req deadmanSwitchFakeRequest
_ = json.NewDecoder(r.Body).Decode(&req) _ = json.NewDecoder(r.Body).Decode(&req)
f.nextSwitchID++ f.nextSwitchID++
@@ -324,8 +372,48 @@ func (f *fakeTerdutServer) handleTeamSubPath(w http.ResponseWriter, r *http.Requ
f.switchDelete[switchID] = true f.switchDelete[switchID] = true
w.WriteHeader(http.StatusNoContent) w.WriteHeader(http.StatusNoContent)
case rest == "/integrations" || strings.HasPrefix(rest, "/integrations/"): default:
f.handleIntegrationSubPath(w, r, id, rest) w.WriteHeader(http.StatusNotFound)
}
}
// handleInviteSubPath answers POST /api/teams/{id}/invites and
// DELETE .../invites/{inviteID} -- split out for the same gocyclo reason as
// handleIntegrationSubPath.
func (f *fakeTerdutServer) handleInviteSubPath(w http.ResponseWriter, r *http.Request, id int64, rest string) {
switch {
case rest == "/invites" && r.Method == http.MethodPost:
var req struct {
Role string `json:"role"`
MaxUses int64 `json:"max_uses"`
}
_ = json.NewDecoder(r.Body).Decode(&req)
f.nextInviteID++
inviteID := f.nextInviteID
inv := tdclient.Invite{
ID: inviteID, TeamID: id, Role: req.Role, MaxUses: req.MaxUses,
ExpiresAt: time.Now().Add(7 * 24 * time.Hour),
URL: fmt.Sprintf("https://terdut.example.invalid/signup?invite=invite-token-%d", inviteID),
}
if f.invites[id] == nil {
f.invites[id] = map[int64]tdclient.Invite{}
}
f.invites[id][inviteID] = inv
writeJSON(w, http.StatusCreated, inv)
case strings.HasPrefix(rest, "/invites/") && r.Method == http.MethodDelete:
inviteID, ok := parseTrailingID(rest, "/invites/")
if !ok {
w.WriteHeader(http.StatusNotFound)
return
}
if _, exists := f.invites[id][inviteID]; !exists {
w.WriteHeader(http.StatusNotFound)
return
}
delete(f.invites[id], inviteID)
f.inviteDelete[inviteID] = true
w.WriteHeader(http.StatusNoContent)
default: default:
w.WriteHeader(http.StatusNotFound) w.WriteHeader(http.StatusNotFound)
@@ -705,14 +793,104 @@ var _ = Describe("TerdutServer Controller", func() {
}) })
}) })
Describe("deletion", func() { Describe("spec.pod", func() {
It("removes the credentials and checkpoint Secrets and the finalizer", func(ctx SpecContext) { It("wires pod-level customization onto the right spot on the Deployment", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer() fake, fakeSrv := newFakeTerdutServer()
_ = fake _ = fake
DeferCleanup(fakeSrv.Close) DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) } reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, dsnSpec()) spec := dsnSpec()
qty := resource.MustParse("250m")
spec.Pod = terdutv1alpha1.PodSpec{
Resources: corev1.ResourceRequirements{Requests: corev1.ResourceList{corev1.ResourceCPU: qty}},
Tolerations: []corev1.Toleration{{Key: "dedicated", Operator: corev1.TolerationOpEqual, Value: "terdut", Effect: corev1.TaintEffectNoSchedule}},
ExtraEnv: []corev1.EnvVar{{Name: "EXTRA_FLAG", Value: "on"}},
ServiceAccountName: "terdut-server-custom",
ExtraVolumes: []corev1.Volume{{Name: "extra-ca", VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}},
ExtraVolumeMounts: []corev1.VolumeMount{{Name: "extra-ca", MountPath: "/etc/extra-ca"}},
}
createServer(ctx, spec)
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service/PDB
var deploy appsv1.Deployment
Expect(k8sClient.Get(ctx, objKey, &deploy)).To(Succeed())
podSpec := deploy.Spec.Template.Spec
Expect(podSpec.Tolerations).To(ConsistOf(spec.Pod.Tolerations))
Expect(podSpec.ServiceAccountName).To(Equal("terdut-server-custom"))
Expect(podSpec.Volumes).To(ConsistOf(spec.Pod.ExtraVolumes))
main := podSpec.Containers[0]
Expect(main.Name).To(Equal("terdut-server"))
Expect(main.Resources).To(Equal(spec.Pod.Resources))
Expect(main.VolumeMounts).To(ConsistOf(spec.Pod.ExtraVolumeMounts))
Expect(main.Env).To(ContainElement(corev1.EnvVar{Name: "EXTRA_FLAG", Value: "on"}))
initContainer := podSpec.InitContainers[0]
Expect(initContainer.Name).To(Equal("wait-for-postgres"))
Expect(initContainer.VolumeMounts).To(BeEmpty(), "extraVolumeMounts must not leak onto wait-for-postgres")
})
})
Describe("spec.pod.disruptionBudget", func() {
It("creates an owned PodDisruptionBudget when set, and deletes it once cleared", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
_ = fake
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
spec := dsnSpec()
minAvail := intstr.FromInt32(1)
spec.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{MinAvailable: &minAvail}
createServer(ctx, spec)
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service/PDB
var pdb policyv1.PodDisruptionBudget
Expect(k8sClient.Get(ctx, objKey, &pdb)).To(Succeed())
Expect(pdb.Spec.Selector.MatchLabels).To(Equal(labelsFor(&terdutv1alpha1.TerdutServer{ObjectMeta: metav1.ObjectMeta{Name: name}})))
Expect(pdb.Spec.MinAvailable).To(Equal(&minAvail))
Expect(pdb.OwnerReferences).To(ContainElement(HaveField("Name", name)))
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
srv.Spec.Pod.DisruptionBudget = nil
Expect(k8sClient.Update(ctx, srv)).To(Succeed())
reconcileOnce(ctx)
err := k8sClient.Get(ctx, objKey, &policyv1.PodDisruptionBudget{})
Expect(apierrors.IsNotFound(err)).To(BeTrue(), "PodDisruptionBudget should be deleted once spec.pod.disruptionBudget is cleared")
})
It("rejects both minAvailable and maxUnavailable set together, and neither set", func(ctx SpecContext) {
bothSet := dsnSpec()
minAvail, maxUnavail := intstr.FromInt32(1), intstr.FromInt32(1)
bothSet.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{MinAvailable: &minAvail, MaxUnavailable: &maxUnavail}
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: name + "-both", Namespace: operatorNamespace},
Spec: bothSet,
})).To(HaveOccurred())
neitherSet := dsnSpec()
neitherSet.Pod.DisruptionBudget = &terdutv1alpha1.PodDisruptionBudgetSpec{}
Expect(k8sClient.Create(ctx, &terdutv1alpha1.TerdutServer{
ObjectMeta: metav1.ObjectMeta{Name: name + "-neither", Namespace: operatorNamespace},
Spec: neitherSet,
})).To(HaveOccurred())
})
})
Describe("deletion", func() {
// runToReady brings a server to Ready against a fresh fake and
// returns the name of the instance credential Secret it minted.
runToReady := func(ctx SpecContext, spec terdutv1alpha1.TerdutServerSpec) string {
_, fakeSrv := newFakeTerdutServer()
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, spec)
reconcileOnce(ctx) reconcileOnce(ctx)
reconcileOnce(ctx) reconcileOnce(ctx)
markDeploymentReady(ctx) markDeploymentReady(ctx)
@@ -720,17 +898,114 @@ var _ = Describe("TerdutServer Controller", func() {
srv := &terdutv1alpha1.TerdutServer{} srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed()) Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
credsName := srv.Status.CredentialsSecretRef.Name return srv.Status.CredentialsSecretRef.Name
}
deleteAndFinalize := func(ctx SpecContext) {
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(k8sClient.Delete(ctx, srv)).To(Succeed()) Expect(k8sClient.Delete(ctx, srv)).To(Succeed())
reconcileOnce(ctx) // runs the finalizer reconcileOnce(ctx) // runs the finalizer
Expect(k8sClient.Get(ctx, objKey, srv)).NotTo(Succeed(),
"the TerdutServer itself should be gone once the finalizer clears")
}
secretExists := func(ctx SpecContext, secretName string) bool {
var s corev1.Secret
err := k8sClient.Get(ctx, types.NamespacedName{Name: secretName, Namespace: operatorNamespace}, &s)
return err == nil
}
err := k8sClient.Get(ctx, objKey, srv) It("keeps the instance credential by default and removes the checkpoint", func(ctx SpecContext) {
Expect(err).To(HaveOccurred(), "the TerdutServer itself should be gone once the finalizer clears") credsName := runToReady(ctx, dsnSpec())
// A checkpoint left behind, as if the best-effort delete had failed.
Expect(k8sClient.Create(ctx, &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: checkpointSecretNameFor(name), Namespace: operatorNamespace},
Data: map[string][]byte{credentialsSecretDataKey: []byte("admin-key-raw")},
})).To(Succeed())
var leftover corev1.Secret deleteAndFinalize(ctx)
err = k8sClient.Get(ctx, types.NamespacedName{Name: credsName, Namespace: operatorNamespace}, &leftover)
Expect(err).To(HaveOccurred(), "the credentials Secret should have been cleaned up by the finalizer") Expect(secretExists(ctx, credsName)).To(BeTrue(),
"the credential is kept so a recreated TerdutServer can adopt it")
Expect(secretExists(ctx, checkpointSecretNameFor(name))).To(BeFalse(),
"the checkpoint is a short-lived admin key and is always removed")
})
It("removes the credentials and checkpoint Secrets and the finalizer when the policy is Delete", func(ctx SpecContext) {
spec := dsnSpec()
spec.Credentials.DeletionPolicy = terdutv1alpha1.CredentialsDelete
credsName := runToReady(ctx, spec)
deleteAndFinalize(ctx)
Expect(secretExists(ctx, credsName)).To(BeFalse(),
"the credentials Secret should have been cleaned up by the finalizer")
})
})
Describe("recreating a TerdutServer against an already-bootstrapped server", func() {
// retainedSecret stands in for what the previous TerdutServer left.
retainedSecret := func(ctx SpecContext, token string) {
Expect(k8sClient.Create(ctx, &corev1.Secret{
ObjectMeta: metav1.ObjectMeta{Name: credentialsSecretNameFor(name), Namespace: operatorNamespace},
Data: map[string][]byte{credentialsSecretDataKey: []byte(token)},
})).To(Succeed())
}
bringUp := func(ctx SpecContext, fakeSrv *httptest.Server) {
DeferCleanup(fakeSrv.Close)
reconciler.NewClient = func(string) *tdclient.Client { return tdclient.New(fakeSrv.URL) }
createServer(ctx, dsnSpec())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // Deployment/Service
markDeploymentReady(ctx)
reconcileOnce(ctx)
}
credsToken := func(ctx SpecContext) string {
var s corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: credentialsSecretNameFor(name), Namespace: operatorNamespace}, &s)).To(Succeed())
return string(s.Data[credentialsSecretDataKey])
}
It("adopts the retained credential when the server still accepts it, without bootstrapping", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
fake.bootstrapped = true // every /api/bootstrap call 403s
fake.accounts[serviceAccountName] = 7
retainedSecret(ctx, "retained-key")
bringUp(ctx, fakeSrv)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
srv := &terdutv1alpha1.TerdutServer{}
Expect(k8sClient.Get(ctx, objKey, srv)).To(Succeed())
Expect(srv.Status.CredentialsSecretRef).NotTo(BeNil())
Expect(srv.Status.CredentialsSecretRef.Name).To(Equal(credentialsSecretNameFor(name)))
Expect(credsToken(ctx)).To(Equal("retained-key"), "the retained key is reused, not replaced")
})
It("ignores a retained credential the server rejects and bootstraps afresh after a database reset", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer() // a fresh database: bootstrap succeeds
fake.rejectedTokens = map[string]bool{"stale-key": true}
retainedSecret(ctx, "stale-key")
bringUp(ctx, fakeSrv)
Expect(readyCondition(ctx).Status).To(Equal(metav1.ConditionTrue))
Expect(credsToken(ctx)).NotTo(Equal("stale-key"), "the stale credential is replaced by a freshly minted one")
})
It("fails closed, naming the Secret to restore, when the server is bootstrapped and the retained credential is rejected", func(ctx SpecContext) {
fake, fakeSrv := newFakeTerdutServer()
fake.bootstrapped = true
fake.rejectedTokens = map[string]bool{"stale-key": true}
retainedSecret(ctx, "stale-key")
bringUp(ctx, fakeSrv)
cond := readyCondition(ctx)
Expect(cond.Status).To(Equal(metav1.ConditionFalse))
Expect(cond.Reason).To(Equal(terdutv1alpha1.ReasonBootstrapStateLost))
Expect(cond.Message).To(ContainSubstring(credentialsSecretNameFor(name)),
"the message names the Secret that would restore it")
Expect(cond.Message).To(ContainSubstring("does not clear the database"))
}) })
}) })
+46 -14
View File
@@ -37,22 +37,47 @@ func (r *TerdutServerReconciler) reconcileDeployment(
_, err := controllerutil.CreateOrUpdate(ctx, r.Client, deploy, func() error { _, err := controllerutil.CreateOrUpdate(ctx, r.Client, deploy, func() error {
replicas := srv.Spec.Replicas replicas := srv.Spec.Replicas
if replicas == 0 { if replicas == 0 {
replicas = 1 // Only reachable for a TerdutServer stored before the
// +kubebuilder:default=2 marker existed -- the API server's own
// CRD defaulting fills this in for anything created or updated
// through it, so a fresh zero value here means a pre-existing
// object that predates the default, not a deliberate "none"
// (there is no way to request zero replicas).
replicas = 2
} }
labels := labelsFor(srv) labels := labelsFor(srv)
deploy.Spec.Replicas = &replicas deploy.Spec.Replicas = &replicas
deploy.Spec.Selector = &metav1.LabelSelector{MatchLabels: labels} deploy.Spec.Selector = &metav1.LabelSelector{MatchLabels: labels}
// Recreate, not RollingUpdate: the sweeper and the notifier are // RollingUpdate, not Recreate: terdut-server v0.36.0 put the sweeper,
// unsynchronised singletons inside terdut-server, and two replicas // the notifier and the migration runner each behind a Postgres
// overlapping during a rollout would both page for the same // advisory lock, and gave incident creation its own conflict
// incident (matches the chart's own deployment.yaml comment). // resolution, so two replicas overlapping during a rollout no longer
deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RecreateDeploymentStrategyType} // double-page, race a migration, or drop a webhook payload (matches
// the chart's own deployment.yaml comment). No explicit
// maxUnavailable/maxSurge: left at the 25%/25% default, which rounds
// to 0/1 at the default replicas: 2 -- already zero-downtime.
deploy.Spec.Strategy = appsv1.DeploymentStrategy{Type: appsv1.RollingUpdateDeploymentStrategyType}
pod := srv.Spec.Pod
deploy.Spec.Template = corev1.PodTemplateSpec{ deploy.Spec.Template = corev1.PodTemplateSpec{
ObjectMeta: metav1.ObjectMeta{Labels: labels}, // pod.Annotations is assigned directly, not merged -- nothing
// else sets pod-template annotations today. If a future change
// needs the controller to set one of its own (e.g. a
// Prometheus-scrape annotation), this needs to become a real
// map merge with a stated precedence rather than silently
// clobbering one side.
ObjectMeta: metav1.ObjectMeta{Labels: labels, Annotations: pod.Annotations},
Spec: corev1.PodSpec{ Spec: corev1.PodSpec{
EnableServiceLinks: new(false), EnableServiceLinks: new(false),
InitContainers: []corev1.Container{waitForPostgresContainer(dbEnv)}, NodeSelector: pod.NodeSelector,
Tolerations: pod.Tolerations,
Affinity: pod.Affinity,
TopologySpreadConstraints: pod.TopologySpreadConstraints,
SecurityContext: pod.SecurityContext,
ServiceAccountName: pod.ServiceAccountName,
ImagePullSecrets: pod.ImagePullSecrets,
InitContainers: []corev1.Container{waitForPostgresContainer(dbEnv)},
Volumes: pod.ExtraVolumes,
Containers: []corev1.Container{{ Containers: []corev1.Container{{
Name: "terdut-server", Name: "terdut-server",
Image: fmt.Sprintf("%s:%s", srv.Spec.Image.Repository, srv.Spec.Image.Tag), Image: fmt.Sprintf("%s:%s", srv.Spec.Image.Repository, srv.Spec.Image.Tag),
@@ -61,9 +86,13 @@ func (r *TerdutServerReconciler) reconcileDeployment(
ContainerPort: servicePort(srv), ContainerPort: servicePort(srv),
Protocol: corev1.ProtocolTCP, Protocol: corev1.ProtocolTCP,
}}, }},
Env: buildEnv(srv, dbEnv), Env: buildEnv(srv, dbEnv),
LivenessProbe: healthzProbe(), EnvFrom: pod.ExtraEnvFrom,
ReadinessProbe: healthzProbe(), VolumeMounts: pod.ExtraVolumeMounts,
Resources: pod.Resources,
SecurityContext: pod.ContainerSecurityContext,
LivenessProbe: healthzProbe(),
ReadinessProbe: healthzProbe(),
}}, }},
}, },
} }
@@ -151,7 +180,10 @@ func healthzProbe() *corev1.Probe {
// field-for-field (confirmed against that source, not reconstructed from // field-for-field (confirmed against that source, not reconstructed from
// DESIGN.md's illustrative YAML alone) — dbEnv (TERDUT_DB_DSN, optionally // DESIGN.md's illustrative YAML alone) — dbEnv (TERDUT_DB_DSN, optionally
// PGPASSWORD) comes from resolveDatabaseEnv, since which of §8's two paths // PGPASSWORD) comes from resolveDatabaseEnv, since which of §8's two paths
// produced it doesn't matter past this point. // produced it doesn't matter past this point. spec.pod.extraEnv is appended
// last, after every fixed var -- this is the one place that owns "what env
// this container gets," so the escape hatch lives here rather than being
// appended separately in reconcileDeployment.
func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.EnvVar { func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.EnvVar {
env := []corev1.EnvVar{{Name: "TERDUT_ADDR", Value: fmt.Sprintf(":%d", servicePort(srv))}} env := []corev1.EnvVar{{Name: "TERDUT_ADDR", Value: fmt.Sprintf(":%d", servicePort(srv))}}
env = append(env, dbEnv...) env = append(env, dbEnv...)
@@ -230,7 +262,7 @@ func buildEnv(srv *terdutv1alpha1.TerdutServer, dbEnv []corev1.EnvVar) []corev1.
} }
} }
return env return append(env, srv.Spec.Pod.ExtraEnv...)
} }
func secretEnvSource(ref *terdutv1alpha1.SecretKeyRef) *corev1.EnvVarSource { func secretEnvSource(ref *terdutv1alpha1.SecretKeyRef) *corev1.EnvVarSource {
+38
View File
@@ -0,0 +1,38 @@
package controller
import (
"context"
policyv1 "k8s.io/api/policy/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
)
// reconcilePodDisruptionBudget creates/updates the PodDisruptionBudget
// spec.pod.disruptionBudget asks for, or deletes a previously-created one
// when the field has been cleared -- the one conditionally-created child
// object in this controller (Deployment/Service are unconditional). Owned
// by srv, same plain-OwnerReference shape as Deployment/Service (DESIGN.md
// §7): same namespace, GC handles it, no finalizer needed.
func (r *TerdutServerReconciler) reconcilePodDisruptionBudget(ctx context.Context, srv *terdutv1alpha1.TerdutServer) error {
pdb := &policyv1.PodDisruptionBudget{ObjectMeta: metav1.ObjectMeta{Name: srv.Name, Namespace: srv.Namespace}}
spec := srv.Spec.Pod.DisruptionBudget
if spec == nil {
if err := r.Delete(ctx, pdb); err != nil && !apierrors.IsNotFound(err) {
return err
}
return nil
}
_, err := controllerutil.CreateOrUpdate(ctx, r.Client, pdb, func() error {
pdb.Spec.Selector = &metav1.LabelSelector{MatchLabels: labelsFor(srv)}
pdb.Spec.MinAvailable = spec.MinAvailable
pdb.Spec.MaxUnavailable = spec.MaxUnavailable
return controllerutil.SetControllerReference(srv, pdb, r.Scheme)
})
return err
}
@@ -131,6 +131,9 @@ func (r *TerdutTeamReconciler) Reconcile(ctx context.Context, req ctrl.Request)
if err := teamClient.SetTeamOIDCGroups(ctx, team.Status.TeamID, team.Spec.OIDC.MemberGroup, team.Spec.OIDC.OwnerGroup); err != nil { if err := teamClient.SetTeamOIDCGroups(ctx, team.Status.TeamID, team.Spec.OIDC.MemberGroup, team.Spec.OIDC.OwnerGroup); err != nil {
return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/oidc-groups: %w", team.Status.TeamID, err) return ctrl.Result{}, fmt.Errorf("PUT /api/teams/%d/oidc-groups: %w", team.Status.TeamID, err)
} }
if err := r.reconcileInvite(ctx, &team, teamClient); err != nil {
return ctrl.Result{}, err
}
meta.SetStatusCondition(&team.Status.Conditions, metav1.Condition{ meta.SetStatusCondition(&team.Status.Conditions, metav1.Condition{
Type: terdutv1alpha1.ConditionReady, Type: terdutv1alpha1.ConditionReady,
@@ -4,6 +4,7 @@ import (
"context" "context"
"fmt" "fmt"
"net/http/httptest" "net/http/httptest"
"time"
. "github.com/onsi/ginkgo/v2" . "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega" . "github.com/onsi/gomega"
@@ -251,6 +252,88 @@ var _ = Describe("TerdutTeam Controller", func() {
}) })
}) })
Describe("spec.invite", func() {
It("mints a link into the TerdutTeam's own namespace, not the operator's", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx) // finalizer
reconcileOnce(ctx) // create+mint+apply
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
team.Spec.Invite.Enabled = true
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.InviteSecretRef).NotTo(BeNil())
var secret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{
Name: team.Status.InviteSecretRef.Name, Namespace: team.Namespace,
}, &secret)).To(Succeed())
Expect(string(secret.Data[inviteSecretURLKey])).To(ContainSubstring("invite="))
Expect(string(secret.Data[inviteSecretInviteIDKey])).To(Equal("1"))
inv := fake.invites[team.Status.TeamID][1]
Expect(inv.Role).To(Equal("member"), "default role")
Expect(inv.MaxUses).To(Equal(int64(1)), "default max uses")
})
It("refreshes a link that's within a day of terdut-server's 7-day TTL", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx)
reconcileOnce(ctx)
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
team.Spec.Invite.Enabled = true
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
firstSecretName := team.Status.InviteSecretRef.Name
var secret corev1.Secret
Expect(k8sClient.Get(ctx, types.NamespacedName{Name: firstSecretName, Namespace: team.Namespace}, &secret)).To(Succeed())
// Simulate the stored link being within the refresh window of
// expiry, the way it genuinely would be six days from now,
// without the test waiting six days.
secret.Data[inviteSecretExpiresAtKey] = []byte(time.Now().Add(12 * time.Hour).Format(time.RFC3339)) // inside inviteRefreshWindow
Expect(k8sClient.Update(ctx, &secret)).To(Succeed())
reconcileOnce(ctx)
Expect(fake.inviteDelete[1]).To(BeTrue(), "the stale invite should have been revoked")
Expect(fake.invites[team.Status.TeamID]).To(HaveKey(int64(2)), "a replacement should have been minted")
})
It("revokes the invite when spec.invite.enabled flips back to false", func(ctx SpecContext) {
createTeam(ctx, "platform", sameNSRef())
reconcileOnce(ctx)
reconcileOnce(ctx)
team := &terdutv1alpha1.TerdutTeam{}
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
team.Spec.Invite.Enabled = true
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
secretName := team.Status.InviteSecretRef.Name
team.Spec.Invite.Enabled = false
Expect(k8sClient.Update(ctx, team)).To(Succeed())
reconcileOnce(ctx)
Expect(fake.inviteDelete[1]).To(BeTrue())
Expect(k8sClient.Get(ctx, teamKey, team)).To(Succeed())
Expect(team.Status.InviteSecretRef).To(BeNil())
var secret corev1.Secret
err := k8sClient.Get(ctx, types.NamespacedName{Name: secretName, Namespace: team.Namespace}, &secret)
Expect(err).To(HaveOccurred(), "the invite Secret should have been deleted")
})
})
Describe("deletion", func() { Describe("deletion", func() {
It("deletes the team server-side and removes the credentials Secret", func(ctx SpecContext) { It("deletes the team server-side and removes the credentials Secret", func(ctx SpecContext) {
createTeam(ctx, "to-delete", sameNSRef()) createTeam(ctx, "to-delete", sameNSRef())
+149
View File
@@ -0,0 +1,149 @@
package controller
import (
"context"
"fmt"
"strconv"
"time"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"sigs.k8s.io/controller-runtime/pkg/client"
"sigs.k8s.io/controller-runtime/pkg/controller/controllerutil"
terdutv1alpha1 "git.ryuvia.com/niklas/terdut-operator/api/v1alpha1"
"git.ryuvia.com/niklas/terdut-operator/internal/tdclient"
)
// Data keys inside the generated invite Secret, following the same naming
// shape as TerdutAlertSource's webhookSecret*Key constants.
const (
inviteSecretURLKey = "url"
inviteSecretInviteIDKey = "inviteID"
inviteSecretExpiresAtKey = "expiresAt"
)
// inviteRefreshWindow is how far ahead of expiry this controller mints a
// replacement link, so a human reading status.inviteSecretRef never finds a
// dead link mid-use. terdut-server's invite TTL is a fixed, unconfigurable
// 7 days (internal/api/signup.go's inviteTTL) -- refreshing a full day
// ahead of that leaves comfortable margin against this controller's own
// 5-minute resync interval ever being delayed.
const inviteRefreshWindow = 24 * time.Hour
func inviteSecretName(team *terdutv1alpha1.TerdutTeam) string {
return team.Name + "-terdut-invite"
}
// reconcileInvite applies spec.invite against teamClient -- this team's own
// team-scoped credential, already owner-equivalent for every /invites route
// (terdut-server's SERVICE-ACCOUNTS.md, ratified not accidental). Mints,
// refreshes ahead of expiry, or revokes, entirely independent of this
// team's own Ready condition: an invite is a convenience for onboarding a
// human, never something anything else in this reconcile waits on.
func (r *TerdutTeamReconciler) reconcileInvite(ctx context.Context, team *terdutv1alpha1.TerdutTeam, teamClient *tdclient.Client) error {
if !team.Spec.Invite.Enabled {
return r.revokeInvite(ctx, team, teamClient)
}
secretName := inviteSecretName(team)
var secret corev1.Secret
err := r.Get(ctx, client.ObjectKey{Namespace: team.Namespace, Name: secretName}, &secret)
switch {
case err == nil:
expiresAt, parseErr := time.Parse(time.RFC3339, string(secret.Data[inviteSecretExpiresAtKey]))
if parseErr == nil && time.Until(expiresAt) > inviteRefreshWindow {
team.Status.InviteSecretRef = &terdutv1alpha1.LocalSecretRef{Name: secretName}
return nil // still fresh, nothing to do this reconcile
}
// Expired, about to expire, or unreadable: mint a replacement.
// Revoke the old row by id first (best-effort) so a leaked old link
// stops working immediately rather than lingering unrevoked until
// its own TTL -- failure here is not fatal, since the replacement
// below is what actually matters.
if oldID, idErr := strconv.ParseInt(string(secret.Data[inviteSecretInviteIDKey]), 10, 64); idErr == nil {
_ = teamClient.RevokeInvite(ctx, team.Status.TeamID, oldID)
}
return r.mintInvite(ctx, team, teamClient, secretName)
case apierrors.IsNotFound(err):
// Low stakes, unlike TerdutAlertSource's webhook URL: nothing
// external holds a durable dependency on one specific invite link
// staying stable the way an Alertmanager config depends on a
// webhook URL -- it's read once by one human and handed out. So
// this silently re-mints rather than failing closed the way
// TerdutAlertSource's ReasonWebhookSecretLost does for its Secret.
return r.mintInvite(ctx, team, teamClient, secretName)
default:
return err
}
}
func (r *TerdutTeamReconciler) mintInvite(
ctx context.Context, team *terdutv1alpha1.TerdutTeam, teamClient *tdclient.Client, secretName string,
) error {
role := team.Spec.Invite.Role
if role == "" {
role = "member"
}
maxUses := team.Spec.Invite.MaxUses
if maxUses == 0 {
maxUses = 1
}
inv, err := teamClient.CreateInvite(ctx, team.Status.TeamID, role, maxUses)
if err != nil {
return fmt.Errorf("POST /api/teams/%d/invites: %w", team.Status.TeamID, err)
}
secret := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: secretName, Namespace: team.Namespace}}
if _, err := controllerutil.CreateOrUpdate(ctx, r.Client, secret, func() error {
secret.Data = map[string][]byte{
inviteSecretURLKey: []byte(inv.URL),
inviteSecretInviteIDKey: []byte(strconv.FormatInt(inv.ID, 10)),
inviteSecretExpiresAtKey: []byte(inv.ExpiresAt.Format(time.RFC3339)),
}
return controllerutil.SetControllerReference(team, secret, r.Scheme)
}); err != nil {
return err
}
team.Status.InviteSecretRef = &terdutv1alpha1.LocalSecretRef{Name: secretName}
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonInviteMinted, terdutv1alpha1.ReasonInviteMinted,
"invite link minted into Secret %q", secretName)
}
return nil
}
// revokeInvite tears down spec.invite's Secret and server-side row when
// spec.invite.enabled is false (or was never set). Same-namespace and
// OwnerReference'd, so deleting the TerdutTeam itself already garbage-
// collects this Secret -- this path exists for the narrower case of
// flipping enabled back to false on an otherwise-live TerdutTeam.
func (r *TerdutTeamReconciler) revokeInvite(ctx context.Context, team *terdutv1alpha1.TerdutTeam, teamClient *tdclient.Client) error {
if team.Status.InviteSecretRef == nil {
return nil
}
name := team.Status.InviteSecretRef.Name
var secret corev1.Secret
err := r.Get(ctx, client.ObjectKey{Namespace: team.Namespace, Name: name}, &secret)
switch {
case err == nil:
if id, idErr := strconv.ParseInt(string(secret.Data[inviteSecretInviteIDKey]), 10, 64); idErr == nil {
if err := teamClient.RevokeInvite(ctx, team.Status.TeamID, id); err != nil {
return fmt.Errorf("DELETE /api/teams/%d/invites/%d: %w", team.Status.TeamID, id, err)
}
}
if err := r.Delete(ctx, &secret); err != nil && !apierrors.IsNotFound(err) {
return err
}
case !apierrors.IsNotFound(err):
return err
}
team.Status.InviteSecretRef = nil
if r.Recorder != nil {
r.Recorder.Eventf(team, nil, corev1.EventTypeNormal, terdutv1alpha1.ReasonInviteRevoked, terdutv1alpha1.ReasonInviteRevoked,
"invite link revoked")
}
return nil
}
+47
View File
@@ -566,3 +566,50 @@ func (c *Client) DeleteIntegration(ctx context.Context, teamID, integrationID in
} }
return c.do(req, nil) return c.do(req, nil)
} }
// Invite is a standing link into a team (POST /api/teams/{teamID}/invites'
// own response shape). URL carries the raw token exactly once, at creation
// -- terdut-server never shows it again (same one-time-shown shape as an
// integration's webhook key) -- so a caller that needs it later has to have
// kept this response, not re-fetched it.
type Invite struct {
ID int64 `json:"id"`
TeamID int64 `json:"team_id"`
Role string `json:"role"`
ExpiresAt time.Time `json:"expires_at"`
MaxUses int64 `json:"max_uses"`
URL string `json:"url,omitempty"`
}
// CreateInvite calls POST /api/teams/{teamID}/invites -- owner-gated
// (requireTeamOwner), so c must hold this team's own team-scoped
// credential, which already satisfies that check via its synthetic owner
// membership (terdut-server's SERVICE-ACCOUNTS.md). No conflict handling
// needed: unlike a team or a service account, an invite has no unique name
// to collide on -- every call mints a brand new row.
func (c *Client) CreateInvite(ctx context.Context, teamID int64, role string, maxUses int64) (*Invite, error) {
req, err := c.newRequest(ctx, http.MethodPost, fmt.Sprintf("/api/teams/%d/invites", teamID),
map[string]any{"role": role, "max_uses": maxUses})
if err != nil {
return nil, err
}
var inv Invite
if err := c.do(req, &inv); err != nil {
return nil, err
}
return &inv, nil
}
// RevokeInvite calls DELETE /api/teams/{teamID}/invites/{inviteID} -- same
// credential requirement as CreateInvite. A 404 (already revoked, or never
// existed) is the caller's to treat as success if it wants to, the same way
// DeleteTeam's own 404 handling works -- this method itself just reports
// whatever terdut-server said.
func (c *Client) RevokeInvite(ctx context.Context, teamID, inviteID int64) error {
req, err := c.newRequest(ctx, http.MethodDelete,
fmt.Sprintf("/api/teams/%d/invites/%d", teamID, inviteID), nil)
if err != nil {
return err
}
return c.do(req, nil)
}