Files
terdut-operator/examples/demo/02-team-platform.yaml
T
Niklas Ye e1103f2b7d Authenticate with a seeded operator key; fold escalation and switches into TerdutTeam
Credentials: the TerdutServer controller generates <name>-operator-key in
the server's own namespace (owned by it) and hands it to the pods as
TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it
at every start. A replaced Secret rolls the pods. The bootstrap handshake,
the checkpoint Secret, per-team service accounts and credentials Secrets,
BootstrapStateLost and credentials.deletionPolicy are gone.

CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule
and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[]
on the team (matched by name, extras removed); team invites are removed.
A team is created under the identity <namespace>/<name> (external_id), so a
retry, a lost status or a deleted team heal by repeating the same call, and
a display name owned by another team is TeamNameTaken instead of an
adoption. The server resolves escalation usernames (UnknownUser condition).
OIDC claim names and trustEmail are spec fields.

Fixes: query values are URL-escaped; every delete treats 404 as success;
deleting a team no longer depends on allowedTeams consent; a switch or
integration deleted on the server is recreated; unnamed switches take the
CR's name.

Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces
and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo
(run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays
cluster-wide, now stated in DESIGN.md section 9.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:22 +02:00

46 lines
1.8 KiB
YAML

# Two teams (this one and 03-team-payments.yaml) so the demo shows
# per-team isolation -- separate incident lists, separate escalation
# ladders, separate alert sources -- rather than one team standing in for
# everything. A team carries its own escalation ladder and dead man's
# switches; alert sources are separate objects (04-alertsource-*.yaml)
# because each owns a webhook Secret.
apiVersion: terdut.ryuvia.com/v1alpha1
kind: TerdutTeam
metadata:
name: terdutteam-platform
spec:
# serverRef.namespace omitted: both this and terdut-operator-demo (01-server.yaml)
# live in whatever namespace you apply this directory into, which is the
# common case and needs no allowedTeams consent on the TerdutServer side
# (DESIGN.md §4.1, §4.6).
serverRef:
name: terdut-operator-demo
displayName: Platform
# No oidc block: this demo is password-login only (01-server.yaml).
# The whole ladder, replaced as one unit. username is required iff kind is
# "user" (CEL validation at apply time). The server resolves it, so a
# username it does not know yet leaves this team at Reason: UnknownUser
# until that person exists -- run-demo.sh creates alice for exactly that.
escalation:
repeatCount: 2
fallbackTopic: platform-fallback
levels:
- timeout: 5m
targets:
- kind: user
username: alice
- timeout: 10m
targets:
- kind: oncall
# Dead man's switches, by name. fire-alerts.sh's "heartbeat" scenario sends
# a matching alert; stop sending it and terdut-server itself opens an
# incident once `timeout` passes with no heartbeat. A switch on the server
# that is not listed here is removed.
deadmanSwitches:
- name: platform-watchdog
matcher: "alertname=PlatformWatchdog"
timeout: 15m
severity: critical