e1103f2b7d
Credentials: the TerdutServer controller generates <name>-operator-key in the server's own namespace (owned by it) and hands it to the pods as TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it at every start. A replaced Secret rolls the pods. The bootstrap handshake, the checkpoint Secret, per-team service accounts and credentials Secrets, BootstrapStateLost and credentials.deletionPolicy are gone. CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[] on the team (matched by name, extras removed); team invites are removed. A team is created under the identity <namespace>/<name> (external_id), so a retry, a lost status or a deleted team heal by repeating the same call, and a display name owned by another team is TeamNameTaken instead of an adoption. The server resolves escalation usernames (UnknownUser condition). OIDC claim names and trustEmail are spec fields. Fixes: query values are URL-escaped; every delete treats 404 as success; deleting a team no longer depends on allowedTeams consent; a switch or integration deleted on the server is recreated; unnamed switches take the CR's name. Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo (run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays cluster-wide, now stated in DESIGN.md section 9. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
46 lines
1.8 KiB
YAML
46 lines
1.8 KiB
YAML
# Two teams (this one and 03-team-payments.yaml) so the demo shows
|
|
# per-team isolation -- separate incident lists, separate escalation
|
|
# ladders, separate alert sources -- rather than one team standing in for
|
|
# everything. A team carries its own escalation ladder and dead man's
|
|
# switches; alert sources are separate objects (04-alertsource-*.yaml)
|
|
# because each owns a webhook Secret.
|
|
apiVersion: terdut.ryuvia.com/v1alpha1
|
|
kind: TerdutTeam
|
|
metadata:
|
|
name: terdutteam-platform
|
|
spec:
|
|
# serverRef.namespace omitted: both this and terdut-operator-demo (01-server.yaml)
|
|
# live in whatever namespace you apply this directory into, which is the
|
|
# common case and needs no allowedTeams consent on the TerdutServer side
|
|
# (DESIGN.md §4.1, §4.6).
|
|
serverRef:
|
|
name: terdut-operator-demo
|
|
displayName: Platform
|
|
# No oidc block: this demo is password-login only (01-server.yaml).
|
|
|
|
# The whole ladder, replaced as one unit. username is required iff kind is
|
|
# "user" (CEL validation at apply time). The server resolves it, so a
|
|
# username it does not know yet leaves this team at Reason: UnknownUser
|
|
# until that person exists -- run-demo.sh creates alice for exactly that.
|
|
escalation:
|
|
repeatCount: 2
|
|
fallbackTopic: platform-fallback
|
|
levels:
|
|
- timeout: 5m
|
|
targets:
|
|
- kind: user
|
|
username: alice
|
|
- timeout: 10m
|
|
targets:
|
|
- kind: oncall
|
|
|
|
# Dead man's switches, by name. fire-alerts.sh's "heartbeat" scenario sends
|
|
# a matching alert; stop sending it and terdut-server itself opens an
|
|
# incident once `timeout` passes with no heartbeat. A switch on the server
|
|
# that is not listed here is removed.
|
|
deadmanSwitches:
|
|
- name: platform-watchdog
|
|
matcher: "alertname=PlatformWatchdog"
|
|
timeout: 15m
|
|
severity: critical
|