Files
terdut-operator/examples/demo/fire-alerts.sh
T
Niklas Ye e1103f2b7d Authenticate with a seeded operator key; fold escalation and switches into TerdutTeam
Credentials: the TerdutServer controller generates <name>-operator-key in
the server's own namespace (owned by it) and hands it to the pods as
TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it
at every start. A replaced Secret rolls the pods. The bootstrap handshake,
the checkpoint Secret, per-team service accounts and credentials Secrets,
BootstrapStateLost and credentials.deletionPolicy are gone.

CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule
and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[]
on the team (matched by name, extras removed); team invites are removed.
A team is created under the identity <namespace>/<name> (external_id), so a
retry, a lost status or a deleted team heal by repeating the same call, and
a display name owned by another team is TeamNameTaken instead of an
adoption. The server resolves escalation usernames (UnknownUser condition).
OIDC claim names and trustEmail are spec fields.

Fixes: query values are URL-escaped; every delete treats 404 as success;
deleting a team no longer depends on allowedTeams consent; a switch or
integration deleted on the server is recreated; unnamed switches take the
CR's name.

Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces
and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo
(run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays
cluster-wide, now stated in DESIGN.md section 9.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:22 +02:00

151 lines
6.1 KiB
Bash
Executable File

#!/usr/bin/env bash
# Sends a synthetic Alertmanager v4 webhook payload at the "platform" or
# "payments" demo team's TerdutAlertSource, so terdut-server opens (or
# resolves) an incident exactly the way it would for a real Alertmanager.
#
# The payload shape here is amPayload/amAlert, read straight out of
# terdut-server's own internal/api/alertmanager.go rather than guessed from
# its docs -- version/status/groupKey/groupLabels, and alerts[] carrying
# status/labels/annotations/startsAt/endsAt/generatorURL/fingerprint.
#
# Why this reads the webhook key out of a kubectl Secret instead of using
# the "url" key already in it: that URL is built from spec.networking.hostname
# (TERDUT_PUBLIC_URL), and nothing in this demo stands up real ingress for
# it (01-server.yaml's own comment) -- so it resolves nowhere. The key
# alone, against whatever you've actually port-forwarded BASE_URL to below,
# is the one part of that URL still usable here.
#
# Usage:
# [CLUSTER=prod-eu] ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
#
# CLUSTER stands in for a Prometheus externalLabel plus `cluster` in
# Alertmanager's group_by (terdut-server's README, "Several clusters, one
# team"): it is put on the alert's labels and on groupLabels, so the incident
# carries it and the web UI shows the cluster chip and the queue's cluster
# filter. It is also part of the group key and the fingerprint, which is what
# keeps the same alert in two clusters from joining one incident. Unset, the
# alert is sent exactly as before.
#
# Prerequisites: kubectl context pointed at the demo namespace, jq, curl,
# and (in another terminal) a running:
# kubectl port-forward svc/terdut-operator-demo 8080:8080
set -euo pipefail
NAMESPACE="${NAMESPACE:-}"
CLUSTER="${CLUSTER:-}"
BASE_URL="${BASE_URL:-http://localhost:8080}"
usage() {
cat >&2 <<'EOF'
usage: fire-alerts.sh <platform|payments> <scenario> [resolve]
scenarios:
high-cpu warning -- CPU usage above 90% for 10 minutes
disk-full critical -- disk usage above 95%
pod-crash error -- a pod crash-looping
heartbeat critical -- the team's dead man's switch heartbeat
(matches the matcher in the deadmanSwitches in 02/03-team-*.yaml -- send this
repeatedly to keep the switch alive, or stop sending it and
watch terdut-server open an incident on its own once
`timeout` passes with no heartbeat. "resolve" is not a valid
third argument for this scenario: a heartbeat is only ever
firing.)
env vars:
NAMESPACE kubectl -n for reading the webhook Secret (required)
BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080)
CLUSTER optional cluster name, e.g. prod-eu: sent as a `cluster` label and
group label, so the UI shows where the incident came from
EOF
exit 1
}
[ $# -ge 2 ] || usage
team="$1" scenario="$2" verb="${3:-fire}"
[ -n "$NAMESPACE" ] || { echo "fire-alerts.sh: set NAMESPACE" >&2; exit 1; }
case "$team" in
platform|payments) ;;
*) usage ;;
esac
case "$scenario" in
high-cpu) alertname=TerdutDemoHighCPU severity=warning summary="CPU usage above 90% for 10 minutes" ;;
disk-full) alertname=TerdutDemoDiskFull severity=critical summary="Disk usage above 95% on /data" ;;
pod-crash) alertname=TerdutDemoPodCrashLooping severity=error summary="Pod web-7f8b9 is crash-looping (5 restarts in 10m)" ;;
heartbeat)
# Must match 02-team-platform.yaml / 03-team-payments.yaml's own
# matcher exactly -- that's what makes this a heartbeat rather than a
# third ordinary alert.
case "$team" in
platform) alertname=PlatformWatchdog ;;
payments) alertname=PaymentsWatchdog ;;
esac
severity=critical summary="demo heartbeat"
[ "$verb" = fire ] || { echo "fire-alerts.sh: heartbeat is only ever fired, never resolved -- just stop sending it" >&2; exit 1; }
;;
*) usage ;;
esac
case "$verb" in
fire) status=firing ;;
resolve) status=resolved ;;
*) usage ;;
esac
secret_name="terdutalertsource-${team}-terdut-webhook"
key="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.key}' | base64 -d)"
[ -n "$key" ] || { echo "fire-alerts.sh: empty key read from Secret $secret_name -- has 04/05-alertsource-*.yaml reconciled yet?" >&2; exit 1; }
# Stable per (team, scenario) so a resolve targets the same alert a fire
# created: terdut-server correlates on (team_id, fingerprint), not on
# anything else in the payload. Real Alertmanager computes this from the
# alert's label set; a fixed string plays the same role here.
fingerprint="demo-${team}-${scenario}${CLUSTER:+-$CLUSTER}"
now="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
if [ "$status" = firing ]; then
ends_at="0001-01-01T00:00:00Z" # Alertmanager's own "not resolved" zero value
else
ends_at="$now"
fi
payload="$(jq -n \
--arg status "$status" \
--arg groupKey "demo:${team}:${scenario}${CLUSTER:+:$CLUSTER}" \
--arg cluster "$CLUSTER" \
--arg alertname "$alertname" \
--arg team "$team" \
--arg severity "$severity" \
--arg summary "$summary" \
--arg startsAt "$now" \
--arg endsAt "$ends_at" \
--arg fingerprint "$fingerprint" \
'{
version: "4",
status: $status,
groupKey: $groupKey,
groupLabels: ({ alertname: $alertname, team: $team } + (if $cluster != "" then { cluster: $cluster } else {} end)),
alerts: [{
status: $status,
labels: ({ alertname: $alertname, severity: $severity, team: $team, instance: "demo" } + (if $cluster != "" then { cluster: $cluster } else {} end)),
annotations: { summary: $summary },
startsAt: $startsAt,
endsAt: $endsAt,
generatorURL: "https://example.com/demo",
fingerprint: $fingerprint
}]
}')"
url="${BASE_URL}/api/integrations/${key}/alertmanager"
echo "POST $url (team=$team scenario=$scenario status=$status${CLUSTER:+ cluster=$CLUSTER})" >&2
code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \
-X POST "$url" -H 'Content-Type: application/json' -d "$payload")"
echo "-> HTTP $code" >&2
cat /tmp/fire-alerts-response.json >&2
echo >&2
if [ "$code" != "200" ]; then
exit 1
fi