e1103f2b7d
Credentials: the TerdutServer controller generates <name>-operator-key in the server's own namespace (owned by it) and hands it to the pods as TERDUT_OPERATOR_KEY; the server creates its instance-scoped account from it at every start. A replaced Secret rolls the pods. The bootstrap handshake, the checkpoint Secret, per-team service accounts and credentials Secrets, BootstrapStateLost and credentials.deletionPolicy are gone. CRDs: TerdutServer, TerdutTeam and TerdutAlertSource. TerdutEscalationRule and TerdutDeadmanSwitch become spec.escalation and spec.deadmanSwitches[] on the team (matched by name, extras removed); team invites are removed. A team is created under the identity <namespace>/<name> (external_id), so a retry, a lost status or a deleted team heal by repeating the same call, and a display name owned by another team is TeamNameTaken instead of an adoption. The server resolves escalation usernames (UnknownUser condition). OIDC claim names and trustEmail are spec fields. Fixes: query values are URL-escaped; every delete treats 404 as success; deleting a team no longer depends on allowedTeams consent; a switch or integration deleted on the server is recreated; unnamed switches take the CR's name. Cleanup: scaffold e2e test, AGENTS.md, devcontainer, unused config/ pieces and Client.Version() removed; DESIGN.md, README, ROADMAP and the demo (run-demo.sh, manifests) rewritten for the new design. Secret RBAC stays cluster-wide, now stated in DESIGN.md section 9. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
151 lines
6.1 KiB
Bash
Executable File
151 lines
6.1 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Sends a synthetic Alertmanager v4 webhook payload at the "platform" or
|
|
# "payments" demo team's TerdutAlertSource, so terdut-server opens (or
|
|
# resolves) an incident exactly the way it would for a real Alertmanager.
|
|
#
|
|
# The payload shape here is amPayload/amAlert, read straight out of
|
|
# terdut-server's own internal/api/alertmanager.go rather than guessed from
|
|
# its docs -- version/status/groupKey/groupLabels, and alerts[] carrying
|
|
# status/labels/annotations/startsAt/endsAt/generatorURL/fingerprint.
|
|
#
|
|
# Why this reads the webhook key out of a kubectl Secret instead of using
|
|
# the "url" key already in it: that URL is built from spec.networking.hostname
|
|
# (TERDUT_PUBLIC_URL), and nothing in this demo stands up real ingress for
|
|
# it (01-server.yaml's own comment) -- so it resolves nowhere. The key
|
|
# alone, against whatever you've actually port-forwarded BASE_URL to below,
|
|
# is the one part of that URL still usable here.
|
|
#
|
|
# Usage:
|
|
# [CLUSTER=prod-eu] ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
|
|
#
|
|
# CLUSTER stands in for a Prometheus externalLabel plus `cluster` in
|
|
# Alertmanager's group_by (terdut-server's README, "Several clusters, one
|
|
# team"): it is put on the alert's labels and on groupLabels, so the incident
|
|
# carries it and the web UI shows the cluster chip and the queue's cluster
|
|
# filter. It is also part of the group key and the fingerprint, which is what
|
|
# keeps the same alert in two clusters from joining one incident. Unset, the
|
|
# alert is sent exactly as before.
|
|
#
|
|
# Prerequisites: kubectl context pointed at the demo namespace, jq, curl,
|
|
# and (in another terminal) a running:
|
|
# kubectl port-forward svc/terdut-operator-demo 8080:8080
|
|
set -euo pipefail
|
|
|
|
NAMESPACE="${NAMESPACE:-}"
|
|
CLUSTER="${CLUSTER:-}"
|
|
BASE_URL="${BASE_URL:-http://localhost:8080}"
|
|
|
|
usage() {
|
|
cat >&2 <<'EOF'
|
|
usage: fire-alerts.sh <platform|payments> <scenario> [resolve]
|
|
|
|
scenarios:
|
|
high-cpu warning -- CPU usage above 90% for 10 minutes
|
|
disk-full critical -- disk usage above 95%
|
|
pod-crash error -- a pod crash-looping
|
|
heartbeat critical -- the team's dead man's switch heartbeat
|
|
(matches the matcher in the deadmanSwitches in 02/03-team-*.yaml -- send this
|
|
repeatedly to keep the switch alive, or stop sending it and
|
|
watch terdut-server open an incident on its own once
|
|
`timeout` passes with no heartbeat. "resolve" is not a valid
|
|
third argument for this scenario: a heartbeat is only ever
|
|
firing.)
|
|
|
|
env vars:
|
|
NAMESPACE kubectl -n for reading the webhook Secret (required)
|
|
BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080)
|
|
CLUSTER optional cluster name, e.g. prod-eu: sent as a `cluster` label and
|
|
group label, so the UI shows where the incident came from
|
|
EOF
|
|
exit 1
|
|
}
|
|
|
|
[ $# -ge 2 ] || usage
|
|
team="$1" scenario="$2" verb="${3:-fire}"
|
|
[ -n "$NAMESPACE" ] || { echo "fire-alerts.sh: set NAMESPACE" >&2; exit 1; }
|
|
|
|
case "$team" in
|
|
platform|payments) ;;
|
|
*) usage ;;
|
|
esac
|
|
|
|
case "$scenario" in
|
|
high-cpu) alertname=TerdutDemoHighCPU severity=warning summary="CPU usage above 90% for 10 minutes" ;;
|
|
disk-full) alertname=TerdutDemoDiskFull severity=critical summary="Disk usage above 95% on /data" ;;
|
|
pod-crash) alertname=TerdutDemoPodCrashLooping severity=error summary="Pod web-7f8b9 is crash-looping (5 restarts in 10m)" ;;
|
|
heartbeat)
|
|
# Must match 02-team-platform.yaml / 03-team-payments.yaml's own
|
|
# matcher exactly -- that's what makes this a heartbeat rather than a
|
|
# third ordinary alert.
|
|
case "$team" in
|
|
platform) alertname=PlatformWatchdog ;;
|
|
payments) alertname=PaymentsWatchdog ;;
|
|
esac
|
|
severity=critical summary="demo heartbeat"
|
|
[ "$verb" = fire ] || { echo "fire-alerts.sh: heartbeat is only ever fired, never resolved -- just stop sending it" >&2; exit 1; }
|
|
;;
|
|
*) usage ;;
|
|
esac
|
|
|
|
case "$verb" in
|
|
fire) status=firing ;;
|
|
resolve) status=resolved ;;
|
|
*) usage ;;
|
|
esac
|
|
|
|
secret_name="terdutalertsource-${team}-terdut-webhook"
|
|
key="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.key}' | base64 -d)"
|
|
[ -n "$key" ] || { echo "fire-alerts.sh: empty key read from Secret $secret_name -- has 04/05-alertsource-*.yaml reconciled yet?" >&2; exit 1; }
|
|
|
|
# Stable per (team, scenario) so a resolve targets the same alert a fire
|
|
# created: terdut-server correlates on (team_id, fingerprint), not on
|
|
# anything else in the payload. Real Alertmanager computes this from the
|
|
# alert's label set; a fixed string plays the same role here.
|
|
fingerprint="demo-${team}-${scenario}${CLUSTER:+-$CLUSTER}"
|
|
|
|
now="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
|
if [ "$status" = firing ]; then
|
|
ends_at="0001-01-01T00:00:00Z" # Alertmanager's own "not resolved" zero value
|
|
else
|
|
ends_at="$now"
|
|
fi
|
|
|
|
payload="$(jq -n \
|
|
--arg status "$status" \
|
|
--arg groupKey "demo:${team}:${scenario}${CLUSTER:+:$CLUSTER}" \
|
|
--arg cluster "$CLUSTER" \
|
|
--arg alertname "$alertname" \
|
|
--arg team "$team" \
|
|
--arg severity "$severity" \
|
|
--arg summary "$summary" \
|
|
--arg startsAt "$now" \
|
|
--arg endsAt "$ends_at" \
|
|
--arg fingerprint "$fingerprint" \
|
|
'{
|
|
version: "4",
|
|
status: $status,
|
|
groupKey: $groupKey,
|
|
groupLabels: ({ alertname: $alertname, team: $team } + (if $cluster != "" then { cluster: $cluster } else {} end)),
|
|
alerts: [{
|
|
status: $status,
|
|
labels: ({ alertname: $alertname, severity: $severity, team: $team, instance: "demo" } + (if $cluster != "" then { cluster: $cluster } else {} end)),
|
|
annotations: { summary: $summary },
|
|
startsAt: $startsAt,
|
|
endsAt: $endsAt,
|
|
generatorURL: "https://example.com/demo",
|
|
fingerprint: $fingerprint
|
|
}]
|
|
}')"
|
|
|
|
url="${BASE_URL}/api/integrations/${key}/alertmanager"
|
|
echo "POST $url (team=$team scenario=$scenario status=$status${CLUSTER:+ cluster=$CLUSTER})" >&2
|
|
code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \
|
|
-X POST "$url" -H 'Content-Type: application/json' -d "$payload")"
|
|
echo "-> HTTP $code" >&2
|
|
cat /tmp/fire-alerts-response.json >&2
|
|
echo >&2
|
|
|
|
if [ "$code" != "200" ]; then
|
|
exit 1
|
|
fi
|