50ce5bcec0
The demo pinned v0.36.0, the floor for replicas: 2, and so showed none of the web UI since: the queue and incident layouts, the rota and escalation pages, the theme toggle, and the cluster chip, filter and page titles (v0.42.0-v0.43.0). It pins v0.43.0 now; the comment keeps v0.36.0 as the floor, which is what the replicas setting actually depends on. fire-alerts.sh takes an optional CLUSTER, standing in for a Prometheus external label plus `cluster` in Alertmanager's group_by (terdut-server's README, "Several clusters, one team"). It goes on the alert's labels and groupLabels, and into the group key and the fingerprint, so the same alert in two clusters is two incidents and not one. Unset, the payload is exactly what it was. run-demo.sh fires its alerts across prod-eu and prod-us, high-cpu in both, so the queue has a chip and a filter to show. run-demo.sh also failed on its second run, though it says it is safe to re-run: it expected HTTP 409 when alice already exists, but a spent invite is answered with 403 "invite link is not usable" before the username is ever checked. It now tries to log alice in first and skips the signup if that works. Checked on the kind cluster: the server rolled to v0.43.0, every CR became Ready and Adopted (server, both teams, both escalation rules, both dead man's switches, both alert sources), and /api/incidents/clusters, /api/incidents?cluster=prod-us and the incident titles came back as expected. No operator code changed, so this needs no operator release. Co-authored-by: Claude <noreply@anthropic.com>
151 lines
6.1 KiB
Bash
Executable File
151 lines
6.1 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Sends a synthetic Alertmanager v4 webhook payload at the "platform" or
|
|
# "payments" demo team's TerdutAlertSource, so terdut-server opens (or
|
|
# resolves) an incident exactly the way it would for a real Alertmanager.
|
|
#
|
|
# The payload shape here is amPayload/amAlert, read straight out of
|
|
# terdut-server's own internal/api/alertmanager.go rather than guessed from
|
|
# its docs -- version/status/groupKey/groupLabels, and alerts[] carrying
|
|
# status/labels/annotations/startsAt/endsAt/generatorURL/fingerprint.
|
|
#
|
|
# Why this reads the webhook key out of a kubectl Secret instead of using
|
|
# the "url" key already in it: that URL is built from spec.networking.hostname
|
|
# (TERDUT_PUBLIC_URL), and nothing in this demo stands up real ingress for
|
|
# it (01-server.yaml's own comment) -- so it resolves nowhere. The key
|
|
# alone, against whatever you've actually port-forwarded BASE_URL to below,
|
|
# is the one part of that URL still usable here.
|
|
#
|
|
# Usage:
|
|
# [CLUSTER=prod-eu] ./fire-alerts.sh <platform|payments> <high-cpu|disk-full|pod-crash|heartbeat> [resolve]
|
|
#
|
|
# CLUSTER stands in for a Prometheus externalLabel plus `cluster` in
|
|
# Alertmanager's group_by (terdut-server's README, "Several clusters, one
|
|
# team"): it is put on the alert's labels and on groupLabels, so the incident
|
|
# carries it and the web UI shows the cluster chip and the queue's cluster
|
|
# filter. It is also part of the group key and the fingerprint, which is what
|
|
# keeps the same alert in two clusters from joining one incident. Unset, the
|
|
# alert is sent exactly as before.
|
|
#
|
|
# Prerequisites: kubectl context pointed at the demo namespace, jq, curl,
|
|
# and (in another terminal) a running:
|
|
# kubectl port-forward svc/terdut-operator-demo 8080:8080
|
|
set -euo pipefail
|
|
|
|
NAMESPACE="${NAMESPACE:-}"
|
|
CLUSTER="${CLUSTER:-}"
|
|
BASE_URL="${BASE_URL:-http://localhost:8080}"
|
|
|
|
usage() {
|
|
cat >&2 <<'EOF'
|
|
usage: fire-alerts.sh <platform|payments> <scenario> [resolve]
|
|
|
|
scenarios:
|
|
high-cpu warning -- CPU usage above 90% for 10 minutes
|
|
disk-full critical -- disk usage above 95%
|
|
pod-crash error -- a pod crash-looping
|
|
heartbeat critical -- the team's dead man's switch heartbeat
|
|
(matches the matcher in 06/07-deadman-*.yaml -- send this
|
|
repeatedly to keep the switch alive, or stop sending it and
|
|
watch terdut-server open an incident on its own once
|
|
`timeout` passes with no heartbeat. "resolve" is not a valid
|
|
third argument for this scenario: a heartbeat is only ever
|
|
firing.)
|
|
|
|
env vars:
|
|
NAMESPACE kubectl -n for reading the webhook Secret (required)
|
|
BASE_URL where the port-forwarded terdut-server is (default http://localhost:8080)
|
|
CLUSTER optional cluster name, e.g. prod-eu: sent as a `cluster` label and
|
|
group label, so the UI shows where the incident came from
|
|
EOF
|
|
exit 1
|
|
}
|
|
|
|
[ $# -ge 2 ] || usage
|
|
team="$1" scenario="$2" verb="${3:-fire}"
|
|
[ -n "$NAMESPACE" ] || { echo "fire-alerts.sh: set NAMESPACE" >&2; exit 1; }
|
|
|
|
case "$team" in
|
|
platform|payments) ;;
|
|
*) usage ;;
|
|
esac
|
|
|
|
case "$scenario" in
|
|
high-cpu) alertname=TerdutDemoHighCPU severity=warning summary="CPU usage above 90% for 10 minutes" ;;
|
|
disk-full) alertname=TerdutDemoDiskFull severity=critical summary="Disk usage above 95% on /data" ;;
|
|
pod-crash) alertname=TerdutDemoPodCrashLooping severity=error summary="Pod web-7f8b9 is crash-looping (5 restarts in 10m)" ;;
|
|
heartbeat)
|
|
# Must match 06-deadman-platform.yaml / 07-deadman-payments.yaml's own
|
|
# matcher exactly -- that's what makes this a heartbeat rather than a
|
|
# third ordinary alert.
|
|
case "$team" in
|
|
platform) alertname=PlatformWatchdog ;;
|
|
payments) alertname=PaymentsWatchdog ;;
|
|
esac
|
|
severity=critical summary="demo heartbeat"
|
|
[ "$verb" = fire ] || { echo "fire-alerts.sh: heartbeat is only ever fired, never resolved -- just stop sending it" >&2; exit 1; }
|
|
;;
|
|
*) usage ;;
|
|
esac
|
|
|
|
case "$verb" in
|
|
fire) status=firing ;;
|
|
resolve) status=resolved ;;
|
|
*) usage ;;
|
|
esac
|
|
|
|
secret_name="terdutalertsource-${team}-terdut-webhook"
|
|
key="$(kubectl -n "$NAMESPACE" get secret "$secret_name" -o jsonpath='{.data.key}' | base64 -d)"
|
|
[ -n "$key" ] || { echo "fire-alerts.sh: empty key read from Secret $secret_name -- has 08/09-alertsource-*.yaml reconciled yet?" >&2; exit 1; }
|
|
|
|
# Stable per (team, scenario) so a resolve targets the same alert a fire
|
|
# created: terdut-server correlates on (team_id, fingerprint), not on
|
|
# anything else in the payload. Real Alertmanager computes this from the
|
|
# alert's label set; a fixed string plays the same role here.
|
|
fingerprint="demo-${team}-${scenario}${CLUSTER:+-$CLUSTER}"
|
|
|
|
now="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
|
if [ "$status" = firing ]; then
|
|
ends_at="0001-01-01T00:00:00Z" # Alertmanager's own "not resolved" zero value
|
|
else
|
|
ends_at="$now"
|
|
fi
|
|
|
|
payload="$(jq -n \
|
|
--arg status "$status" \
|
|
--arg groupKey "demo:${team}:${scenario}${CLUSTER:+:$CLUSTER}" \
|
|
--arg cluster "$CLUSTER" \
|
|
--arg alertname "$alertname" \
|
|
--arg team "$team" \
|
|
--arg severity "$severity" \
|
|
--arg summary "$summary" \
|
|
--arg startsAt "$now" \
|
|
--arg endsAt "$ends_at" \
|
|
--arg fingerprint "$fingerprint" \
|
|
'{
|
|
version: "4",
|
|
status: $status,
|
|
groupKey: $groupKey,
|
|
groupLabels: ({ alertname: $alertname, team: $team } + (if $cluster != "" then { cluster: $cluster } else {} end)),
|
|
alerts: [{
|
|
status: $status,
|
|
labels: ({ alertname: $alertname, severity: $severity, team: $team, instance: "demo" } + (if $cluster != "" then { cluster: $cluster } else {} end)),
|
|
annotations: { summary: $summary },
|
|
startsAt: $startsAt,
|
|
endsAt: $endsAt,
|
|
generatorURL: "https://example.com/demo",
|
|
fingerprint: $fingerprint
|
|
}]
|
|
}')"
|
|
|
|
url="${BASE_URL}/api/integrations/${key}/alertmanager"
|
|
echo "POST $url (team=$team scenario=$scenario status=$status${CLUSTER:+ cluster=$CLUSTER})" >&2
|
|
code="$(curl -sS -o /tmp/fire-alerts-response.json -w '%{http_code}' \
|
|
-X POST "$url" -H 'Content-Type: application/json' -d "$payload")"
|
|
echo "-> HTTP $code" >&2
|
|
cat /tmp/fire-alerts-response.json >&2
|
|
echo >&2
|
|
|
|
if [ "$code" != "200" ]; then
|
|
exit 1
|
|
fi
|