The README was 1,240 lines of reference material and still described a SQLite quick start. It is now a short tour (highlights, screenshots of the web UI, an accurate quick start against Postgres), and each topic has its own page under docs/ with an index: deployment, configuration, Alertmanager, incidents, notifications, escalation, dead man's switches, single sign-on, web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it described. The "Upgrading to ..." sections for an unreleased product are dropped. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
4.1 KiB
Dead man's switches
Detecting that alerts have stopped arriving. Back to the README and the documentation index.
Everything above assumes alerts arrive. If Prometheus stops evaluating, or Alertmanager cannot reach this server, nothing arrives — and silence looks exactly like everything being fine. A dead man's switch inverts the handling for one designated alert so that silence is the signal:
- receiving it opens no incident, and
- the absence of it does.
kube-prometheus-stack already ships the alert for this. Watchdog is
expr: vector(1), so it fires permanently and is re-sent forever; it is worth
nothing unless something downstream notices it stop. It is the usual first switch.
Switches belong to a team, which decides which of its own alerts are
heartbeats and how long a silence has to last. Each switch is a row of its
own — a name, one matcher, a timeout and a severity — so switches in one team
can have different deadlines. An owner adds and removes them on Team →
Switches, which lists each with a status (healthy, dead, or
dormant until its first heartbeat), when it was last heard from, and when it
last opened an incident; a matcher that several clusters satisfy is broken down
per cluster. The API is POST/DELETE /api/teams/{teamID}/deadman/switches. A
missed heartbeat opens an incident in the team whose integration received it.
Removing a switch stops the watching; an incident it already opened stays open
until somebody resolves it.
A new team watches nothing until its owner (or terdut-operator, from a
TerdutTeam) adds a switch: inheriting an install-wide heartbeat would page a
new team about a source it has never heard of.
A matcher is a set of exact label conditions, one of which must be the
alertname, , one matcher per switch, , between the label conditions:
alertname=Watchdog,cluster=prod
The unit of monitoring is the fingerprint, not the alert name. Two clusters
sending the same Watchdog are two independent switches, so a healthy one can
never mask a dead one.
The lifecycle
A switch is dormant until its first heartbeat arrives. A configured matcher that has never been heard from opens nothing, so a fresh deploy or a restored database does not page. It also means a matcher that never matches anything is silently inert.
Once armed, the sweeper declares it dead when either the heartbeat has not
been refreshed within the switch's timeout_seconds, or Alertmanager explicitly
resolved it — the sender saying the heartbeat stopped needs no further waiting.
That opens an incident at the switch's severity, assigned and paged like any
other, and marks the heartbeat alert "resolution_source": "deadman" so the
alert list stops claiming a dead switch is firing.
It recovers when the heartbeat starts arriving again: the incident resolves
with "resolution_source": "recovered" and the all-clear goes to whoever was
paged.
Resolving the incident by hand sticks, the same way it does for an alert-backed one. While the switch stays silent nothing new opens — so a decommissioned source is a one-time page rather than a nag. The switch re-arms on the next heartbeat: come back and die again, and that is a new incident.
Two things to know
A switch's timeout must be shorter than the repeat_interval of the
route carrying the heartbeat, which is the exact opposite of
TERDUT_STALE_AFTER. Inheriting a default repeat_interval of 4h gives you a
switch that takes four hours to notice anything, so give the heartbeat
its own route. Matched alerts are exempt from
stale-alert expiry — a heartbeat answers to its own timeout and nothing else.
A dead man's switch incident has no member alerts:
GET /api/incidents/{id}/alerts returns an empty list. There is no alert
describing the problem, because the problem is that no alert arrived. What
happened is on the timeline instead, as a deadman_silent event carrying the age
of the last heartbeat, and the heartbeat's labels are on the incident's
group_labels.