Files
terdut-server/docs/dead-mans-switch.md
T
Niklas Ye 44b2eb2cc3 Rewrite the README as highlights with screenshots; move the detail into docs/
The README was 1,240 lines of reference material and still described a
SQLite quick start. It is now a short tour (highlights, screenshots of the
web UI, an accurate quick start against Postgres), and each topic has its
own page under docs/ with an index: deployment, configuration, Alertmanager,
incidents, notifications, escalation, dead man's switches, single sign-on,
web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a
proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it
described. The "Upgrading to ..." sections for an unreleased product are
dropped.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:13 +02:00

4.1 KiB

Dead man's switches

Detecting that alerts have stopped arriving. Back to the README and the documentation index.

Everything above assumes alerts arrive. If Prometheus stops evaluating, or Alertmanager cannot reach this server, nothing arrives — and silence looks exactly like everything being fine. A dead man's switch inverts the handling for one designated alert so that silence is the signal:

  • receiving it opens no incident, and
  • the absence of it does.

kube-prometheus-stack already ships the alert for this. Watchdog is expr: vector(1), so it fires permanently and is re-sent forever; it is worth nothing unless something downstream notices it stop. It is the usual first switch.

Switches belong to a team, which decides which of its own alerts are heartbeats and how long a silence has to last. Each switch is a row of its own — a name, one matcher, a timeout and a severity — so switches in one team can have different deadlines. An owner adds and removes them on Team → Switches, which lists each with a status (healthy, dead, or dormant until its first heartbeat), when it was last heard from, and when it last opened an incident; a matcher that several clusters satisfy is broken down per cluster. The API is POST/DELETE /api/teams/{teamID}/deadman/switches. A missed heartbeat opens an incident in the team whose integration received it. Removing a switch stops the watching; an incident it already opened stays open until somebody resolves it.

A new team watches nothing until its owner (or terdut-operator, from a TerdutTeam) adds a switch: inheriting an install-wide heartbeat would page a new team about a source it has never heard of.

A matcher is a set of exact label conditions, one of which must be the alertname, , one matcher per switch, , between the label conditions:

alertname=Watchdog,cluster=prod

The unit of monitoring is the fingerprint, not the alert name. Two clusters sending the same Watchdog are two independent switches, so a healthy one can never mask a dead one.

The lifecycle

A switch is dormant until its first heartbeat arrives. A configured matcher that has never been heard from opens nothing, so a fresh deploy or a restored database does not page. It also means a matcher that never matches anything is silently inert.

Once armed, the sweeper declares it dead when either the heartbeat has not been refreshed within the switch's timeout_seconds, or Alertmanager explicitly resolved it — the sender saying the heartbeat stopped needs no further waiting. That opens an incident at the switch's severity, assigned and paged like any other, and marks the heartbeat alert "resolution_source": "deadman" so the alert list stops claiming a dead switch is firing.

It recovers when the heartbeat starts arriving again: the incident resolves with "resolution_source": "recovered" and the all-clear goes to whoever was paged.

Resolving the incident by hand sticks, the same way it does for an alert-backed one. While the switch stays silent nothing new opens — so a decommissioned source is a one-time page rather than a nag. The switch re-arms on the next heartbeat: come back and die again, and that is a new incident.

Two things to know

A switch's timeout must be shorter than the repeat_interval of the route carrying the heartbeat, which is the exact opposite of TERDUT_STALE_AFTER. Inheriting a default repeat_interval of 4h gives you a switch that takes four hours to notice anything, so give the heartbeat its own route. Matched alerts are exempt from stale-alert expiry — a heartbeat answers to its own timeout and nothing else.

A dead man's switch incident has no member alerts: GET /api/incidents/{id}/alerts returns an empty list. There is no alert describing the problem, because the problem is that no alert arrived. What happened is on the timeline instead, as a deadman_silent event carrying the age of the last heartbeat, and the heartbeat's labels are on the incident's group_labels.