# Dead man's switches _Detecting that alerts have stopped arriving._ Back to the [README](../README.md) and the [documentation index](./README.md). Everything above assumes alerts arrive. If Prometheus stops evaluating, or Alertmanager cannot reach this server, nothing arrives — and silence looks exactly like everything being fine. A dead man's switch inverts the handling for one designated alert so that silence is the signal: - **receiving** it opens no incident, and - the **absence** of it does. kube-prometheus-stack already ships the alert for this. `Watchdog` is `expr: vector(1)`, so it fires permanently and is re-sent forever; it is worth nothing unless something downstream notices it stop. It is the usual first switch. **Switches belong to a team**, which decides which of its own alerts are heartbeats and how long a silence has to last. Each **switch** is a row of its own — a name, one matcher, a timeout and a severity — so switches in one team can have different deadlines. An owner adds and removes them on **Team → Switches**, which lists each with a status (**healthy**, **dead**, or **dormant** until its first heartbeat), when it was last heard from, and when it last opened an incident; a matcher that several clusters satisfy is broken down per cluster. The API is `POST`/`DELETE /api/teams/{teamID}/deadman/switches`. A missed heartbeat opens an incident in the team whose integration received it. Removing a switch stops the watching; an incident it already opened stays open until somebody resolves it. A new team watches nothing until its owner (or terdut-operator, from a `TerdutTeam`) adds a switch: inheriting an install-wide heartbeat would page a new team about a source it has never heard of. A matcher is a set of exact label conditions, one of which must be the `alertname`, , one matcher per switch, `,` between the label conditions: ``` alertname=Watchdog,cluster=prod ``` **The unit of monitoring is the fingerprint, not the alert name.** Two clusters sending the same `Watchdog` are two independent switches, so a healthy one can never mask a dead one. ## The lifecycle A switch is **dormant** until its first heartbeat arrives. A configured matcher that has never been heard from opens nothing, so a fresh deploy or a restored database does not page. It also means a matcher that never matches anything is silently inert. Once armed, the sweeper declares it **dead** when either the heartbeat has not been refreshed within the switch's `timeout_seconds`, or Alertmanager explicitly resolved it — the sender saying the heartbeat stopped needs no further waiting. That opens an incident at the switch's `severity`, assigned and paged like any other, and marks the heartbeat alert `"resolution_source": "deadman"` so the alert list stops claiming a dead switch is firing. It **recovers** when the heartbeat starts arriving again: the incident resolves with `"resolution_source": "recovered"` and the all-clear goes to whoever was paged. Resolving the incident by hand sticks, the same way it does for an alert-backed one. While the switch stays silent nothing new opens — so a decommissioned source is a one-time page rather than a nag. The switch **re-arms** on the next heartbeat: come back and die again, and that is a new incident. ## Two things to know A switch's timeout must be **shorter** than the `repeat_interval` of the route carrying the heartbeat, which is the exact opposite of `TERDUT_STALE_AFTER`. Inheriting a default `repeat_interval` of 4h gives you a switch that takes four hours to notice anything, so give the heartbeat [its own route](./alertmanager.md#alertmanager-configuration). Matched alerts are exempt from stale-alert expiry — a heartbeat answers to its own timeout and nothing else. A dead man's switch incident has **no member alerts**: `GET /api/incidents/{id}/alerts` returns an empty list. There is no alert describing the problem, because the problem is that no alert arrived. What happened is on the timeline instead, as a `deadman_silent` event carrying the age of the last heartbeat, and the heartbeat's labels are on the incident's `group_labels`.