Files
terdut-server/docs/dead-mans-switch.md
T
Niklas Ye 44b2eb2cc3 Rewrite the README as highlights with screenshots; move the detail into docs/
The README was 1,240 lines of reference material and still described a
SQLite quick start. It is now a short tour (highlights, screenshots of the
web UI, an accurate quick start against Postgres), and each topic has its
own page under docs/ with an index: deployment, configuration, Alertmanager,
incidents, notifications, escalation, dead man's switches, single sign-on,
web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a
proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it
described. The "Upgrading to ..." sections for an unreleased product are
dropped.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:13 +02:00

82 lines
4.1 KiB
Markdown

# Dead man's switches
_Detecting that alerts have stopped arriving._ Back to the [README](../README.md) and the [documentation index](./README.md).
Everything above assumes alerts arrive. If Prometheus stops evaluating, or
Alertmanager cannot reach this server, nothing arrives — and silence looks
exactly like everything being fine. A dead man's switch inverts the handling for
one designated alert so that silence is the signal:
- **receiving** it opens no incident, and
- the **absence** of it does.
kube-prometheus-stack already ships the alert for this. `Watchdog` is
`expr: vector(1)`, so it fires permanently and is re-sent forever; it is worth
nothing unless something downstream notices it stop. It is the usual first switch.
**Switches belong to a team**, which decides which of its own alerts are
heartbeats and how long a silence has to last. Each **switch** is a row of its
own — a name, one matcher, a timeout and a severity — so switches in one team
can have different deadlines. An owner adds and removes them on **Team →
Switches**, which lists each with a status (**healthy**, **dead**, or
**dormant** until its first heartbeat), when it was last heard from, and when it
last opened an incident; a matcher that several clusters satisfy is broken down
per cluster. The API is `POST`/`DELETE /api/teams/{teamID}/deadman/switches`. A
missed heartbeat opens an incident in the team whose integration received it.
Removing a switch stops the watching; an incident it already opened stays open
until somebody resolves it.
A new team watches nothing until its owner (or terdut-operator, from a
`TerdutTeam`) adds a switch: inheriting an install-wide heartbeat would page a
new team about a source it has never heard of.
A matcher is a set of exact label conditions, one of which must be the
`alertname`, , one matcher per switch, `,` between the label conditions:
```
alertname=Watchdog,cluster=prod
```
**The unit of monitoring is the fingerprint, not the alert name.** Two clusters
sending the same `Watchdog` are two independent switches, so a healthy one can
never mask a dead one.
## The lifecycle
A switch is **dormant** until its first heartbeat arrives. A configured matcher
that has never been heard from opens nothing, so a fresh deploy or a restored
database does not page. It also means a matcher that never matches anything is
silently inert.
Once armed, the sweeper declares it **dead** when either the heartbeat has not
been refreshed within the switch's `timeout_seconds`, or Alertmanager explicitly
resolved it — the sender saying the heartbeat stopped needs no further waiting.
That opens an incident at the switch's `severity`, assigned and paged like any
other, and marks the heartbeat alert `"resolution_source": "deadman"` so the
alert list stops claiming a dead switch is firing.
It **recovers** when the heartbeat starts arriving again: the incident resolves
with `"resolution_source": "recovered"` and the all-clear goes to whoever was
paged.
Resolving the incident by hand sticks, the same way it does for an alert-backed
one. While the switch stays silent nothing new opens — so a decommissioned
source is a one-time page rather than a nag. The switch **re-arms** on the next
heartbeat: come back and die again, and that is a new incident.
## Two things to know
A switch's timeout must be **shorter** than the `repeat_interval` of the
route carrying the heartbeat, which is the exact opposite of
`TERDUT_STALE_AFTER`. Inheriting a default `repeat_interval` of 4h gives you a
switch that takes four hours to notice anything, so give the heartbeat
[its own route](./alertmanager.md#alertmanager-configuration). Matched alerts are exempt from
stale-alert expiry — a heartbeat answers to its own timeout and nothing else.
A dead man's switch incident has **no member alerts**:
`GET /api/incidents/{id}/alerts` returns an empty list. There is no alert
describing the problem, because the problem is that no alert arrived. What
happened is on the timeline instead, as a `deadman_silent` event carrying the age
of the last heartbeat, and the heartbeat's labels are on the incident's
`group_labels`.