44b2eb2cc3
The README was 1,240 lines of reference material and still described a SQLite quick start. It is now a short tour (highlights, screenshots of the web UI, an accurate quick start against Postgres), and each topic has its own page under docs/ with an index: deployment, configuration, Alertmanager, incidents, notifications, escalation, dead man's switches, single sign-on, web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it described. The "Upgrading to ..." sections for an unreleased product are dropped. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
82 lines
4.1 KiB
Markdown
82 lines
4.1 KiB
Markdown
# Dead man's switches
|
|
|
|
_Detecting that alerts have stopped arriving._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
|
|
|
Everything above assumes alerts arrive. If Prometheus stops evaluating, or
|
|
Alertmanager cannot reach this server, nothing arrives — and silence looks
|
|
exactly like everything being fine. A dead man's switch inverts the handling for
|
|
one designated alert so that silence is the signal:
|
|
|
|
- **receiving** it opens no incident, and
|
|
- the **absence** of it does.
|
|
|
|
kube-prometheus-stack already ships the alert for this. `Watchdog` is
|
|
`expr: vector(1)`, so it fires permanently and is re-sent forever; it is worth
|
|
nothing unless something downstream notices it stop. It is the usual first switch.
|
|
|
|
**Switches belong to a team**, which decides which of its own alerts are
|
|
heartbeats and how long a silence has to last. Each **switch** is a row of its
|
|
own — a name, one matcher, a timeout and a severity — so switches in one team
|
|
can have different deadlines. An owner adds and removes them on **Team →
|
|
Switches**, which lists each with a status (**healthy**, **dead**, or
|
|
**dormant** until its first heartbeat), when it was last heard from, and when it
|
|
last opened an incident; a matcher that several clusters satisfy is broken down
|
|
per cluster. The API is `POST`/`DELETE /api/teams/{teamID}/deadman/switches`. A
|
|
missed heartbeat opens an incident in the team whose integration received it.
|
|
Removing a switch stops the watching; an incident it already opened stays open
|
|
until somebody resolves it.
|
|
|
|
A new team watches nothing until its owner (or terdut-operator, from a
|
|
`TerdutTeam`) adds a switch: inheriting an install-wide heartbeat would page a
|
|
new team about a source it has never heard of.
|
|
|
|
A matcher is a set of exact label conditions, one of which must be the
|
|
`alertname`, , one matcher per switch, `,` between the label conditions:
|
|
|
|
```
|
|
alertname=Watchdog,cluster=prod
|
|
```
|
|
|
|
**The unit of monitoring is the fingerprint, not the alert name.** Two clusters
|
|
sending the same `Watchdog` are two independent switches, so a healthy one can
|
|
never mask a dead one.
|
|
|
|
## The lifecycle
|
|
|
|
A switch is **dormant** until its first heartbeat arrives. A configured matcher
|
|
that has never been heard from opens nothing, so a fresh deploy or a restored
|
|
database does not page. It also means a matcher that never matches anything is
|
|
silently inert.
|
|
|
|
Once armed, the sweeper declares it **dead** when either the heartbeat has not
|
|
been refreshed within the switch's `timeout_seconds`, or Alertmanager explicitly
|
|
resolved it — the sender saying the heartbeat stopped needs no further waiting.
|
|
That opens an incident at the switch's `severity`, assigned and paged like any
|
|
other, and marks the heartbeat alert `"resolution_source": "deadman"` so the
|
|
alert list stops claiming a dead switch is firing.
|
|
|
|
It **recovers** when the heartbeat starts arriving again: the incident resolves
|
|
with `"resolution_source": "recovered"` and the all-clear goes to whoever was
|
|
paged.
|
|
|
|
Resolving the incident by hand sticks, the same way it does for an alert-backed
|
|
one. While the switch stays silent nothing new opens — so a decommissioned
|
|
source is a one-time page rather than a nag. The switch **re-arms** on the next
|
|
heartbeat: come back and die again, and that is a new incident.
|
|
|
|
## Two things to know
|
|
|
|
A switch's timeout must be **shorter** than the `repeat_interval` of the
|
|
route carrying the heartbeat, which is the exact opposite of
|
|
`TERDUT_STALE_AFTER`. Inheriting a default `repeat_interval` of 4h gives you a
|
|
switch that takes four hours to notice anything, so give the heartbeat
|
|
[its own route](./alertmanager.md#alertmanager-configuration). Matched alerts are exempt from
|
|
stale-alert expiry — a heartbeat answers to its own timeout and nothing else.
|
|
|
|
A dead man's switch incident has **no member alerts**:
|
|
`GET /api/incidents/{id}/alerts` returns an empty list. There is no alert
|
|
describing the problem, because the problem is that no alert arrived. What
|
|
happened is on the timeline instead, as a `deadman_silent` event carrying the age
|
|
of the last heartbeat, and the heartbeat's labels are on the incident's
|
|
`group_labels`.
|