44b2eb2cc3
The README was 1,240 lines of reference material and still described a SQLite quick start. It is now a short tour (highlights, screenshots of the web UI, an accurate quick start against Postgres), and each topic has its own page under docs/ with an index: deployment, configuration, Alertmanager, incidents, notifications, escalation, dead man's switches, single sign-on, web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it described. The "Upgrading to ..." sections for an unreleased product are dropped. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
114 lines
5.7 KiB
Markdown
114 lines
5.7 KiB
Markdown
# Alerts and incidents
|
|
|
|
_How alerts become incidents and how incidents are worked, notified, escalated and expired._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
|
|
|
There are two objects, and the difference between them is the whole design.
|
|
|
|
**An alert is Alertmanager's record.** It has two states, `firing` and
|
|
`resolved`, one row per fingerprint, and no human ever writes to it. The API
|
|
exposes alerts read-only.
|
|
|
|
**An incident is the work item.** It goes `triggered → acknowledged → resolved`,
|
|
carries an assignee, a snooze, notes and a timeline, and is the only thing people
|
|
act on. Many alerts belong to one incident.
|
|
|
|
## Correlation uses Alertmanager's `groupKey`
|
|
|
|
Alertmanager has already grouped alerts according to the `group_by` routing tree
|
|
you configured, and it sends the resulting `groupKey` and `groupLabels` on every
|
|
webhook. Incidents adopt that answer rather than re-grouping alerts a second
|
|
time — if you want different correlation, change `group_by` in
|
|
`alertmanager.yml` and terdut follows.
|
|
|
|
At most one incident is open per `groupKey` at a time. Alerts firing in a group
|
|
that already has an open incident join it. The incident's `severity` is a
|
|
high-water mark — the highest `severity` label any of its alerts has carried — so
|
|
an incident that hit `critical` still reads as critical after the critical alert
|
|
clears.
|
|
|
|
## Several clusters, one team
|
|
|
|
A team with one Alertmanager per Kubernetes cluster, each posting to its own
|
|
source, needs two settings or the clusters run together.
|
|
|
|
1. Give every alert a `cluster` label at the source. In Prometheus that is
|
|
`externalLabels: {cluster: prod-eu}` (kube-prometheus-stack:
|
|
`prometheus.prometheusSpec.externalLabels`).
|
|
2. Add `cluster` to `group_by` in `alertmanager.yml`.
|
|
|
|
The second one is the one that matters. Incidents are matched on the team and
|
|
Alertmanager's `groupKey`, and the `groupKey` does not include external labels:
|
|
without `cluster` in `group_by`, the same alert in two clusters has the same
|
|
key and joins one incident. With it, each cluster gets its own, `cluster` is in
|
|
the incident's `group_labels`, and the web UI shows it as a coloured chip on the
|
|
queue, the incident and the alert list, instead of leaving it in the title.
|
|
An alert that is not grouped by `cluster` still shows the chip on the alert
|
|
list, which reads the label from the alert itself.
|
|
|
|
The queue has a cluster dropdown once there are two or more values to choose
|
|
between. It filters on the incident's `cluster` group label
|
|
(`GET /api/incidents?cluster=...`), so it only sees incidents grouped by it.
|
|
|
|
## An incident opens only on a new occurrence
|
|
|
|
An incident opens when an alert **transitions into firing**: a fingerprint that
|
|
was never seen, an alert with a newer `startsAt`, or a resolved alert that
|
|
started again. The unchanged firing notifications Alertmanager re-sends every
|
|
`repeat_interval` are none of those, and open nothing.
|
|
|
|
This is what makes closing an incident by hand mean something. Without the rule,
|
|
`POST /api/incidents/{id}/resolve` would be undone by the next re-send of an
|
|
alert that never stopped firing.
|
|
|
|
## Leaving the open state
|
|
|
|
- **Automatically**, once every alert under the incident has stopped firing —
|
|
whether by a resolved webhook or by the sweeper's
|
|
[stale-alert expiry](#stale-alert-expiry). The incident gets
|
|
`"resolution_source": "alerts"`.
|
|
- **By hand**, via `POST /api/incidents/{id}/resolve`
|
|
(`"resolution_source": "manual"`). This is **terminal**: a later occurrence in
|
|
that group opens a *new* incident rather than reopening this one. If the alert
|
|
underneath never stops firing, the incident stays closed — that is what
|
|
resolving by hand asserts.
|
|
- **On recovery**, for a [dead man's switch](./dead-mans-switch.md) incident whose
|
|
heartbeat started arriving again (`"resolution_source": "recovered"`). These
|
|
incidents have no member alerts, so the automatic cascade above cannot reach
|
|
them.
|
|
|
|
To quieten an incident you expect to come back, snooze it instead
|
|
(`POST /api/incidents/{id}/snooze`). A snooze hides the incident from the default
|
|
list without closing it, and expires by simply falling into the past.
|
|
|
|
## On-call assignment
|
|
|
|
A new incident is assigned to whoever holds today's schedule entry at the moment
|
|
it opens (`GET /api/schedule/current`). If nobody is scheduled it opens
|
|
unassigned. Reassign with `POST /api/incidents/{id}/assign`.
|
|
|
|
One person holds a given day, so `POST /api/schedule` refuses a date somebody
|
|
already has: taking a shift off the person expecting to be paged for it should
|
|
not be something a plain call does by accident. Pass `"replace": true` to take
|
|
them anyway. Either way the whole request is one transaction — a week where some
|
|
days are free and some are taken moves as a unit, and a failure leaves the rota
|
|
exactly as it was rather than with a hole in it.
|
|
|
|
## Stale alert expiry
|
|
|
|
A resolved webhook is the only signal that an alert has stopped firing, so a
|
|
notification that is dropped, silenced, or lost to a restart would otherwise pin
|
|
that alert as firing forever. A background sweeper resolves firing alerts that
|
|
Alertmanager has stopped refreshing, using either signal:
|
|
|
|
- the `endsAt` watermark on the last notification has passed, or
|
|
- no webhook has refreshed the alert within `TERDUT_STALE_AFTER`.
|
|
|
|
Alertmanager re-sends firing notifications every `repeat_interval`, which is what
|
|
keeps a live alert fresh — so `TERDUT_STALE_AFTER` must be comfortably larger
|
|
than your `repeat_interval` (default 4h), or live alerts will be resolved
|
|
prematurely. Alerts resolved this way are marked `"resolution_source": "expiry"`
|
|
to distinguish them from a real Alertmanager resolve (`"alertmanager"`).
|
|
|
|
An expiry cascades: once it leaves an incident with nothing firing under it, the
|
|
incident resolves too, in the same sweep.
|