# Alerts and incidents _How alerts become incidents and how incidents are worked, notified, escalated and expired._ Back to the [README](../README.md) and the [documentation index](./README.md). There are two objects, and the difference between them is the whole design. **An alert is Alertmanager's record.** It has two states, `firing` and `resolved`, one row per fingerprint, and no human ever writes to it. The API exposes alerts read-only. **An incident is the work item.** It goes `triggered → acknowledged → resolved`, carries an assignee, a snooze, notes and a timeline, and is the only thing people act on. Many alerts belong to one incident. ## Correlation uses Alertmanager's `groupKey` Alertmanager has already grouped alerts according to the `group_by` routing tree you configured, and it sends the resulting `groupKey` and `groupLabels` on every webhook. Incidents adopt that answer rather than re-grouping alerts a second time — if you want different correlation, change `group_by` in `alertmanager.yml` and terdut follows. At most one incident is open per `groupKey` at a time. Alerts firing in a group that already has an open incident join it. The incident's `severity` is a high-water mark — the highest `severity` label any of its alerts has carried — so an incident that hit `critical` still reads as critical after the critical alert clears. ## Several clusters, one team A team with one Alertmanager per Kubernetes cluster, each posting to its own source, needs two settings or the clusters run together. 1. Give every alert a `cluster` label at the source. In Prometheus that is `externalLabels: {cluster: prod-eu}` (kube-prometheus-stack: `prometheus.prometheusSpec.externalLabels`). 2. Add `cluster` to `group_by` in `alertmanager.yml`. The second one is the one that matters. Incidents are matched on the team and Alertmanager's `groupKey`, and the `groupKey` does not include external labels: without `cluster` in `group_by`, the same alert in two clusters has the same key and joins one incident. With it, each cluster gets its own, `cluster` is in the incident's `group_labels`, and the web UI shows it as a coloured chip on the queue, the incident and the alert list, instead of leaving it in the title. An alert that is not grouped by `cluster` still shows the chip on the alert list, which reads the label from the alert itself. The queue has a cluster dropdown once there are two or more values to choose between. It filters on the incident's `cluster` group label (`GET /api/incidents?cluster=...`), so it only sees incidents grouped by it. ## An incident opens only on a new occurrence An incident opens when an alert **transitions into firing**: a fingerprint that was never seen, an alert with a newer `startsAt`, or a resolved alert that started again. The unchanged firing notifications Alertmanager re-sends every `repeat_interval` are none of those, and open nothing. This is what makes closing an incident by hand mean something. Without the rule, `POST /api/incidents/{id}/resolve` would be undone by the next re-send of an alert that never stopped firing. ## Leaving the open state - **Automatically**, once every alert under the incident has stopped firing — whether by a resolved webhook or by the sweeper's [stale-alert expiry](#stale-alert-expiry). The incident gets `"resolution_source": "alerts"`. - **By hand**, via `POST /api/incidents/{id}/resolve` (`"resolution_source": "manual"`). This is **terminal**: a later occurrence in that group opens a *new* incident rather than reopening this one. If the alert underneath never stops firing, the incident stays closed — that is what resolving by hand asserts. - **On recovery**, for a [dead man's switch](./dead-mans-switch.md) incident whose heartbeat started arriving again (`"resolution_source": "recovered"`). These incidents have no member alerts, so the automatic cascade above cannot reach them. To quieten an incident you expect to come back, snooze it instead (`POST /api/incidents/{id}/snooze`). A snooze hides the incident from the default list without closing it, and expires by simply falling into the past. ## On-call assignment A new incident is assigned to whoever holds today's schedule entry at the moment it opens (`GET /api/schedule/current`). If nobody is scheduled it opens unassigned. Reassign with `POST /api/incidents/{id}/assign`. One person holds a given day, so `POST /api/schedule` refuses a date somebody already has: taking a shift off the person expecting to be paged for it should not be something a plain call does by accident. Pass `"replace": true` to take them anyway. Either way the whole request is one transaction — a week where some days are free and some are taken moves as a unit, and a failure leaves the rota exactly as it was rather than with a hole in it. ## Stale alert expiry A resolved webhook is the only signal that an alert has stopped firing, so a notification that is dropped, silenced, or lost to a restart would otherwise pin that alert as firing forever. A background sweeper resolves firing alerts that Alertmanager has stopped refreshing, using either signal: - the `endsAt` watermark on the last notification has passed, or - no webhook has refreshed the alert within `TERDUT_STALE_AFTER`. Alertmanager re-sends firing notifications every `repeat_interval`, which is what keeps a live alert fresh — so `TERDUT_STALE_AFTER` must be comfortably larger than your `repeat_interval` (default 4h), or live alerts will be resolved prematurely. Alerts resolved this way are marked `"resolution_source": "expiry"` to distinguish them from a real Alertmanager resolve (`"alertmanager"`). An expiry cascades: once it leaves an incident with nothing firing under it, the incident resolves too, in the same sweep.