Files
terdut-server/docs/incidents.md
T
Niklas Ye 44b2eb2cc3 Rewrite the README as highlights with screenshots; move the detail into docs/
The README was 1,240 lines of reference material and still described a
SQLite quick start. It is now a short tour (highlights, screenshots of the
web UI, an accurate quick start against Postgres), and each topic has its
own page under docs/ with an index: deployment, configuration, Alertmanager,
incidents, notifications, escalation, dead man's switches, single sign-on,
web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a
proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it
described. The "Upgrading to ..." sections for an unreleased product are
dropped.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:13 +02:00

5.7 KiB

Alerts and incidents

How alerts become incidents and how incidents are worked, notified, escalated and expired. Back to the README and the documentation index.

There are two objects, and the difference between them is the whole design.

An alert is Alertmanager's record. It has two states, firing and resolved, one row per fingerprint, and no human ever writes to it. The API exposes alerts read-only.

An incident is the work item. It goes triggered → acknowledged → resolved, carries an assignee, a snooze, notes and a timeline, and is the only thing people act on. Many alerts belong to one incident.

Correlation uses Alertmanager's groupKey

Alertmanager has already grouped alerts according to the group_by routing tree you configured, and it sends the resulting groupKey and groupLabels on every webhook. Incidents adopt that answer rather than re-grouping alerts a second time — if you want different correlation, change group_by in alertmanager.yml and terdut follows.

At most one incident is open per groupKey at a time. Alerts firing in a group that already has an open incident join it. The incident's severity is a high-water mark — the highest severity label any of its alerts has carried — so an incident that hit critical still reads as critical after the critical alert clears.

Several clusters, one team

A team with one Alertmanager per Kubernetes cluster, each posting to its own source, needs two settings or the clusters run together.

  1. Give every alert a cluster label at the source. In Prometheus that is externalLabels: {cluster: prod-eu} (kube-prometheus-stack: prometheus.prometheusSpec.externalLabels).
  2. Add cluster to group_by in alertmanager.yml.

The second one is the one that matters. Incidents are matched on the team and Alertmanager's groupKey, and the groupKey does not include external labels: without cluster in group_by, the same alert in two clusters has the same key and joins one incident. With it, each cluster gets its own, cluster is in the incident's group_labels, and the web UI shows it as a coloured chip on the queue, the incident and the alert list, instead of leaving it in the title. An alert that is not grouped by cluster still shows the chip on the alert list, which reads the label from the alert itself.

The queue has a cluster dropdown once there are two or more values to choose between. It filters on the incident's cluster group label (GET /api/incidents?cluster=...), so it only sees incidents grouped by it.

An incident opens only on a new occurrence

An incident opens when an alert transitions into firing: a fingerprint that was never seen, an alert with a newer startsAt, or a resolved alert that started again. The unchanged firing notifications Alertmanager re-sends every repeat_interval are none of those, and open nothing.

This is what makes closing an incident by hand mean something. Without the rule, POST /api/incidents/{id}/resolve would be undone by the next re-send of an alert that never stopped firing.

Leaving the open state

  • Automatically, once every alert under the incident has stopped firing — whether by a resolved webhook or by the sweeper's stale-alert expiry. The incident gets "resolution_source": "alerts".
  • By hand, via POST /api/incidents/{id}/resolve ("resolution_source": "manual"). This is terminal: a later occurrence in that group opens a new incident rather than reopening this one. If the alert underneath never stops firing, the incident stays closed — that is what resolving by hand asserts.
  • On recovery, for a dead man's switch incident whose heartbeat started arriving again ("resolution_source": "recovered"). These incidents have no member alerts, so the automatic cascade above cannot reach them.

To quieten an incident you expect to come back, snooze it instead (POST /api/incidents/{id}/snooze). A snooze hides the incident from the default list without closing it, and expires by simply falling into the past.

On-call assignment

A new incident is assigned to whoever holds today's schedule entry at the moment it opens (GET /api/schedule/current). If nobody is scheduled it opens unassigned. Reassign with POST /api/incidents/{id}/assign.

One person holds a given day, so POST /api/schedule refuses a date somebody already has: taking a shift off the person expecting to be paged for it should not be something a plain call does by accident. Pass "replace": true to take them anyway. Either way the whole request is one transaction — a week where some days are free and some are taken moves as a unit, and a failure leaves the rota exactly as it was rather than with a hole in it.

Stale alert expiry

A resolved webhook is the only signal that an alert has stopped firing, so a notification that is dropped, silenced, or lost to a restart would otherwise pin that alert as firing forever. A background sweeper resolves firing alerts that Alertmanager has stopped refreshing, using either signal:

  • the endsAt watermark on the last notification has passed, or
  • no webhook has refreshed the alert within TERDUT_STALE_AFTER.

Alertmanager re-sends firing notifications every repeat_interval, which is what keeps a live alert fresh — so TERDUT_STALE_AFTER must be comfortably larger than your repeat_interval (default 4h), or live alerts will be resolved prematurely. Alerts resolved this way are marked "resolution_source": "expiry" to distinguish them from a real Alertmanager resolve ("alertmanager").

An expiry cascades: once it leaves an incident with nothing firing under it, the incident resolves too, in the same sweep.