• Turn incoming alerts into incidents
    Release / build (amd64, linux) (push) Failing after 11s
    Release / build (amd64, darwin) (push) Failing after 12s
    Release / build (arm64, darwin) (push) Failing after 11s
    Release / build (arm64, linux) (push) Failing after 11s
    Release / release (push) Has been skipped
    Release / chart (push) Failing after 13s
    Release / docker (push) Failing after 19s

    niklas released this 2026-07-30 15:02:13 +00:00 | 110 commits to main since this release

    The alerts row was both Alertmanager's record and the human work queue, and
    the two have different owners. The webhook upsert rewrites that row on every
    notification; acknowledgement, comments and archiving were columns on it that
    the upsert happened not to touch. So an alert that resolved and re-fired days
    later still read as acknowledged by whoever acked the first occurrence — the
    ack outlived the thing it referred to. Nothing recorded transitions either:
    rows are mutated in place, so there was no timeline and no way to compute how
    long anything took.

    Alerts are now read-only signal records with two states, and incidents are
    the work item: triggered, acknowledged or resolved, with an assignee, a
    snooze, notes and an append-only timeline. Many alerts map to one incident,
    and a new occurrence opens a new incident, which is what makes a stale ack
    impossible rather than merely unlikely.

    Correlation uses Alertmanager's own groupKey. It already grouped the alerts
    according to the group_by routing tree the operator configured and sends the
    result on every webhook, where it was being discarded; adopting it means
    changing group_by in alertmanager.yml changes correlation here, with no
    second grouping scheme to configure and keep in sync.

    An incident opens only when an alert transitions into firing — an unseen
    fingerprint, a newer startsAt, or a resolved alert starting again. The
    unchanged notifications Alertmanager re-sends every repeat_interval are none
    of those. That rule is what lets manual resolution be terminal: without it,
    closing an incident by hand would be undone by the next re-send of an alert
    that never stopped firing, and the button would be a lie. Snooze covers the
    "not now" case instead. Incidents otherwise resolve by cascade, once every
    alert under them has stopped firing, whether by webhook or by expiry.

    New incidents are assigned to whoever holds today's schedule entry. The
    schedule table has existed since the first release with nothing reading it.

    Also here, following from the split:

    • Incident severity is a high-water mark over its alerts, never lowered.
      An incident that hit critical was a critical incident, and downgrading a
      live one would demote it in the queue while the work is still open.
    • /api/stats/incidents reports MTTA and MTTR, null rather than zero until
      there is something to average. Neither was computable before.
    • Alert archiving becomes sweeper-only housekeeping; the archive people
      interact with is the incident's.

    Breaking: the alert acknowledge, archive and comment endpoints are gone, and
    the alert object drops the acknowledgement fields and gains incident_id. The
    README maps each removed endpoint to its replacement. Migration 008 backfills
    an incident per existing alert, archived ones included so no comment is
    orphaned, carrying acknowledgements across and turning comments into timeline
    notes.

    Both documented alert contracts are untouched: received_at still advances on
    every accepted payload, re-sends included, and resolution_source still says
    how much to trust ends_at. The upsert is byte-for-byte what it was, now
    running inside the ingest transaction.

    Downloads