-
Expire stale firing alerts
Release / build (amd64, darwin) (push) Failing after 12sRelease / build (arm64, darwin) (push) Failing after 11sRelease / build (arm64, linux) (push) Failing after 11sRelease / release (push) Has been skippedRelease / docker (push) Failing after 19sRelease / build (amd64, linux) (push) Failing after 12sRelease / chart (push) Failing after 9sreleased this
2026-07-28 09:49:39 +00:00 | 113 commits to main since this releaseA resolved webhook was the only path out of the firing state, so a
notification that was dropped, silenced, or lost to a restart pinned an
alert as firing forever — Prometheus showed it resolved while
terdut-server kept listing it. The archiver only ever touched resolved
alerts, and both the list and stats queries compared status with plain
equality, so a stale row was indistinguishable from a live one.A sweeper pass now resolves firing alerts on either of two signals: the
ends_at watermark Alertmanager sets on outgoing firing notifications has
passed (plus a grace period for clock skew), or no webhook has refreshed
the alert within TERDUT_STALE_AFTER (default 6h, above Alertmanager's 4h
repeat_interval). Such alerts get resolution_source = 'expiry',
distinguishing them from a real 'alertmanager' resolve.Two related webhook bugs fixed alongside:
- The upsert had no ordering guard, so a retried firing notification
arriving after the resolved one resurrected the alert. Payloads for
an older alert instance are now discarded: a stale retry carries the
same startsAt, a genuine re-fire a newer one. - archived_at was never cleared on re-fire, leaving a re-fired alert
archived and invisible in the default list.
Stats now exclude archived alerts to match the default list view; this
lowers historical firing/resolved totals.The chart exposes both sweeper durations via sweeper.staleAfter and
sweeper.archiveAfter.Downloads
- The upsert had no ordering guard, so a retried firing notification