42e846f876
Release / build (amd64, darwin) (push) Failing after 12s
Release / build (arm64, darwin) (push) Failing after 11s
Release / build (arm64, linux) (push) Failing after 11s
Release / release (push) Has been skipped
Release / docker (push) Failing after 19s
Release / build (amd64, linux) (push) Failing after 12s
Release / chart (push) Failing after 9s
A resolved webhook was the only path out of the firing state, so a
notification that was dropped, silenced, or lost to a restart pinned an
alert as firing forever — Prometheus showed it resolved while
terdut-server kept listing it. The archiver only ever touched resolved
alerts, and both the list and stats queries compared status with plain
equality, so a stale row was indistinguishable from a live one.
A sweeper pass now resolves firing alerts on either of two signals: the
ends_at watermark Alertmanager sets on outgoing firing notifications has
passed (plus a grace period for clock skew), or no webhook has refreshed
the alert within TERDUT_STALE_AFTER (default 6h, above Alertmanager's 4h
repeat_interval). Such alerts get resolution_source = 'expiry',
distinguishing them from a real 'alertmanager' resolve.
Two related webhook bugs fixed alongside:
- The upsert had no ordering guard, so a retried firing notification
arriving after the resolved one resurrected the alert. Payloads for
an older alert instance are now discarded: a stale retry carries the
same startsAt, a genuine re-fire a newer one.
- archived_at was never cleared on re-fire, leaving a re-fired alert
archived and invisible in the default list.
Stats now exclude archived alerts to match the default list view; this
lowers historical firing/resolved totals.
The chart exposes both sweeper durations via sweeper.staleAfter and
sweeper.archiveAfter.
42 lines
984 B
Go
42 lines
984 B
Go
package config
|
|
|
|
import (
|
|
"os"
|
|
"time"
|
|
)
|
|
|
|
type Config struct {
|
|
Addr string
|
|
DBPath string
|
|
ArchiveAfter time.Duration
|
|
|
|
// StaleAfter is how long a firing alert may go without a refreshing webhook
|
|
// before the sweeper treats it as resolved. It must exceed Alertmanager's
|
|
// repeat_interval (default 4h), which is what refreshes the alert.
|
|
StaleAfter time.Duration
|
|
}
|
|
|
|
func Load() Config {
|
|
addr := os.Getenv("TERDUT_ADDR")
|
|
if addr == "" {
|
|
addr = ":8080"
|
|
}
|
|
dbPath := os.Getenv("TERDUT_DB_PATH")
|
|
if dbPath == "" {
|
|
dbPath = "terdut.db"
|
|
}
|
|
archiveAfter := 7 * 24 * time.Hour
|
|
if s := os.Getenv("TERDUT_ARCHIVE_AFTER"); s != "" {
|
|
if d, err := time.ParseDuration(s); err == nil {
|
|
archiveAfter = d
|
|
}
|
|
}
|
|
staleAfter := 6 * time.Hour
|
|
if s := os.Getenv("TERDUT_STALE_AFTER"); s != "" {
|
|
if d, err := time.ParseDuration(s); err == nil {
|
|
staleAfter = d
|
|
}
|
|
}
|
|
return Config{Addr: addr, DBPath: dbPath, ArchiveAfter: archiveAfter, StaleAfter: staleAfter}
|
|
}
|