• v0.9.0 14c24f8fda

    Notice when the Watchdog alert stops arriving
    Release / release (push) Has been skipped
    Release / build (amd64, linux) (push) Has been skipped
    Release / build (arm64, darwin) (push) Has been skipped
    Release / test (push) Failing after 5s
    Release / build (arm64, linux) (push) Has been skipped
    Release / docker (push) Has been skipped
    Release / chart (push) Has been skipped
    Release / build (amd64, darwin) (push) Has been skipped

    niklas released this 2026-08-08 19:28:55 +00:00 | 99 commits to main since this release

    Everything this server does assumes alerts arrive. If Prometheus stops
    evaluating, or Alertmanager cannot reach us, nothing arrives — and
    silence is indistinguishable from everything being fine. The cluster has
    shipped the alert for exactly this case all along: Watchdog is
    expr: vector(1), so it fires permanently and is re-sent forever, and it
    is worth nothing unless something downstream notices it stop. Nothing
    did. It arrived, opened no incident because a repeat_interval re-send is
    not a new occurrence, and when the monitoring stack died the sweeper
    quietly expired it and paged nobody.

    So the handling is inverted for a configurable set of alerts: receiving
    one opens no incident, and the absence of one does. TERDUT_DEADMAN_MATCHERS
    selects them as label matchers, defaulting to alertname=Watchdog.

    The unit of monitoring is the fingerprint rather than the alert name. Two
    clusters sending the same Watchdog are two independent switches, so a
    healthy one can never mask a dead one. Every matcher must name an
    alertname, which keeps the sweeper's candidate query on alerts_name_idx
    instead of JSON-extracting labels from every row, and leaves matching
    with a single implementation.

    A switch is dormant until its first heartbeat: a matcher nothing has ever
    sent opens nothing, so a fresh deploy or a restored database does not
    page. Resolving the incident by hand sticks, exactly as it does for an
    alert-backed one, so a decommissioned source is a one-time page rather
    than a nag; the switch re-arms only when the heartbeat comes back, and
    dying again is a new incident.

    The incident has no member alerts on purpose. Linking the heartbeat would
    have the settled-incident cascade close it on the very sweep that opened
    it, and there is no alert describing the problem anyway — the problem is
    that no alert arrived. What happened is on the timeline instead, and
    recovery is the only automatic way out.

    One narrow exemption in the ingest guard makes recovery possible at all.
    A heartbeat we declared dead is marked resolved, and the one that proves
    us wrong carries the unchanged startsAt of an alert that never stopped
    firing — so "resolution is terminal within an instance" would discard it
    forever and a switch could die exactly once. The exemption is scoped to
    resolution_source = 'deadman', which is the only resolution this server
    infers from silence on a timeout of its own, so nothing another writer
    set can be undone by a stale retry. Matched alerts are also held back
    from the generic staleness expiry, which would otherwise resolve a
    heartbeat as 'expiry' long before its own tighter deadline.

    The timeout points the opposite way to TERDUT_STALE_AFTER: staleness is a
    generous grace period around a repeat_interval you do not control, while
    this is a deadline you set deliberately and configure the heartbeat's
    route to beat. Inheriting a 4h or 12h repeat_interval gives a dead man's
    switch with a twelve hour fuse, so the README spells out the route the
    heartbeat needs.

    Downloads