# Terminal Duty (terdut-server) Incident management server for teams using Prometheus Alertmanager. - Receives Alertmanager webhooks directly — no adapter needed - Turns alerts into **incidents**, correlated by Alertmanager's own `groupKey` - Incident workflow: acknowledge, assign, snooze, note, resolve, with a full timeline - On-call schedule management, with new incidents auto-assigned to whoever is on call - Alert and incident statistics, including MTTA and MTTR - REST API with per-user API key authentication - Single binary, SQLite storage — trivial to self-host --- ## Quick start **Prerequisites:** Go 1.21+ ```bash git clone https://git.ryuvia.com/niklas/terdut-server cd terdut-server go run ./cmd/terdut ``` The server starts on `:8080` with a `terdut.db` file in the working directory. ### Create the first user ```bash curl -X POST http://localhost:8080/api/bootstrap \ -H "Content-Type: application/json" \ -d '{"username": "admin", "email": "admin@example.com"}' ``` Save the `api_key.key` value from the response — it is shown **once only**. Use it as a bearer token for all subsequent requests: ```bash export KEY= curl -H "Authorization: Bearer $KEY" http://localhost:8080/api/users ``` ### Docker ```bash docker build -t terdut-server . docker run -p 8080:8080 -v $(pwd)/data:/data \ -e TERDUT_DB_PATH=/data/terdut.db \ terdut-server ``` ### Kubernetes A Helm chart is published from this repository as an OCI artifact, versioned in lockstep with the app — chart `x.y.z` is always app `vx.y.z`: ```bash helm upgrade --install terdut-server oci://git.ryuvia.com/niklas/terdut-server \ --version 0.9.2 \ --namespace terdut-server --create-namespace \ --set networking.hostname=terdut.example.com ``` The chart expects a [Gateway API](https://gateway-api.sigs.k8s.io/) Gateway named `envoy-main` in the `envoy-gateway-system` namespace to already exist — it renders an `HTTPRoute` against it rather than an `Ingress`. TLS is terminated at the gateway, so the server itself never sees a certificate. | Value | Default | Description | |---|---|---| | `networking.hostname` | `terdut.example.com` | Hostname the `HTTPRoute` serves | | `networking.listener` | `""` | Gateway listener (`sectionName`) to bind to. Empty attaches to every matching listener, **including plaintext HTTP** — set it to the HTTPS listener's name to serve TLS only | | `networking.servicePort` | `8080` | Port the route forwards to; keep in sync with `service.port` | | `bootstrap.enabled` | `true` | Runs a post-install hook that creates the first user and stores its API key in the `-admin-key` Secret. Already-bootstrapped servers are left alone | | `backupSidecar.enabled` | `true` | Adds an idle `python` sidecar and the [k8up](https://k8up.io/) annotations that dump the database through it | The API key travels in an `Authorization: Bearer` header, so set `networking.listener` whenever the hostname is reachable outside a trusted network. #### Backups The server image is `FROM scratch` — the binary and nothing else — so there is no interpreter to run a database dump in, and the database runs in WAL mode, where a file-level copy of the volume is not crash-consistent. The chart therefore ships an idle `python:*-alpine` sidecar that shares the data volume, and points k8up's `backupcommand` at it with `k8up.io/backupcommand-container`. Without that annotation k8up execs into `.spec.containers[0]` and the dump fails. The dump is buffered and sanity-checked before its first byte reaches stdout, because k8up streams stdout straight into Restic: a dump that dies partway is otherwise stored as a silently truncated snapshot that k8up still reports as successful. Set `backupSidecar.enabled=false` if you back the volume up some other way. --- ## Configuration | Variable | Default | Description | |---|---|---| | `TERDUT_ADDR` | `:8080` | TCP address to listen on | | `TERDUT_DB_PATH` | `terdut.db` | Path to the SQLite database file | | `TERDUT_ARCHIVE_AFTER` | `168h` (7d) | How long a resolved alert or incident stays in the default list before being auto-archived | | `TERDUT_STALE_AFTER` | `6h` | How long a firing alert may go without a refreshing webhook before it is treated as resolved — **must exceed your Alertmanager `repeat_interval`** | | `TERDUT_DEADMAN_MATCHERS` | `alertname=Watchdog` | Which alerts are [dead man's switches](#dead-mans-switch). `;` separates matchers, `,` the label conditions within one, `=` is exact equality. Every matcher must name an `alertname` | | `TERDUT_DEADMAN_TIMEOUT` | `15m` | How long a heartbeat may go unheard before its switch is declared dead — **must be shorter than the `repeat_interval` of the route carrying it**. `0` disables dead man's switch handling | | `TERDUT_DEADMAN_SEVERITY` | `critical` | Severity a dead man's switch incident opens at | | `TERDUT_NTFY_URL` | — | ntfy server to publish push notifications to. Empty disables notifications entirely | | `TERDUT_NTFY_TOKEN` | — | Bearer token for an access-controlled ntfy | | `TERDUT_NTFY_FALLBACK_TOPIC` | — | Topic used when nobody is on call | | `TERDUT_PUBLIC_URL` | — | Base URL a phone uses to reach this server, for the link and Acknowledge button inside a notification | | `TERDUT_NOTIFY_REPEAT` | `15m` | How long an incident may sit unacknowledged before it is paged again. `0` notifies once and never repeats | Durations use Go syntax (`30m`, `12h`, `168h`). An unparseable value falls back to the default. Note that `TERDUT_STALE_AFTER` and `TERDUT_DEADMAN_TIMEOUT` point in opposite directions. Staleness is a generous grace period around a `repeat_interval` you do not control; a dead man's switch is a deadline you set deliberately, and the heartbeat's route is configured to beat faster than it. In the Helm chart the two sweeper durations are set via `sweeper.staleAfter` and `sweeper.archiveAfter`, dead man's switches via the `deadman.*` values, and notifications via the `notify.*` values. --- ## Alertmanager configuration Add terdut-server as a webhook receiver in your `alertmanager.yml`: ```yaml receivers: - name: terdut webhook_configs: - url: http://terdut-server:8080/api/alertmanager/webhook send_resolved: true route: receiver: terdut ``` The webhook endpoint requires no authentication. If you use the [dead man's switch](#dead-mans-switch) — and the default configuration does — give the heartbeat a route of its own, because the deadline is only as tight as the interval feeding it: ```yaml route: receiver: terdut repeat_interval: 4h routes: - matchers: [ 'alertname = "Watchdog"' ] receiver: terdut group_wait: 0s group_interval: 1m repeat_interval: 1m ``` That delivers a heartbeat every **2 minutes**, not every minute. Alertmanager only reconsiders a group every `group_interval`, and at exactly one elapsed interval `repeat_interval` has not *quite* passed, so the send slips to the next tick — equal values give 2×. Two minutes against the 15 minute default is seven heartbeats per window, which is the point; use `group_interval: 30s` if you want the numbers to mean what they say. kube-prometheus-stack users get the `Watchdog` alert (`expr: vector(1)`) for free; it just needs routing to terdut rather than to `null`. --- ## Alerts and incidents There are two objects, and the difference between them is the whole design. **An alert is Alertmanager's record.** It has two states, `firing` and `resolved`, one row per fingerprint, and no human ever writes to it. The API exposes alerts read-only. **An incident is the work item.** It goes `triggered → acknowledged → resolved`, carries an assignee, a snooze, notes and a timeline, and is the only thing people act on. Many alerts belong to one incident. ### Correlation uses Alertmanager's `groupKey` Alertmanager has already grouped alerts according to the `group_by` routing tree you configured, and it sends the resulting `groupKey` and `groupLabels` on every webhook. Incidents adopt that answer rather than re-grouping alerts a second time — if you want different correlation, change `group_by` in `alertmanager.yml` and terdut follows. At most one incident is open per `groupKey` at a time. Alerts firing in a group that already has an open incident join it. The incident's `severity` is a high-water mark — the highest `severity` label any of its alerts has carried — so an incident that hit `critical` still reads as critical after the critical alert clears. ### An incident opens only on a new occurrence An incident opens when an alert **transitions into firing**: a fingerprint that was never seen, an alert with a newer `startsAt`, or a resolved alert that started again. The unchanged firing notifications Alertmanager re-sends every `repeat_interval` are none of those, and open nothing. This is what makes closing an incident by hand mean something. Without the rule, `POST /api/incidents/{id}/resolve` would be undone by the next re-send of an alert that never stopped firing. ### Leaving the open state - **Automatically**, once every alert under the incident has stopped firing — whether by a resolved webhook or by the sweeper's [stale-alert expiry](#stale-alert-expiry). The incident gets `"resolution_source": "alerts"`. - **By hand**, via `POST /api/incidents/{id}/resolve` (`"resolution_source": "manual"`). This is **terminal**: a later occurrence in that group opens a *new* incident rather than reopening this one. If the alert underneath never stops firing, the incident stays closed — that is what resolving by hand asserts. - **On recovery**, for a [dead man's switch](#dead-mans-switch) incident whose heartbeat started arriving again (`"resolution_source": "recovered"`). These incidents have no member alerts, so the automatic cascade above cannot reach them. To quieten an incident you expect to come back, snooze it instead (`POST /api/incidents/{id}/snooze`). A snooze hides the incident from the default list without closing it, and expires by simply falling into the past. ### On-call assignment A new incident is assigned to whoever holds today's schedule entry at the moment it opens (`GET /api/schedule/current`). If nobody is scheduled it opens unassigned. Reassign with `POST /api/incidents/{id}/assign`. One person holds a given day, so `POST /api/schedule` refuses a date somebody already has: taking a shift off the person expecting to be paged for it should not be something a plain call does by accident. Pass `"replace": true` to take them anyway. Either way the whole request is one transaction — a week where some days are free and some are taken moves as a unit, and a failure leaves the rota exactly as it was rather than with a hole in it. ### Push notifications With `TERDUT_NTFY_URL` set, an incident that opens is pushed to the on-call person's phone through [ntfy](https://ntfy.sh). Set each user's topic with `PUT /api/users/{id}/notify`; a user with no topic falls back to `TERDUT_NTFY_FALLBACK_TOPIC`, as does an incident that opens with nobody on call. If neither yields a topic, nothing is queued. Three things get pushed: - **triggered** — an incident opened. Priority follows severity (`critical` maps to ntfy's max priority, the one that overrides the phone's quiet settings). - **reminder** — the incident is still `triggered` after `TERDUT_NOTIFY_REPEAT`. Repeats until somebody acts. Acknowledging, snoozing, resolving or archiving all stop it — snooze is the mute button. - **resolved** — every alert under the incident stopped firing. Only sent to whoever was paged in the first place, and only for the automatic cascade: resolving by hand pushes nothing, since the person who did it already knows. Notifications carry an **Acknowledge** button that acknowledges the incident without opening anything. It POSTs to `/api/notify/ack/{token}`, an unauthenticated route authorised by the 256-bit token in its path — minted fresh per notification, scoped to one incident and one action, and valid for 24 hours. A real API key is never put in a notification, because the message is stored on the ntfy server and cached on the device. The token is **not** consumed by use. Acknowledging is idempotent, so a token stays valid for its full 24 hours and a second tap is a no-op that reports the incident's current state rather than an error — which is what you want when a tap is retried on a flaky mobile connection. What bounds it is scope, not a use count: one incident, one action, one day. Expired tokens are purged by the sweeper. Two consequences worth planning for: - `/api/notify/ack/{token}` **must stay publicly reachable**, or the button will not work when the responder is off your network. - Notifications sent to the fallback topic carry **no** Acknowledge button. The topic is shared, and a button on it would let any subscriber acknowledge as somebody else. Delivery is a queue, not an inline call: the webhook writes a row and a background notifier sends it within 30 seconds, retrying with exponential backoff up to 8 attempts. Nothing about ingestion blocks on ntfy being reachable. Every delivery is recorded on the incident's timeline: a `notified` event once ntfy accepts the publish, and a `notify_failed` event when a notification exhausts its retries. Written from the result rather than at enqueue, so the timeline says what actually happened — and a page that never landed is visible instead of looking the same as one that did. ### Stale alert expiry A resolved webhook is the only signal that an alert has stopped firing, so a notification that is dropped, silenced, or lost to a restart would otherwise pin that alert as firing forever. A background sweeper resolves firing alerts that Alertmanager has stopped refreshing, using either signal: - the `endsAt` watermark on the last notification has passed, or - no webhook has refreshed the alert within `TERDUT_STALE_AFTER`. Alertmanager re-sends firing notifications every `repeat_interval`, which is what keeps a live alert fresh — so `TERDUT_STALE_AFTER` must be comfortably larger than your `repeat_interval` (default 4h), or live alerts will be resolved prematurely. Alerts resolved this way are marked `"resolution_source": "expiry"` to distinguish them from a real Alertmanager resolve (`"alertmanager"`). An expiry cascades: once it leaves an incident with nothing firing under it, the incident resolves too, in the same sweep. ### Dead man's switch Everything above assumes alerts arrive. If Prometheus stops evaluating, or Alertmanager cannot reach this server, nothing arrives — and silence looks exactly like everything being fine. A dead man's switch inverts the handling for one designated alert so that silence is the signal: - **receiving** it opens no incident, and - the **absence** of it does. kube-prometheus-stack already ships the alert for this. `Watchdog` is `expr: vector(1)`, so it fires permanently and is re-sent forever; it is worth nothing unless something downstream notices it stop. That is what `TERDUT_DEADMAN_MATCHERS` defaults to. A matcher is a set of exact label conditions, one of which must be the `alertname`: ``` TERDUT_DEADMAN_MATCHERS="alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat" ``` **The unit of monitoring is the fingerprint, not the alert name.** Two clusters sending the same `Watchdog` are two independent switches, so a healthy one can never mask a dead one. #### The lifecycle A switch is **dormant** until its first heartbeat arrives. A configured matcher that has never been heard from opens nothing, so a fresh deploy or a restored database does not page. It also means a matcher that never matches anything is silently inert — check the startup log line, which lists the matchers that survived parsing. Once armed, the sweeper declares it **dead** when either the heartbeat has not been refreshed within `TERDUT_DEADMAN_TIMEOUT`, or Alertmanager explicitly resolved it — the sender saying the heartbeat stopped needs no further waiting. That opens an incident at `TERDUT_DEADMAN_SEVERITY`, assigned and paged like any other, and marks the heartbeat alert `"resolution_source": "deadman"` so the alert list stops claiming a dead switch is firing. It **recovers** when the heartbeat starts arriving again: the incident resolves with `"resolution_source": "recovered"` and the all-clear goes to whoever was paged. Resolving the incident by hand sticks, the same way it does for an alert-backed one. While the switch stays silent nothing new opens — so a decommissioned source is a one-time page rather than a nag. The switch **re-arms** on the next heartbeat: come back and die again, and that is a new incident. #### Two things to know `TERDUT_DEADMAN_TIMEOUT` must be **shorter** than the `repeat_interval` of the route carrying the heartbeat, which is the exact opposite of `TERDUT_STALE_AFTER`. Inheriting a default `repeat_interval` of 4h gives you a switch that takes four hours to notice anything, so give the heartbeat [its own route](#alertmanager-configuration). Matched alerts are exempt from stale-alert expiry — a heartbeat answers to its own timeout and nothing else. A dead man's switch incident has **no member alerts**: `GET /api/incidents/{id}/alerts` returns an empty list. There is no alert describing the problem, because the problem is that no alert arrived. What happened is on the timeline instead, as a `deadman_silent` event carrying the age of the last heartbeat, and the heartbeat's labels are on the incident's `group_labels`. --- ## API reference ### Authentication All endpoints except `/api/bootstrap` and `/api/alertmanager/webhook` require: ``` Authorization: Bearer ``` ### Users | Method | Path | Description | |---|---|---| | `POST` | `/api/bootstrap` | Create first user + API key (only works on empty DB) | | `GET` | `/api/users` | List users | | `POST` | `/api/users` | Create user `{"username","email"}` | | `DELETE` | `/api/users/{id}` | Delete user (cascades to keys) | | `PUT` | `/api/users/{id}/notify` | Set push notification target `{"ntfy_topic"}` — empty string clears it | | `POST` | `/api/users/{id}/api-keys` | Issue API key `{"name"}` — key shown once | | `DELETE` | `/api/users/{id}/api-keys/{keyID}` | Revoke API key | ### Alert ingestion | Method | Path | Description | |---|---|---| | `POST` | `/api/alertmanager/webhook` | Alertmanager v4 webhook receiver (no auth) | ### Notifications | Method | Path | Description | |---|---|---| | `POST` | `/api/notify/ack/{token}` | Acknowledge an incident from a push notification's Acknowledge button. No auth: the token in the path is the credential — one incident, one action, 24 hours, idempotent. Must stay publicly reachable | ### Incidents | Method | Path | Description | |---|---|---| | `GET` | `/api/incidents` | List incidents. Filters: `?status=triggered\|acknowledged\|resolved`, `?severity=`, `?assigned_to=`, `?archived=true`, `?snoozed=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?sort=severity`, `?limit=` (default 50, max 500) | | `GET` | `/api/incidents/{id}` | Get single incident, with its alerts inline | | `GET` | `/api/incidents/{id}/alerts` | Alerts under this incident | | `GET` | `/api/incidents/{id}/timeline` | Full event history, chronological | | `POST` | `/api/incidents/{id}/acknowledge` | Acknowledge (stamps authed user + time) | | `DELETE` | `/api/incidents/{id}/acknowledge` | Clear acknowledgement, back to `triggered` | | `POST` | `/api/incidents/{id}/resolve` | Close by hand — **terminal**, see above | | `POST` | `/api/incidents/{id}/assign` | Reassign `{"user_id"}` | | `POST` | `/api/incidents/{id}/snooze` | Hide until `{"until": RFC3339}` or `{"duration": "2h"}` | | `DELETE` | `/api/incidents/{id}/snooze` | Un-snooze | | `POST` | `/api/incidents/{id}/archive` | Archive (hides from the default list) | | `DELETE` | `/api/incidents/{id}/archive` | Un-archive | | `POST` | `/api/incidents/{id}/notes` | Add a note `{"content"}` | | `DELETE` | `/api/incidents/{id}/notes/{eventID}` | Delete own note | With no `?status=` filter, `GET /api/incidents` returns **open** incidents only — the queue an on-call person wants. Currently snoozed and archived incidents are excluded unless asked for. Actions that only make sense on an open incident return `409` once it is resolved. Notes are ordinary timeline events of type `note`; only they are deletable, and only by their author. The rest of the timeline is a record of what happened. #### The incident object | Field | Type | Notes | |---|---|---| | `id` | integer | Server-assigned | | `group_key` | string | Alertmanager's `groupKey` — opaque, treat as an identifier | | `title` | string | Rendered from `groupLabels` | | `group_labels` | object | String→string, as sent by Alertmanager | | `status` | string | `"triggered"`, `"acknowledged"` or `"resolved"` | | `severity` | string | *optional* — high-water mark across the incident's alerts; never lowered | | `triggered_at` | timestamp | When the incident opened | | `acknowledged_by_id` / `acknowledged_by` / `acknowledged_at` | | *optional* — user id, username, time | | `assigned_to_id` / `assigned_to` | | *optional* — user id, username | | `snoozed_until` | timestamp | *optional* — a value in the past reads as not snoozed | | `resolved_at` | timestamp | *optional* | | `resolution_source` | string | *optional* — `"alerts"`, `"manual"` or `"recovered"` | | `archived_at` | timestamp | *optional* | | `alerts` | array | Only on `GET /api/incidents/{id}` | Treat `resolution_source` as an open set, as with the alert field of the same name: degrade unknown values to "resolved, reason unknown". #### The timeline event object | Field | Type | Notes | |---|---|---| | `id` | integer | | | `incident_id` | integer | | | `type` | string | See below — treat as an open set | | `user_id` / `username` | | *optional* — absent when the server acted rather than a person | | `alert_id` | integer | *optional* — the alert an `alert_added` / `alert_resolved` event refers to | | `detail` | string | *optional* — the note body, the snooze deadline, etc. | | `created_at` | timestamp | | Types written today: `triggered`, `alert_added`, `alert_resolved`, `acknowledged`, `unacknowledged`, `assigned`, `snoozed`, `unsnoozed`, `resolved`, `note`, `notified`, `notify_failed`, `deadman_silent`. On an `assigned` event `user_id` is the **assignee**, not the actor. New types may be added; render unknown ones generically rather than dropping them. On `notified` and `notify_failed`, `detail` carries the notification kind (`triggered` | `reminder` | `resolved`), and on a failure the reason after it. `user_id` is who was paged — absent means the page went to the shared fallback topic and so belongs to nobody. The topic itself is never written to the timeline: it is a shared secret with the ntfy server, and every API key can read this. ### Alerts Alerts are read-only. Everything a person does happens on the incident. | Method | Path | Description | |---|---|---| | `GET` | `/api/alerts` | List alerts. Filters: `?status=firing\|resolved`, `?name=`, `?incident_id=`, `?archived=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?limit=` (default 50, max 500) | | `GET` | `/api/alerts/{id}` | Get single alert | Archived alerts are hidden from `GET /api/alerts` unless `?archived=true` is passed; alert archiving is automatic housekeeping by the sweeper, not a user action. Resolved alerts carry `resolution_source`: `"alertmanager"` for a real resolved webhook, `"expiry"` when the sweeper inferred it (see [Stale alert expiry](#stale-alert-expiry)), `"deadman"` for a heartbeat declared dead (see [Dead man's switch](#dead-mans-switch)). #### The alert object Returned by `GET /api/alerts` (as an array) and `GET /api/alerts/{id}`. Timestamps are RFC 3339 in UTC. Fields marked *optional* are omitted entirely when unset, so clients must treat them as nullable. | Field | Type | Notes | |---|---|---| | `id` | integer | Server-assigned; stable for the life of the row | | `fingerprint` | string | Alertmanager's fingerprint — the upsert key | | `name` | string | From the `alertname` label | | `status` | string | `"firing"` or `"resolved"` | | `labels` | object | String→string, as sent by Alertmanager | | `annotations` | object | String→string, as sent by Alertmanager | | `starts_at` | timestamp | When the alert instance began, **per Prometheus** | | `ends_at` | timestamp | *optional* — absent while no end is known | | `generator_url` | string | Link back to the originating Prometheus | | `received_at` | timestamp | When the server last accepted a webhook for this alert — see below | | `incident_id` | integer | *optional* — the most recent incident this alert belongs to | | `resolution_source` | string | *optional* — `"alertmanager"`, `"expiry"` or `"deadman"` | | `archived_at` | timestamp | *optional* — set while archived | ##### `received_at` is a liveness heartbeat `starts_at` comes from Prometheus and **never changes** for the lifetime of an alert instance. It says when the problem began, not whether it is still happening — an alert that started twelve days ago looks identical whether Alertmanager refreshed it a minute ago or went silent a week ago. `received_at` is the field that answers "is this still live". It is set to the server's clock on **every accepted webhook** for that fingerprint, including the unchanged firing notifications Alertmanager re-sends every `repeat_interval`. Clients may rely on this: - **A firing alert whose `received_at` is advancing is still being refreshed.** Stale-dating it against `repeat_interval` is a valid liveness check, and it is what the built-in sweeper does (see [Stale alert expiry](#stale-alert-expiry)). - **`received_at` tracks accepted payloads, not delivery attempts.** A retry that describes an older instance than the stored one is discarded, and a discarded payload does not move `received_at`. - **It stops advancing once the alert resolves,** because Alertmanager stops re-sending. On an alert resolved by the sweeper (`"resolution_source": "expiry"`) it therefore marks the last time Alertmanager was actually heard from, which is earlier than `ends_at`. `GET /api/alerts` is ordered by `received_at` descending — most recently refreshed first — and the `?from=` / `?to=` filters on both the alert and stats endpoints select on `received_at`, not `starts_at`. ##### `resolution_source` says how much to trust `ends_at` An alert can leave the firing state two ways, and `resolution_source` records which happened. Clients may rely on this: - **Absent while firing.** It is set only on resolve, and a re-fire under the same fingerprint clears it again, so its presence always agrees with `"status": "resolved"`. - **`"alertmanager"` — a real resolved webhook arrived.** `ends_at` is the end time Alertmanager reported. It is an observed value and can be displayed as fact. - **`"expiry"` — the sweeper inferred the resolve** because Alertmanager stopped refreshing the alert (see [Stale alert expiry](#stale-alert-expiry)). Nothing ever reported an end, so **`ends_at` is approximate**: it is either the stale `endsAt` watermark from the last notification, or — when that notification carried none — the time the sweep ran, which lags the last real contact by up to `TERDUT_STALE_AFTER` plus a sweep interval. Treat it as "no later than", not as when the problem stopped. On these alerts `received_at` is the more truthful signal: it marks the last time Alertmanager was actually heard from. Surfacing the distinction is worthwhile, since `"expiry"` can also mean the alert is still firing and the notification path broke. - **`"deadman"` — a heartbeat was declared dead** (see [Dead man's switch](#dead-mans-switch)). Like `"expiry"`, an inference from silence rather than an observed end, so `ends_at` is approximate — but a much tighter one, bounded by `TERDUT_DEADMAN_TIMEOUT`. It is also the one resolution a re-fire under the same `starts_at` can undo, since the switch coming back is exactly the evidence that the inference was wrong. Treat the value as an open set and tolerate ones you do not recognise — new sources may be added, and unknown values should degrade to "resolved, reason unknown" rather than being rejected. ### On-call schedule | Method | Path | Description | |---|---|---| | `POST` | `/api/schedule` | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...], "replace"}` — all-or-nothing | | `GET` | `/api/schedule` | List entries. Filters: `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD` | | `GET` | `/api/schedule/current` | Today's on-call user (UTC), 404 if none | | `DELETE` | `/api/schedule/{id}` | Remove schedule entry | ### Statistics All stat endpoints accept optional `?from=YYYY-MM-DD` and `?to=YYYY-MM-DD`, and exclude archived rows to match the default list views. Alert stats filter on `received_at`; incident stats filter on `triggered_at`. | Method | Path | Description | |---|---|---| | `GET` | `/api/stats/incidents` | `{total, triggered, acknowledged, resolved, mtta_seconds, mttr_seconds}` | | `GET` | `/api/stats/alerts` | `{total, firing, resolved}` counts | | `GET` | `/api/stats/alerts/top` | Most frequent alert names. `?limit=` (default 10, max 100) | | `GET` | `/api/stats/alerts/by-hour` | Count per hour-of-day (UTC), all 24 slots returned | | `GET` | `/api/stats/alerts/by-day` | Count per day-of-week, all 7 slots with names returned | `mtta_seconds` (time to acknowledge) and `mttr_seconds` (time to resolve) are averages over incidents that have actually been acknowledged or resolved, and are **null** until there are any — null means "no data", not zero. --- ## Upgrading to incidents The incidents release moves the workflow off alerts, which is a **breaking API change**. These endpoints are gone: | Removed | Replacement | |---|---| | `POST`/`DELETE` `/api/alerts/{id}/acknowledge` | `POST`/`DELETE` `/api/incidents/{id}/acknowledge` | | `POST`/`DELETE` `/api/alerts/{id}/archive` | `POST`/`DELETE` `/api/incidents/{id}/archive` (alert archiving is now sweeper-only) | | `GET`/`POST` `/api/alerts/{id}/comments` | `GET /api/incidents/{id}/timeline`, `POST /api/incidents/{id}/notes` | | `DELETE /api/alerts/{id}/comments/{commentID}` | `DELETE /api/incidents/{id}/notes/{eventID}` | The alert object also drops `acknowledged_by_id`, `acknowledged_by` and `acknowledged_at`, and gains `incident_id`. Migration `008_incidents.sql` runs automatically on start and preserves existing data: every alert gets a backfilled incident carrying its acknowledgement, and comments become timeline notes. Backfilled incidents have a `group_key` of `backfill:` — there is no historical `groupKey` to correlate on, so they are one-per-alert rather than grouped. Nothing about the two documented alert contracts changes: `received_at` is still advanced on every accepted webhook, and `resolution_source` still means what it did. ## Upgrading to dead man's switches Dead man's switch handling is **on by default**, watching `alertname=Watchdog` with a 15 minute timeout. If you already route `Watchdog` to this server, the behaviour of that alert changes on upgrade, in both directions: - it stops opening incidents when it arrives, and - it starts opening one when it stops arriving. **Check your `repeat_interval` before upgrading.** The switch pages whenever a heartbeat has not been refreshed within `TERDUT_DEADMAN_TIMEOUT`, so a `Watchdog` route inheriting a 4h or 12h `repeat_interval` will page constantly against the 15 minute default. Either give the heartbeat [its own fast route](#alertmanager-configuration) — the point of the feature — or set `TERDUT_DEADMAN_TIMEOUT` above your current `repeat_interval` until you have. `TERDUT_DEADMAN_TIMEOUT=0` turns the whole thing off. There is no migration and no schema change. An existing open incident from a `Watchdog` that arrived under the old behaviour is unaffected; resolve it by hand. --- ## Development ```bash go test ./... # run all tests go build ./... # compile all packages go run ./cmd/terdut # run locally ``` `make fmt lint test helm-lint` is the gate. It mirrors `.gitea/workflows/ci.yaml` step for step, so a green run here means a green pipeline — with one deliberate exception: `make test` adds `-race`, which CI does not. The sweeper, the notifier goroutine and the dead man's switch sweep all touch the same single database connection, and a race between them would surface as a flaky incident in production rather than as a red build. ## Releasing ``` push to main → ci.yaml gofmt, go vet, go test push tag vX.Y.Z → release.yaml same gate, then publish: git.ryuvia.com/niklas/terdut-server:vX.Y.Z oci://git.ryuvia.com/niklas/terdut-server X.Y.Z PR to Ryuvia/charts → bump the wrapper chart to X.Y.Z; on merge Flux reconciles and the release rolls out ``` Both artifacts go to the **personal** Gitea namespace rather than `ryuvia`, because Gitea scopes package visibility to the owner with no per-package override — so `ryuvia/*` is private because the org is. Publishing to `niklas` keeps them anonymously pullable, which is why no pull secret is needed in the cluster. Same reasoning, and the same choice, as riksdata and rd-web. Saying **"Release"** runs all three rows: the `release` skill commits, pushes, tags, waits for the pipeline, and opens the `Ryuvia/charts` PR, stopping before the merge. See `~/.claude/skills/release/`, or `.release.conf` here for this repo's part of it. The chart is published **only** from the tag, by the `chart` job. There used to be a second publisher on every `charts/**` push to main, and the two raced for the same chart version with different answers — chart 0.9.0 went out reading `appVersion: "latest"` that way. One publisher, triggered by the tag (`766f439`). The cost is that a chart-only change has no version of its own and rides the next app tag. Both workflows are thin drivers over the Makefile: `ci.yaml` runs `make fmt lint test` and `make helm-lint`, `release.yaml` adds `make binaries`, `make push`, `make helm-package` and `make helm-push`. That is deliberate — it is what makes a green local gate and a green pipeline the same code rather than two descriptions of it, and it is how riksdata and rd-web have always worked. `make push` builds and pushes in one step, unlike those two, because the image is `linux/amd64,linux/arm64` and buildx cannot load a multi-platform result into the local image store. `make build` stays single-platform and local-only. Both refuse `VERSION=dev`: publishing is one command, so it is also one command to run by accident. Publishing happens by pushing a tag. Two things the release process needs to know about this repo: - **The pipeline has no image scan**, unlike riksdata and rd-web. A green release run is not evidence the image is CVE-clean. - **The wrapper chart's `values.yaml` has two `tag:` lines** — the app image and the python backup sidecar — so `chart-bump` is given `--image` to say which one moves. The wrapper chart must have **its own `version:` bumped in the same commit**. Flux reconciles with `reconcileStrategy: ChartVersion`, so a chart whose version did not change produces no new artifact and the change is never deployed — with no error anywhere.