289eca8076
CI / test (push) Successful in 2m15s
The module path, the container image, the Helm chart and the CI pipeline all named GitHub. They now name the Gitea instance everything else already runs on. The workflows are rewritten rather than translated. Gitea's runner image is ubuntu:22.04, whose nodejs is Node 12, so no JS action runs there at all -- actions/checkout@v4 dies with a SyntaxError before it does anything. Every step is shell, checkout is a plain clone (this repo is public, so it needs no credential), and the jobs that need docker or helm run in host mode because the dind bridge a `container:` job gets cannot reach github.com or get.helm.sh. Two consequences worth naming: - upload-artifact/download-artifact are also JS actions, and there is no artifact store here, so the job that builds the binaries is the job that publishes them. Nothing is passed between jobs. - setup-qemu-action is gone with the rest, and the runner has no binfmt registration. The Dockerfile's builder stage now runs on $BUILDPLATFORM and cross-compiles from TARGETARCH instead, which is what keeps the arm64 image buildable -- and makes it native rather than emulated. The chart moves from a GitHub Pages index to an OCI artifact in Gitea's registry. Publishing stays tag-only for the reason recorded in release.yaml: a workflow triggered by the branch push cannot know the version it is about to be tagged with. The GitHub repository is left in place and untouched. Nothing pushes to it any more, but its existing release downloads and chart index keep resolving.
664 lines
32 KiB
Markdown
664 lines
32 KiB
Markdown
# Terminal Duty (terdut-server)
|
||
|
||
Incident management server for teams using Prometheus Alertmanager.
|
||
|
||
- Receives Alertmanager webhooks directly — no adapter needed
|
||
- Turns alerts into **incidents**, correlated by Alertmanager's own `groupKey`
|
||
- Incident workflow: acknowledge, assign, snooze, note, resolve, with a full timeline
|
||
- On-call schedule management, with new incidents auto-assigned to whoever is on call
|
||
- Alert and incident statistics, including MTTA and MTTR
|
||
- REST API with per-user API key authentication
|
||
- Single binary, SQLite storage — trivial to self-host
|
||
|
||
---
|
||
|
||
## Quick start
|
||
|
||
**Prerequisites:** Go 1.21+
|
||
|
||
```bash
|
||
git clone https://git.ryuvia.com/niklas/terdut-server
|
||
cd terdut-server
|
||
go run ./cmd/terdut
|
||
```
|
||
|
||
The server starts on `:8080` with a `terdut.db` file in the working directory.
|
||
|
||
### Create the first user
|
||
|
||
```bash
|
||
curl -X POST http://localhost:8080/api/bootstrap \
|
||
-H "Content-Type: application/json" \
|
||
-d '{"username": "admin", "email": "admin@example.com"}'
|
||
```
|
||
|
||
Save the `api_key.key` value from the response — it is shown **once only**.
|
||
|
||
Use it as a bearer token for all subsequent requests:
|
||
|
||
```bash
|
||
export KEY=<your-key>
|
||
curl -H "Authorization: Bearer $KEY" http://localhost:8080/api/users
|
||
```
|
||
|
||
### Docker
|
||
|
||
```bash
|
||
docker build -t terdut-server .
|
||
docker run -p 8080:8080 -v $(pwd)/data:/data \
|
||
-e TERDUT_DB_PATH=/data/terdut.db \
|
||
terdut-server
|
||
```
|
||
|
||
### Kubernetes
|
||
|
||
A Helm chart is published from this repository as an OCI artifact, versioned in lockstep
|
||
with the app — chart `x.y.z` is always app `vx.y.z`:
|
||
|
||
```bash
|
||
helm upgrade --install terdut-server oci://git.ryuvia.com/niklas/terdut-server \
|
||
--version 0.9.0 \
|
||
--namespace terdut-server --create-namespace \
|
||
--set networking.hostname=terdut.example.com
|
||
```
|
||
|
||
The chart expects a [Gateway API](https://gateway-api.sigs.k8s.io/) Gateway named `envoy-main` in
|
||
the `envoy-gateway-system` namespace to already exist — it renders an `HTTPRoute` against it rather
|
||
than an `Ingress`. TLS is terminated at the gateway, so the server itself never sees a certificate.
|
||
|
||
| Value | Default | Description |
|
||
|---|---|---|
|
||
| `networking.hostname` | `terdut.example.com` | Hostname the `HTTPRoute` serves |
|
||
| `networking.listener` | `""` | Gateway listener (`sectionName`) to bind to. Empty attaches to every matching listener, **including plaintext HTTP** — set it to the HTTPS listener's name to serve TLS only |
|
||
| `networking.servicePort` | `8080` | Port the route forwards to; keep in sync with `service.port` |
|
||
| `bootstrap.enabled` | `true` | Runs a post-install hook that creates the first user and stores its API key in the `<release>-admin-key` Secret. Already-bootstrapped servers are left alone |
|
||
| `backupSidecar.enabled` | `true` | Adds an idle `python` sidecar and the [k8up](https://k8up.io/) annotations that dump the database through it |
|
||
|
||
The API key travels in an `Authorization: Bearer` header, so set `networking.listener` whenever the
|
||
hostname is reachable outside a trusted network.
|
||
|
||
#### Backups
|
||
|
||
The server image is `FROM scratch` — the binary and nothing else — so there is no interpreter to
|
||
run a database dump in, and the database runs in WAL mode, where a file-level copy of the volume is
|
||
not crash-consistent. The chart therefore ships an idle `python:*-alpine` sidecar that shares the
|
||
data volume, and points k8up's `backupcommand` at it with `k8up.io/backupcommand-container`. Without
|
||
that annotation k8up execs into `.spec.containers[0]` and the dump fails.
|
||
|
||
The dump is buffered and sanity-checked before its first byte reaches stdout, because k8up streams
|
||
stdout straight into Restic: a dump that dies partway is otherwise stored as a silently truncated
|
||
snapshot that k8up still reports as successful.
|
||
|
||
Set `backupSidecar.enabled=false` if you back the volume up some other way.
|
||
|
||
---
|
||
|
||
## Configuration
|
||
|
||
| Variable | Default | Description |
|
||
|---|---|---|
|
||
| `TERDUT_ADDR` | `:8080` | TCP address to listen on |
|
||
| `TERDUT_DB_PATH` | `terdut.db` | Path to the SQLite database file |
|
||
| `TERDUT_ARCHIVE_AFTER` | `168h` (7d) | How long a resolved alert or incident stays in the default list before being auto-archived |
|
||
| `TERDUT_STALE_AFTER` | `6h` | How long a firing alert may go without a refreshing webhook before it is treated as resolved — **must exceed your Alertmanager `repeat_interval`** |
|
||
| `TERDUT_DEADMAN_MATCHERS` | `alertname=Watchdog` | Which alerts are [dead man's switches](#dead-mans-switch). `;` separates matchers, `,` the label conditions within one, `=` is exact equality. Every matcher must name an `alertname` |
|
||
| `TERDUT_DEADMAN_TIMEOUT` | `15m` | How long a heartbeat may go unheard before its switch is declared dead — **must be shorter than the `repeat_interval` of the route carrying it**. `0` disables dead man's switch handling |
|
||
| `TERDUT_DEADMAN_SEVERITY` | `critical` | Severity a dead man's switch incident opens at |
|
||
| `TERDUT_NTFY_URL` | — | ntfy server to publish push notifications to. Empty disables notifications entirely |
|
||
| `TERDUT_NTFY_TOKEN` | — | Bearer token for an access-controlled ntfy |
|
||
| `TERDUT_NTFY_FALLBACK_TOPIC` | — | Topic used when nobody is on call |
|
||
| `TERDUT_PUBLIC_URL` | — | Base URL a phone uses to reach this server, for the link and Acknowledge button inside a notification |
|
||
| `TERDUT_NOTIFY_REPEAT` | `15m` | How long an incident may sit unacknowledged before it is paged again. `0` notifies once and never repeats |
|
||
|
||
Durations use Go syntax (`30m`, `12h`, `168h`). An unparseable value falls back to the default.
|
||
|
||
Note that `TERDUT_STALE_AFTER` and `TERDUT_DEADMAN_TIMEOUT` point in opposite directions. Staleness
|
||
is a generous grace period around a `repeat_interval` you do not control; a dead man's switch is a
|
||
deadline you set deliberately, and the heartbeat's route is configured to beat faster than it.
|
||
|
||
In the Helm chart the two sweeper durations are set via `sweeper.staleAfter` and `sweeper.archiveAfter`, dead man's switches via the `deadman.*` values, and notifications via the `notify.*` values.
|
||
|
||
---
|
||
|
||
## Alertmanager configuration
|
||
|
||
Add terdut-server as a webhook receiver in your `alertmanager.yml`:
|
||
|
||
```yaml
|
||
receivers:
|
||
- name: terdut
|
||
webhook_configs:
|
||
- url: http://terdut-server:8080/api/alertmanager/webhook
|
||
send_resolved: true
|
||
|
||
route:
|
||
receiver: terdut
|
||
```
|
||
|
||
The webhook endpoint requires no authentication.
|
||
|
||
If you use the [dead man's switch](#dead-mans-switch) — and the default configuration does — give
|
||
the heartbeat a route of its own, because the deadline is only as tight as the interval feeding it:
|
||
|
||
```yaml
|
||
route:
|
||
receiver: terdut
|
||
repeat_interval: 4h
|
||
routes:
|
||
- matchers: [ 'alertname = "Watchdog"' ]
|
||
receiver: terdut
|
||
group_wait: 0s
|
||
group_interval: 1m
|
||
repeat_interval: 1m
|
||
```
|
||
|
||
That delivers a heartbeat every **2 minutes**, not every minute. Alertmanager only reconsiders a
|
||
group every `group_interval`, and at exactly one elapsed interval `repeat_interval` has not *quite*
|
||
passed, so the send slips to the next tick — equal values give 2×. Two minutes against the 15 minute
|
||
default is seven heartbeats per window, which is the point; use `group_interval: 30s` if you want
|
||
the numbers to mean what they say.
|
||
|
||
kube-prometheus-stack users get the `Watchdog` alert (`expr: vector(1)`) for free; it just needs
|
||
routing to terdut rather than to `null`.
|
||
|
||
---
|
||
|
||
## Alerts and incidents
|
||
|
||
There are two objects, and the difference between them is the whole design.
|
||
|
||
**An alert is Alertmanager's record.** It has two states, `firing` and
|
||
`resolved`, one row per fingerprint, and no human ever writes to it. The API
|
||
exposes alerts read-only.
|
||
|
||
**An incident is the work item.** It goes `triggered → acknowledged → resolved`,
|
||
carries an assignee, a snooze, notes and a timeline, and is the only thing people
|
||
act on. Many alerts belong to one incident.
|
||
|
||
### Correlation uses Alertmanager's `groupKey`
|
||
|
||
Alertmanager has already grouped alerts according to the `group_by` routing tree
|
||
you configured, and it sends the resulting `groupKey` and `groupLabels` on every
|
||
webhook. Incidents adopt that answer rather than re-grouping alerts a second
|
||
time — if you want different correlation, change `group_by` in
|
||
`alertmanager.yml` and terdut follows.
|
||
|
||
At most one incident is open per `groupKey` at a time. Alerts firing in a group
|
||
that already has an open incident join it. The incident's `severity` is a
|
||
high-water mark — the highest `severity` label any of its alerts has carried — so
|
||
an incident that hit `critical` still reads as critical after the critical alert
|
||
clears.
|
||
|
||
### An incident opens only on a new occurrence
|
||
|
||
An incident opens when an alert **transitions into firing**: a fingerprint that
|
||
was never seen, an alert with a newer `startsAt`, or a resolved alert that
|
||
started again. The unchanged firing notifications Alertmanager re-sends every
|
||
`repeat_interval` are none of those, and open nothing.
|
||
|
||
This is what makes closing an incident by hand mean something. Without the rule,
|
||
`POST /api/incidents/{id}/resolve` would be undone by the next re-send of an
|
||
alert that never stopped firing.
|
||
|
||
### Leaving the open state
|
||
|
||
- **Automatically**, once every alert under the incident has stopped firing —
|
||
whether by a resolved webhook or by the sweeper's
|
||
[stale-alert expiry](#stale-alert-expiry). The incident gets
|
||
`"resolution_source": "alerts"`.
|
||
- **By hand**, via `POST /api/incidents/{id}/resolve`
|
||
(`"resolution_source": "manual"`). This is **terminal**: a later occurrence in
|
||
that group opens a *new* incident rather than reopening this one. If the alert
|
||
underneath never stops firing, the incident stays closed — that is what
|
||
resolving by hand asserts.
|
||
- **On recovery**, for a [dead man's switch](#dead-mans-switch) incident whose
|
||
heartbeat started arriving again (`"resolution_source": "recovered"`). These
|
||
incidents have no member alerts, so the automatic cascade above cannot reach
|
||
them.
|
||
|
||
To quieten an incident you expect to come back, snooze it instead
|
||
(`POST /api/incidents/{id}/snooze`). A snooze hides the incident from the default
|
||
list without closing it, and expires by simply falling into the past.
|
||
|
||
### On-call assignment
|
||
|
||
A new incident is assigned to whoever holds today's schedule entry at the moment
|
||
it opens (`GET /api/schedule/current`). If nobody is scheduled it opens
|
||
unassigned. Reassign with `POST /api/incidents/{id}/assign`.
|
||
|
||
One person holds a given day, so `POST /api/schedule` refuses a date somebody
|
||
already has: taking a shift off the person expecting to be paged for it should
|
||
not be something a plain call does by accident. Pass `"replace": true` to take
|
||
them anyway. Either way the whole request is one transaction — a week where some
|
||
days are free and some are taken moves as a unit, and a failure leaves the rota
|
||
exactly as it was rather than with a hole in it.
|
||
|
||
### Push notifications
|
||
|
||
With `TERDUT_NTFY_URL` set, an incident that opens is pushed to the on-call
|
||
person's phone through [ntfy](https://ntfy.sh). Set each user's topic with
|
||
`PUT /api/users/{id}/notify`; a user with no topic falls back to
|
||
`TERDUT_NTFY_FALLBACK_TOPIC`, as does an incident that opens with nobody on call.
|
||
If neither yields a topic, nothing is queued.
|
||
|
||
Three things get pushed:
|
||
|
||
- **triggered** — an incident opened. Priority follows severity (`critical` maps
|
||
to ntfy's max priority, the one that overrides the phone's quiet settings).
|
||
- **reminder** — the incident is still `triggered` after `TERDUT_NOTIFY_REPEAT`.
|
||
Repeats until somebody acts. Acknowledging, snoozing, resolving or archiving
|
||
all stop it — snooze is the mute button.
|
||
- **resolved** — every alert under the incident stopped firing. Only sent to
|
||
whoever was paged in the first place, and only for the automatic cascade:
|
||
resolving by hand pushes nothing, since the person who did it already knows.
|
||
|
||
Notifications carry an **Acknowledge** button that acknowledges the incident
|
||
without opening anything. It POSTs to `/api/notify/ack/{token}`, an
|
||
unauthenticated route authorised by the 256-bit token in its path — minted fresh
|
||
per notification, scoped to one incident and one action, and valid for 24 hours.
|
||
A real API key is never put in a notification, because the message is stored on
|
||
the ntfy server and cached on the device.
|
||
|
||
The token is **not** consumed by use. Acknowledging is idempotent, so a token
|
||
stays valid for its full 24 hours and a second tap is a no-op that reports the
|
||
incident's current state rather than an error — which is what you want when a
|
||
tap is retried on a flaky mobile connection. What bounds it is scope, not a use
|
||
count: one incident, one action, one day. Expired tokens are purged by the
|
||
sweeper.
|
||
|
||
Two consequences worth planning for:
|
||
|
||
- `/api/notify/ack/{token}` **must stay publicly reachable**, or the button will
|
||
not work when the responder is off your network.
|
||
- Notifications sent to the fallback topic carry **no** Acknowledge button. The
|
||
topic is shared, and a button on it would let any subscriber acknowledge as
|
||
somebody else.
|
||
|
||
Delivery is a queue, not an inline call: the webhook writes a row and a
|
||
background notifier sends it within 30 seconds, retrying with exponential
|
||
backoff up to 8 attempts. Nothing about ingestion blocks on ntfy being reachable.
|
||
|
||
Every delivery is recorded on the incident's timeline: a `notified` event once
|
||
ntfy accepts the publish, and a `notify_failed` event when a notification
|
||
exhausts its retries. Written from the result rather than at enqueue, so the
|
||
timeline says what actually happened — and a page that never landed is visible
|
||
instead of looking the same as one that did.
|
||
|
||
### Stale alert expiry
|
||
|
||
A resolved webhook is the only signal that an alert has stopped firing, so a
|
||
notification that is dropped, silenced, or lost to a restart would otherwise pin
|
||
that alert as firing forever. A background sweeper resolves firing alerts that
|
||
Alertmanager has stopped refreshing, using either signal:
|
||
|
||
- the `endsAt` watermark on the last notification has passed, or
|
||
- no webhook has refreshed the alert within `TERDUT_STALE_AFTER`.
|
||
|
||
Alertmanager re-sends firing notifications every `repeat_interval`, which is what
|
||
keeps a live alert fresh — so `TERDUT_STALE_AFTER` must be comfortably larger
|
||
than your `repeat_interval` (default 4h), or live alerts will be resolved
|
||
prematurely. Alerts resolved this way are marked `"resolution_source": "expiry"`
|
||
to distinguish them from a real Alertmanager resolve (`"alertmanager"`).
|
||
|
||
An expiry cascades: once it leaves an incident with nothing firing under it, the
|
||
incident resolves too, in the same sweep.
|
||
|
||
### Dead man's switch
|
||
|
||
Everything above assumes alerts arrive. If Prometheus stops evaluating, or
|
||
Alertmanager cannot reach this server, nothing arrives — and silence looks
|
||
exactly like everything being fine. A dead man's switch inverts the handling for
|
||
one designated alert so that silence is the signal:
|
||
|
||
- **receiving** it opens no incident, and
|
||
- the **absence** of it does.
|
||
|
||
kube-prometheus-stack already ships the alert for this. `Watchdog` is
|
||
`expr: vector(1)`, so it fires permanently and is re-sent forever; it is worth
|
||
nothing unless something downstream notices it stop. That is what
|
||
`TERDUT_DEADMAN_MATCHERS` defaults to.
|
||
|
||
A matcher is a set of exact label conditions, one of which must be the
|
||
`alertname`:
|
||
|
||
```
|
||
TERDUT_DEADMAN_MATCHERS="alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat"
|
||
```
|
||
|
||
**The unit of monitoring is the fingerprint, not the alert name.** Two clusters
|
||
sending the same `Watchdog` are two independent switches, so a healthy one can
|
||
never mask a dead one.
|
||
|
||
#### The lifecycle
|
||
|
||
A switch is **dormant** until its first heartbeat arrives. A configured matcher
|
||
that has never been heard from opens nothing, so a fresh deploy or a restored
|
||
database does not page. It also means a matcher that never matches anything is
|
||
silently inert — check the startup log line, which lists the matchers that
|
||
survived parsing.
|
||
|
||
Once armed, the sweeper declares it **dead** when either the heartbeat has not
|
||
been refreshed within `TERDUT_DEADMAN_TIMEOUT`, or Alertmanager explicitly
|
||
resolved it — the sender saying the heartbeat stopped needs no further waiting.
|
||
That opens an incident at `TERDUT_DEADMAN_SEVERITY`, assigned and paged like any
|
||
other, and marks the heartbeat alert `"resolution_source": "deadman"` so the
|
||
alert list stops claiming a dead switch is firing.
|
||
|
||
It **recovers** when the heartbeat starts arriving again: the incident resolves
|
||
with `"resolution_source": "recovered"` and the all-clear goes to whoever was
|
||
paged.
|
||
|
||
Resolving the incident by hand sticks, the same way it does for an alert-backed
|
||
one. While the switch stays silent nothing new opens — so a decommissioned
|
||
source is a one-time page rather than a nag. The switch **re-arms** on the next
|
||
heartbeat: come back and die again, and that is a new incident.
|
||
|
||
#### Two things to know
|
||
|
||
`TERDUT_DEADMAN_TIMEOUT` must be **shorter** than the `repeat_interval` of the
|
||
route carrying the heartbeat, which is the exact opposite of
|
||
`TERDUT_STALE_AFTER`. Inheriting a default `repeat_interval` of 4h gives you a
|
||
switch that takes four hours to notice anything, so give the heartbeat
|
||
[its own route](#alertmanager-configuration). Matched alerts are exempt from
|
||
stale-alert expiry — a heartbeat answers to its own timeout and nothing else.
|
||
|
||
A dead man's switch incident has **no member alerts**:
|
||
`GET /api/incidents/{id}/alerts` returns an empty list. There is no alert
|
||
describing the problem, because the problem is that no alert arrived. What
|
||
happened is on the timeline instead, as a `deadman_silent` event carrying the age
|
||
of the last heartbeat, and the heartbeat's labels are on the incident's
|
||
`group_labels`.
|
||
|
||
---
|
||
|
||
## API reference
|
||
|
||
### Authentication
|
||
|
||
All endpoints except `/api/bootstrap` and `/api/alertmanager/webhook` require:
|
||
|
||
```
|
||
Authorization: Bearer <api-key>
|
||
```
|
||
|
||
### Users
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `POST` | `/api/bootstrap` | Create first user + API key (only works on empty DB) |
|
||
| `GET` | `/api/users` | List users |
|
||
| `POST` | `/api/users` | Create user `{"username","email"}` |
|
||
| `DELETE` | `/api/users/{id}` | Delete user (cascades to keys) |
|
||
| `PUT` | `/api/users/{id}/notify` | Set push notification target `{"ntfy_topic"}` — empty string clears it |
|
||
| `POST` | `/api/users/{id}/api-keys` | Issue API key `{"name"}` — key shown once |
|
||
| `DELETE` | `/api/users/{id}/api-keys/{keyID}` | Revoke API key |
|
||
|
||
### Alert ingestion
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `POST` | `/api/alertmanager/webhook` | Alertmanager v4 webhook receiver (no auth) |
|
||
|
||
### Notifications
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `POST` | `/api/notify/ack/{token}` | Acknowledge an incident from a push notification's Acknowledge button. No auth: the token in the path is the credential — one incident, one action, 24 hours, idempotent. Must stay publicly reachable |
|
||
|
||
### Incidents
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `GET` | `/api/incidents` | List incidents. Filters: `?status=triggered\|acknowledged\|resolved`, `?severity=`, `?assigned_to=<user id>`, `?archived=true`, `?snoozed=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?sort=severity`, `?limit=` (default 50, max 500) |
|
||
| `GET` | `/api/incidents/{id}` | Get single incident, with its alerts inline |
|
||
| `GET` | `/api/incidents/{id}/alerts` | Alerts under this incident |
|
||
| `GET` | `/api/incidents/{id}/timeline` | Full event history, chronological |
|
||
| `POST` | `/api/incidents/{id}/acknowledge` | Acknowledge (stamps authed user + time) |
|
||
| `DELETE` | `/api/incidents/{id}/acknowledge` | Clear acknowledgement, back to `triggered` |
|
||
| `POST` | `/api/incidents/{id}/resolve` | Close by hand — **terminal**, see above |
|
||
| `POST` | `/api/incidents/{id}/assign` | Reassign `{"user_id"}` |
|
||
| `POST` | `/api/incidents/{id}/snooze` | Hide until `{"until": RFC3339}` or `{"duration": "2h"}` |
|
||
| `DELETE` | `/api/incidents/{id}/snooze` | Un-snooze |
|
||
| `POST` | `/api/incidents/{id}/archive` | Archive (hides from the default list) |
|
||
| `DELETE` | `/api/incidents/{id}/archive` | Un-archive |
|
||
| `POST` | `/api/incidents/{id}/notes` | Add a note `{"content"}` |
|
||
| `DELETE` | `/api/incidents/{id}/notes/{eventID}` | Delete own note |
|
||
|
||
With no `?status=` filter, `GET /api/incidents` returns **open** incidents only —
|
||
the queue an on-call person wants. Currently snoozed and archived incidents are
|
||
excluded unless asked for. Actions that only make sense on an open incident
|
||
return `409` once it is resolved.
|
||
|
||
Notes are ordinary timeline events of type `note`; only they are deletable, and
|
||
only by their author. The rest of the timeline is a record of what happened.
|
||
|
||
#### The incident object
|
||
|
||
| Field | Type | Notes |
|
||
|---|---|---|
|
||
| `id` | integer | Server-assigned |
|
||
| `group_key` | string | Alertmanager's `groupKey` — opaque, treat as an identifier |
|
||
| `title` | string | Rendered from `groupLabels` |
|
||
| `group_labels` | object | String→string, as sent by Alertmanager |
|
||
| `status` | string | `"triggered"`, `"acknowledged"` or `"resolved"` |
|
||
| `severity` | string | *optional* — high-water mark across the incident's alerts; never lowered |
|
||
| `triggered_at` | timestamp | When the incident opened |
|
||
| `acknowledged_by_id` / `acknowledged_by` / `acknowledged_at` | | *optional* — user id, username, time |
|
||
| `assigned_to_id` / `assigned_to` | | *optional* — user id, username |
|
||
| `snoozed_until` | timestamp | *optional* — a value in the past reads as not snoozed |
|
||
| `resolved_at` | timestamp | *optional* |
|
||
| `resolution_source` | string | *optional* — `"alerts"`, `"manual"` or `"recovered"` |
|
||
| `archived_at` | timestamp | *optional* |
|
||
| `alerts` | array | Only on `GET /api/incidents/{id}` |
|
||
|
||
Treat `resolution_source` as an open set, as with the alert field of the same
|
||
name: degrade unknown values to "resolved, reason unknown".
|
||
|
||
#### The timeline event object
|
||
|
||
| Field | Type | Notes |
|
||
|---|---|---|
|
||
| `id` | integer | |
|
||
| `incident_id` | integer | |
|
||
| `type` | string | See below — treat as an open set |
|
||
| `user_id` / `username` | | *optional* — absent when the server acted rather than a person |
|
||
| `alert_id` | integer | *optional* — the alert an `alert_added` / `alert_resolved` event refers to |
|
||
| `detail` | string | *optional* — the note body, the snooze deadline, etc. |
|
||
| `created_at` | timestamp | |
|
||
|
||
Types written today: `triggered`, `alert_added`, `alert_resolved`,
|
||
`acknowledged`, `unacknowledged`, `assigned`, `snoozed`, `unsnoozed`, `resolved`,
|
||
`note`, `notified`, `notify_failed`, `deadman_silent`. On an `assigned` event
|
||
`user_id` is the **assignee**, not the actor. New types may be added; render
|
||
unknown ones generically rather than dropping them.
|
||
|
||
On `notified` and `notify_failed`, `detail` carries the notification kind
|
||
(`triggered` | `reminder` | `resolved`), and on a failure the reason after it.
|
||
`user_id` is who was paged — absent means the page went to the shared fallback
|
||
topic and so belongs to nobody. The topic itself is never written to the
|
||
timeline: it is a shared secret with the ntfy server, and every API key can read
|
||
this.
|
||
|
||
### Alerts
|
||
|
||
Alerts are read-only. Everything a person does happens on the incident.
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `GET` | `/api/alerts` | List alerts. Filters: `?status=firing\|resolved`, `?name=`, `?incident_id=`, `?archived=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?limit=` (default 50, max 500) |
|
||
| `GET` | `/api/alerts/{id}` | Get single alert |
|
||
|
||
Archived alerts are hidden from `GET /api/alerts` unless `?archived=true` is
|
||
passed; alert archiving is automatic housekeeping by the sweeper, not a user
|
||
action. Resolved alerts carry `resolution_source`: `"alertmanager"` for a real
|
||
resolved webhook, `"expiry"` when the sweeper inferred it (see
|
||
[Stale alert expiry](#stale-alert-expiry)), `"deadman"` for a heartbeat declared
|
||
dead (see [Dead man's switch](#dead-mans-switch)).
|
||
|
||
#### The alert object
|
||
|
||
Returned by `GET /api/alerts` (as an array) and `GET /api/alerts/{id}`.
|
||
Timestamps are RFC 3339 in UTC. Fields marked *optional* are omitted entirely
|
||
when unset, so clients must treat them as nullable.
|
||
|
||
| Field | Type | Notes |
|
||
|---|---|---|
|
||
| `id` | integer | Server-assigned; stable for the life of the row |
|
||
| `fingerprint` | string | Alertmanager's fingerprint — the upsert key |
|
||
| `name` | string | From the `alertname` label |
|
||
| `status` | string | `"firing"` or `"resolved"` |
|
||
| `labels` | object | String→string, as sent by Alertmanager |
|
||
| `annotations` | object | String→string, as sent by Alertmanager |
|
||
| `starts_at` | timestamp | When the alert instance began, **per Prometheus** |
|
||
| `ends_at` | timestamp | *optional* — absent while no end is known |
|
||
| `generator_url` | string | Link back to the originating Prometheus |
|
||
| `received_at` | timestamp | When the server last accepted a webhook for this alert — see below |
|
||
| `incident_id` | integer | *optional* — the most recent incident this alert belongs to |
|
||
| `resolution_source` | string | *optional* — `"alertmanager"`, `"expiry"` or `"deadman"` |
|
||
| `archived_at` | timestamp | *optional* — set while archived |
|
||
|
||
##### `received_at` is a liveness heartbeat
|
||
|
||
`starts_at` comes from Prometheus and **never changes** for the lifetime of an
|
||
alert instance. It says when the problem began, not whether it is still
|
||
happening — an alert that started twelve days ago looks identical whether
|
||
Alertmanager refreshed it a minute ago or went silent a week ago.
|
||
|
||
`received_at` is the field that answers "is this still live". It is set to the
|
||
server's clock on **every accepted webhook** for that fingerprint, including the
|
||
unchanged firing notifications Alertmanager re-sends every `repeat_interval`.
|
||
Clients may rely on this:
|
||
|
||
- **A firing alert whose `received_at` is advancing is still being refreshed.**
|
||
Stale-dating it against `repeat_interval` is a valid liveness check, and it is
|
||
what the built-in sweeper does (see
|
||
[Stale alert expiry](#stale-alert-expiry)).
|
||
- **`received_at` tracks accepted payloads, not delivery attempts.** A retry
|
||
that describes an older instance than the stored one is discarded, and a
|
||
discarded payload does not move `received_at`.
|
||
- **It stops advancing once the alert resolves,** because Alertmanager stops
|
||
re-sending. On an alert resolved by the sweeper
|
||
(`"resolution_source": "expiry"`) it therefore marks the last time
|
||
Alertmanager was actually heard from, which is earlier than `ends_at`.
|
||
|
||
`GET /api/alerts` is ordered by `received_at` descending — most recently
|
||
refreshed first — and the `?from=` / `?to=` filters on both the alert and stats
|
||
endpoints select on `received_at`, not `starts_at`.
|
||
|
||
##### `resolution_source` says how much to trust `ends_at`
|
||
|
||
An alert can leave the firing state two ways, and `resolution_source` records
|
||
which happened. Clients may rely on this:
|
||
|
||
- **Absent while firing.** It is set only on resolve, and a re-fire under the
|
||
same fingerprint clears it again, so its presence always agrees with
|
||
`"status": "resolved"`.
|
||
- **`"alertmanager"` — a real resolved webhook arrived.** `ends_at` is the end
|
||
time Alertmanager reported. It is an observed value and can be displayed as
|
||
fact.
|
||
- **`"expiry"` — the sweeper inferred the resolve** because Alertmanager stopped
|
||
refreshing the alert (see [Stale alert expiry](#stale-alert-expiry)). Nothing
|
||
ever reported an end, so **`ends_at` is approximate**: it is either the stale
|
||
`endsAt` watermark from the last notification, or — when that notification
|
||
carried none — the time the sweep ran, which lags the last real contact by up
|
||
to `TERDUT_STALE_AFTER` plus a sweep interval. Treat it as "no later than",
|
||
not as when the problem stopped.
|
||
|
||
On these alerts `received_at` is the more truthful signal: it marks the last
|
||
time Alertmanager was actually heard from. Surfacing the distinction is
|
||
worthwhile, since `"expiry"` can also mean the alert is still firing and the
|
||
notification path broke.
|
||
|
||
- **`"deadman"` — a heartbeat was declared dead** (see
|
||
[Dead man's switch](#dead-mans-switch)). Like `"expiry"`, an inference from
|
||
silence rather than an observed end, so `ends_at` is approximate — but a much
|
||
tighter one, bounded by `TERDUT_DEADMAN_TIMEOUT`. It is also the one resolution
|
||
a re-fire under the same `starts_at` can undo, since the switch coming back is
|
||
exactly the evidence that the inference was wrong.
|
||
|
||
Treat the value as an open set and tolerate ones you do not recognise — new
|
||
sources may be added, and unknown values should degrade to "resolved, reason
|
||
unknown" rather than being rejected.
|
||
|
||
### On-call schedule
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `POST` | `/api/schedule` | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...], "replace"}` — all-or-nothing |
|
||
| `GET` | `/api/schedule` | List entries. Filters: `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD` |
|
||
| `GET` | `/api/schedule/current` | Today's on-call user (UTC), 404 if none |
|
||
| `DELETE` | `/api/schedule/{id}` | Remove schedule entry |
|
||
|
||
### Statistics
|
||
|
||
All stat endpoints accept optional `?from=YYYY-MM-DD` and `?to=YYYY-MM-DD`, and exclude archived rows to match the default list views. Alert stats filter on `received_at`; incident stats filter on `triggered_at`.
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `GET` | `/api/stats/incidents` | `{total, triggered, acknowledged, resolved, mtta_seconds, mttr_seconds}` |
|
||
| `GET` | `/api/stats/alerts` | `{total, firing, resolved}` counts |
|
||
| `GET` | `/api/stats/alerts/top` | Most frequent alert names. `?limit=` (default 10, max 100) |
|
||
| `GET` | `/api/stats/alerts/by-hour` | Count per hour-of-day (UTC), all 24 slots returned |
|
||
| `GET` | `/api/stats/alerts/by-day` | Count per day-of-week, all 7 slots with names returned |
|
||
|
||
`mtta_seconds` (time to acknowledge) and `mttr_seconds` (time to resolve) are
|
||
averages over incidents that have actually been acknowledged or resolved, and are
|
||
**null** until there are any — null means "no data", not zero.
|
||
|
||
---
|
||
|
||
## Upgrading to incidents
|
||
|
||
The incidents release moves the workflow off alerts, which is a **breaking API
|
||
change**. These endpoints are gone:
|
||
|
||
| Removed | Replacement |
|
||
|---|---|
|
||
| `POST`/`DELETE` `/api/alerts/{id}/acknowledge` | `POST`/`DELETE` `/api/incidents/{id}/acknowledge` |
|
||
| `POST`/`DELETE` `/api/alerts/{id}/archive` | `POST`/`DELETE` `/api/incidents/{id}/archive` (alert archiving is now sweeper-only) |
|
||
| `GET`/`POST` `/api/alerts/{id}/comments` | `GET /api/incidents/{id}/timeline`, `POST /api/incidents/{id}/notes` |
|
||
| `DELETE /api/alerts/{id}/comments/{commentID}` | `DELETE /api/incidents/{id}/notes/{eventID}` |
|
||
|
||
The alert object also drops `acknowledged_by_id`, `acknowledged_by` and
|
||
`acknowledged_at`, and gains `incident_id`.
|
||
|
||
Migration `008_incidents.sql` runs automatically on start and preserves existing
|
||
data: every alert gets a backfilled incident carrying its acknowledgement, and
|
||
comments become timeline notes. Backfilled incidents have a `group_key` of
|
||
`backfill:<fingerprint>` — there is no historical `groupKey` to correlate on, so
|
||
they are one-per-alert rather than grouped.
|
||
|
||
Nothing about the two documented alert contracts changes: `received_at` is still
|
||
advanced on every accepted webhook, and `resolution_source` still means what it
|
||
did.
|
||
|
||
## Upgrading to dead man's switches
|
||
|
||
Dead man's switch handling is **on by default**, watching `alertname=Watchdog`
|
||
with a 15 minute timeout. If you already route `Watchdog` to this server, the
|
||
behaviour of that alert changes on upgrade, in both directions:
|
||
|
||
- it stops opening incidents when it arrives, and
|
||
- it starts opening one when it stops arriving.
|
||
|
||
**Check your `repeat_interval` before upgrading.** The switch pages whenever a
|
||
heartbeat has not been refreshed within `TERDUT_DEADMAN_TIMEOUT`, so a `Watchdog`
|
||
route inheriting a 4h or 12h `repeat_interval` will page constantly against the
|
||
15 minute default. Either give the heartbeat
|
||
[its own fast route](#alertmanager-configuration) — the point of the feature — or
|
||
set `TERDUT_DEADMAN_TIMEOUT` above your current `repeat_interval` until you have.
|
||
`TERDUT_DEADMAN_TIMEOUT=0` turns the whole thing off.
|
||
|
||
There is no migration and no schema change. An existing open incident from a
|
||
`Watchdog` that arrived under the old behaviour is unaffected; resolve it by hand.
|
||
|
||
---
|
||
|
||
## Development
|
||
|
||
```bash
|
||
go test ./... # run all tests
|
||
go build ./... # compile all packages
|
||
go run ./cmd/terdut # run locally
|
||
```
|