Files
terdut-server/README.md
T
Niklas Ye 17ee290d90 docs: the ack token is scoped, not single-use
The handler never deletes the token: it stays valid until expires_at and
is purged by the sweeper, so a second tap is an idempotent no-op rather
than a rejection. Caught by pressing Acknowledge twice against the live
server. What bounds the token is scope -- one incident, one action, one
day -- not a use count.
2026-08-07 11:41:15 +02:00

24 KiB

Terminal Duty (terdut-server)

Incident management server for teams using Prometheus Alertmanager.

  • Receives Alertmanager webhooks directly — no adapter needed
  • Turns alerts into incidents, correlated by Alertmanager's own groupKey
  • Incident workflow: acknowledge, assign, snooze, note, resolve, with a full timeline
  • On-call schedule management, with new incidents auto-assigned to whoever is on call
  • Alert and incident statistics, including MTTA and MTTR
  • REST API with per-user API key authentication
  • Single binary, SQLite storage — trivial to self-host

Quick start

Prerequisites: Go 1.21+

git clone https://github.com/yeniklas/terdut-server
cd terdut-server
go run ./cmd/terdut

The server starts on :8080 with a terdut.db file in the working directory.

Create the first user

curl -X POST http://localhost:8080/api/bootstrap \
  -H "Content-Type: application/json" \
  -d '{"username": "admin", "email": "admin@example.com"}'

Save the api_key.key value from the response — it is shown once only.

Use it as a bearer token for all subsequent requests:

export KEY=<your-key>
curl -H "Authorization: Bearer $KEY" http://localhost:8080/api/users

Docker

docker build -t terdut-server .
docker run -p 8080:8080 -v $(pwd)/data:/data \
  -e TERDUT_DB_PATH=/data/terdut.db \
  terdut-server

Kubernetes

A Helm chart is published from this repository:

helm repo add terdut-server https://yeniklas.github.io/terdut-server
helm upgrade --install terdut-server terdut-server/terdut-server \
  --namespace terdut-server --create-namespace \
  --set networking.hostname=terdut.example.com

The chart expects a Gateway API Gateway named envoy-main in the envoy-gateway-system namespace to already exist — it renders an HTTPRoute against it rather than an Ingress. TLS is terminated at the gateway, so the server itself never sees a certificate.

Value Default Description
networking.hostname terdut.example.com Hostname the HTTPRoute serves
networking.listener "" Gateway listener (sectionName) to bind to. Empty attaches to every matching listener, including plaintext HTTP — set it to the HTTPS listener's name to serve TLS only
networking.servicePort 8080 Port the route forwards to; keep in sync with service.port
bootstrap.enabled true Runs a post-install hook that creates the first user and stores its API key in the <release>-admin-key Secret. Already-bootstrapped servers are left alone
backupSidecar.enabled true Adds an idle python sidecar and the k8up annotations that dump the database through it

The API key travels in an Authorization: Bearer header, so set networking.listener whenever the hostname is reachable outside a trusted network.

Backups

The server image is FROM scratch — the binary and nothing else — so there is no interpreter to run a database dump in, and the database runs in WAL mode, where a file-level copy of the volume is not crash-consistent. The chart therefore ships an idle python:*-alpine sidecar that shares the data volume, and points k8up's backupcommand at it with k8up.io/backupcommand-container. Without that annotation k8up execs into .spec.containers[0] and the dump fails.

The dump is buffered and sanity-checked before its first byte reaches stdout, because k8up streams stdout straight into Restic: a dump that dies partway is otherwise stored as a silently truncated snapshot that k8up still reports as successful.

Set backupSidecar.enabled=false if you back the volume up some other way.


Configuration

Variable Default Description
TERDUT_ADDR :8080 TCP address to listen on
TERDUT_DB_PATH terdut.db Path to the SQLite database file
TERDUT_ARCHIVE_AFTER 168h (7d) How long a resolved alert or incident stays in the default list before being auto-archived
TERDUT_STALE_AFTER 6h How long a firing alert may go without a refreshing webhook before it is treated as resolved — must exceed your Alertmanager repeat_interval
TERDUT_NTFY_URL — ntfy server to publish push notifications to. Empty disables notifications entirely
TERDUT_NTFY_TOKEN — Bearer token for an access-controlled ntfy
TERDUT_NTFY_FALLBACK_TOPIC — Topic used when nobody is on call
TERDUT_PUBLIC_URL — Base URL a phone uses to reach this server, for the link and Acknowledge button inside a notification
TERDUT_NOTIFY_REPEAT 15m How long an incident may sit unacknowledged before it is paged again. 0 notifies once and never repeats

Durations use Go syntax (30m, 12h, 168h). An unparseable value falls back to the default.

In the Helm chart the two sweeper durations are set via sweeper.staleAfter and sweeper.archiveAfter, and notifications via the notify.* values.


Alertmanager configuration

Add terdut-server as a webhook receiver in your alertmanager.yml:

receivers:
  - name: terdut
    webhook_configs:
      - url: http://terdut-server:8080/api/alertmanager/webhook
        send_resolved: true

route:
  receiver: terdut

The webhook endpoint requires no authentication.


Alerts and incidents

There are two objects, and the difference between them is the whole design.

An alert is Alertmanager's record. It has two states, firing and resolved, one row per fingerprint, and no human ever writes to it. The API exposes alerts read-only.

An incident is the work item. It goes triggered → acknowledged → resolved, carries an assignee, a snooze, notes and a timeline, and is the only thing people act on. Many alerts belong to one incident.

Correlation uses Alertmanager's groupKey

Alertmanager has already grouped alerts according to the group_by routing tree you configured, and it sends the resulting groupKey and groupLabels on every webhook. Incidents adopt that answer rather than re-grouping alerts a second time — if you want different correlation, change group_by in alertmanager.yml and terdut follows.

At most one incident is open per groupKey at a time. Alerts firing in a group that already has an open incident join it. The incident's severity is a high-water mark — the highest severity label any of its alerts has carried — so an incident that hit critical still reads as critical after the critical alert clears.

An incident opens only on a new occurrence

An incident opens when an alert transitions into firing: a fingerprint that was never seen, an alert with a newer startsAt, or a resolved alert that started again. The unchanged firing notifications Alertmanager re-sends every repeat_interval are none of those, and open nothing.

This is what makes closing an incident by hand mean something. Without the rule, POST /api/incidents/{id}/resolve would be undone by the next re-send of an alert that never stopped firing.

Leaving the open state

  • Automatically, once every alert under the incident has stopped firing — whether by a resolved webhook or by the sweeper's stale-alert expiry. The incident gets "resolution_source": "alerts".
  • By hand, via POST /api/incidents/{id}/resolve ("resolution_source": "manual"). This is terminal: a later occurrence in that group opens a new incident rather than reopening this one. If the alert underneath never stops firing, the incident stays closed — that is what resolving by hand asserts.

To quieten an incident you expect to come back, snooze it instead (POST /api/incidents/{id}/snooze). A snooze hides the incident from the default list without closing it, and expires by simply falling into the past.

On-call assignment

A new incident is assigned to whoever holds today's schedule entry at the moment it opens (GET /api/schedule/current). If nobody is scheduled it opens unassigned. Reassign with POST /api/incidents/{id}/assign.

Push notifications

With TERDUT_NTFY_URL set, an incident that opens is pushed to the on-call person's phone through ntfy. Set each user's topic with PUT /api/users/{id}/notify; a user with no topic falls back to TERDUT_NTFY_FALLBACK_TOPIC, as does an incident that opens with nobody on call. If neither yields a topic, nothing is queued.

Three things get pushed:

  • triggered — an incident opened. Priority follows severity (critical maps to ntfy's max priority, the one that overrides the phone's quiet settings).
  • reminder — the incident is still triggered after TERDUT_NOTIFY_REPEAT. Repeats until somebody acts. Acknowledging, snoozing, resolving or archiving all stop it — snooze is the mute button.
  • resolved — every alert under the incident stopped firing. Only sent to whoever was paged in the first place, and only for the automatic cascade: resolving by hand pushes nothing, since the person who did it already knows.

Notifications carry an Acknowledge button that acknowledges the incident without opening anything. It POSTs to /api/notify/ack/{token}, an unauthenticated route authorised by the 256-bit token in its path — minted fresh per notification, scoped to one incident and one action, and valid for 24 hours. A real API key is never put in a notification, because the message is stored on the ntfy server and cached on the device.

The token is not consumed by use. Acknowledging is idempotent, so a token stays valid for its full 24 hours and a second tap is a no-op that reports the incident's current state rather than an error — which is what you want when a tap is retried on a flaky mobile connection. What bounds it is scope, not a use count: one incident, one action, one day. Expired tokens are purged by the sweeper.

Two consequences worth planning for:

  • /api/notify/ack/{token} must stay publicly reachable, or the button will not work when the responder is off your network.
  • Notifications sent to the fallback topic carry no Acknowledge button. The topic is shared, and a button on it would let any subscriber acknowledge as somebody else.

Delivery is a queue, not an inline call: the webhook writes a row and a background notifier sends it within 30 seconds, retrying with exponential backoff up to 8 attempts. Nothing about ingestion blocks on ntfy being reachable.

Stale alert expiry

A resolved webhook is the only signal that an alert has stopped firing, so a notification that is dropped, silenced, or lost to a restart would otherwise pin that alert as firing forever. A background sweeper resolves firing alerts that Alertmanager has stopped refreshing, using either signal:

  • the endsAt watermark on the last notification has passed, or
  • no webhook has refreshed the alert within TERDUT_STALE_AFTER.

Alertmanager re-sends firing notifications every repeat_interval, which is what keeps a live alert fresh — so TERDUT_STALE_AFTER must be comfortably larger than your repeat_interval (default 4h), or live alerts will be resolved prematurely. Alerts resolved this way are marked "resolution_source": "expiry" to distinguish them from a real Alertmanager resolve ("alertmanager").

An expiry cascades: once it leaves an incident with nothing firing under it, the incident resolves too, in the same sweep.


API reference

Authentication

All endpoints except /api/bootstrap and /api/alertmanager/webhook require:

Authorization: Bearer <api-key>

Users

Method Path Description
POST /api/bootstrap Create first user + API key (only works on empty DB)
GET /api/users List users
POST /api/users Create user {"username","email"}
DELETE /api/users/{id} Delete user (cascades to keys)
PUT /api/users/{id}/notify Set push notification target {"ntfy_topic"} — empty string clears it
POST /api/users/{id}/api-keys Issue API key {"name"} — key shown once
DELETE /api/users/{id}/api-keys/{keyID} Revoke API key

Alert ingestion

Method Path Description
POST /api/alertmanager/webhook Alertmanager v4 webhook receiver (no auth)

Notifications

Method Path Description
POST /api/notify/ack/{token} Acknowledge an incident from a push notification's Acknowledge button. No auth: the token in the path is the credential — one incident, one action, 24 hours, idempotent. Must stay publicly reachable

Incidents

Method Path Description
GET /api/incidents List incidents. Filters: ?status=triggered|acknowledged|resolved, ?severity=, ?assigned_to=<user id>, ?archived=true, ?snoozed=true, ?from=YYYY-MM-DD, ?to=YYYY-MM-DD, ?sort=severity, ?limit= (default 50, max 500)
GET /api/incidents/{id} Get single incident, with its alerts inline
GET /api/incidents/{id}/alerts Alerts under this incident
GET /api/incidents/{id}/timeline Full event history, chronological
POST /api/incidents/{id}/acknowledge Acknowledge (stamps authed user + time)
DELETE /api/incidents/{id}/acknowledge Clear acknowledgement, back to triggered
POST /api/incidents/{id}/resolve Close by hand — terminal, see above
POST /api/incidents/{id}/assign Reassign {"user_id"}
POST /api/incidents/{id}/snooze Hide until {"until": RFC3339} or {"duration": "2h"}
DELETE /api/incidents/{id}/snooze Un-snooze
POST /api/incidents/{id}/archive Archive (hides from the default list)
DELETE /api/incidents/{id}/archive Un-archive
POST /api/incidents/{id}/notes Add a note {"content"}
DELETE /api/incidents/{id}/notes/{eventID} Delete own note

With no ?status= filter, GET /api/incidents returns open incidents only — the queue an on-call person wants. Currently snoozed and archived incidents are excluded unless asked for. Actions that only make sense on an open incident return 409 once it is resolved.

Notes are ordinary timeline events of type note; only they are deletable, and only by their author. The rest of the timeline is a record of what happened.

The incident object

Field Type Notes
id integer Server-assigned
group_key string Alertmanager's groupKey — opaque, treat as an identifier
title string Rendered from groupLabels
group_labels object String→string, as sent by Alertmanager
status string "triggered", "acknowledged" or "resolved"
severity string optional — high-water mark across the incident's alerts; never lowered
triggered_at timestamp When the incident opened
acknowledged_by_id / acknowledged_by / acknowledged_at optional — user id, username, time
assigned_to_id / assigned_to optional — user id, username
snoozed_until timestamp optional — a value in the past reads as not snoozed
resolved_at timestamp optional
resolution_source string optional — "alerts" or "manual"
archived_at timestamp optional
alerts array Only on GET /api/incidents/{id}

Treat resolution_source as an open set, as with the alert field of the same name: degrade unknown values to "resolved, reason unknown".

The timeline event object

Field Type Notes
id integer
incident_id integer
type string See below — treat as an open set
user_id / username optional — absent when the server acted rather than a person
alert_id integer optional — the alert an alert_added / alert_resolved event refers to
detail string optional — the note body, the snooze deadline, etc.
created_at timestamp

Types written today: triggered, alert_added, alert_resolved, acknowledged, unacknowledged, assigned, snoozed, unsnoozed, resolved, note. On an assigned event user_id is the assignee, not the actor. New types may be added; render unknown ones generically rather than dropping them.

Alerts

Alerts are read-only. Everything a person does happens on the incident.

Method Path Description
GET /api/alerts List alerts. Filters: ?status=firing|resolved, ?name=, ?incident_id=, ?archived=true, ?from=YYYY-MM-DD, ?to=YYYY-MM-DD, ?limit= (default 50, max 500)
GET /api/alerts/{id} Get single alert

Archived alerts are hidden from GET /api/alerts unless ?archived=true is passed; alert archiving is automatic housekeeping by the sweeper, not a user action. Resolved alerts carry resolution_source: "alertmanager" for a real resolved webhook, "expiry" when the sweeper inferred it (see Stale alert expiry).

The alert object

Returned by GET /api/alerts (as an array) and GET /api/alerts/{id}. Timestamps are RFC 3339 in UTC. Fields marked optional are omitted entirely when unset, so clients must treat them as nullable.

Field Type Notes
id integer Server-assigned; stable for the life of the row
fingerprint string Alertmanager's fingerprint — the upsert key
name string From the alertname label
status string "firing" or "resolved"
labels object String→string, as sent by Alertmanager
annotations object String→string, as sent by Alertmanager
starts_at timestamp When the alert instance began, per Prometheus
ends_at timestamp optional — absent while no end is known
generator_url string Link back to the originating Prometheus
received_at timestamp When the server last accepted a webhook for this alert — see below
incident_id integer optional — the most recent incident this alert belongs to
resolution_source string optional — "alertmanager" or "expiry"
archived_at timestamp optional — set while archived
received_at is a liveness heartbeat

starts_at comes from Prometheus and never changes for the lifetime of an alert instance. It says when the problem began, not whether it is still happening — an alert that started twelve days ago looks identical whether Alertmanager refreshed it a minute ago or went silent a week ago.

received_at is the field that answers "is this still live". It is set to the server's clock on every accepted webhook for that fingerprint, including the unchanged firing notifications Alertmanager re-sends every repeat_interval. Clients may rely on this:

  • A firing alert whose received_at is advancing is still being refreshed. Stale-dating it against repeat_interval is a valid liveness check, and it is what the built-in sweeper does (see Stale alert expiry).
  • received_at tracks accepted payloads, not delivery attempts. A retry that describes an older instance than the stored one is discarded, and a discarded payload does not move received_at.
  • It stops advancing once the alert resolves, because Alertmanager stops re-sending. On an alert resolved by the sweeper ("resolution_source": "expiry") it therefore marks the last time Alertmanager was actually heard from, which is earlier than ends_at.

GET /api/alerts is ordered by received_at descending — most recently refreshed first — and the ?from= / ?to= filters on both the alert and stats endpoints select on received_at, not starts_at.

resolution_source says how much to trust ends_at

An alert can leave the firing state two ways, and resolution_source records which happened. Clients may rely on this:

  • Absent while firing. It is set only on resolve, and a re-fire under the same fingerprint clears it again, so its presence always agrees with "status": "resolved".

  • "alertmanager" — a real resolved webhook arrived. ends_at is the end time Alertmanager reported. It is an observed value and can be displayed as fact.

  • "expiry" — the sweeper inferred the resolve because Alertmanager stopped refreshing the alert (see Stale alert expiry). Nothing ever reported an end, so ends_at is approximate: it is either the stale endsAt watermark from the last notification, or — when that notification carried none — the time the sweep ran, which lags the last real contact by up to TERDUT_STALE_AFTER plus a sweep interval. Treat it as "no later than", not as when the problem stopped.

    On these alerts received_at is the more truthful signal: it marks the last time Alertmanager was actually heard from. Surfacing the distinction is worthwhile, since "expiry" can also mean the alert is still firing and the notification path broke.

Treat the value as an open set and tolerate ones you do not recognise — new sources may be added, and unknown values should degrade to "resolved, reason unknown" rather than being rejected.

On-call schedule

Method Path Description
POST /api/schedule Assign user to dates {"user_id", "dates":["YYYY-MM-DD",...]} — all-or-nothing
GET /api/schedule List entries. Filters: ?from=YYYY-MM-DD, ?to=YYYY-MM-DD
GET /api/schedule/current Today's on-call user (UTC), 404 if none
DELETE /api/schedule/{id} Remove schedule entry

Statistics

All stat endpoints accept optional ?from=YYYY-MM-DD and ?to=YYYY-MM-DD, and exclude archived rows to match the default list views. Alert stats filter on received_at; incident stats filter on triggered_at.

Method Path Description
GET /api/stats/incidents {total, triggered, acknowledged, resolved, mtta_seconds, mttr_seconds}
GET /api/stats/alerts {total, firing, resolved} counts
GET /api/stats/alerts/top Most frequent alert names. ?limit= (default 10, max 100)
GET /api/stats/alerts/by-hour Count per hour-of-day (UTC), all 24 slots returned
GET /api/stats/alerts/by-day Count per day-of-week, all 7 slots with names returned

mtta_seconds (time to acknowledge) and mttr_seconds (time to resolve) are averages over incidents that have actually been acknowledged or resolved, and are null until there are any — null means "no data", not zero.


Upgrading to incidents

The incidents release moves the workflow off alerts, which is a breaking API change. These endpoints are gone:

Removed Replacement
POST/DELETE /api/alerts/{id}/acknowledge POST/DELETE /api/incidents/{id}/acknowledge
POST/DELETE /api/alerts/{id}/archive POST/DELETE /api/incidents/{id}/archive (alert archiving is now sweeper-only)
GET/POST /api/alerts/{id}/comments GET /api/incidents/{id}/timeline, POST /api/incidents/{id}/notes
DELETE /api/alerts/{id}/comments/{commentID} DELETE /api/incidents/{id}/notes/{eventID}

The alert object also drops acknowledged_by_id, acknowledged_by and acknowledged_at, and gains incident_id.

Migration 008_incidents.sql runs automatically on start and preserves existing data: every alert gets a backfilled incident carrying its acknowledgement, and comments become timeline notes. Backfilled incidents have a group_key of backfill:<fingerprint> — there is no historical groupKey to correlate on, so they are one-per-alert rather than grouped.

Nothing about the two documented alert contracts changes: received_at is still advanced on every accepted webhook, and resolution_source still means what it did.


Development

go test ./...        # run all tests
go build ./...       # compile all packages
go run ./cmd/terdut  # run locally