niklas 392eab4127
CI / chart (push) Successful in 1s
CI / test (push) Successful in 9s
CI / security (push) Successful in 16s
Update README.md
2026-10-09 13:38:19 +00:00
2026-10-09 13:38:19 +00:00

terdut-server

Incident management for teams that already run Prometheus Alertmanager. Point Alertmanager at it, and alerts become incidents that get assigned to whoever is on call, paged, escalated when nobody answers, and tracked to resolution. One binary, one Postgres.

The incident queue with an incident open

Highlights

  • Alertmanager-native. Receives Alertmanager webhooks directly, with no adapter, and groups alerts into incidents by Alertmanager's own groupKey. An incident opens on a new occurrence, not on every re-send. Details
  • A real incident workflow. Acknowledge, assign, snooze, add notes, resolve and archive, with a full timeline of who did what and when. Several clusters can feed one team without cross-talk. Details
  • On-call rota. Each team keeps its own rota, and new incidents go to whoever is on call. Web UI
  • Pages that escalate. Push notifications through ntfy, with an Acknowledge button right in the notification, and escalation ladders that move on to the next level when nobody answers. Notifications · Escalation
  • Notices when the alerts stop. Dead man's switches turn the absence of a heartbeat such as Watchdog into an incident. Details
  • Teams. Every team owns its queue, rota, escalation, alert sources and switches; people see only the teams they belong to.
  • Stats. Incident counts, mean time to acknowledge and resolve, and alert frequency by name, hour and day. Web UI
  • Single sign-on. OpenID Connect with group-to-team and administrator mapping, and a device flow so a terminal client can sign in through the browser. Details
  • Built for the phone first. The web UI is served by the same binary, follows the system's dark mode, and can be added to the home screen.
  • An API for everything. A REST API with API keys and service accounts for automation. Reference
  • Easy to run. A scratch container image, a Helm chart, and terdut-operator if you want teams and escalation as Kubernetes objects. Deployment

Screenshots

An incident with a note on its timeline The queue in dark mode
Work an incident: acknowledge, assign, snooze, note, resolve Dark mode, following the system
The on-call page An escalation ladder
On call: who holds the pager now, and the week ahead Escalation: who is paged next, and when
Statistics The alert feed
Stats: MTTA, MTTR and what fires most Alerts: the raw feed behind the incidents

On a phone the queue and the incident page are the same interface, with a sticky action bar:

The queue on a phone An incident on a phone

Quick start

You need Go 1.25+ and a Postgres 14+ the server can reach.

git clone https://git.ryuvia.com/niklas/terdut-server
cd terdut-server
export TERDUT_DB_DSN='postgres://terdut:secret@localhost:5432/terdut?sslmode=disable'
go run ./cmd/terdut

The server creates its schema on startup and listens on :8080. Create the first user, an administrator, while no user exists yet:

curl -X POST http://localhost:8080/api/bootstrap \
  -H "Content-Type: application/json" \
  -d '{"username": "admin", "email": "admin@example.com", "password": "<at least 10 characters>"}'

Sign in at http://localhost:8080 with that username and password. The response also carries an API key, shown once, for scripts:

curl -H "Authorization: Bearer $KEY" http://localhost:8080/api/users

Then create a team, add an alert source to get a webhook URL, and point Alertmanager's webhook receiver at it: see Alertmanager configuration.

Documentation

The documentation index lists everything. The main pages:

Deployment Docker, Helm chart, database, backups, the operator
Configuration Environment variables and settings
Alertmanager configuration Routes, keys, webhooks
Alerts and incidents Correlation, lifecycle, on-call assignment
Single sign-on OIDC and the terminal device flow
API reference Every endpoint
Development and releasing Tests, CI gate, release pipeline
  • terdut-tui: a terminal client for the same server.
  • terdut-operator: a Kubernetes operator that runs the server and manages teams, escalation, switches and alert sources as objects.

License

See LICENSE.

S
Description
Incident management server for teams using Prometheus Alertmanager
Readme GPL-3.0 4.7 MiB
v0.43.0 Latest
2026-10-08 16:34:08 +00:00
Languages
Go 69.2%
JavaScript 22.2%
CSS 6%
Makefile 1.3%
HTML 1%
Other 0.3%