The README was 1,240 lines of reference material and still described a SQLite quick start. It is now a short tour (highlights, screenshots of the web UI, an accurate quick start against Postgres), and each topic has its own page under docs/ with an index: deployment, configuration, Alertmanager, incidents, notifications, escalation, dead man's switches, single sign-on, web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it described. The "Upgrading to ..." sections for an unreleased product are dropped. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
5.2 KiB
Terminal Duty (terdut-server)
Incident management for teams that already run Prometheus Alertmanager. Point Alertmanager at it, and alerts become incidents that get assigned to whoever is on call, paged, escalated when nobody answers, and tracked to resolution. One binary, one Postgres.
Highlights
- Alertmanager-native. Receives Alertmanager webhooks directly, with no adapter, and groups
alerts into incidents by Alertmanager's own
groupKey. An incident opens on a new occurrence, not on every re-send. Details - A real incident workflow. Acknowledge, assign, snooze, add notes, resolve and archive, with a full timeline of who did what and when. Several clusters can feed one team without cross-talk. Details
- On-call rota. Each team keeps its own rota, and new incidents go to whoever is on call. Web UI
- Pages that escalate. Push notifications through ntfy, with an Acknowledge button right in the notification, and escalation ladders that move on to the next level when nobody answers. Notifications · Escalation
- Notices when the alerts stop. Dead man's switches turn the absence of a heartbeat such as
Watchdoginto an incident. Details - Teams. Every team owns its queue, rota, escalation, alert sources and switches; people see only the teams they belong to.
- Stats. Incident counts, mean time to acknowledge and resolve, and alert frequency by name, hour and day. Web UI
- Single sign-on. OpenID Connect with group-to-team and administrator mapping, and a device flow so a terminal client can sign in through the browser. Details
- Built for the phone first. The web UI is served by the same binary, follows the system's dark mode, and can be added to the home screen.
- An API for everything. A REST API with API keys and service accounts for automation. Reference
- Easy to run. A scratch container image, a Helm chart, and terdut-operator if you want teams and escalation as Kubernetes objects. Deployment
Screenshots
On a phone the queue and the incident page are the same interface, with a sticky action bar:
Quick start
You need Go 1.25+ and a Postgres 14+ the server can reach.
git clone https://git.ryuvia.com/niklas/terdut-server
cd terdut-server
export TERDUT_DB_DSN='postgres://terdut:secret@localhost:5432/terdut?sslmode=disable'
go run ./cmd/terdut
The server creates its schema on startup and listens on :8080. Create the first user, an
administrator, while no user exists yet:
curl -X POST http://localhost:8080/api/bootstrap \
-H "Content-Type: application/json" \
-d '{"username": "admin", "email": "admin@example.com", "password": "<at least 10 characters>"}'
Sign in at http://localhost:8080 with that username and password. The response also carries an API key, shown once, for scripts:
curl -H "Authorization: Bearer $KEY" http://localhost:8080/api/users
Then create a team, add an alert source to get a webhook URL, and point Alertmanager's webhook receiver at it: see Alertmanager configuration.
Documentation
The documentation index lists everything. The main pages:
| Deployment | Docker, Helm chart, database, backups, the operator |
| Configuration | Environment variables and settings |
| Alertmanager configuration | Routes, keys, webhooks |
| Alerts and incidents | Correlation, lifecycle, on-call assignment |
| Single sign-on | OIDC and the terminal device flow |
| API reference | Every endpoint |
| Development and releasing | Tests, CI gate, release pipeline |
Related
- terdut-tui: a terminal client for the same server.
- terdut-operator: a Kubernetes operator that runs the server and manages teams, escalation, switches and alert sources as objects.
License
See LICENSE.








