# Terminal Duty (terdut-server) Incident management for teams that already run Prometheus Alertmanager. Point Alertmanager at it, and alerts become incidents that get assigned to whoever is on call, paged, escalated when nobody answers, and tracked to resolution. One binary, one Postgres.  ## Highlights - **Alertmanager-native.** Receives Alertmanager webhooks directly, with no adapter, and groups alerts into incidents by Alertmanager's own `groupKey`. An incident opens on a new occurrence, not on every re-send. [Details](./docs/incidents.md) - **A real incident workflow.** Acknowledge, assign, snooze, add notes, resolve and archive, with a full timeline of who did what and when. Several clusters can feed one team without cross-talk. [Details](./docs/incidents.md) - **On-call rota.** Each team keeps its own rota, and new incidents go to whoever is on call. [Web UI](./docs/web-ui.md) - **Pages that escalate.** Push notifications through [ntfy](https://ntfy.sh), with an Acknowledge button right in the notification, and escalation ladders that move on to the next level when nobody answers. [Notifications](./docs/notifications.md) ยท [Escalation](./docs/escalation.md) - **Notices when the alerts stop.** Dead man's switches turn the absence of a heartbeat such as `Watchdog` into an incident. [Details](./docs/dead-mans-switch.md) - **Teams.** Every team owns its queue, rota, escalation, alert sources and switches; people see only the teams they belong to. - **Stats.** Incident counts, mean time to acknowledge and resolve, and alert frequency by name, hour and day. [Web UI](./docs/web-ui.md) - **Single sign-on.** OpenID Connect with group-to-team and administrator mapping, and a device flow so a terminal client can sign in through the browser. [Details](./docs/single-sign-on.md) - **Built for the phone first.** The web UI is served by the same binary, follows the system's dark mode, and can be added to the home screen. - **An API for everything.** A REST API with API keys and service accounts for automation. [Reference](./docs/api.md) - **Easy to run.** A scratch container image, a Helm chart, and [terdut-operator](https://git.ryuvia.com/niklas/terdut-operator) if you want teams and escalation as Kubernetes objects. [Deployment](./docs/deployment.md) ## Screenshots | | | |---|---| |  |  | | **Work an incident**: acknowledge, assign, snooze, note, resolve | **Dark mode**, following the system | |  |  | | **On call**: who holds the pager now, and the week ahead | **Escalation**: who is paged next, and when | |  |  | | **Stats**: MTTA, MTTR and what fires most | **Alerts**: the raw feed behind the incidents | On a phone the queue and the incident page are the same interface, with a sticky action bar: