110 lines
5.2 KiB
Markdown
110 lines
5.2 KiB
Markdown
# terdut-server
|
|
|
|
Incident management for teams that already run Prometheus Alertmanager. Point Alertmanager at it,
|
|
and alerts become incidents that get assigned to whoever is on call, paged, escalated when nobody
|
|
answers, and tracked to resolution. One binary, one Postgres.
|
|
|
|

|
|
|
|
## Highlights
|
|
|
|
- **Alertmanager-native.** Receives Alertmanager webhooks directly, with no adapter, and groups
|
|
alerts into incidents by Alertmanager's own `groupKey`. An incident opens on a new occurrence,
|
|
not on every re-send. [Details](./docs/incidents.md)
|
|
- **A real incident workflow.** Acknowledge, assign, snooze, add notes, resolve and archive, with a
|
|
full timeline of who did what and when. Several clusters can feed one team without
|
|
cross-talk. [Details](./docs/incidents.md)
|
|
- **On-call rota.** Each team keeps its own rota, and new incidents go to whoever is on call.
|
|
[Web UI](./docs/web-ui.md)
|
|
- **Pages that escalate.** Push notifications through [ntfy](https://ntfy.sh), with an
|
|
Acknowledge button right in the notification, and escalation ladders that move on to the next
|
|
level when nobody answers. [Notifications](./docs/notifications.md) ·
|
|
[Escalation](./docs/escalation.md)
|
|
- **Notices when the alerts stop.** Dead man's switches turn the absence of a heartbeat such as
|
|
`Watchdog` into an incident. [Details](./docs/dead-mans-switch.md)
|
|
- **Teams.** Every team owns its queue, rota, escalation, alert sources and switches; people see
|
|
only the teams they belong to.
|
|
- **Stats.** Incident counts, mean time to acknowledge and resolve, and alert frequency by name,
|
|
hour and day. [Web UI](./docs/web-ui.md)
|
|
- **Single sign-on.** OpenID Connect with group-to-team and administrator mapping, and a device
|
|
flow so a terminal client can sign in through the browser. [Details](./docs/single-sign-on.md)
|
|
- **Built for the phone first.** The web UI is served by the same binary, follows the system's
|
|
dark mode, and can be added to the home screen.
|
|
- **An API for everything.** A REST API with API keys and service accounts for automation.
|
|
[Reference](./docs/api.md)
|
|
- **Easy to run.** A scratch container image, a Helm chart, and
|
|
[terdut-operator](https://git.ryuvia.com/niklas/terdut-operator) if you want teams and
|
|
escalation as Kubernetes objects. [Deployment](./docs/deployment.md)
|
|
|
|
## Screenshots
|
|
|
|
| | |
|
|
|---|---|
|
|
|  |  |
|
|
| **Work an incident**: acknowledge, assign, snooze, note, resolve | **Dark mode**, following the system |
|
|
|  |  |
|
|
| **On call**: who holds the pager now, and the week ahead | **Escalation**: who is paged next, and when |
|
|
|  |  |
|
|
| **Stats**: MTTA, MTTR and what fires most | **Alerts**: the raw feed behind the incidents |
|
|
|
|
On a phone the queue and the incident page are the same interface, with a sticky action bar:
|
|
|
|
<p>
|
|
<img src="./docs/images/mobile-queue.png" alt="The queue on a phone" width="240">
|
|
<img src="./docs/images/mobile-incident.png" alt="An incident on a phone" width="240">
|
|
</p>
|
|
|
|
## Quick start
|
|
|
|
You need Go 1.25+ and a Postgres 14+ the server can reach.
|
|
|
|
```bash
|
|
git clone https://git.ryuvia.com/niklas/terdut-server
|
|
cd terdut-server
|
|
export TERDUT_DB_DSN='postgres://terdut:secret@localhost:5432/terdut?sslmode=disable'
|
|
go run ./cmd/terdut
|
|
```
|
|
|
|
The server creates its schema on startup and listens on `:8080`. Create the first user, an
|
|
administrator, while no user exists yet:
|
|
|
|
```bash
|
|
curl -X POST http://localhost:8080/api/bootstrap \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"username": "admin", "email": "admin@example.com", "password": "<at least 10 characters>"}'
|
|
```
|
|
|
|
Sign in at <http://localhost:8080> with that username and password. The response also carries an
|
|
API key, shown **once**, for scripts:
|
|
|
|
```bash
|
|
curl -H "Authorization: Bearer $KEY" http://localhost:8080/api/users
|
|
```
|
|
|
|
Then create a team, add an alert source to get a webhook URL, and point Alertmanager's webhook
|
|
receiver at it: see [Alertmanager configuration](./docs/alertmanager.md).
|
|
|
|
## Documentation
|
|
|
|
The [documentation index](./docs/README.md) lists everything. The main pages:
|
|
|
|
| | |
|
|
|---|---|
|
|
| [Deployment](./docs/deployment.md) | Docker, Helm chart, database, backups, the operator |
|
|
| [Configuration](./docs/configuration.md) | Environment variables and settings |
|
|
| [Alertmanager configuration](./docs/alertmanager.md) | Routes, keys, webhooks |
|
|
| [Alerts and incidents](./docs/incidents.md) | Correlation, lifecycle, on-call assignment |
|
|
| [Single sign-on](./docs/single-sign-on.md) | OIDC and the terminal device flow |
|
|
| [API reference](./docs/api.md) | Every endpoint |
|
|
| [Development and releasing](./docs/development.md) | Tests, CI gate, release pipeline |
|
|
|
|
## Related
|
|
|
|
- [terdut-tui](https://git.ryuvia.com/niklas/terdut-tui): a terminal client for the same server.
|
|
- [terdut-operator](https://git.ryuvia.com/niklas/terdut-operator): a Kubernetes operator that
|
|
runs the server and manages teams, escalation, switches and alert sources as objects.
|
|
|
|
## License
|
|
|
|
See [LICENSE](./LICENSE).
|