5b4683febf
Team membership from single sign-on used to come from one env var,
TERDUT_OIDC_GROUP_MAPPINGS, matched against a team by name and creating
the team if none existed. That put the decision in the server's
environment rather than the team's own hands, needed a restart to
change, and let a typo in a team name silently create a stray team.
Each team now carries its own oidc_member_group and oidc_owner_group,
set by its owner (or an administrator) from the Members tab, or PUT
/api/teams/{teamID}/oidc-groups. The "highest role wins" rule
TERDUT_OIDC_GROUP_MAPPINGS used to apply across mappings now applies
across one team's own two fields: being in both makes somebody an
owner. The sync no longer creates a team by name; a group only ever
grants into a team that already exists.
This is a breaking change for anyone already using
TERDUT_OIDC_GROUP_MAPPINGS, deliberately not auto-migrated: an
OIDC-sourced membership is dropped at a user's next sign-in until its
team's owner re-sets the group. The README's OIDC section spells out
the migration and the risk of a visible access gap during it.
TERDUT_OIDC_ADMIN_GROUP and TERDUT_OIDC_ALLOWED_GROUPS are untouched --
only team membership moved. terdut-tui needs no change: it only reads
GET /api/teams and GET /api/teams/{id}/members, and neither response
shape moved.
1281 lines
72 KiB
Markdown
1281 lines
72 KiB
Markdown
# Terminal Duty (terdut-server)
|
||
|
||
Incident management server for teams using Prometheus Alertmanager.
|
||
|
||
- Receives Alertmanager webhooks directly — no adapter needed
|
||
- Turns alerts into **incidents**, correlated by Alertmanager's own `groupKey`
|
||
- Incident workflow: acknowledge, assign, snooze, note, resolve, with a full timeline
|
||
- On-call schedule management, with new incidents auto-assigned to whoever is on call
|
||
- Alert and incident statistics, including MTTA and MTTR
|
||
- Web UI for phones and desktops, served by the same binary
|
||
- REST API with per-user API key authentication
|
||
- Single binary plus a Postgres — straightforward to self-host
|
||
|
||
---
|
||
|
||
## Quick start
|
||
|
||
**Prerequisites:** Go 1.21+
|
||
|
||
```bash
|
||
git clone https://git.ryuvia.com/niklas/terdut-server
|
||
cd terdut-server
|
||
go run ./cmd/terdut
|
||
```
|
||
|
||
The server starts on `:8080` with a `terdut.db` file in the working directory.
|
||
|
||
### Create the first user
|
||
|
||
```bash
|
||
curl -X POST http://localhost:8080/api/bootstrap \
|
||
-H "Content-Type: application/json" \
|
||
-d '{"username": "admin", "email": "admin@example.com", "password": "<at least 10 characters>"}'
|
||
```
|
||
|
||
Save the `api_key.key` value from the response — it is shown **once only**. The
|
||
`password` is optional and is what signs you in to the [web UI](#web-ui).
|
||
|
||
Use it as a bearer token for all subsequent requests:
|
||
|
||
```bash
|
||
export KEY=<your-key>
|
||
curl -H "Authorization: Bearer $KEY" http://localhost:8080/api/users
|
||
```
|
||
|
||
### Web UI
|
||
|
||
The server serves a web UI at `/`: the incident queue, each incident's alerts
|
||
and timeline with every action (acknowledge, assign, snooze, note, resolve,
|
||
archive), who is on call, the alert feed, and an *Account* tab for your own
|
||
password and the ntfy topic your pages go to. It is built for a phone first. On a phone
|
||
it navigates through a hamburger menu and has a sticky action bar, it follows the
|
||
system's dark mode, and it can be added to the home screen. From 900px wide it switches
|
||
to a sidebar with the queue and the incident side by side. The Stats page shows
|
||
incident counts, MTTA and MTTR, and alert frequency by name, hour and day over a
|
||
chosen range.
|
||
|
||
You sign in with a username and password. Users have no password until one is
|
||
set, and a user without one can only use API keys:
|
||
|
||
```bash
|
||
# an admin sets someone's first password with their API key
|
||
curl -X PUT http://localhost:8080/api/users/2/password \
|
||
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
|
||
-d '{"password": "<at least 10 characters>"}'
|
||
```
|
||
|
||
After that, users change it themselves under *Account*. Changing your own
|
||
password requires the current one.
|
||
|
||
How a browser stays signed in:
|
||
|
||
- A successful login sets an `HttpOnly`, `SameSite=Lax` session cookie. It lasts
|
||
30 days and slides forward while it is used, so an on-call phone stays signed
|
||
in.
|
||
- The cookie is marked `Secure` when `TERDUT_PUBLIC_URL` starts with `https://`,
|
||
so set it to the HTTPS address. TLS terminates at the gateway and the server
|
||
itself only ever sees plain HTTP.
|
||
- Requests authenticated by the cookie are checked for cross-origin use (Go's
|
||
`http.CrossOriginProtection`). That is the CSRF guard. Bearer-key clients are
|
||
not affected.
|
||
- Setting a password signs that user out everywhere else.
|
||
- Ten failed logins for one username within 15 minutes lock that username for
|
||
the rest of the window.
|
||
|
||
With `TERDUT_PUBLIC_URL` set, tapping a push notification opens the incident in
|
||
the web UI (`/incidents/{id}`).
|
||
|
||
A **Team** tab holds everything a team owns, in five sub-sections with a URL
|
||
each and a strip across the top to move between them: the on-call rota
|
||
(`/team/rota`), the membership (`/team/members`), the escalation ladder
|
||
(`/team/escalation`), the alert sources with their keys (`/team/sources`) and
|
||
the dead man's switches (`/team/deadman`). `/team` itself is an overview — who
|
||
is on call today, how many members and owners, how many ladder levels, how many
|
||
keys and how many switches — so a page fetches only what it shows. An owner
|
||
edits it; a member sees the same pages read-only, because the server refuses
|
||
their writes anyway. Somebody in more than one team picks between them above
|
||
the strip, since the choice changes the subject of all five.
|
||
|
||
The rota is a month at a time, one coloured initial per day with a legend
|
||
underneath, and it says how many days are left uncovered — the question a rota
|
||
is read for is who holds which stretch, and a run of one colour answers it
|
||
where a list of dates does not. An owner taps a day to hand it to somebody or
|
||
empty it, and fills a whole shift from the range form folded in below.
|
||
|
||
The **Admin** tab appears only for a system administrator, and holds what
|
||
belongs to the whole server rather than to one team. It has three sub-sections,
|
||
each with a URL of its own and a strip across the top to move between them:
|
||
every team (`/admin/teams`), every user (`/admin/users`), and the settings that
|
||
used to be environment variables (`/admin/settings`). `/admin` itself is an
|
||
overview — how many of each, and what each section is for. Adding somebody is
|
||
minting them an invite link into a team, rather than creating a bare account:
|
||
the person who accepts it picks their own password, so one never passes through
|
||
an administrator, and the link carries the team, so they land somewhere with a
|
||
queue in it. That happens on the team's own page, since an invite is a fact
|
||
about a team; the user list points there rather than asking which team beside a
|
||
form.
|
||
|
||
A name in the team list opens **that team's page**, at `/admin/teams/{id}`: when it
|
||
was created, how many are in it and how much is open, a field to rename it, the
|
||
members with their roles, the invites into it, and deletion. The member list is the
|
||
one thing there that needed a new endpoint — `GET /api/teams/{id}/members` is
|
||
member-only and answers `404` to an administrator who is not in the team, which is
|
||
the rule and not an oversight, so the page reads `GET /api/admin/teams/{id}` instead.
|
||
An administrator still sees none of that team's incidents, alerts or rota.
|
||
|
||
A name in the user list opens **that person's page**, at `/admin/users/{id}`: their
|
||
email and when they joined, where their notifications go, whether they are an
|
||
administrator, whether the account is disabled, the teams they are in with their
|
||
role in each, a password field for a first or forgotten one, and deletion. It is
|
||
the one place membership is edited from the person's side — the Team tab answers
|
||
"who is in this team", and answering "which teams is this person in" there means
|
||
visiting each team in turn.
|
||
|
||
### Single sign-on (OIDC)
|
||
|
||
terdut can sign people in through any OpenID Connect provider; the examples use
|
||
[Authentik](https://goauthentik.io/). Groups at the provider decide who may sign
|
||
in, which teams they belong to and whether they administer the install, much as
|
||
Grafana's OAuth role and org mapping does. Password login keeps working alongside
|
||
it unless you turn it off.
|
||
|
||
**At the provider**, create an OAuth2/OpenID provider and an application for it:
|
||
a *confidential* client, redirect URI `<TERDUT_PUBLIC_URL>/api/oidc/callback`, and
|
||
the `openid`, `profile` and `email` scopes. The issuer is the application's, e.g.
|
||
`https://auth.example.com/application/o/terdut/`. Then set:
|
||
|
||
```sh
|
||
TERDUT_PUBLIC_URL=https://terdut.example.com
|
||
TERDUT_OIDC_ISSUER=https://auth.example.com/application/o/terdut/
|
||
TERDUT_OIDC_CLIENT_ID=terdut
|
||
TERDUT_OIDC_CLIENT_SECRET=...
|
||
TERDUT_OIDC_ALLOWED_GROUPS=terdut-users,terdut-admins
|
||
TERDUT_OIDC_ADMIN_GROUP=terdut-admins
|
||
```
|
||
|
||
Which team a group grants is not server-wide config: each team names its own
|
||
group(s), set by that team's own owner (or an administrator) from its Members
|
||
tab, or `PUT /api/teams/{teamID}/oidc-groups {"member_group":"sre","owner_group":"sre-leads"}`.
|
||
A team must already exist before a group can grant access to it — the sync
|
||
never creates one.
|
||
|
||
The web UI's sign-in page shows a "Sign in with <name>" button (a plain link to
|
||
`/api/oidc/login`) above the password form, or instead of it when
|
||
`TERDUT_PASSWORD_LOGIN=false`; it asks `GET /api/auth/config` what the server offers
|
||
(`password_login`, `oidc.enabled`, `oidc.name`). A refused sign-in comes back to that
|
||
page with the reason spelled out. Access the groups grant is badged **SSO** on the
|
||
Team, Admin and per-user pages, with its edit and remove controls disabled, and the
|
||
Account page does not offer to set a password nobody could use.
|
||
|
||
**What a sign-in does**
|
||
|
||
1. *Who.* The provider's `(issuer, subject)` is the identity. The first time, a
|
||
user is found by email — only when the provider marks it verified, or
|
||
`TERDUT_OIDC_TRUST_EMAIL` is set — or created with no password. A username taken
|
||
by somebody else gets a numeric suffix (`alice-2`). Username and email follow the
|
||
provider at each sign-in. Authentik reports `email_verified` as false unless
|
||
configured otherwise, so linking existing users usually needs
|
||
`TERDUT_OIDC_TRUST_EMAIL=true`.
|
||
2. *Whether.* With `TERDUT_OIDC_ALLOWED_GROUPS` set, somebody in none of them is
|
||
refused and nothing is created.
|
||
3. *What.* The administrator flag follows `TERDUT_OIDC_ADMIN_GROUP`. Team roles
|
||
follow each team's own `oidc_member_group`/`oidc_owner_group`; where both of a
|
||
team's groups match, the owner group wins.
|
||
|
||
**Managed access.** What the sync grants is marked as managed by single sign-on,
|
||
and only that is ever changed by it. It is added at sign-in, and removed at the
|
||
next sign-in after the group is gone, even if that leaves a team without an owner
|
||
(an administrator can always repair a team) — the provider is the source of truth
|
||
for what it grants, so the last-owner and last-administrator guards do not apply.
|
||
Memberships and administrators added by hand are left alone; the exception is a
|
||
hand-added member whose team's own group grants a *higher* role, who is raised and
|
||
from then on managed. Editing managed access by hand (`POST` or `DELETE` on a
|
||
team's members, revoking an SSO-granted administrator) is refused with `409`, since
|
||
the next sign-in would undo it.
|
||
|
||
> **Upgrading past migration 013: reconfigure every team's groups.**
|
||
> `TERDUT_OIDC_GROUP_MAPPINGS` is gone, and the sync no longer creates a team by
|
||
> name. Group-to-team-role mapping is now each team's own setting — an owner sets
|
||
> it from the Members tab, or `PUT /api/teams/{teamID}/oidc-groups`. Until a team's
|
||
> owner does that, an OIDC-sourced membership in it is dropped at that user's next
|
||
> SSO sign-in, the same as any other loss of group access. Set every team's groups
|
||
> before affected users next sign in, to avoid a visible gap in access.
|
||
|
||
**How fast changes arrive.** Groups are read only at sign-in. A session made by an
|
||
SSO sign-in has a hard ceiling (`TERDUT_OIDC_SESSION_MAX_AGE`, default 12h) that
|
||
sliding never extends, so a change at the provider reaches terdut within that time.
|
||
Password sessions are unaffected.
|
||
|
||
> **API keys are not revoked when somebody is removed at the provider.** terdut
|
||
> holds no refresh token and never asks the provider again, so a person removed
|
||
> from every allowed group loses their sessions within `TERDUT_OIDC_SESSION_MAX_AGE`
|
||
> and cannot sign in again, but keeps any API key they made (the TUI and scripts use
|
||
> them) until an administrator disables the user in terdut.
|
||
|
||
**Signing in from a terminal.** A client with no browser of its own, such as the
|
||
TUI over SSH, signs in with a device code, run by terdut itself so the terminal
|
||
never talks to the provider:
|
||
|
||
1. The terminal calls `POST /api/oidc/device` and shows the person a link
|
||
(`<TERDUT_PUBLIC_URL>/device?code=XXXX-XXXX`) and the code.
|
||
2. On any device the person opens the link, signs in (by the provider or by
|
||
password, whatever the login page offers), sees the code and the account, and
|
||
presses **Approve**. Only a browser session can approve; an API key cannot.
|
||
3. The terminal polls `POST /api/oidc/device/token` every 5 seconds and is given the
|
||
ordinary `terdut_session` cookie once. A person who signs in through the provider
|
||
gets the same `TERDUT_OIDC_SESSION_MAX_AGE` ceiling on the terminal's session as
|
||
on their browser's.
|
||
|
||
A login expires after 10 minutes. `GET /api/auth/config` reports `device_login`.
|
||
|
||
**If the provider is down**, terdut still starts (discovery is fetched on first
|
||
use) and password login is the way in. With `TERDUT_PASSWORD_LOGIN=false` that way
|
||
is closed: set it back to `true`. The first administrator comes from the bootstrap
|
||
endpoint, and stays a manual administrator that no group can revoke; on an SSO-only
|
||
install set `bootstrap.enabled: false` in the chart if you don't want that account,
|
||
or keep it and never give it a password.
|
||
|
||
### Docker
|
||
|
||
```bash
|
||
docker build -t terdut-server .
|
||
docker run -p 8080:8080 \
|
||
-e TERDUT_DB_DSN='postgres://terdut:secret@host.docker.internal:5432/terdut?sslmode=disable' \
|
||
terdut-server
|
||
```
|
||
|
||
The server creates its own schema on startup and needs a reachable Postgres; it stores nothing on
|
||
disk, so there is no volume to mount.
|
||
|
||
### Kubernetes
|
||
|
||
A Helm chart is published from this repository as an OCI artifact, versioned in lockstep
|
||
with the app — chart `x.y.z` is always app `vx.y.z`:
|
||
|
||
```bash
|
||
helm upgrade --install terdut-server oci://git.ryuvia.com/niklas/terdut-server \
|
||
--version 0.9.2 \
|
||
--namespace terdut-server --create-namespace \
|
||
--set networking.hostname=terdut.example.com
|
||
```
|
||
|
||
The chart expects a [Gateway API](https://gateway-api.sigs.k8s.io/) Gateway named `envoy-main` in
|
||
the `envoy-gateway-system` namespace to already exist — it renders an `HTTPRoute` against it rather
|
||
than an `Ingress`. TLS is terminated at the gateway, so the server itself never sees a certificate.
|
||
|
||
| Value | Default | Description |
|
||
|---|---|---|
|
||
| `networking.hostname` | `terdut.example.com` | Hostname the `HTTPRoute` serves |
|
||
| `networking.listener` | `""` | Gateway listener (`sectionName`) to bind to. Empty attaches to every matching listener, **including plaintext HTTP** — set it to the HTTPS listener's name to serve TLS only |
|
||
| `networking.servicePort` | `8080` | Port the route forwards to; keep in sync with `service.port` |
|
||
| `bootstrap.enabled` | `true` | Runs a post-install hook that creates the first user and stores its API key in the `<release>-admin-key` Secret. Already-bootstrapped servers are left alone |
|
||
| `database.dsn` | `""` | **Required.** Postgres DSN, with no password in it. The chart provisions no database |
|
||
| `database.passwordSecret.name` | `""` | Secret supplying `PGPASSWORD`. With the Zalando postgres operator, the Secret it generates for the role |
|
||
| `database.passwordSecret.key` | `password` | Key within that Secret |
|
||
|
||
The API key travels in an `Authorization: Bearer` header, so set `networking.listener` whenever the
|
||
hostname is reachable outside a trusted network.
|
||
|
||
#### The database
|
||
|
||
The chart provisions no database: it takes a DSN and expects a Postgres that already exists. In this
|
||
cluster the wrapper chart declares an `acid.zalan.do/v1 postgresql` CR; anywhere else, any reachable
|
||
Postgres 14+ will do.
|
||
|
||
The DSN carries no password. pgx falls back to libpq's environment variables for whatever the DSN
|
||
leaves out, so the password arrives as `PGPASSWORD` from a Secret and never appears in values, in
|
||
the rendered manifest or in `kubectl describe pod`. With the postgres operator that Secret is the
|
||
one it generates for the role, so a rebuild mints a new password with nothing to keep in sync —
|
||
the same wiring miniflux uses.
|
||
|
||
The server migrates its own schema on startup, so a new database only has to exist and be writable.
|
||
|
||
#### Backups
|
||
|
||
Postgres is backed up where it runs, not from here. The database pod carries a
|
||
[k8up](https://k8up.io/) `k8up.io/backupcommand` annotation that streams a `pg_dump`, the same way
|
||
gitea and immich do in this cluster.
|
||
|
||
This used to be the app's problem: the SQLite database lived on a PVC beside the server, the image
|
||
is `FROM scratch` with no interpreter to dump it, and WAL mode makes a file-level copy of the volume
|
||
non-crash-consistent — so the chart shipped an idle `python:*-alpine` sidecar purely to give k8up
|
||
somewhere to exec. The sidecar, the PVC and the `backupSidecar` values are all gone.
|
||
|
||
---
|
||
|
||
## Configuration
|
||
|
||
Two kinds of setting, split by who changes them and how often.
|
||
|
||
**Where the server is plugged in** stays in the environment: the listen address,
|
||
the database DSN, the ntfy URL and token, the public URL. They are needed before
|
||
the database is open, and two of them are credentials.
|
||
|
||
**How the server behaves** lives in the database and is edited by an
|
||
administrator in the web UI or through `PUT /api/admin/settings`, taking effect
|
||
on the next sweep rather than at the next restart. The variables below marked
|
||
**seed** are the value each of those starts from: written once, on first start,
|
||
and never overwritten afterwards — a redeploy cannot put a chart's default back
|
||
over an administrator's edit.
|
||
|
||
| Variable | Default | Description |
|
||
|---|---|---|
|
||
| `TERDUT_ADDR` | `:8080` | TCP address to listen on |
|
||
| `TERDUT_DB_DSN` | — | **Required.** Postgres connection string, e.g. `postgres://terdut:secret@localhost:5432/terdut?sslmode=require` |
|
||
| `TERDUT_ARCHIVE_AFTER` | `168h` (7d) | **seed.** How long a resolved alert or incident stays in the default list before being auto-archived |
|
||
| `TERDUT_STALE_AFTER` | `6h` | **seed.** How long a firing alert may go without a refreshing webhook before it is treated as resolved — **must exceed your Alertmanager `repeat_interval`** |
|
||
| `TERDUT_DEADMAN_MATCHERS` | `alertname=Watchdog` | The **default** matchers a team starts with — switches are per team now, and this seeds teams that have no configuration of their own. `;` separates matchers, `,` the label conditions within one, `=` is exact equality. Every matcher must name an `alertname` |
|
||
| `TERDUT_DEADMAN_TIMEOUT` | `15m` | How long a heartbeat may go unheard before its switch is declared dead — **must be shorter than the `repeat_interval` of the route carrying it**. `0` disables dead man's switch handling |
|
||
| `TERDUT_DEADMAN_SEVERITY` | `critical` | Severity a dead man's switch incident opens at |
|
||
| `TERDUT_NTFY_URL` | — | ntfy server to publish push notifications to. Empty disables notifications entirely |
|
||
| `TERDUT_NTFY_TOKEN` | — | Bearer token for an access-controlled ntfy |
|
||
| `TERDUT_NTFY_FALLBACK_TOPIC` | — | Topic used when nobody is on call |
|
||
| `TERDUT_PUBLIC_URL` | — | Base URL a phone uses to reach this server: the notification's link into the web UI, its Acknowledge button, and whether the session cookie is `Secure` |
|
||
| `TERDUT_NOTIFY_REPEAT` | `15m` | **seed.** How long an incident may sit unacknowledged before it is paged again. `0` notifies once and never repeats |
|
||
| `TERDUT_PASSWORD_LOGIN` | `true` | `false` refuses password login and password sign-up (`403`), leaving single sign-on the only way in. Refused at startup unless SSO is configured |
|
||
| `TERDUT_OIDC_ISSUER` | — | Turns single sign-on on. The provider's issuer URL; discovery is read from `<issuer>/.well-known/openid-configuration`. See [Single sign-on](#single-sign-on-oidc) |
|
||
| `TERDUT_OIDC_CLIENT_ID` / `TERDUT_OIDC_CLIENT_SECRET` | — | **Required with an issuer.** The confidential client registered at the provider. Keep the secret in a Secret, not in values |
|
||
| `TERDUT_OIDC_NAME` | `SSO` | What the sign-in button calls the provider |
|
||
| `TERDUT_OIDC_SCOPES` | `openid profile email` | Scopes requested, comma or space separated. Authentik puts `groups` behind `profile` |
|
||
| `TERDUT_OIDC_USERNAME_CLAIM` / `_EMAIL_CLAIM` / `_GROUPS_CLAIM` | `preferred_username` / `email` / `groups` | ID token claims read for the username, email and groups |
|
||
| `TERDUT_OIDC_TRUST_EMAIL` | `false` | Link a first sign-in to an existing local user by email even if the provider does not mark the address verified |
|
||
| `TERDUT_OIDC_ALLOWED_GROUPS` | — | Comma-separated. Only people in one of these may sign in. Empty admits everybody the provider authenticates |
|
||
| `TERDUT_OIDC_ADMIN_GROUP` | — | Members are system administrators |
|
||
| `TERDUT_OIDC_SESSION_MAX_AGE` | `12h` | Hard ceiling on a session made by an SSO sign-in |
|
||
|
||
Durations use Go syntax (`30m`, `12h`, `168h`). An unparseable value falls back to the default.
|
||
|
||
Note that `TERDUT_STALE_AFTER` and `TERDUT_DEADMAN_TIMEOUT` point in opposite directions. Staleness
|
||
is a generous grace period around a `repeat_interval` you do not control; a dead man's switch is a
|
||
deadline you set deliberately, and the heartbeat's route is configured to beat faster than it.
|
||
|
||
In the Helm chart the two sweeper durations are set via `sweeper.staleAfter` and `sweeper.archiveAfter`, dead man's switches via the `deadman.*` values, notifications via the `notify.*` values, and single sign-on via `oidc.*` and `passwordLogin`.
|
||
|
||
---
|
||
|
||
## Alertmanager configuration
|
||
|
||
Alerts arrive on a team's **integration key**, which says both that the sender
|
||
may post and which team the alerts belong to. Mint one as an owner of the team:
|
||
|
||
```bash
|
||
curl -X POST https://terdut.example.com/api/teams/1/integrations \
|
||
-H "Authorization: Bearer $TERDUT_API_KEY" \
|
||
-H 'Content-Type: application/json' \
|
||
-d '{"name":"prod alertmanager"}'
|
||
```
|
||
|
||
The response carries the key and the full URL **once**; only a SHA-256 hash is
|
||
stored. Put it in your `alertmanager.yml`:
|
||
|
||
```yaml
|
||
receivers:
|
||
- name: terdut
|
||
webhook_configs:
|
||
- url: http://terdut-server:8080/api/integrations/<key>/alertmanager
|
||
send_resolved: true
|
||
|
||
route:
|
||
receiver: terdut
|
||
```
|
||
|
||
The whole URL is a credential, so treat it like one. Alertmanager 0.26 and
|
||
later can read it from a file with `url_file:` instead, which keeps it out of
|
||
your configuration repository:
|
||
|
||
```yaml
|
||
- url_file: /etc/alertmanager/secrets/terdut-webhook-url/url
|
||
send_resolved: true
|
||
```
|
||
|
||
The webhook endpoint requires no authentication.
|
||
|
||
If you use the [dead man's switch](#dead-mans-switch) — and the default configuration does — give
|
||
the heartbeat a route of its own, because the deadline is only as tight as the interval feeding it:
|
||
|
||
```yaml
|
||
route:
|
||
receiver: terdut
|
||
repeat_interval: 4h
|
||
routes:
|
||
- matchers: [ 'alertname = "Watchdog"' ]
|
||
receiver: terdut
|
||
group_wait: 0s
|
||
group_interval: 1m
|
||
repeat_interval: 1m
|
||
```
|
||
|
||
That delivers a heartbeat every **2 minutes**, not every minute. Alertmanager only reconsiders a
|
||
group every `group_interval`, and at exactly one elapsed interval `repeat_interval` has not *quite*
|
||
passed, so the send slips to the next tick — equal values give 2×. Two minutes against the 15 minute
|
||
default is seven heartbeats per window, which is the point; use `group_interval: 30s` if you want
|
||
the numbers to mean what they say.
|
||
|
||
kube-prometheus-stack users get the `Watchdog` alert (`expr: vector(1)`) for free; it just needs
|
||
routing to terdut rather than to `null`.
|
||
|
||
---
|
||
|
||
## Alerts and incidents
|
||
|
||
There are two objects, and the difference between them is the whole design.
|
||
|
||
**An alert is Alertmanager's record.** It has two states, `firing` and
|
||
`resolved`, one row per fingerprint, and no human ever writes to it. The API
|
||
exposes alerts read-only.
|
||
|
||
**An incident is the work item.** It goes `triggered → acknowledged → resolved`,
|
||
carries an assignee, a snooze, notes and a timeline, and is the only thing people
|
||
act on. Many alerts belong to one incident.
|
||
|
||
### Correlation uses Alertmanager's `groupKey`
|
||
|
||
Alertmanager has already grouped alerts according to the `group_by` routing tree
|
||
you configured, and it sends the resulting `groupKey` and `groupLabels` on every
|
||
webhook. Incidents adopt that answer rather than re-grouping alerts a second
|
||
time — if you want different correlation, change `group_by` in
|
||
`alertmanager.yml` and terdut follows.
|
||
|
||
At most one incident is open per `groupKey` at a time. Alerts firing in a group
|
||
that already has an open incident join it. The incident's `severity` is a
|
||
high-water mark — the highest `severity` label any of its alerts has carried — so
|
||
an incident that hit `critical` still reads as critical after the critical alert
|
||
clears.
|
||
|
||
### An incident opens only on a new occurrence
|
||
|
||
An incident opens when an alert **transitions into firing**: a fingerprint that
|
||
was never seen, an alert with a newer `startsAt`, or a resolved alert that
|
||
started again. The unchanged firing notifications Alertmanager re-sends every
|
||
`repeat_interval` are none of those, and open nothing.
|
||
|
||
This is what makes closing an incident by hand mean something. Without the rule,
|
||
`POST /api/incidents/{id}/resolve` would be undone by the next re-send of an
|
||
alert that never stopped firing.
|
||
|
||
### Leaving the open state
|
||
|
||
- **Automatically**, once every alert under the incident has stopped firing —
|
||
whether by a resolved webhook or by the sweeper's
|
||
[stale-alert expiry](#stale-alert-expiry). The incident gets
|
||
`"resolution_source": "alerts"`.
|
||
- **By hand**, via `POST /api/incidents/{id}/resolve`
|
||
(`"resolution_source": "manual"`). This is **terminal**: a later occurrence in
|
||
that group opens a *new* incident rather than reopening this one. If the alert
|
||
underneath never stops firing, the incident stays closed — that is what
|
||
resolving by hand asserts.
|
||
- **On recovery**, for a [dead man's switch](#dead-mans-switch) incident whose
|
||
heartbeat started arriving again (`"resolution_source": "recovered"`). These
|
||
incidents have no member alerts, so the automatic cascade above cannot reach
|
||
them.
|
||
|
||
To quieten an incident you expect to come back, snooze it instead
|
||
(`POST /api/incidents/{id}/snooze`). A snooze hides the incident from the default
|
||
list without closing it, and expires by simply falling into the past.
|
||
|
||
### On-call assignment
|
||
|
||
A new incident is assigned to whoever holds today's schedule entry at the moment
|
||
it opens (`GET /api/schedule/current`). If nobody is scheduled it opens
|
||
unassigned. Reassign with `POST /api/incidents/{id}/assign`.
|
||
|
||
One person holds a given day, so `POST /api/schedule` refuses a date somebody
|
||
already has: taking a shift off the person expecting to be paged for it should
|
||
not be something a plain call does by accident. Pass `"replace": true` to take
|
||
them anyway. Either way the whole request is one transaction — a week where some
|
||
days are free and some are taken moves as a unit, and a failure leaves the rota
|
||
exactly as it was rather than with a hole in it.
|
||
|
||
### Push notifications
|
||
|
||
With `TERDUT_NTFY_URL` set, an incident that opens is pushed to the on-call
|
||
person's phone through [ntfy](https://ntfy.sh). Everybody sets their own topic
|
||
under *Account* in the web UI, where a **Send a test push** button proves it
|
||
before an incident has to; `PUT /api/users/{id}/notify` is the same thing over
|
||
the API, and an administrator may set somebody else's. A user with no topic
|
||
falls back to `TERDUT_NTFY_FALLBACK_TOPIC`, as does an incident that opens with
|
||
nobody on call. If neither yields a topic, nothing is queued.
|
||
|
||
The **server** is the install's one ntfy, from `TERDUT_NTFY_URL`, and is not
|
||
something a user picks. Only the topic is per-person.
|
||
|
||
A topic is a shared secret with the ntfy server: anyone who knows it can both
|
||
read the pages and publish to it, so an unguessable one is worth the trouble.
|
||
That is also why the topic never appears in an incident's timeline, which every
|
||
API key can read.
|
||
|
||
Three things get pushed:
|
||
|
||
- **triggered** — an incident opened. Priority follows severity (`critical` maps
|
||
to ntfy's max priority, the one that overrides the phone's quiet settings).
|
||
- **reminder** — the incident is still `triggered` after `TERDUT_NOTIFY_REPEAT`.
|
||
Repeats until somebody acts. Acknowledging, snoozing, resolving or archiving
|
||
all stop it — snooze is the mute button.
|
||
- **resolved** — every alert under the incident stopped firing. Only sent to
|
||
whoever was paged in the first place, and only for the automatic cascade:
|
||
resolving by hand pushes nothing, since the person who did it already knows.
|
||
|
||
Notifications carry an **Acknowledge** button that acknowledges the incident
|
||
without opening anything. It POSTs to `/api/notify/ack/{token}`, an
|
||
unauthenticated route authorised by the 256-bit token in its path — minted fresh
|
||
per notification, scoped to one incident and one action, and valid for 24 hours.
|
||
A real API key is never put in a notification, because the message is stored on
|
||
the ntfy server and cached on the device.
|
||
|
||
The token is **not** consumed by use. Acknowledging is idempotent, so a token
|
||
stays valid for its full 24 hours and a second tap is a no-op that reports the
|
||
incident's current state rather than an error — which is what you want when a
|
||
tap is retried on a flaky mobile connection. What bounds it is scope, not a use
|
||
count: one incident, one action, one day. Expired tokens are purged by the
|
||
sweeper.
|
||
|
||
Two consequences worth planning for:
|
||
|
||
- `/api/notify/ack/{token}` **must stay publicly reachable**, or the button will
|
||
not work when the responder is off your network.
|
||
- Notifications sent to the fallback topic carry **no** Acknowledge button. The
|
||
topic is shared, and a button on it would let any subscriber acknowledge as
|
||
somebody else.
|
||
|
||
Delivery is a queue, not an inline call: the webhook writes a row and a
|
||
background notifier sends it within 30 seconds, retrying with exponential
|
||
backoff up to 8 attempts. Nothing about ingestion blocks on ntfy being reachable.
|
||
|
||
Every delivery is recorded on the incident's timeline: a `notified` event once
|
||
ntfy accepts the publish, and a `notify_failed` event when a notification
|
||
exhausts its retries. Written from the result rather than at enqueue, so the
|
||
timeline says what actually happened — and a page that never landed is visible
|
||
instead of looking the same as one that did.
|
||
|
||
### Escalation
|
||
|
||
Without a ladder, an unacknowledged incident re-pages the same topic every
|
||
`notify_repeat` forever. That is a louder version of the same silence: if the
|
||
person on call is asleep, out of signal, or has left, nothing else happens.
|
||
|
||
A team can configure an ordered ladder instead. Each level has a timeout and a
|
||
set of targets, and a target is either a named person or **whoever the team's
|
||
rota says is on call today** — the target that keeps working when the rota
|
||
changes and nobody remembers to edit the policy.
|
||
|
||
```
|
||
level 1 5m oncall the rota gets first refusal
|
||
level 2 5m user:bob then a named second
|
||
then repeat_count more rounds
|
||
then the team's fallback topic, once
|
||
```
|
||
|
||
When a level's timeout passes with the incident still `triggered`, the next
|
||
level is paged. Off the end of the ladder the whole thing runs again
|
||
`repeat_count` times, and after that the team's `fallback_topic` is paged once
|
||
as the end of the line. The incident stays open throughout: running out of
|
||
people to wake is not the same as somebody answering.
|
||
|
||
**Acknowledging or resolving stops it**, which is the point — continuing to wake
|
||
people after somebody has said "I have this" is how a tool teaches people to
|
||
mute it. **Snoozing pauses it**: a deliberate "not now" holds the ladder where
|
||
it is, and it resumes when the snooze runs out.
|
||
|
||
Every step is on the incident's timeline with the level and the names it woke,
|
||
so somebody reading it afterwards can tell why their phone rang at 04:00. A
|
||
level whose targets are all unreachable — no ntfy topic, a disabled account, an
|
||
empty rota — is recorded as `nobody reachable` and the ladder moves on rather
|
||
than stalling on a rung that cannot ring.
|
||
|
||
**Reminders and escalation never both run.** A team with a ladder gets
|
||
escalation; a team without keeps the reminder behaviour exactly as it was. Two
|
||
pages for one silence is the surest way to get a tool muted.
|
||
|
||
The ladder's `fallback_topic` is per team, unlike `TERDUT_NTFY_FALLBACK_TOPIC`,
|
||
which is the install-wide topic used when an incident opens with nobody on call.
|
||
They answer different questions: one is "nobody was scheduled", the other is
|
||
"everybody scheduled has been tried".
|
||
|
||
### Stale alert expiry
|
||
|
||
A resolved webhook is the only signal that an alert has stopped firing, so a
|
||
notification that is dropped, silenced, or lost to a restart would otherwise pin
|
||
that alert as firing forever. A background sweeper resolves firing alerts that
|
||
Alertmanager has stopped refreshing, using either signal:
|
||
|
||
- the `endsAt` watermark on the last notification has passed, or
|
||
- no webhook has refreshed the alert within `TERDUT_STALE_AFTER`.
|
||
|
||
Alertmanager re-sends firing notifications every `repeat_interval`, which is what
|
||
keeps a live alert fresh — so `TERDUT_STALE_AFTER` must be comfortably larger
|
||
than your `repeat_interval` (default 4h), or live alerts will be resolved
|
||
prematurely. Alerts resolved this way are marked `"resolution_source": "expiry"`
|
||
to distinguish them from a real Alertmanager resolve (`"alertmanager"`).
|
||
|
||
An expiry cascades: once it leaves an incident with nothing firing under it, the
|
||
incident resolves too, in the same sweep.
|
||
|
||
### Dead man's switch
|
||
|
||
Everything above assumes alerts arrive. If Prometheus stops evaluating, or
|
||
Alertmanager cannot reach this server, nothing arrives — and silence looks
|
||
exactly like everything being fine. A dead man's switch inverts the handling for
|
||
one designated alert so that silence is the signal:
|
||
|
||
- **receiving** it opens no incident, and
|
||
- the **absence** of it does.
|
||
|
||
kube-prometheus-stack already ships the alert for this. `Watchdog` is
|
||
`expr: vector(1)`, so it fires permanently and is re-sent forever; it is worth
|
||
nothing unless something downstream notices it stop. That is what
|
||
`TERDUT_DEADMAN_MATCHERS` defaults to.
|
||
|
||
**Switches belong to a team**, which decides which of its own alerts are
|
||
heartbeats and how long a silence has to last. Each **switch** is a row of its
|
||
own — a name, one matcher, a timeout and a severity — so switches in one team
|
||
can have different deadlines. An owner adds and removes them on **Team →
|
||
Switches**, which lists each with a status (**healthy**, **dead**, or
|
||
**dormant** until its first heartbeat), when it was last heard from, and when it
|
||
last opened an incident; a matcher that several clusters satisfy is broken down
|
||
per cluster. The API is `POST`/`DELETE /api/teams/{teamID}/deadman/switches`. A
|
||
missed heartbeat opens an incident in the team whose integration received it.
|
||
Removing a switch stops the watching; an incident it already opened stays open
|
||
until somebody resolves it.
|
||
|
||
The environment variables are the starting point, not the setting: the **first**
|
||
time the server starts, every team is given a switch per default matcher from
|
||
them, once. After that a team's switches are its own — an owner's edit or
|
||
deletion is never put back by a redeploy. A team created later starts watching
|
||
nothing until its owner says otherwise — inheriting an install-wide heartbeat
|
||
would page a new team about a source it has never heard of.
|
||
|
||
A matcher is a set of exact label conditions, one of which must be the
|
||
`alertname`, in the format the environment variable uses (one matcher per switch; the
|
||
variable takes several, separated by `;`):
|
||
|
||
```
|
||
alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat
|
||
```
|
||
|
||
**The unit of monitoring is the fingerprint, not the alert name.** Two clusters
|
||
sending the same `Watchdog` are two independent switches, so a healthy one can
|
||
never mask a dead one.
|
||
|
||
#### The lifecycle
|
||
|
||
A switch is **dormant** until its first heartbeat arrives. A configured matcher
|
||
that has never been heard from opens nothing, so a fresh deploy or a restored
|
||
database does not page. It also means a matcher that never matches anything is
|
||
silently inert — check the startup log line, which lists the matchers that
|
||
survived parsing.
|
||
|
||
Once armed, the sweeper declares it **dead** when either the heartbeat has not
|
||
been refreshed within `TERDUT_DEADMAN_TIMEOUT`, or Alertmanager explicitly
|
||
resolved it — the sender saying the heartbeat stopped needs no further waiting.
|
||
That opens an incident at `TERDUT_DEADMAN_SEVERITY`, assigned and paged like any
|
||
other, and marks the heartbeat alert `"resolution_source": "deadman"` so the
|
||
alert list stops claiming a dead switch is firing.
|
||
|
||
It **recovers** when the heartbeat starts arriving again: the incident resolves
|
||
with `"resolution_source": "recovered"` and the all-clear goes to whoever was
|
||
paged.
|
||
|
||
Resolving the incident by hand sticks, the same way it does for an alert-backed
|
||
one. While the switch stays silent nothing new opens — so a decommissioned
|
||
source is a one-time page rather than a nag. The switch **re-arms** on the next
|
||
heartbeat: come back and die again, and that is a new incident.
|
||
|
||
#### Two things to know
|
||
|
||
`TERDUT_DEADMAN_TIMEOUT` must be **shorter** than the `repeat_interval` of the
|
||
route carrying the heartbeat, which is the exact opposite of
|
||
`TERDUT_STALE_AFTER`. Inheriting a default `repeat_interval` of 4h gives you a
|
||
switch that takes four hours to notice anything, so give the heartbeat
|
||
[its own route](#alertmanager-configuration). Matched alerts are exempt from
|
||
stale-alert expiry — a heartbeat answers to its own timeout and nothing else.
|
||
|
||
A dead man's switch incident has **no member alerts**:
|
||
`GET /api/incidents/{id}/alerts` returns an empty list. There is no alert
|
||
describing the problem, because the problem is that no alert arrived. What
|
||
happened is on the timeline instead, as a `deadman_silent` event carrying the age
|
||
of the last heartbeat, and the heartbeat's labels are on the incident's
|
||
`group_labels`.
|
||
|
||
---
|
||
|
||
## API reference
|
||
|
||
### Authentication
|
||
|
||
All endpoints except `/api/bootstrap`, `/api/integrations/{key}/alertmanager`,
|
||
`/api/notify/ack/{token}`, `/api/login`, `/api/logout`, `/api/auth/config`,
|
||
`/api/oidc/login`, `/api/oidc/callback`, `/api/oidc/device` and `/api/oidc/device/token`
|
||
require either an API key:
|
||
|
||
```
|
||
Authorization: Bearer <api-key>
|
||
```
|
||
|
||
or the web UI's session cookie. A request that carries an `Authorization` header
|
||
is judged on that header alone.
|
||
|
||
Two kinds of user exist. An **administrator** manages accounts: creating and
|
||
deleting users, setting anybody's password, minting keys for anybody, and
|
||
granting the flag itself. Everybody else works incidents — acknowledging,
|
||
assigning, snoozing, resolving, noting — and manages their own account and
|
||
nobody else's. An API key carries exactly the rights of the user it belongs to.
|
||
|
||
**Getting an account.** The first one comes from `/api/bootstrap`. After that
|
||
it depends on `signup_mode`, an administrator setting:
|
||
|
||
- `invite_only` (the default) — a team owner mints a link with
|
||
`POST /api/teams/{teamID}/invites`, and the person who opens it picks a
|
||
username and password and lands in that team with the role the link carries.
|
||
Links are single-use unless told otherwise, expire after seven days, and can
|
||
be revoked before that.
|
||
- `open` — anybody who can reach the server can create an account, and must
|
||
name a team, which they then own.
|
||
|
||
Invites are **links, not email**: this server has no SMTP, and adding it to send
|
||
one message would be a subsystem to run, secure and monitor. Send the link
|
||
however you already talk to the person.
|
||
|
||
A domain-restricted third mode was considered and dropped: with no email there
|
||
is nothing to verify an address against, so it would only check the domain of a
|
||
string somebody typed.
|
||
|
||
The first user, from `/api/bootstrap`, is an administrator. Users created
|
||
afterwards are not, until an administrator says so. An install always keeps at
|
||
least one: the last administrator can be neither deleted nor demoted, and
|
||
nobody can delete or demote themselves.
|
||
|
||
Endpoints that require the flag answer `403` with
|
||
`{"error":"administrator access required"}`.
|
||
|
||
**Teams** are the unit of tenancy, and are a separate axis from the administrator
|
||
flag. A team owns its incidents, alerts, schedule and integrations, and a user
|
||
sees exactly the teams they belong to. Within a team an **owner** configures it
|
||
(schedule, integrations, membership) and a **member** works its incidents.
|
||
|
||
An administrator crosses that line in one direction only. They **configure any
|
||
team** without being in it — every owner-only endpoint accepts the flag, because
|
||
otherwise a team whose last owner left could never be repaired. They do **not
|
||
read any team**: the queue, the alerts and the incidents are filtered by real
|
||
membership, so an administrator sees a team's work only by joining it, which is
|
||
a membership change and shows up as one. Administration is about accounts and
|
||
the shape of a team, not about reading other people's incidents.
|
||
|
||
Anything belonging to a team you are not in answers `404`, not `403`: whether an
|
||
incident exists is itself something only its team should learn.
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `GET` | `/api/auth/config` | How to sign in: `{"password_login", "oidc": {"enabled","name"}, "device_login"}`. No session needed |
|
||
| `POST` | `/api/login` | `{"username","password"}` → sets the session cookie, returns `{user, has_password}`. `429` after too many failures; `403` when `TERDUT_PASSWORD_LOGIN=false` |
|
||
| `GET` | `/api/oidc/login` | Starts a single sign-on sign-in: redirects the browser to the provider. `?next=/path` is where to land afterwards; only a path on this server is honoured. Only exists when SSO is configured |
|
||
| `POST` | `/api/oidc/device` | Starts a device login: returns `{device_code, user_code, verification_url, interval, expires_in}`. Only exists when SSO is configured |
|
||
| `POST` | `/api/oidc/device/token` | `{"device_code"}` → `202 {"status":"pending"}`, then `200` with the session cookie once approved (once only). `410` with `{"error":"expired"}` or `{"error":"denied"}`; `429 {"error":"slow_down"}` if polled faster than `interval` |
|
||
| `POST` | `/api/oidc/device/approve` | **session** — `{"user_code"}`. Approves a pending device login as the caller. `403` for an API key; `404` for an unknown, expired or already decided code |
|
||
| `POST` | `/api/oidc/device/deny` | **session** — `{"user_code"}`. Refuses it |
|
||
| `GET` | `/api/oidc/callback` | Where the provider sends the browser back. Sets the session cookie and redirects to `/`, or to `/?sso_error=<code>` — one of `denied`, `expired`, `failed`, `unavailable`, `not_allowed`, `no_email`, `email_conflict`, `disabled` |
|
||
| `POST` | `/api/logout` | Ends the session and clears the cookie |
|
||
| `GET` | `/api/me` | The caller: `{user, has_password}` |
|
||
|
||
### Users
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
**admin** marks an endpoint that requires the administrator flag; **self or
|
||
admin** marks one you may use on your own account and an administrator may use
|
||
on anybody's.
|
||
|
||
| Method | Path | Who | Description |
|
||
|---|---|---|---|
|
||
| `GET` | `/api/signup` | — | Whether sign-up is open, and whether `?invite=` is usable. No session needed: the caller has no account yet |
|
||
| `POST` | `/api/signup` | — | Create an account `{"username","email","password","invite"?,"team_name"?}` and sign in. `403` without a usable invite when the mode is invite-only |
|
||
| `POST` | `/api/bootstrap` | — | Create first user + API key `{"username","email","password"?}` (only works on empty DB). The user is an administrator |
|
||
| `GET` | `/api/users` | any | List users. Open to everybody: the queue's assignment control and the schedule both have to name people |
|
||
| `GET` | `/api/users/{id}/teams` | self or admin | The teams that user is in, each with their role. `/api/teams` is always about the caller; this one answers it about somebody else, for the admin page's per-user view. `404` for a user who does not exist, so "no teams" and "no such person" are distinguishable |
|
||
| `POST` | `/api/users` | **admin** | Create user `{"username","email"}`. Not an administrator |
|
||
| `DELETE` | `/api/users/{id}` | **admin** | Delete user (cascades to keys). `409` for yourself or the last administrator |
|
||
| `PUT` | `/api/users/{id}/admin` | **admin** | Grant or revoke the administrator flag `{"is_admin"}`. `409` for yourself, the last administrator, or an administrator granted by single sign-on |
|
||
| `PUT` | `/api/users/{id}/disabled` | **admin** | Take an account out of use, or put it back `{"disabled"}`. `409` for yourself or the last administrator |
|
||
| `PUT` | `/api/users/{id}/notify` | self or admin | Set push notification target `{"ntfy_topic"}` — empty string clears it |
|
||
| `PUT` | `/api/users/{id}/password` | self or admin | Set web UI password `{"password","current_password"}`. `current_password` is required only when changing your own existing password. Ends the user's other sessions |
|
||
| `POST` | `/api/users/{id}/api-keys` | self or admin | Issue API key `{"name"}` — key shown once |
|
||
| `DELETE` | `/api/users/{id}/api-keys/{keyID}` | self or admin | Revoke API key |
|
||
|
||
### Administration
|
||
|
||
| Method | Path | Who | Description |
|
||
|---|---|---|---|
|
||
| `GET` | `/api/admin/teams` | **admin** | Every team on the server, with its member and open-incident counts. `/api/teams` answers "what am I in"; this answers "what is there" |
|
||
| `GET` | `/api/admin/teams/{teamID}` | **admin** | One team and who is in it: `{"team", "members"}`. `404` for a team that does not exist. `GET /api/teams/{teamID}/members` is **member**-only and still `404`s an administrator from outside the team — reading a team's shape and reading its work are different questions, so they are different endpoints |
|
||
| `GET` | `/api/admin/settings` | **admin** | The editable settings with their bounds, plus the environment-configured ones, read-only. Never credentials |
|
||
| `PUT` | `/api/admin/settings` | **admin** | Change one or more `{"key": seconds}`, or `{"signup_mode": "open"\|"invite_only"}`. `400` for an unknown key or a value outside its bounds |
|
||
|
||
### Alert ingestion
|
||
|
||
Alerts arrive on a team's integration key. The key is both the credential and the
|
||
routing: it says that the sender may post, and which team the alerts belong to.
|
||
Create one with `POST /api/teams/{teamID}/integrations`, which returns the key
|
||
and the full URL once and stores only a SHA-256 hash.
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `POST` | `/api/integrations/{key}/alertmanager` | Alertmanager v4 webhook receiver for the key's team. `401` for an unknown key |
|
||
|
||
This is the only way in. The pre-teams `POST /api/alertmanager/webhook` took no
|
||
credential at all — anything able to reach the port could open an incident —
|
||
and was removed in v0.13.0 once senders had moved onto keys.
|
||
|
||
### Teams
|
||
|
||
**owner** below means an owner of that team *or* a system administrator, who
|
||
passes every one of these without being a member — see
|
||
[Authentication](#authentication). **member** means membership and nothing else: an
|
||
administrator who is not in the team gets the same `404` as anybody else.
|
||
|
||
| Method | Path | Who | Description |
|
||
|---|---|---|---|
|
||
| `GET` | `/api/teams` | any | The caller's own teams, each with their role |
|
||
| `POST` | `/api/teams` | any | Create a team `{"name"}`; the creator becomes its first owner |
|
||
| `PUT` | `/api/teams/{teamID}` | **owner** | Rename it `{"name"}`. `409` if the name is taken |
|
||
| `DELETE` | `/api/teams/{teamID}` | **owner** | Delete a team and everything under it. `409` while it has open incidents |
|
||
| `GET` | `/api/teams/{teamID}/members` | member | Who is in the team, with `status` (`oncall` if the rota has them today, `unpageable` when a page to them would go nowhere — even if they are on call — else `reachable`), `on_call`, `next_shift` (first rota day after today), `pageable` and `problem` (`has no ntfy topic` / `account is disabled`; never the topic itself) and `last_active_at` (their newest session or API-key use). Every member sees the same list |
|
||
| `POST` | `/api/teams/{teamID}/members` | **owner** | Add a member, or change their role `{"user_id","role"}`. `409` when it would demote the last owner, or the membership is managed by single sign-on |
|
||
| `DELETE` | `/api/teams/{teamID}/members/{userID}` | **owner** | Remove a member. `409` for the last owner, or a membership managed by single sign-on |
|
||
| `GET` | `/api/teams/{teamID}/oidc-groups` | member | Which groups control this team's membership: `{"member_group","owner_group"}`. An empty string means no group grants that role here |
|
||
| `PUT` | `/api/teams/{teamID}/oidc-groups` | **owner** | Set them. An empty string clears a binding |
|
||
| `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys. Each carries `status` (`active` if its key posted within 24h, `quiet` if it has but not lately, `never`), `last_used_at` (last webhook, usable or not), `last_alert_at` (when an alert last arrived on it) and `alerts_24h` (distinct alerts it refreshed in the last day). Alerts delivered before the source was recorded (migration 010) have none, so the last two fill in as Alertmanager re-sends them |
|
||
| `PATCH` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Rename `{"name"}`. The key does not change |
|
||
| `POST` | `/api/teams/{teamID}/integrations` | **owner** | Mint an integration `{"name","kind"}` — key and URL shown once |
|
||
| `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration. Alerts it delivered stay, unattributed |
|
||
| `GET` | `/api/teams/{teamID}/invites` | **owner** | The team's invite links, with their uses and expiry. Never the tokens |
|
||
| `POST` | `/api/teams/{teamID}/invites` | **owner** | Mint one `{"role","max_uses"}` — the full URL is returned once |
|
||
| `DELETE` | `/api/teams/{teamID}/invites/{inviteID}` | **owner** | Revoke a link before it expires |
|
||
| `GET` | `/api/teams/{teamID}/escalation` | member | The team's [escalation ladder](#escalation) `{repeat_count, fallback_topic, levels[], last_escalated_at?, last_escalated_incident_id?}`. Empty levels means the team has none. Each level also carries `status` (`ready`, `escalating` when an unanswered incident has climbed to it, `unreachable` when nobody on it could be woken), `waiting` (ids of the open incidents on it) and, per target, `username` (who it means today — the person on call, for a rota target), `reachable` and `problem`. The extra fields are output only; `PUT` takes the plain shape |
|
||
| `PUT` | `/api/teams/{teamID}/escalation` | **owner** | Replace it wholesale. `400` for a level with no targets or no timeout — a rung that pages nobody is a silence with a number on it |
|
||
| `GET` | `/api/teams/{teamID}/deadman/switches` | member | The team's [dead man's switches](#dead-mans-switch), each `{id, name, matcher, timeout_seconds, severity, status, last_heartbeat_at, last_triggered_at, open_incident_id, sources[]}`. `status` is `healthy`, `dead` or `dormant`; `sources` has one entry per heartbeat fingerprint. Empty when the team watches nothing |
|
||
| `POST` | `/api/teams/{teamID}/deadman/switches` | **owner** | Add one: `{name?, matcher, timeout_seconds, severity?}`. `400` when the matcher names no `alertname` or holds several, or the timeout is not positive — a switch that silently watches nothing is the failure this feature exists to prevent |
|
||
| `DELETE` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Stop watching. An incident it opened stays open. `404` for a switch of another team |
|
||
|
||
### Notifications
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `POST` | `/api/notify/ack/{token}` | Acknowledge an incident from a push notification's Acknowledge button. No auth: the token in the path is the credential — one incident, one action, 24 hours, idempotent. Must stay publicly reachable |
|
||
|
||
### Incidents
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `GET` | `/api/incidents` | List incidents. Filters: `?status=triggered\|acknowledged\|resolved`, `?severity=`, `?assigned_to=<user id>`, `?archived=true`, `?snoozed=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?sort=severity`, `?limit=` (default 50, max 500) |
|
||
| `GET` | `/api/incidents/{id}` | Get single incident, with its alerts inline |
|
||
| `GET` | `/api/incidents/{id}/alerts` | Alerts under this incident |
|
||
| `GET` | `/api/incidents/{id}/timeline` | Full event history, chronological |
|
||
| `POST` | `/api/incidents/{id}/acknowledge` | Acknowledge (stamps authed user + time) |
|
||
| `DELETE` | `/api/incidents/{id}/acknowledge` | Clear acknowledgement, back to `triggered` |
|
||
| `POST` | `/api/incidents/{id}/resolve` | Close by hand — **terminal**, see above |
|
||
| `POST` | `/api/incidents/{id}/assign` | Reassign `{"user_id"}` |
|
||
| `POST` | `/api/incidents/{id}/snooze` | Hide until `{"until": RFC3339}` or `{"duration": "2h"}` |
|
||
| `DELETE` | `/api/incidents/{id}/snooze` | Un-snooze |
|
||
| `POST` | `/api/incidents/{id}/archive` | Archive (hides from the default list) |
|
||
| `DELETE` | `/api/incidents/{id}/archive` | Un-archive |
|
||
| `POST` | `/api/incidents/{id}/notes` | Add a note `{"content"}` |
|
||
| `DELETE` | `/api/incidents/{id}/notes/{eventID}` | Delete own note |
|
||
|
||
With no `?status=` filter, `GET /api/incidents` returns **open** incidents only —
|
||
the queue an on-call person wants. Currently snoozed and archived incidents are
|
||
excluded unless asked for. Actions that only make sense on an open incident
|
||
return `409` once it is resolved.
|
||
|
||
Notes are ordinary timeline events of type `note`; only they are deletable, and
|
||
only by their author. The rest of the timeline is a record of what happened.
|
||
|
||
#### The incident object
|
||
|
||
| Field | Type | Notes |
|
||
|---|---|---|
|
||
| `id` | integer | Server-assigned |
|
||
| `group_key` | string | Alertmanager's `groupKey` — opaque, treat as an identifier |
|
||
| `title` | string | Rendered from `groupLabels` |
|
||
| `group_labels` | object | String→string, as sent by Alertmanager |
|
||
| `status` | string | `"triggered"`, `"acknowledged"` or `"resolved"` |
|
||
| `severity` | string | *optional* — high-water mark across the incident's alerts; never lowered |
|
||
| `triggered_at` | timestamp | When the incident opened |
|
||
| `acknowledged_by_id` / `acknowledged_by` / `acknowledged_at` | | *optional* — user id, username, time |
|
||
| `assigned_to_id` / `assigned_to` | | *optional* — user id, username |
|
||
| `snoozed_until` | timestamp | *optional* — a value in the past reads as not snoozed |
|
||
| `resolved_at` | timestamp | *optional* |
|
||
| `resolution_source` | string | *optional* — `"alerts"`, `"manual"` or `"recovered"` |
|
||
| `archived_at` | timestamp | *optional* |
|
||
| `alerts` | array | Only on `GET /api/incidents/{id}` |
|
||
|
||
Treat `resolution_source` as an open set, as with the alert field of the same
|
||
name: degrade unknown values to "resolved, reason unknown".
|
||
|
||
#### The timeline event object
|
||
|
||
| Field | Type | Notes |
|
||
|---|---|---|
|
||
| `id` | integer | |
|
||
| `incident_id` | integer | |
|
||
| `type` | string | See below — treat as an open set |
|
||
| `user_id` / `username` | | *optional* — absent when the server acted rather than a person |
|
||
| `alert_id` | integer | *optional* — the alert an `alert_added` / `alert_resolved` event refers to |
|
||
| `detail` | string | *optional* — the note body, the snooze deadline, etc. |
|
||
| `created_at` | timestamp | |
|
||
|
||
Types written today: `triggered`, `alert_added`, `alert_resolved`,
|
||
`acknowledged`, `unacknowledged`, `assigned`, `snoozed`, `unsnoozed`, `resolved`,
|
||
`note`, `notified`, `notify_failed`, `deadman_silent`. On an `assigned` event
|
||
`user_id` is the **assignee**, not the actor. New types may be added; render
|
||
unknown ones generically rather than dropping them.
|
||
|
||
On `notified` and `notify_failed`, `detail` carries the notification kind
|
||
(`triggered` | `reminder` | `resolved`), and on a failure the reason after it.
|
||
`user_id` is who was paged — absent means the page went to the shared fallback
|
||
topic and so belongs to nobody. The topic itself is never written to the
|
||
timeline: it is a shared secret with the ntfy server, and every API key can read
|
||
this.
|
||
|
||
### Alerts
|
||
|
||
Alerts are read-only. Everything a person does happens on the incident.
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `GET` | `/api/alerts` | List alerts. Filters: `?status=firing\|resolved`, `?name=`, `?incident_id=`, `?archived=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?limit=` (default 50, max 500) |
|
||
| `GET` | `/api/alerts/{id}` | Get single alert |
|
||
|
||
Archived alerts are hidden from `GET /api/alerts` unless `?archived=true` is
|
||
passed; alert archiving is automatic housekeeping by the sweeper, not a user
|
||
action. Resolved alerts carry `resolution_source`: `"alertmanager"` for a real
|
||
resolved webhook, `"expiry"` when the sweeper inferred it (see
|
||
[Stale alert expiry](#stale-alert-expiry)), `"deadman"` for a heartbeat declared
|
||
dead (see [Dead man's switch](#dead-mans-switch)).
|
||
|
||
#### The alert object
|
||
|
||
Returned by `GET /api/alerts` (as an array) and `GET /api/alerts/{id}`.
|
||
Timestamps are RFC 3339 in UTC. Fields marked *optional* are omitted entirely
|
||
when unset, so clients must treat them as nullable.
|
||
|
||
| Field | Type | Notes |
|
||
|---|---|---|
|
||
| `id` | integer | Server-assigned; stable for the life of the row |
|
||
| `fingerprint` | string | Alertmanager's fingerprint — the upsert key |
|
||
| `name` | string | From the `alertname` label |
|
||
| `status` | string | `"firing"` or `"resolved"` |
|
||
| `labels` | object | String→string, as sent by Alertmanager |
|
||
| `annotations` | object | String→string, as sent by Alertmanager |
|
||
| `starts_at` | timestamp | When the alert instance began, **per Prometheus** |
|
||
| `ends_at` | timestamp | *optional* — absent while no end is known |
|
||
| `generator_url` | string | Link back to the originating Prometheus |
|
||
| `received_at` | timestamp | When the server last accepted a webhook for this alert — see below |
|
||
| `incident_id` | integer | *optional* — the most recent incident this alert belongs to |
|
||
| `resolution_source` | string | *optional* — `"alertmanager"`, `"expiry"` or `"deadman"` |
|
||
| `archived_at` | timestamp | *optional* — set while archived |
|
||
|
||
##### `received_at` is a liveness heartbeat
|
||
|
||
`starts_at` comes from Prometheus and **never changes** for the lifetime of an
|
||
alert instance. It says when the problem began, not whether it is still
|
||
happening — an alert that started twelve days ago looks identical whether
|
||
Alertmanager refreshed it a minute ago or went silent a week ago.
|
||
|
||
`received_at` is the field that answers "is this still live". It is set to the
|
||
server's clock on **every accepted webhook** for that fingerprint, including the
|
||
unchanged firing notifications Alertmanager re-sends every `repeat_interval`.
|
||
Clients may rely on this:
|
||
|
||
- **A firing alert whose `received_at` is advancing is still being refreshed.**
|
||
Stale-dating it against `repeat_interval` is a valid liveness check, and it is
|
||
what the built-in sweeper does (see
|
||
[Stale alert expiry](#stale-alert-expiry)).
|
||
- **`received_at` tracks accepted payloads, not delivery attempts.** A retry
|
||
that describes an older instance than the stored one is discarded, and a
|
||
discarded payload does not move `received_at`.
|
||
- **It stops advancing once the alert resolves,** because Alertmanager stops
|
||
re-sending. On an alert resolved by the sweeper
|
||
(`"resolution_source": "expiry"`) it therefore marks the last time
|
||
Alertmanager was actually heard from, which is earlier than `ends_at`.
|
||
|
||
`GET /api/alerts` is ordered by `received_at` descending — most recently
|
||
refreshed first — and the `?from=` / `?to=` filters on both the alert and stats
|
||
endpoints select on `received_at`, not `starts_at`.
|
||
|
||
##### `resolution_source` says how much to trust `ends_at`
|
||
|
||
An alert can leave the firing state two ways, and `resolution_source` records
|
||
which happened. Clients may rely on this:
|
||
|
||
- **Absent while firing.** It is set only on resolve, and a re-fire under the
|
||
same fingerprint clears it again, so its presence always agrees with
|
||
`"status": "resolved"`.
|
||
- **`"alertmanager"` — a real resolved webhook arrived.** `ends_at` is the end
|
||
time Alertmanager reported. It is an observed value and can be displayed as
|
||
fact.
|
||
- **`"expiry"` — the sweeper inferred the resolve** because Alertmanager stopped
|
||
refreshing the alert (see [Stale alert expiry](#stale-alert-expiry)). Nothing
|
||
ever reported an end, so **`ends_at` is approximate**: it is either the stale
|
||
`endsAt` watermark from the last notification, or — when that notification
|
||
carried none — the time the sweep ran, which lags the last real contact by up
|
||
to `TERDUT_STALE_AFTER` plus a sweep interval. Treat it as "no later than",
|
||
not as when the problem stopped.
|
||
|
||
On these alerts `received_at` is the more truthful signal: it marks the last
|
||
time Alertmanager was actually heard from. Surfacing the distinction is
|
||
worthwhile, since `"expiry"` can also mean the alert is still firing and the
|
||
notification path broke.
|
||
|
||
- **`"deadman"` — a heartbeat was declared dead** (see
|
||
[Dead man's switch](#dead-mans-switch)). Like `"expiry"`, an inference from
|
||
silence rather than an observed end, so `ends_at` is approximate — but a much
|
||
tighter one, bounded by `TERDUT_DEADMAN_TIMEOUT`. It is also the one resolution
|
||
a re-fire under the same `starts_at` can undo, since the switch coming back is
|
||
exactly the evidence that the inference was wrong.
|
||
|
||
Treat the value as an open set and tolerate ones you do not recognise — new
|
||
sources may be added, and unknown values should degrade to "resolved, reason
|
||
unknown" rather than being rejected.
|
||
|
||
### On-call schedule
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
Each team keeps its own rota, so two teams can have two different people on call
|
||
on the same day. The person taking a shift has to be in the team — paging
|
||
somebody who cannot open the incident is worse than paging nobody.
|
||
|
||
| Method | Path | Who | Description |
|
||
|---|---|---|---|
|
||
| `POST` | `/api/teams/{teamID}/schedule` | **owner** | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...], "replace"}` — all-or-nothing |
|
||
| `GET` | `/api/teams/{teamID}/schedule` | member | List entries. Filters: `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD` |
|
||
| `DELETE` | `/api/teams/{teamID}/schedule/{id}` | **owner** | Remove schedule entry |
|
||
| `GET` | `/api/schedule/current` | any | Who is on call today (UTC) in **every** team the caller is in — one entry per team, `[]` when nobody anywhere |
|
||
|
||
### Statistics
|
||
|
||
Every figure counts the caller's own teams only: a report that counted other
|
||
teams' incidents would leak their volume, and their alert names through the
|
||
top-alerts list, and would not be a number about the reader's work anyway.
|
||
|
||
All stat endpoints accept optional `?from=YYYY-MM-DD` and `?to=YYYY-MM-DD`, and exclude archived rows to match the default list views. Alert stats filter on `received_at`; incident stats filter on `triggered_at`.
|
||
|
||
| Method | Path | Description |
|
||
|---|---|---|
|
||
| `GET` | `/api/stats/incidents` | `{total, triggered, acknowledged, resolved, mtta_seconds, mttr_seconds}` |
|
||
| `GET` | `/api/stats/alerts` | `{total, firing, resolved}` counts |
|
||
| `GET` | `/api/stats/alerts/top` | Most frequent alert names. `?limit=` (default 10, max 100) |
|
||
| `GET` | `/api/stats/alerts/by-hour` | Count per hour-of-day (UTC), all 24 slots returned |
|
||
| `GET` | `/api/stats/alerts/by-day` | Count per day-of-week, all 7 slots with names returned |
|
||
|
||
`mtta_seconds` (time to acknowledge) and `mttr_seconds` (time to resolve) are
|
||
averages over incidents that have actually been acknowledged or resolved, and are
|
||
**null** until there are any — null means "no data", not zero.
|
||
|
||
---
|
||
|
||
## Upgrading to teams
|
||
|
||
Everything that existed before teams moves into one team called **Default**, and
|
||
every existing user becomes an owner of it. The upgrade is a no-op for the
|
||
people using it: the same queue, the same schedule, the same incidents, with a
|
||
name on them.
|
||
|
||
What changes, and will need attention:
|
||
|
||
- **Alert ingestion moved.** Mint a key with
|
||
`POST /api/teams/{teamID}/integrations` and point Alertmanager at the URL it
|
||
returns. In v0.12.0 the old `POST /api/alertmanager/webhook` still worked,
|
||
deprecated, routing everything to the oldest team; **v0.13.0 removes it**, so
|
||
upgrade straight from v0.11.x to v0.13.0 only after the senders are moved.
|
||
- **The schedule endpoints moved** under `/api/teams/{teamID}/schedule`, and
|
||
editing the rota is now an owner's job. `GET /api/schedule/current` stayed
|
||
where it was but now returns an **array** — one entry per team with somebody
|
||
on call — instead of a single object or a 404. This is a breaking API change
|
||
for anything that reads it, terdut-tui included.
|
||
- **Uniqueness is per team now.** Two teams can legitimately see the same alert
|
||
fingerprint, the same Alertmanager groupKey, and put somebody on call on the
|
||
same date.
|
||
|
||
**Dead man's switches moved too.** `TERDUT_DEADMAN_MATCHERS`, `_TIMEOUT` and
|
||
`_SEVERITY` are no longer the setting; they are the default each existing team
|
||
is seeded with at startup, after which an owner manages them per team through
|
||
`/api/teams/{teamID}/deadman/switches` and a redeploy never overwrites that.
|
||
|
||
Nothing else about an incident changes, and incidents never move between teams:
|
||
an alert belongs to whichever team's key it arrived on.
|
||
|
||
## Upgrading to roles
|
||
|
||
Before this release every authenticated caller could create and delete users,
|
||
set anybody's password and mint anybody's API keys. That is now the
|
||
administrator flag, and the migration **makes every existing user an
|
||
administrator** — they already held those powers, so nobody's access changes on
|
||
upgrade and demotion is a deliberate act afterwards. Promoting only the first
|
||
user would have silently stripped the rest, and could leave an install whose
|
||
only administrator is an account nobody has a password for.
|
||
|
||
Users created after the upgrade are not administrators. Hand the flag out with:
|
||
|
||
```bash
|
||
curl -X PUT https://terdut.example.com/api/users/7/admin \
|
||
-H "Authorization: Bearer $TERDUT_API_KEY" \
|
||
-H 'Content-Type: application/json' \
|
||
-d '{"is_admin": true}'
|
||
```
|
||
|
||
Nothing in the API changed shape, so terdut-tui needs no new version — but a
|
||
non-administrator now gets `403` where a `200` used to come back.
|
||
|
||
## Upgrading from SQLite
|
||
|
||
Versions up to v0.10.2 stored everything in a SQLite file. From v0.11.1 the server needs
|
||
`TERDUT_DB_DSN` and keeps nothing on disk.
|
||
|
||
The copy was done by `scripts/sqlite-to-postgres.go`, which **was deleted in v0.13.0** along
|
||
with the SQLite driver it was the last user of. It is still in the history — check out the
|
||
`v0.12.0` tag to get it:
|
||
|
||
```bash
|
||
git show v0.12.0:scripts/sqlite-to-postgres.go > sqlite-to-postgres.go
|
||
```
|
||
|
||
The cutover is ordered, and the server must not be running while the copy happens: stop the
|
||
old version, let the new binary build the schema against an empty Postgres, run the script
|
||
with `-sqlite` and `-dsn`, then start the new version for good. On Kubernetes step three runs
|
||
as a Job with the same image against the PVC before it is removed.
|
||
|
||
The copy preserves every id, so incidents keep their numbers and the timeline, alert
|
||
membership, outbox and ack tokens all still point where they did. It refuses a target that
|
||
already has rows, so a second run cannot double-insert.
|
||
|
||
## Upgrading to incidents
|
||
|
||
The incidents release moves the workflow off alerts, which is a **breaking API
|
||
change**. These endpoints are gone:
|
||
|
||
| Removed | Replacement |
|
||
|---|---|
|
||
| `POST`/`DELETE` `/api/alerts/{id}/acknowledge` | `POST`/`DELETE` `/api/incidents/{id}/acknowledge` |
|
||
| `POST`/`DELETE` `/api/alerts/{id}/archive` | `POST`/`DELETE` `/api/incidents/{id}/archive` (alert archiving is now sweeper-only) |
|
||
| `GET`/`POST` `/api/alerts/{id}/comments` | `GET /api/incidents/{id}/timeline`, `POST /api/incidents/{id}/notes` |
|
||
| `DELETE /api/alerts/{id}/comments/{commentID}` | `DELETE /api/incidents/{id}/notes/{eventID}` |
|
||
|
||
The alert object also drops `acknowledged_by_id`, `acknowledged_by` and
|
||
`acknowledged_at`, and gains `incident_id`.
|
||
|
||
Migration `008_incidents.sql` runs automatically on start and preserves existing
|
||
data: every alert gets a backfilled incident carrying its acknowledgement, and
|
||
comments become timeline notes. Backfilled incidents have a `group_key` of
|
||
`backfill:<fingerprint>` — there is no historical `groupKey` to correlate on, so
|
||
they are one-per-alert rather than grouped.
|
||
|
||
Nothing about the two documented alert contracts changes: `received_at` is still
|
||
advanced on every accepted webhook, and `resolution_source` still means what it
|
||
did.
|
||
|
||
## Upgrading to dead man's switches
|
||
|
||
Dead man's switch handling is **on by default**, watching `alertname=Watchdog`
|
||
with a 15 minute timeout. If you already route `Watchdog` to this server, the
|
||
behaviour of that alert changes on upgrade, in both directions:
|
||
|
||
- it stops opening incidents when it arrives, and
|
||
- it starts opening one when it stops arriving.
|
||
|
||
**Check your `repeat_interval` before upgrading.** The switch pages whenever a
|
||
heartbeat has not been refreshed within `TERDUT_DEADMAN_TIMEOUT`, so a `Watchdog`
|
||
route inheriting a 4h or 12h `repeat_interval` will page constantly against the
|
||
15 minute default. Either give the heartbeat
|
||
[its own fast route](#alertmanager-configuration) — the point of the feature — or
|
||
set `TERDUT_DEADMAN_TIMEOUT` above your current `repeat_interval` until you have.
|
||
`TERDUT_DEADMAN_TIMEOUT=0` turns the whole thing off.
|
||
|
||
There is no migration and no schema change. An existing open incident from a
|
||
`Watchdog` that arrived under the old behaviour is unaffected; resolve it by hand.
|
||
|
||
---
|
||
|
||
## Development
|
||
|
||
```bash
|
||
make test-db # start a local Postgres for the tests (podman or docker)
|
||
make test # run all tests
|
||
go build ./... # compile all packages
|
||
go run ./cmd/terdut # run locally (needs TERDUT_DB_DSN)
|
||
```
|
||
|
||
The tests need a real Postgres, because the server does — there is no in-memory Postgres the
|
||
way there was an in-memory SQLite. `TERDUT_TEST_DSN` says where it is, `make test-db` starts
|
||
one on port 5433 and prints the DSN, and `make test-db-stop` removes it. Each test gets its
|
||
own schema on that server, so tests cannot see each other's rows. An unset `TERDUT_TEST_DSN`
|
||
fails the suite rather than skipping it: a run that quietly tests nothing is worse than one
|
||
that does not run.
|
||
|
||
`make fmt lint test helm-lint` is the gate. It mirrors `.gitea/workflows/ci.yaml` step for
|
||
step, so a green run here means a green pipeline — with one deliberate exception: `make test`
|
||
adds `-race`, which CI does not. The sweeper, the notifier goroutine and the dead man's switch
|
||
sweep all run concurrently against the same database, and a race between them would surface as
|
||
a flaky incident in production rather than as a red build.
|
||
|
||
The web UI lives in `internal/web/static/` as plain HTML, CSS and ES modules,
|
||
embedded into the binary with `go:embed`. It has no build step and no npm, so
|
||
editing a file and restarting the server is the whole loop.
|
||
|
||
## Releasing
|
||
|
||
```
|
||
push or PR → ci.yaml gofmt, go vet, go test -race
|
||
govulncheck, gitleaks
|
||
helm lint + render
|
||
push tag vX.Y.Z → release.yaml the same gate, then publish:
|
||
git.ryuvia.com/niklas/terdut-server:vX.Y.Z
|
||
oci://git.ryuvia.com/niklas/terdut-server X.Y.Z
|
||
then trivy-scan the pushed image
|
||
PR to Ryuvia/charts → bump the wrapper chart to X.Y.Z; on merge
|
||
Flux reconciles and the release rolls out
|
||
```
|
||
|
||
Both artifacts go to the **personal** Gitea namespace rather than `ryuvia`, because Gitea
|
||
scopes package visibility to the owner with no per-package override — so `ryuvia/*` is private
|
||
because the org is. Publishing to `niklas` keeps them anonymously pullable, which is why no
|
||
pull secret is needed in the cluster. Same reasoning, and the same choice, as riksdata and
|
||
rd-web.
|
||
|
||
Saying **"Release"** runs all three rows: the `release` skill commits, pushes, tags, waits for
|
||
the pipeline, and opens the `Ryuvia/charts` PR, stopping before the merge. See
|
||
`~/.claude/skills/release/`, or `.release.conf` here for this repo's part of it.
|
||
|
||
The chart is published **only** from the tag, by the `chart` job. There used to be a second
|
||
publisher on every `charts/**` push to main, and the two raced for the same chart version with
|
||
different answers — chart 0.9.0 went out reading `appVersion: "latest"` that way. One
|
||
publisher, triggered by the tag (`766f439`). The cost is that a chart-only change has no
|
||
version of its own and rides the next app tag.
|
||
|
||
Both workflows are thin drivers over the Makefile: `ci.yaml` runs `make fmt lint test` and
|
||
`make helm-lint`, `release.yaml` adds `make binaries`, `make push`, `make helm-package` and
|
||
`make helm-push`. That is deliberate — it is what makes a green local gate and a green
|
||
pipeline the same code rather than two descriptions of it, and it is how riksdata and rd-web
|
||
have always worked.
|
||
|
||
`make push` builds and pushes in one step, unlike those two, because the image is
|
||
`linux/amd64,linux/arm64` and buildx cannot load a multi-platform result into the local image
|
||
store. `make build` stays single-platform and local-only. Both refuse `VERSION=dev`:
|
||
publishing is one command, so it is also one command to run by accident. Publishing happens
|
||
by pushing a tag.
|
||
|
||
Two things the release process needs to know about this repo:
|
||
|
||
- **The image scan runs after publishing**, like riksdata's and rd-web's: trivy cannot read
|
||
a locally built image on this runner, so it pulls the pushed one. A red `scan-image` means
|
||
do not bump the wrapper chart to that version — it cannot unpublish anything. The image is
|
||
`FROM scratch`, so trivy sees exactly one target, the Go binary and its module graph.
|
||
- **The wrapper chart's `values.yaml` has two `tag:` lines** — the app image and the python
|
||
backup sidecar — so `chart-bump` is given `--image` to say which one moves. The sidecar is
|
||
on its way out with SQLite: once the wrapper chart drops it and declares a `postgresql` CR
|
||
instead, there is one `tag:` line again, and `--image` becomes belt and braces.
|
||
|
||
The wrapper chart must have **its own `version:` bumped in the same commit**. Flux reconciles
|
||
with `reconcileStrategy: ChartVersion`, so a chart whose version did not change produces no
|
||
new artifact and the change is never deployed — with no error anywhere.
|