Files
terdut-server/docs/api.md
T
Niklas Ye 44b2eb2cc3 Rewrite the README as highlights with screenshots; move the detail into docs/
The README was 1,240 lines of reference material and still described a
SQLite quick start. It is now a short tour (highlights, screenshots of the
web UI, an accurate quick start against Postgres), and each topic has its
own page under docs/ with an index: deployment, configuration, Alertmanager,
incidents, notifications, escalation, dead man's switches, single sign-on,
web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a
proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it
described. The "Upgrading to ..." sections for an unreleased product are
dropped.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:13 +02:00

437 lines
30 KiB
Markdown

# API reference
_The REST API._ Back to the [README](../README.md) and the [documentation index](./README.md).
## Authentication
All endpoints except `/api/bootstrap`, `/api/integrations/{key}/alertmanager`,
`/api/notify/ack/{token}`, `/api/login`, `/api/logout`, `/api/auth/config`,
`/api/version`, `/api/oidc/login`, `/api/oidc/callback`, `/api/oidc/device` and
`/api/oidc/device/token` require either an API key:
```
Authorization: Bearer <api-key>
```
or the web UI's session cookie. A request that carries an `Authorization` header
is judged on that header alone.
Two kinds of user exist. An **administrator** manages accounts: creating and
deleting users, setting anybody's password, minting keys for anybody, and
granting the flag itself. Everybody else works incidents — acknowledging,
assigning, snoozing, resolving, noting — and manages their own account and
nobody else's. An API key carries exactly the rights of the user it belongs to.
A third principal, the **service account**, exists for automation (a
Kubernetes operator, most likely) that needs to manage teams, escalation
policies, dead man's switches and integrations without impersonating a human.
It is not a user — it never signs in, never appears in a team's member list,
and never holds the administrator flag — and its key is prefixed `tdsa_` so it
reads as one at a glance in a log line. See [Service accounts](#service-accounts).
**Getting an account.** The first one comes from `/api/bootstrap`. After that
it depends on `signup_mode`, an administrator setting:
- `invite_only` (the default) — a team owner mints a link with
`POST /api/teams/{teamID}/invites`, and the person who opens it picks a
username and password and lands in that team with the role the link carries.
Links are single-use unless told otherwise, expire after seven days, and can
be revoked before that.
- `open` — anybody who can reach the server can create an account, and must
name a team, which they then own.
Invites are **links, not email**: this server has no SMTP, and adding it to send
one message would be a subsystem to run, secure and monitor. Send the link
however you already talk to the person.
A domain-restricted third mode was considered and dropped: with no email there
is nothing to verify an address against, so it would only check the domain of a
string somebody typed.
The first user, from `/api/bootstrap`, is an administrator. Users created
afterwards are not, until an administrator says so. An install always keeps at
least one: the last administrator can be neither deleted nor demoted, and
nobody can delete or demote themselves.
Endpoints that require the flag answer `403` with
`{"error":"administrator access required"}`.
**Teams** are the unit of tenancy, and are a separate axis from the administrator
flag. A team owns its incidents, alerts, schedule and integrations, and a user
sees exactly the teams they belong to. Within a team an **owner** configures it
(schedule, integrations, membership) and a **member** works its incidents.
An administrator crosses that line in one direction only. They **configure any
team** without being in it — every owner-only endpoint accepts the flag, because
otherwise a team whose last owner left could never be repaired. They do **not
read any team**: the queue, the alerts and the incidents are filtered by real
membership, so an administrator sees a team's work only by joining it, which is
a membership change and shows up as one. Administration is about accounts and
the shape of a team, not about reading other people's incidents.
Anything belonging to a team you are not in answers `404`, not `403`: whether an
incident exists is itself something only its team should learn.
**Operator mode** (`TERDUT_OPERATOR_MODE`, see [Configuration](./configuration.md#configuration))
declares this install gitops-managed. When it is on, a session or a user's own
API key gets `403 {"error": "...", "reason": "operator_managed"}` on every
write this page marks **owner**-gated under Teams below (creating, renaming
or deleting a team; its OIDC group binding; its escalation ladder; its dead
man's switches; its integrations) — a service account's writes are unaffected.
Team membership and invites are deliberately excluded: they are never
gitops-managed, in operator mode or out of it. `GET /api/auth/config` reports
`operator_mode` so a client can grey those sections out before a write is ever
attempted.
| Method | Path | Description |
|---|---|---|
| `GET` | `/api/auth/config` | How to sign in: `{"password_login", "oidc": {"enabled","name"}, "device_login", "operator_mode"}`. No session needed |
| `GET` | `/api/version` | `{"version"}` — this build's version string. No session needed, the same as `/healthz` |
| `POST` | `/api/login` | `{"username","password"}` → sets the session cookie, returns `{user, has_password}`. `429` after too many failures; `403` when `TERDUT_PASSWORD_LOGIN=false` |
| `GET` | `/api/oidc/login` | Starts a single sign-on sign-in: redirects the browser to the provider. `?next=/path` is where to land afterwards; only a path on this server is honoured. Only exists when SSO is configured |
| `POST` | `/api/oidc/device` | Starts a device login: returns `{device_code, user_code, verification_url, interval, expires_in}`. Only exists when SSO is configured |
| `POST` | `/api/oidc/device/token` | `{"device_code"}` → `202 {"status":"pending"}`, then `200` with the session cookie once approved (once only). `410` with `{"error":"expired"}` or `{"error":"denied"}`; `429 {"error":"slow_down"}` if polled faster than `interval` |
| `POST` | `/api/oidc/device/approve` | **session** — `{"user_code"}`. Approves a pending device login as the caller. `403` for an API key; `404` for an unknown, expired or already decided code |
| `POST` | `/api/oidc/device/deny` | **session** — `{"user_code"}`. Refuses it |
| `GET` | `/api/oidc/callback` | Where the provider sends the browser back. Sets the session cookie and redirects to `/`, or to `/?sso_error=<code>` — one of `denied`, `expired`, `failed`, `unavailable`, `not_allowed`, `no_email`, `email_conflict`, `disabled`, `not_bootstrapped` (no user exists on this install yet — sign in again once something has called `/api/bootstrap`) |
| `POST` | `/api/logout` | Ends the session and clears the cookie |
| `GET` | `/api/me` | The caller: `{user, has_password}` |
## Users
| Method | Path | Description |
|---|---|---|
**admin** marks an endpoint that requires the administrator flag; **self or
admin** marks one you may use on your own account and an administrator may use
on anybody's.
| Method | Path | Who | Description |
|---|---|---|---|
| `GET` | `/api/signup` | — | Whether sign-up is open, and whether `?invite=` is usable. No session needed: the caller has no account yet |
| `POST` | `/api/signup` | — | Create an account `{"username","email","password","invite"?,"team_name"?}` and sign in. `403` without a usable invite when the mode is invite-only |
| `POST` | `/api/bootstrap` | — | Create first user + API key `{"username","email","password"?}` (only works on empty DB). The user is an administrator |
| `GET` | `/api/users` | any | List users. Open to everybody: the queue's assignment control and the schedule both have to name people |
| `GET` | `/api/users/{id}/teams` | self or admin | The teams that user is in, each with their role. `/api/teams` is always about the caller; this one answers it about somebody else, for the admin page's per-user view. `404` for a user who does not exist, so "no teams" and "no such person" are distinguishable |
| `POST` | `/api/users` | **admin** | Create user `{"username","email"}`. Not an administrator |
| `DELETE` | `/api/users/{id}` | **admin** | Delete user (cascades to keys). `409` for yourself or the last administrator |
| `PUT` | `/api/users/{id}/admin` | **admin** | Grant or revoke the administrator flag `{"is_admin"}`. `409` for yourself, the last administrator, or an administrator granted by single sign-on |
| `PUT` | `/api/users/{id}/disabled` | **admin** | Take an account out of use, or put it back `{"disabled"}`. `409` for yourself or the last administrator |
| `PUT` | `/api/users/{id}/notify` | self or admin | Set push notification target `{"ntfy_topic"}` — empty string clears it |
| `PUT` | `/api/users/{id}/password` | self or admin | Set web UI password `{"password","current_password"}`. `current_password` is required only when changing your own existing password. Ends the user's other sessions |
| `POST` | `/api/users/{id}/api-keys` | self or admin | Issue API key `{"name"}` — key shown once |
| `DELETE` | `/api/users/{id}/api-keys/{keyID}` | self or admin | Revoke API key |
## Administration
| Method | Path | Who | Description |
|---|---|---|---|
| `GET` | `/api/admin/teams` | **admin** | Every team on the server, with its member and open-incident counts. `/api/teams` answers "what am I in"; this answers "what is there" |
| `GET` | `/api/admin/teams/{teamID}` | **admin** | One team and who is in it: `{"team", "members"}`. `404` for a team that does not exist. `GET /api/teams/{teamID}/members` is **member**-only and still `404`s an administrator from outside the team — reading a team's shape and reading its work are different questions, so they are different endpoints |
| `GET` | `/api/admin/settings` | **admin** | The editable settings with their bounds, plus the environment-configured ones, read-only. Never credentials |
| `PUT` | `/api/admin/settings` | **admin** | Change one or more `{"key": seconds}`, or `{"signup_mode": "open"\|"invite_only"}`. `400` for an unknown key or a value outside its bounds |
## Service accounts
A service account is a scoped, non-human credential for automation — not a
`users` row, so it never signs in, is never a team member, and never carries
the administrator flag. Two scopes:
- **instance** — the same reach system administration has over teams: create
one, and mint a **team**-scoped account against any of them. There is no
cap on how many instance-scoped accounts exist, but ordinarily there is one,
belonging to whatever is provisioning this install end to end.
- **team** — owner-equivalent for that one team, and nothing else: every
**owner**-gated endpoint under [Teams](#teams), membership and invites
included. Nothing narrower is enforced server-side; what actually keeps
membership out of automation's hands is that no operator built against this
scope should ever call those two endpoints — see
[operator mode](#authentication) and [`SERVICE-ACCOUNTS.md`](../SERVICE-ACCOUNTS.md)'s note on this.
A key is shown once, at creation or rotation, and only its hash is stored —
the same handling as a user's API key. Losing it means minting a new one;
there is no way to recover a raw key from the server.
| Method | Path | Who | Description |
|---|---|---|---|
| `GET` | `/api/service-accounts` | **admin** | Every service account. Pass `?name=` instead to look one up by its exact name — open to **any** authenticated caller (human or service account), since it returns no key material and is how an account finds its own id |
| `POST` | `/api/service-accounts` | owner\* | Create one and mint its first key `{"name","scope","team_id"?}` (`team_id` required for `scope:"team"`, absent for `scope:"instance"`). Returns `{"service_account", "key"}` — `key.key` shown once |
| `POST` | `/api/service-accounts/{id}/keys` | owner\* | Mint an additional key `{"name"}` — rotation without recreating the account. Shown once |
| `DELETE` | `/api/service-accounts/{id}/keys/{keyID}` | owner\* | Revoke one key |
\* For an **instance**-scoped account: a system administrator only. For a
**team**-scoped account: a system administrator, that team's own human owner,
an instance-scoped service account (minting a narrower credential for a team
it just created), or — for the two key endpoints only — the account rotating
or revoking its own key, which is not a privilege escalation, the same
reasoning a user's own API keys rest on.
## Alert ingestion
Alerts arrive on a team's integration key. The key is both the credential and the
routing: it says that the sender may post, and which team the alerts belong to.
Create one with `POST /api/teams/{teamID}/integrations`, which returns the key
and the full URL once and stores only a SHA-256 hash.
| Method | Path | Description |
|---|---|---|
| `POST` | `/api/integrations/{key}/alertmanager` | Alertmanager v4 webhook receiver for the key's team. `401` for an unknown key |
This is the only way in. The pre-teams `POST /api/alertmanager/webhook` took no
credential at all — anything able to reach the port could open an incident —
and was removed in v0.13.0 once senders had moved onto keys.
## Teams
**owner** below means an owner of that team, a system administrator (who
passes every one of these without being a member), or that team's own
team-scoped [service account](#service-accounts) — including membership and
invites, technically, though no automation this scope was designed for
(a Kubernetes operator's CRDs, see [`SERVICE-ACCOUNTS.md`](../SERVICE-ACCOUNTS.md)) ever models team
membership or would call those two. See [Authentication](#authentication).
**member** means membership and nothing else: an administrator who is not in
the team gets the same `404` as anybody else.
| Method | Path | Who | Description |
|---|---|---|---|
| `GET` | `/api/teams` | any | The caller's own teams, each with their role |
| `POST` | `/api/teams` | any | Create a team `{"name"}`; a human creator becomes its first owner. An instance-scoped [service account](#service-accounts) may also create one, and it gets no owner at all — expected for a team an operator is about to hand a team-scoped credential to, not an orphaned team a human made |
| `PUT` | `/api/teams/{teamID}` | **owner** | Rename it `{"name"}`. `409` if the name is taken |
| `DELETE` | `/api/teams/{teamID}` | **owner** | Delete a team and everything under it. `409` while it has open incidents |
| `GET` | `/api/teams/{teamID}/members` | member | Who is in the team, with `status` (`oncall` if the rota has them today, `unpageable` when a page to them would go nowhere — even if they are on call — else `reachable`), `on_call`, `next_shift` (first rota day after today), `pageable` and `problem` (`has no ntfy topic` / `account is disabled`; never the topic itself) and `last_active_at` (their newest session or API-key use). Every member sees the same list |
| `POST` | `/api/teams/{teamID}/members` | **owner** | Add a member, or change their role `{"user_id","role"}`. `409` when it would demote the last owner, or the membership is managed by single sign-on |
| `DELETE` | `/api/teams/{teamID}/members/{userID}` | **owner** | Remove a member. `409` for the last owner, or a membership managed by single sign-on |
| `GET` | `/api/teams/{teamID}/oidc-groups` | member | Which groups control this team's membership: `{"member_group","owner_group"}`. An empty string means no group grants that role here |
| `PUT` | `/api/teams/{teamID}/oidc-groups` | **owner** | Set them. An empty string clears a binding |
| `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys. Each carries `status` (`active` if its key posted within 24h, `quiet` if it has but not lately, `never`), `last_used_at` (last webhook, usable or not), `last_alert_at` (when an alert last arrived on it) and `alerts_24h` (distinct alerts it refreshed in the last day). Alerts delivered before the source was recorded (migration 010) have none, so the last two fill in as Alertmanager re-sends them |
| `PATCH` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Rename `{"name"}`. The key does not change |
| `POST` | `/api/teams/{teamID}/integrations` | **owner** | Mint an integration `{"name","kind"}` — key and URL shown once |
| `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration. Alerts it delivered stay, unattributed |
| `GET` | `/api/teams/{teamID}/invites` | **owner** | The team's invite links, with their uses and expiry. Never the tokens |
| `POST` | `/api/teams/{teamID}/invites` | **owner** | Mint one `{"role","max_uses"}` — the full URL is returned once |
| `DELETE` | `/api/teams/{teamID}/invites/{inviteID}` | **owner** | Revoke a link before it expires |
| `GET` | `/api/teams/{teamID}/escalation` | member | The team's [escalation ladder](./escalation.md#escalation) `{repeat_count, fallback_topic, levels[], last_escalated_at?, last_escalated_incident_id?}`. Empty levels means the team has none. Each level also carries `status` (`ready`, `escalating` when an unanswered incident has climbed to it, `unreachable` when nobody on it could be woken), `waiting` (ids of the open incidents on it) and, per target, `username` (who it means today — the person on call, for a rota target), `reachable` and `problem`. The extra fields are output only; `PUT` takes the plain shape |
| `PUT` | `/api/teams/{teamID}/escalation` | **owner** | Replace it wholesale. `400` for a level with no targets or no timeout — a rung that pages nobody is a silence with a number on it |
| `GET` | `/api/teams/{teamID}/deadman/switches` | member | The team's [dead man's switches](./dead-mans-switch.md), each `{id, name, matcher, timeout_seconds, severity, status, last_heartbeat_at, last_triggered_at, open_incident_id, sources[]}`. `status` is `healthy`, `dead` or `dormant`; `sources` has one entry per heartbeat fingerprint. Empty when the team watches nothing |
| `POST` | `/api/teams/{teamID}/deadman/switches` | **owner** | Add one: `{name?, matcher, timeout_seconds, severity?}`. `400` when the matcher names no `alertname` or holds several, or the timeout is not positive — a switch that silently watches nothing is the failure this feature exists to prevent |
| `PUT` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Replace one in place, same body and validation as create. Its id is unchanged — for an automated caller reconciling a spec change, unlike delete-and-recreate |
| `DELETE` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Stop watching. An incident it opened stays open. `404` for a switch of another team |
## Notifications
| Method | Path | Description |
|---|---|---|
| `POST` | `/api/notify/ack/{token}` | Acknowledge an incident from a push notification's Acknowledge button. No auth: the token in the path is the credential — one incident, one action, 24 hours, idempotent. Must stay publicly reachable |
## Incidents
| Method | Path | Description |
|---|---|---|
| `GET` | `/api/incidents` | List incidents. Filters: `?status=triggered\|acknowledged\|resolved`, `?severity=`, `?assigned_to=<user id>`, `?archived=true`, `?snoozed=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?sort=severity`, `?cluster=<value of the cluster group label>`, `?limit=` (default 50, max 500) |
| `GET` | `/api/incidents/clusters` | The distinct `cluster` values on the caller's incidents from the last 90 days, sorted (`?team_id=` narrows it). An empty array when nothing carries the label |
| `GET` | `/api/incidents/{id}` | Get single incident, with its alerts inline |
| `GET` | `/api/incidents/{id}/alerts` | Alerts under this incident |
| `GET` | `/api/incidents/{id}/timeline` | Full event history, chronological |
| `POST` | `/api/incidents/{id}/acknowledge` | Acknowledge (stamps authed user + time) |
| `DELETE` | `/api/incidents/{id}/acknowledge` | Clear acknowledgement, back to `triggered` |
| `POST` | `/api/incidents/{id}/resolve` | Close by hand — **terminal**, see above |
| `POST` | `/api/incidents/{id}/assign` | Reassign `{"user_id"}` |
| `POST` | `/api/incidents/{id}/snooze` | Hide until `{"until": RFC3339}` or `{"duration": "2h"}` |
| `DELETE` | `/api/incidents/{id}/snooze` | Un-snooze |
| `POST` | `/api/incidents/{id}/archive` | Archive (hides from the default list) |
| `DELETE` | `/api/incidents/{id}/archive` | Un-archive |
| `POST` | `/api/incidents/{id}/notes` | Add a note `{"content"}` |
| `DELETE` | `/api/incidents/{id}/notes/{eventID}` | Delete own note |
With no `?status=` filter, `GET /api/incidents` returns **open** incidents only —
the queue an on-call person wants. Currently snoozed and archived incidents are
excluded unless asked for. Actions that only make sense on an open incident
return `409` once it is resolved.
Notes are ordinary timeline events of type `note`; only they are deletable, and
only by their author. The rest of the timeline is a record of what happened.
### The incident object
| Field | Type | Notes |
|---|---|---|
| `id` | integer | Server-assigned |
| `group_key` | string | Alertmanager's `groupKey` — opaque, treat as an identifier |
| `title` | string | Rendered from `groupLabels` |
| `group_labels` | object | String→string, as sent by Alertmanager |
| `status` | string | `"triggered"`, `"acknowledged"` or `"resolved"` |
| `severity` | string | *optional* — high-water mark across the incident's alerts; never lowered |
| `triggered_at` | timestamp | When the incident opened |
| `acknowledged_by_id` / `acknowledged_by` / `acknowledged_at` | | *optional* — user id, username, time |
| `assigned_to_id` / `assigned_to` | | *optional* — user id, username |
| `snoozed_until` | timestamp | *optional* — a value in the past reads as not snoozed |
| `resolved_at` | timestamp | *optional* |
| `resolution_source` | string | *optional* — `"alerts"`, `"manual"` or `"recovered"` |
| `archived_at` | timestamp | *optional* |
| `alerts` | array | Only on `GET /api/incidents/{id}` |
Treat `resolution_source` as an open set, as with the alert field of the same
name: degrade unknown values to "resolved, reason unknown".
### The timeline event object
| Field | Type | Notes |
|---|---|---|
| `id` | integer | |
| `incident_id` | integer | |
| `type` | string | See below — treat as an open set |
| `user_id` / `username` | | *optional* — absent when the server acted rather than a person |
| `alert_id` | integer | *optional* — the alert an `alert_added` / `alert_resolved` event refers to |
| `detail` | string | *optional* — the note body, the snooze deadline, etc. |
| `created_at` | timestamp | |
Types written today: `triggered`, `alert_added`, `alert_resolved`,
`acknowledged`, `unacknowledged`, `assigned`, `archived`, `unarchived`, `snoozed`,
`unsnoozed`, `resolved`, `note`, `notified`, `notify_failed`, `deadman_silent`. On an
`assigned` event `user_id` is the **assignee**, not the actor; the actor is in
`actor_user_id`/`actor_username` or `actor_service_account_id`/`actor_service_account_name`
(absent on assignments made before they were recorded). New types may be added; render
unknown ones generically rather than dropping them.
On `notified` and `notify_failed`, `detail` carries the notification kind
(`triggered` | `reminder` | `resolved`), and on a failure the reason after it.
`user_id` is who was paged — absent means the page went to the shared fallback
topic and so belongs to nobody. The topic itself is never written to the
timeline: it is a shared secret with the ntfy server, and every API key can read
this.
## Alerts
Alerts are read-only. Everything a person does happens on the incident.
| Method | Path | Description |
|---|---|---|
| `GET` | `/api/alerts` | List alerts. Filters: `?status=firing\|resolved`, `?name=`, `?incident_id=`, `?archived=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?limit=` (default 50, max 500) |
| `GET` | `/api/alerts/{id}` | Get single alert |
Archived alerts are hidden from `GET /api/alerts` unless `?archived=true` is
passed; alert archiving is automatic housekeeping by the sweeper, not a user
action. Resolved alerts carry `resolution_source`: `"alertmanager"` for a real
resolved webhook, `"expiry"` when the sweeper inferred it (see
[Stale alert expiry](./incidents.md#stale-alert-expiry)), `"deadman"` for a heartbeat declared
dead (see [Dead man's switch](./dead-mans-switch.md)).
### The alert object
Returned by `GET /api/alerts` (as an array) and `GET /api/alerts/{id}`.
Timestamps are RFC 3339 in UTC. Fields marked *optional* are omitted entirely
when unset, so clients must treat them as nullable.
| Field | Type | Notes |
|---|---|---|
| `id` | integer | Server-assigned; stable for the life of the row |
| `fingerprint` | string | Alertmanager's fingerprint — the upsert key |
| `name` | string | From the `alertname` label |
| `status` | string | `"firing"` or `"resolved"` |
| `labels` | object | String→string, as sent by Alertmanager |
| `annotations` | object | String→string, as sent by Alertmanager |
| `starts_at` | timestamp | When the alert instance began, **per Prometheus** |
| `ends_at` | timestamp | *optional* — absent while no end is known |
| `generator_url` | string | Link back to the originating Prometheus |
| `received_at` | timestamp | When the server last accepted a webhook for this alert — see below |
| `incident_id` | integer | *optional* — the most recent incident this alert belongs to |
| `resolution_source` | string | *optional* — `"alertmanager"`, `"expiry"` or `"deadman"` |
| `archived_at` | timestamp | *optional* — set while archived |
#### `received_at` is a liveness heartbeat
`starts_at` comes from Prometheus and **never changes** for the lifetime of an
alert instance. It says when the problem began, not whether it is still
happening — an alert that started twelve days ago looks identical whether
Alertmanager refreshed it a minute ago or went silent a week ago.
`received_at` is the field that answers "is this still live". It is set to the
server's clock on **every accepted webhook** for that fingerprint, including the
unchanged firing notifications Alertmanager re-sends every `repeat_interval`.
Clients may rely on this:
- **A firing alert whose `received_at` is advancing is still being refreshed.**
Stale-dating it against `repeat_interval` is a valid liveness check, and it is
what the built-in sweeper does (see
[Stale alert expiry](./incidents.md#stale-alert-expiry)).
- **`received_at` tracks accepted payloads, not delivery attempts.** A retry
that describes an older instance than the stored one is discarded, and a
discarded payload does not move `received_at`.
- **It stops advancing once the alert resolves,** because Alertmanager stops
re-sending. On an alert resolved by the sweeper
(`"resolution_source": "expiry"`) it therefore marks the last time
Alertmanager was actually heard from, which is earlier than `ends_at`.
`GET /api/alerts` is ordered by `received_at` descending — most recently
refreshed first — and the `?from=` / `?to=` filters on both the alert and stats
endpoints select on `received_at`, not `starts_at`.
#### `resolution_source` says how much to trust `ends_at`
An alert can leave the firing state two ways, and `resolution_source` records
which happened. Clients may rely on this:
- **Absent while firing.** It is set only on resolve, and a re-fire under the
same fingerprint clears it again, so its presence always agrees with
`"status": "resolved"`.
- **`"alertmanager"` — a real resolved webhook arrived.** `ends_at` is the end
time Alertmanager reported. It is an observed value and can be displayed as
fact.
- **`"expiry"` — the sweeper inferred the resolve** because Alertmanager stopped
refreshing the alert (see [Stale alert expiry](./incidents.md#stale-alert-expiry)). Nothing
ever reported an end, so **`ends_at` is approximate**: it is either the stale
`endsAt` watermark from the last notification, or — when that notification
carried none — the time the sweep ran, which lags the last real contact by up
to `TERDUT_STALE_AFTER` plus a sweep interval. Treat it as "no later than",
not as when the problem stopped.
On these alerts `received_at` is the more truthful signal: it marks the last
time Alertmanager was actually heard from. Surfacing the distinction is
worthwhile, since `"expiry"` can also mean the alert is still firing and the
notification path broke.
- **`"deadman"` — a heartbeat was declared dead** (see
[Dead man's switch](./dead-mans-switch.md)). Like `"expiry"`, an inference from
silence rather than an observed end, so `ends_at` is approximate — but a much
tighter one, bounded by the switch's timeout. It is also the one resolution
a re-fire under the same `starts_at` can undo, since the switch coming back is
exactly the evidence that the inference was wrong.
Treat the value as an open set and tolerate ones you do not recognise — new
sources may be added, and unknown values should degrade to "resolved, reason
unknown" rather than being rejected.
## On-call schedule
| Method | Path | Description |
|---|---|---|
Each team keeps its own rota, so two teams can have two different people on call
on the same day. The person taking a shift has to be in the team — paging
somebody who cannot open the incident is worse than paging nobody.
| Method | Path | Who | Description |
|---|---|---|---|
| `POST` | `/api/teams/{teamID}/schedule` | **owner** | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...], "replace"}` — all-or-nothing |
| `GET` | `/api/teams/{teamID}/schedule` | member | List entries. Filters: `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD` |
| `DELETE` | `/api/teams/{teamID}/schedule/{id}` | **owner** | Remove schedule entry |
| `GET` | `/api/schedule/current` | any | Who is on call today (UTC) in **every** team the caller is in — one entry per team, `[]` when nobody anywhere |
## Statistics
Every figure counts the caller's own teams only: a report that counted other
teams' incidents would leak their volume, and their alert names through the
top-alerts list, and would not be a number about the reader's work anyway.
All stat endpoints accept optional `?from=YYYY-MM-DD` and `?to=YYYY-MM-DD`, and exclude archived rows to match the default list views. Alert stats filter on `received_at`; incident stats filter on `triggered_at`.
| Method | Path | Description |
|---|---|---|
| `GET` | `/api/stats/incidents` | `{total, triggered, acknowledged, resolved, mtta_seconds, mttr_seconds}` |
| `GET` | `/api/stats/alerts` | `{total, firing, resolved}` counts |
| `GET` | `/api/stats/alerts/top` | Most frequent alert names. `?limit=` (default 10, max 100) |
| `GET` | `/api/stats/alerts/by-hour` | Count per hour-of-day (UTC), all 24 slots returned |
| `GET` | `/api/stats/alerts/by-day` | Count per day-of-week, all 7 slots with names returned |
`mtta_seconds` (time to acknowledge) and `mttr_seconds` (time to resolve) are
averages over incidents that have actually been acknowledged or resolved, and are
**null** until there are any — null means "no data", not zero.