44b2eb2cc3
The README was 1,240 lines of reference material and still described a SQLite quick start. It is now a short tour (highlights, screenshots of the web UI, an accurate quick start against Postgres), and each topic has its own page under docs/ with an index: deployment, configuration, Alertmanager, incidents, notifications, escalation, dead man's switches, single sign-on, web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it described. The "Upgrading to ..." sections for an unreleased product are dropped. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
437 lines
30 KiB
Markdown
437 lines
30 KiB
Markdown
# API reference
|
|
|
|
_The REST API._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
|
|
|
## Authentication
|
|
|
|
All endpoints except `/api/bootstrap`, `/api/integrations/{key}/alertmanager`,
|
|
`/api/notify/ack/{token}`, `/api/login`, `/api/logout`, `/api/auth/config`,
|
|
`/api/version`, `/api/oidc/login`, `/api/oidc/callback`, `/api/oidc/device` and
|
|
`/api/oidc/device/token` require either an API key:
|
|
|
|
```
|
|
Authorization: Bearer <api-key>
|
|
```
|
|
|
|
or the web UI's session cookie. A request that carries an `Authorization` header
|
|
is judged on that header alone.
|
|
|
|
Two kinds of user exist. An **administrator** manages accounts: creating and
|
|
deleting users, setting anybody's password, minting keys for anybody, and
|
|
granting the flag itself. Everybody else works incidents — acknowledging,
|
|
assigning, snoozing, resolving, noting — and manages their own account and
|
|
nobody else's. An API key carries exactly the rights of the user it belongs to.
|
|
|
|
A third principal, the **service account**, exists for automation (a
|
|
Kubernetes operator, most likely) that needs to manage teams, escalation
|
|
policies, dead man's switches and integrations without impersonating a human.
|
|
It is not a user — it never signs in, never appears in a team's member list,
|
|
and never holds the administrator flag — and its key is prefixed `tdsa_` so it
|
|
reads as one at a glance in a log line. See [Service accounts](#service-accounts).
|
|
|
|
**Getting an account.** The first one comes from `/api/bootstrap`. After that
|
|
it depends on `signup_mode`, an administrator setting:
|
|
|
|
- `invite_only` (the default) — a team owner mints a link with
|
|
`POST /api/teams/{teamID}/invites`, and the person who opens it picks a
|
|
username and password and lands in that team with the role the link carries.
|
|
Links are single-use unless told otherwise, expire after seven days, and can
|
|
be revoked before that.
|
|
- `open` — anybody who can reach the server can create an account, and must
|
|
name a team, which they then own.
|
|
|
|
Invites are **links, not email**: this server has no SMTP, and adding it to send
|
|
one message would be a subsystem to run, secure and monitor. Send the link
|
|
however you already talk to the person.
|
|
|
|
A domain-restricted third mode was considered and dropped: with no email there
|
|
is nothing to verify an address against, so it would only check the domain of a
|
|
string somebody typed.
|
|
|
|
The first user, from `/api/bootstrap`, is an administrator. Users created
|
|
afterwards are not, until an administrator says so. An install always keeps at
|
|
least one: the last administrator can be neither deleted nor demoted, and
|
|
nobody can delete or demote themselves.
|
|
|
|
Endpoints that require the flag answer `403` with
|
|
`{"error":"administrator access required"}`.
|
|
|
|
**Teams** are the unit of tenancy, and are a separate axis from the administrator
|
|
flag. A team owns its incidents, alerts, schedule and integrations, and a user
|
|
sees exactly the teams they belong to. Within a team an **owner** configures it
|
|
(schedule, integrations, membership) and a **member** works its incidents.
|
|
|
|
An administrator crosses that line in one direction only. They **configure any
|
|
team** without being in it — every owner-only endpoint accepts the flag, because
|
|
otherwise a team whose last owner left could never be repaired. They do **not
|
|
read any team**: the queue, the alerts and the incidents are filtered by real
|
|
membership, so an administrator sees a team's work only by joining it, which is
|
|
a membership change and shows up as one. Administration is about accounts and
|
|
the shape of a team, not about reading other people's incidents.
|
|
|
|
Anything belonging to a team you are not in answers `404`, not `403`: whether an
|
|
incident exists is itself something only its team should learn.
|
|
|
|
**Operator mode** (`TERDUT_OPERATOR_MODE`, see [Configuration](./configuration.md#configuration))
|
|
declares this install gitops-managed. When it is on, a session or a user's own
|
|
API key gets `403 {"error": "...", "reason": "operator_managed"}` on every
|
|
write this page marks **owner**-gated under Teams below (creating, renaming
|
|
or deleting a team; its OIDC group binding; its escalation ladder; its dead
|
|
man's switches; its integrations) — a service account's writes are unaffected.
|
|
Team membership and invites are deliberately excluded: they are never
|
|
gitops-managed, in operator mode or out of it. `GET /api/auth/config` reports
|
|
`operator_mode` so a client can grey those sections out before a write is ever
|
|
attempted.
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
| `GET` | `/api/auth/config` | How to sign in: `{"password_login", "oidc": {"enabled","name"}, "device_login", "operator_mode"}`. No session needed |
|
|
| `GET` | `/api/version` | `{"version"}` — this build's version string. No session needed, the same as `/healthz` |
|
|
| `POST` | `/api/login` | `{"username","password"}` → sets the session cookie, returns `{user, has_password}`. `429` after too many failures; `403` when `TERDUT_PASSWORD_LOGIN=false` |
|
|
| `GET` | `/api/oidc/login` | Starts a single sign-on sign-in: redirects the browser to the provider. `?next=/path` is where to land afterwards; only a path on this server is honoured. Only exists when SSO is configured |
|
|
| `POST` | `/api/oidc/device` | Starts a device login: returns `{device_code, user_code, verification_url, interval, expires_in}`. Only exists when SSO is configured |
|
|
| `POST` | `/api/oidc/device/token` | `{"device_code"}` → `202 {"status":"pending"}`, then `200` with the session cookie once approved (once only). `410` with `{"error":"expired"}` or `{"error":"denied"}`; `429 {"error":"slow_down"}` if polled faster than `interval` |
|
|
| `POST` | `/api/oidc/device/approve` | **session** — `{"user_code"}`. Approves a pending device login as the caller. `403` for an API key; `404` for an unknown, expired or already decided code |
|
|
| `POST` | `/api/oidc/device/deny` | **session** — `{"user_code"}`. Refuses it |
|
|
| `GET` | `/api/oidc/callback` | Where the provider sends the browser back. Sets the session cookie and redirects to `/`, or to `/?sso_error=<code>` — one of `denied`, `expired`, `failed`, `unavailable`, `not_allowed`, `no_email`, `email_conflict`, `disabled`, `not_bootstrapped` (no user exists on this install yet — sign in again once something has called `/api/bootstrap`) |
|
|
| `POST` | `/api/logout` | Ends the session and clears the cookie |
|
|
| `GET` | `/api/me` | The caller: `{user, has_password}` |
|
|
|
|
## Users
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
**admin** marks an endpoint that requires the administrator flag; **self or
|
|
admin** marks one you may use on your own account and an administrator may use
|
|
on anybody's.
|
|
|
|
| Method | Path | Who | Description |
|
|
|---|---|---|---|
|
|
| `GET` | `/api/signup` | — | Whether sign-up is open, and whether `?invite=` is usable. No session needed: the caller has no account yet |
|
|
| `POST` | `/api/signup` | — | Create an account `{"username","email","password","invite"?,"team_name"?}` and sign in. `403` without a usable invite when the mode is invite-only |
|
|
| `POST` | `/api/bootstrap` | — | Create first user + API key `{"username","email","password"?}` (only works on empty DB). The user is an administrator |
|
|
| `GET` | `/api/users` | any | List users. Open to everybody: the queue's assignment control and the schedule both have to name people |
|
|
| `GET` | `/api/users/{id}/teams` | self or admin | The teams that user is in, each with their role. `/api/teams` is always about the caller; this one answers it about somebody else, for the admin page's per-user view. `404` for a user who does not exist, so "no teams" and "no such person" are distinguishable |
|
|
| `POST` | `/api/users` | **admin** | Create user `{"username","email"}`. Not an administrator |
|
|
| `DELETE` | `/api/users/{id}` | **admin** | Delete user (cascades to keys). `409` for yourself or the last administrator |
|
|
| `PUT` | `/api/users/{id}/admin` | **admin** | Grant or revoke the administrator flag `{"is_admin"}`. `409` for yourself, the last administrator, or an administrator granted by single sign-on |
|
|
| `PUT` | `/api/users/{id}/disabled` | **admin** | Take an account out of use, or put it back `{"disabled"}`. `409` for yourself or the last administrator |
|
|
| `PUT` | `/api/users/{id}/notify` | self or admin | Set push notification target `{"ntfy_topic"}` — empty string clears it |
|
|
| `PUT` | `/api/users/{id}/password` | self or admin | Set web UI password `{"password","current_password"}`. `current_password` is required only when changing your own existing password. Ends the user's other sessions |
|
|
| `POST` | `/api/users/{id}/api-keys` | self or admin | Issue API key `{"name"}` — key shown once |
|
|
| `DELETE` | `/api/users/{id}/api-keys/{keyID}` | self or admin | Revoke API key |
|
|
|
|
## Administration
|
|
|
|
| Method | Path | Who | Description |
|
|
|---|---|---|---|
|
|
| `GET` | `/api/admin/teams` | **admin** | Every team on the server, with its member and open-incident counts. `/api/teams` answers "what am I in"; this answers "what is there" |
|
|
| `GET` | `/api/admin/teams/{teamID}` | **admin** | One team and who is in it: `{"team", "members"}`. `404` for a team that does not exist. `GET /api/teams/{teamID}/members` is **member**-only and still `404`s an administrator from outside the team — reading a team's shape and reading its work are different questions, so they are different endpoints |
|
|
| `GET` | `/api/admin/settings` | **admin** | The editable settings with their bounds, plus the environment-configured ones, read-only. Never credentials |
|
|
| `PUT` | `/api/admin/settings` | **admin** | Change one or more `{"key": seconds}`, or `{"signup_mode": "open"\|"invite_only"}`. `400` for an unknown key or a value outside its bounds |
|
|
|
|
## Service accounts
|
|
|
|
A service account is a scoped, non-human credential for automation — not a
|
|
`users` row, so it never signs in, is never a team member, and never carries
|
|
the administrator flag. Two scopes:
|
|
|
|
- **instance** — the same reach system administration has over teams: create
|
|
one, and mint a **team**-scoped account against any of them. There is no
|
|
cap on how many instance-scoped accounts exist, but ordinarily there is one,
|
|
belonging to whatever is provisioning this install end to end.
|
|
- **team** — owner-equivalent for that one team, and nothing else: every
|
|
**owner**-gated endpoint under [Teams](#teams), membership and invites
|
|
included. Nothing narrower is enforced server-side; what actually keeps
|
|
membership out of automation's hands is that no operator built against this
|
|
scope should ever call those two endpoints — see
|
|
[operator mode](#authentication) and [`SERVICE-ACCOUNTS.md`](../SERVICE-ACCOUNTS.md)'s note on this.
|
|
|
|
A key is shown once, at creation or rotation, and only its hash is stored —
|
|
the same handling as a user's API key. Losing it means minting a new one;
|
|
there is no way to recover a raw key from the server.
|
|
|
|
| Method | Path | Who | Description |
|
|
|---|---|---|---|
|
|
| `GET` | `/api/service-accounts` | **admin** | Every service account. Pass `?name=` instead to look one up by its exact name — open to **any** authenticated caller (human or service account), since it returns no key material and is how an account finds its own id |
|
|
| `POST` | `/api/service-accounts` | owner\* | Create one and mint its first key `{"name","scope","team_id"?}` (`team_id` required for `scope:"team"`, absent for `scope:"instance"`). Returns `{"service_account", "key"}` — `key.key` shown once |
|
|
| `POST` | `/api/service-accounts/{id}/keys` | owner\* | Mint an additional key `{"name"}` — rotation without recreating the account. Shown once |
|
|
| `DELETE` | `/api/service-accounts/{id}/keys/{keyID}` | owner\* | Revoke one key |
|
|
|
|
\* For an **instance**-scoped account: a system administrator only. For a
|
|
**team**-scoped account: a system administrator, that team's own human owner,
|
|
an instance-scoped service account (minting a narrower credential for a team
|
|
it just created), or — for the two key endpoints only — the account rotating
|
|
or revoking its own key, which is not a privilege escalation, the same
|
|
reasoning a user's own API keys rest on.
|
|
|
|
## Alert ingestion
|
|
|
|
Alerts arrive on a team's integration key. The key is both the credential and the
|
|
routing: it says that the sender may post, and which team the alerts belong to.
|
|
Create one with `POST /api/teams/{teamID}/integrations`, which returns the key
|
|
and the full URL once and stores only a SHA-256 hash.
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
| `POST` | `/api/integrations/{key}/alertmanager` | Alertmanager v4 webhook receiver for the key's team. `401` for an unknown key |
|
|
|
|
This is the only way in. The pre-teams `POST /api/alertmanager/webhook` took no
|
|
credential at all — anything able to reach the port could open an incident —
|
|
and was removed in v0.13.0 once senders had moved onto keys.
|
|
|
|
## Teams
|
|
|
|
**owner** below means an owner of that team, a system administrator (who
|
|
passes every one of these without being a member), or that team's own
|
|
team-scoped [service account](#service-accounts) — including membership and
|
|
invites, technically, though no automation this scope was designed for
|
|
(a Kubernetes operator's CRDs, see [`SERVICE-ACCOUNTS.md`](../SERVICE-ACCOUNTS.md)) ever models team
|
|
membership or would call those two. See [Authentication](#authentication).
|
|
**member** means membership and nothing else: an administrator who is not in
|
|
the team gets the same `404` as anybody else.
|
|
|
|
| Method | Path | Who | Description |
|
|
|---|---|---|---|
|
|
| `GET` | `/api/teams` | any | The caller's own teams, each with their role |
|
|
| `POST` | `/api/teams` | any | Create a team `{"name"}`; a human creator becomes its first owner. An instance-scoped [service account](#service-accounts) may also create one, and it gets no owner at all — expected for a team an operator is about to hand a team-scoped credential to, not an orphaned team a human made |
|
|
| `PUT` | `/api/teams/{teamID}` | **owner** | Rename it `{"name"}`. `409` if the name is taken |
|
|
| `DELETE` | `/api/teams/{teamID}` | **owner** | Delete a team and everything under it. `409` while it has open incidents |
|
|
| `GET` | `/api/teams/{teamID}/members` | member | Who is in the team, with `status` (`oncall` if the rota has them today, `unpageable` when a page to them would go nowhere — even if they are on call — else `reachable`), `on_call`, `next_shift` (first rota day after today), `pageable` and `problem` (`has no ntfy topic` / `account is disabled`; never the topic itself) and `last_active_at` (their newest session or API-key use). Every member sees the same list |
|
|
| `POST` | `/api/teams/{teamID}/members` | **owner** | Add a member, or change their role `{"user_id","role"}`. `409` when it would demote the last owner, or the membership is managed by single sign-on |
|
|
| `DELETE` | `/api/teams/{teamID}/members/{userID}` | **owner** | Remove a member. `409` for the last owner, or a membership managed by single sign-on |
|
|
| `GET` | `/api/teams/{teamID}/oidc-groups` | member | Which groups control this team's membership: `{"member_group","owner_group"}`. An empty string means no group grants that role here |
|
|
| `PUT` | `/api/teams/{teamID}/oidc-groups` | **owner** | Set them. An empty string clears a binding |
|
|
| `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys. Each carries `status` (`active` if its key posted within 24h, `quiet` if it has but not lately, `never`), `last_used_at` (last webhook, usable or not), `last_alert_at` (when an alert last arrived on it) and `alerts_24h` (distinct alerts it refreshed in the last day). Alerts delivered before the source was recorded (migration 010) have none, so the last two fill in as Alertmanager re-sends them |
|
|
| `PATCH` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Rename `{"name"}`. The key does not change |
|
|
| `POST` | `/api/teams/{teamID}/integrations` | **owner** | Mint an integration `{"name","kind"}` — key and URL shown once |
|
|
| `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration. Alerts it delivered stay, unattributed |
|
|
| `GET` | `/api/teams/{teamID}/invites` | **owner** | The team's invite links, with their uses and expiry. Never the tokens |
|
|
| `POST` | `/api/teams/{teamID}/invites` | **owner** | Mint one `{"role","max_uses"}` — the full URL is returned once |
|
|
| `DELETE` | `/api/teams/{teamID}/invites/{inviteID}` | **owner** | Revoke a link before it expires |
|
|
| `GET` | `/api/teams/{teamID}/escalation` | member | The team's [escalation ladder](./escalation.md#escalation) `{repeat_count, fallback_topic, levels[], last_escalated_at?, last_escalated_incident_id?}`. Empty levels means the team has none. Each level also carries `status` (`ready`, `escalating` when an unanswered incident has climbed to it, `unreachable` when nobody on it could be woken), `waiting` (ids of the open incidents on it) and, per target, `username` (who it means today — the person on call, for a rota target), `reachable` and `problem`. The extra fields are output only; `PUT` takes the plain shape |
|
|
| `PUT` | `/api/teams/{teamID}/escalation` | **owner** | Replace it wholesale. `400` for a level with no targets or no timeout — a rung that pages nobody is a silence with a number on it |
|
|
| `GET` | `/api/teams/{teamID}/deadman/switches` | member | The team's [dead man's switches](./dead-mans-switch.md), each `{id, name, matcher, timeout_seconds, severity, status, last_heartbeat_at, last_triggered_at, open_incident_id, sources[]}`. `status` is `healthy`, `dead` or `dormant`; `sources` has one entry per heartbeat fingerprint. Empty when the team watches nothing |
|
|
| `POST` | `/api/teams/{teamID}/deadman/switches` | **owner** | Add one: `{name?, matcher, timeout_seconds, severity?}`. `400` when the matcher names no `alertname` or holds several, or the timeout is not positive — a switch that silently watches nothing is the failure this feature exists to prevent |
|
|
| `PUT` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Replace one in place, same body and validation as create. Its id is unchanged — for an automated caller reconciling a spec change, unlike delete-and-recreate |
|
|
| `DELETE` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Stop watching. An incident it opened stays open. `404` for a switch of another team |
|
|
|
|
## Notifications
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
| `POST` | `/api/notify/ack/{token}` | Acknowledge an incident from a push notification's Acknowledge button. No auth: the token in the path is the credential — one incident, one action, 24 hours, idempotent. Must stay publicly reachable |
|
|
|
|
## Incidents
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
| `GET` | `/api/incidents` | List incidents. Filters: `?status=triggered\|acknowledged\|resolved`, `?severity=`, `?assigned_to=<user id>`, `?archived=true`, `?snoozed=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?sort=severity`, `?cluster=<value of the cluster group label>`, `?limit=` (default 50, max 500) |
|
|
| `GET` | `/api/incidents/clusters` | The distinct `cluster` values on the caller's incidents from the last 90 days, sorted (`?team_id=` narrows it). An empty array when nothing carries the label |
|
|
| `GET` | `/api/incidents/{id}` | Get single incident, with its alerts inline |
|
|
| `GET` | `/api/incidents/{id}/alerts` | Alerts under this incident |
|
|
| `GET` | `/api/incidents/{id}/timeline` | Full event history, chronological |
|
|
| `POST` | `/api/incidents/{id}/acknowledge` | Acknowledge (stamps authed user + time) |
|
|
| `DELETE` | `/api/incidents/{id}/acknowledge` | Clear acknowledgement, back to `triggered` |
|
|
| `POST` | `/api/incidents/{id}/resolve` | Close by hand — **terminal**, see above |
|
|
| `POST` | `/api/incidents/{id}/assign` | Reassign `{"user_id"}` |
|
|
| `POST` | `/api/incidents/{id}/snooze` | Hide until `{"until": RFC3339}` or `{"duration": "2h"}` |
|
|
| `DELETE` | `/api/incidents/{id}/snooze` | Un-snooze |
|
|
| `POST` | `/api/incidents/{id}/archive` | Archive (hides from the default list) |
|
|
| `DELETE` | `/api/incidents/{id}/archive` | Un-archive |
|
|
| `POST` | `/api/incidents/{id}/notes` | Add a note `{"content"}` |
|
|
| `DELETE` | `/api/incidents/{id}/notes/{eventID}` | Delete own note |
|
|
|
|
With no `?status=` filter, `GET /api/incidents` returns **open** incidents only —
|
|
the queue an on-call person wants. Currently snoozed and archived incidents are
|
|
excluded unless asked for. Actions that only make sense on an open incident
|
|
return `409` once it is resolved.
|
|
|
|
Notes are ordinary timeline events of type `note`; only they are deletable, and
|
|
only by their author. The rest of the timeline is a record of what happened.
|
|
|
|
### The incident object
|
|
|
|
| Field | Type | Notes |
|
|
|---|---|---|
|
|
| `id` | integer | Server-assigned |
|
|
| `group_key` | string | Alertmanager's `groupKey` — opaque, treat as an identifier |
|
|
| `title` | string | Rendered from `groupLabels` |
|
|
| `group_labels` | object | String→string, as sent by Alertmanager |
|
|
| `status` | string | `"triggered"`, `"acknowledged"` or `"resolved"` |
|
|
| `severity` | string | *optional* — high-water mark across the incident's alerts; never lowered |
|
|
| `triggered_at` | timestamp | When the incident opened |
|
|
| `acknowledged_by_id` / `acknowledged_by` / `acknowledged_at` | | *optional* — user id, username, time |
|
|
| `assigned_to_id` / `assigned_to` | | *optional* — user id, username |
|
|
| `snoozed_until` | timestamp | *optional* — a value in the past reads as not snoozed |
|
|
| `resolved_at` | timestamp | *optional* |
|
|
| `resolution_source` | string | *optional* — `"alerts"`, `"manual"` or `"recovered"` |
|
|
| `archived_at` | timestamp | *optional* |
|
|
| `alerts` | array | Only on `GET /api/incidents/{id}` |
|
|
|
|
Treat `resolution_source` as an open set, as with the alert field of the same
|
|
name: degrade unknown values to "resolved, reason unknown".
|
|
|
|
### The timeline event object
|
|
|
|
| Field | Type | Notes |
|
|
|---|---|---|
|
|
| `id` | integer | |
|
|
| `incident_id` | integer | |
|
|
| `type` | string | See below — treat as an open set |
|
|
| `user_id` / `username` | | *optional* — absent when the server acted rather than a person |
|
|
| `alert_id` | integer | *optional* — the alert an `alert_added` / `alert_resolved` event refers to |
|
|
| `detail` | string | *optional* — the note body, the snooze deadline, etc. |
|
|
| `created_at` | timestamp | |
|
|
|
|
Types written today: `triggered`, `alert_added`, `alert_resolved`,
|
|
`acknowledged`, `unacknowledged`, `assigned`, `archived`, `unarchived`, `snoozed`,
|
|
`unsnoozed`, `resolved`, `note`, `notified`, `notify_failed`, `deadman_silent`. On an
|
|
`assigned` event `user_id` is the **assignee**, not the actor; the actor is in
|
|
`actor_user_id`/`actor_username` or `actor_service_account_id`/`actor_service_account_name`
|
|
(absent on assignments made before they were recorded). New types may be added; render
|
|
unknown ones generically rather than dropping them.
|
|
|
|
On `notified` and `notify_failed`, `detail` carries the notification kind
|
|
(`triggered` | `reminder` | `resolved`), and on a failure the reason after it.
|
|
`user_id` is who was paged — absent means the page went to the shared fallback
|
|
topic and so belongs to nobody. The topic itself is never written to the
|
|
timeline: it is a shared secret with the ntfy server, and every API key can read
|
|
this.
|
|
|
|
## Alerts
|
|
|
|
Alerts are read-only. Everything a person does happens on the incident.
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
| `GET` | `/api/alerts` | List alerts. Filters: `?status=firing\|resolved`, `?name=`, `?incident_id=`, `?archived=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?limit=` (default 50, max 500) |
|
|
| `GET` | `/api/alerts/{id}` | Get single alert |
|
|
|
|
Archived alerts are hidden from `GET /api/alerts` unless `?archived=true` is
|
|
passed; alert archiving is automatic housekeeping by the sweeper, not a user
|
|
action. Resolved alerts carry `resolution_source`: `"alertmanager"` for a real
|
|
resolved webhook, `"expiry"` when the sweeper inferred it (see
|
|
[Stale alert expiry](./incidents.md#stale-alert-expiry)), `"deadman"` for a heartbeat declared
|
|
dead (see [Dead man's switch](./dead-mans-switch.md)).
|
|
|
|
### The alert object
|
|
|
|
Returned by `GET /api/alerts` (as an array) and `GET /api/alerts/{id}`.
|
|
Timestamps are RFC 3339 in UTC. Fields marked *optional* are omitted entirely
|
|
when unset, so clients must treat them as nullable.
|
|
|
|
| Field | Type | Notes |
|
|
|---|---|---|
|
|
| `id` | integer | Server-assigned; stable for the life of the row |
|
|
| `fingerprint` | string | Alertmanager's fingerprint — the upsert key |
|
|
| `name` | string | From the `alertname` label |
|
|
| `status` | string | `"firing"` or `"resolved"` |
|
|
| `labels` | object | String→string, as sent by Alertmanager |
|
|
| `annotations` | object | String→string, as sent by Alertmanager |
|
|
| `starts_at` | timestamp | When the alert instance began, **per Prometheus** |
|
|
| `ends_at` | timestamp | *optional* — absent while no end is known |
|
|
| `generator_url` | string | Link back to the originating Prometheus |
|
|
| `received_at` | timestamp | When the server last accepted a webhook for this alert — see below |
|
|
| `incident_id` | integer | *optional* — the most recent incident this alert belongs to |
|
|
| `resolution_source` | string | *optional* — `"alertmanager"`, `"expiry"` or `"deadman"` |
|
|
| `archived_at` | timestamp | *optional* — set while archived |
|
|
|
|
#### `received_at` is a liveness heartbeat
|
|
|
|
`starts_at` comes from Prometheus and **never changes** for the lifetime of an
|
|
alert instance. It says when the problem began, not whether it is still
|
|
happening — an alert that started twelve days ago looks identical whether
|
|
Alertmanager refreshed it a minute ago or went silent a week ago.
|
|
|
|
`received_at` is the field that answers "is this still live". It is set to the
|
|
server's clock on **every accepted webhook** for that fingerprint, including the
|
|
unchanged firing notifications Alertmanager re-sends every `repeat_interval`.
|
|
Clients may rely on this:
|
|
|
|
- **A firing alert whose `received_at` is advancing is still being refreshed.**
|
|
Stale-dating it against `repeat_interval` is a valid liveness check, and it is
|
|
what the built-in sweeper does (see
|
|
[Stale alert expiry](./incidents.md#stale-alert-expiry)).
|
|
- **`received_at` tracks accepted payloads, not delivery attempts.** A retry
|
|
that describes an older instance than the stored one is discarded, and a
|
|
discarded payload does not move `received_at`.
|
|
- **It stops advancing once the alert resolves,** because Alertmanager stops
|
|
re-sending. On an alert resolved by the sweeper
|
|
(`"resolution_source": "expiry"`) it therefore marks the last time
|
|
Alertmanager was actually heard from, which is earlier than `ends_at`.
|
|
|
|
`GET /api/alerts` is ordered by `received_at` descending — most recently
|
|
refreshed first — and the `?from=` / `?to=` filters on both the alert and stats
|
|
endpoints select on `received_at`, not `starts_at`.
|
|
|
|
#### `resolution_source` says how much to trust `ends_at`
|
|
|
|
An alert can leave the firing state two ways, and `resolution_source` records
|
|
which happened. Clients may rely on this:
|
|
|
|
- **Absent while firing.** It is set only on resolve, and a re-fire under the
|
|
same fingerprint clears it again, so its presence always agrees with
|
|
`"status": "resolved"`.
|
|
- **`"alertmanager"` — a real resolved webhook arrived.** `ends_at` is the end
|
|
time Alertmanager reported. It is an observed value and can be displayed as
|
|
fact.
|
|
- **`"expiry"` — the sweeper inferred the resolve** because Alertmanager stopped
|
|
refreshing the alert (see [Stale alert expiry](./incidents.md#stale-alert-expiry)). Nothing
|
|
ever reported an end, so **`ends_at` is approximate**: it is either the stale
|
|
`endsAt` watermark from the last notification, or — when that notification
|
|
carried none — the time the sweep ran, which lags the last real contact by up
|
|
to `TERDUT_STALE_AFTER` plus a sweep interval. Treat it as "no later than",
|
|
not as when the problem stopped.
|
|
|
|
On these alerts `received_at` is the more truthful signal: it marks the last
|
|
time Alertmanager was actually heard from. Surfacing the distinction is
|
|
worthwhile, since `"expiry"` can also mean the alert is still firing and the
|
|
notification path broke.
|
|
|
|
- **`"deadman"` — a heartbeat was declared dead** (see
|
|
[Dead man's switch](./dead-mans-switch.md)). Like `"expiry"`, an inference from
|
|
silence rather than an observed end, so `ends_at` is approximate — but a much
|
|
tighter one, bounded by the switch's timeout. It is also the one resolution
|
|
a re-fire under the same `starts_at` can undo, since the switch coming back is
|
|
exactly the evidence that the inference was wrong.
|
|
|
|
Treat the value as an open set and tolerate ones you do not recognise — new
|
|
sources may be added, and unknown values should degrade to "resolved, reason
|
|
unknown" rather than being rejected.
|
|
|
|
## On-call schedule
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
Each team keeps its own rota, so two teams can have two different people on call
|
|
on the same day. The person taking a shift has to be in the team — paging
|
|
somebody who cannot open the incident is worse than paging nobody.
|
|
|
|
| Method | Path | Who | Description |
|
|
|---|---|---|---|
|
|
| `POST` | `/api/teams/{teamID}/schedule` | **owner** | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...], "replace"}` — all-or-nothing |
|
|
| `GET` | `/api/teams/{teamID}/schedule` | member | List entries. Filters: `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD` |
|
|
| `DELETE` | `/api/teams/{teamID}/schedule/{id}` | **owner** | Remove schedule entry |
|
|
| `GET` | `/api/schedule/current` | any | Who is on call today (UTC) in **every** team the caller is in — one entry per team, `[]` when nobody anywhere |
|
|
|
|
## Statistics
|
|
|
|
Every figure counts the caller's own teams only: a report that counted other
|
|
teams' incidents would leak their volume, and their alert names through the
|
|
top-alerts list, and would not be a number about the reader's work anyway.
|
|
|
|
All stat endpoints accept optional `?from=YYYY-MM-DD` and `?to=YYYY-MM-DD`, and exclude archived rows to match the default list views. Alert stats filter on `received_at`; incident stats filter on `triggered_at`.
|
|
|
|
| Method | Path | Description |
|
|
|---|---|---|
|
|
| `GET` | `/api/stats/incidents` | `{total, triggered, acknowledged, resolved, mtta_seconds, mttr_seconds}` |
|
|
| `GET` | `/api/stats/alerts` | `{total, firing, resolved}` counts |
|
|
| `GET` | `/api/stats/alerts/top` | Most frequent alert names. `?limit=` (default 10, max 100) |
|
|
| `GET` | `/api/stats/alerts/by-hour` | Count per hour-of-day (UTC), all 24 slots returned |
|
|
| `GET` | `/api/stats/alerts/by-day` | Count per day-of-week, all 7 slots with names returned |
|
|
|
|
`mtta_seconds` (time to acknowledge) and `mttr_seconds` (time to resolve) are
|
|
averages over incidents that have actually been acknowledged or resolved, and are
|
|
**null** until there are any — null means "no data", not zero.
|