Files
terdut-server/docs/api.md
T
Niklas Ye 44b2eb2cc3 Rewrite the README as highlights with screenshots; move the detail into docs/
The README was 1,240 lines of reference material and still described a
SQLite quick start. It is now a short tour (highlights, screenshots of the
web UI, an accurate quick start against Postgres), and each topic has its
own page under docs/ with an index: deployment, configuration, Alertmanager,
incidents, notifications, escalation, dead man's switches, single sign-on,
web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a
proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it
described. The "Upgrading to ..." sections for an unreleased product are
dropped.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:13 +02:00

30 KiB

API reference

The REST API. Back to the README and the documentation index.

Authentication

All endpoints except /api/bootstrap, /api/integrations/{key}/alertmanager, /api/notify/ack/{token}, /api/login, /api/logout, /api/auth/config, /api/version, /api/oidc/login, /api/oidc/callback, /api/oidc/device and /api/oidc/device/token require either an API key:

Authorization: Bearer <api-key>

or the web UI's session cookie. A request that carries an Authorization header is judged on that header alone.

Two kinds of user exist. An administrator manages accounts: creating and deleting users, setting anybody's password, minting keys for anybody, and granting the flag itself. Everybody else works incidents — acknowledging, assigning, snoozing, resolving, noting — and manages their own account and nobody else's. An API key carries exactly the rights of the user it belongs to.

A third principal, the service account, exists for automation (a Kubernetes operator, most likely) that needs to manage teams, escalation policies, dead man's switches and integrations without impersonating a human. It is not a user — it never signs in, never appears in a team's member list, and never holds the administrator flag — and its key is prefixed tdsa_ so it reads as one at a glance in a log line. See Service accounts.

Getting an account. The first one comes from /api/bootstrap. After that it depends on signup_mode, an administrator setting:

  • invite_only (the default) — a team owner mints a link with POST /api/teams/{teamID}/invites, and the person who opens it picks a username and password and lands in that team with the role the link carries. Links are single-use unless told otherwise, expire after seven days, and can be revoked before that.
  • open — anybody who can reach the server can create an account, and must name a team, which they then own.

Invites are links, not email: this server has no SMTP, and adding it to send one message would be a subsystem to run, secure and monitor. Send the link however you already talk to the person.

A domain-restricted third mode was considered and dropped: with no email there is nothing to verify an address against, so it would only check the domain of a string somebody typed.

The first user, from /api/bootstrap, is an administrator. Users created afterwards are not, until an administrator says so. An install always keeps at least one: the last administrator can be neither deleted nor demoted, and nobody can delete or demote themselves.

Endpoints that require the flag answer 403 with {"error":"administrator access required"}.

Teams are the unit of tenancy, and are a separate axis from the administrator flag. A team owns its incidents, alerts, schedule and integrations, and a user sees exactly the teams they belong to. Within a team an owner configures it (schedule, integrations, membership) and a member works its incidents.

An administrator crosses that line in one direction only. They configure any team without being in it — every owner-only endpoint accepts the flag, because otherwise a team whose last owner left could never be repaired. They do not read any team: the queue, the alerts and the incidents are filtered by real membership, so an administrator sees a team's work only by joining it, which is a membership change and shows up as one. Administration is about accounts and the shape of a team, not about reading other people's incidents.

Anything belonging to a team you are not in answers 404, not 403: whether an incident exists is itself something only its team should learn.

Operator mode (TERDUT_OPERATOR_MODE, see Configuration) declares this install gitops-managed. When it is on, a session or a user's own API key gets 403 {"error": "...", "reason": "operator_managed"} on every write this page marks owner-gated under Teams below (creating, renaming or deleting a team; its OIDC group binding; its escalation ladder; its dead man's switches; its integrations) — a service account's writes are unaffected. Team membership and invites are deliberately excluded: they are never gitops-managed, in operator mode or out of it. GET /api/auth/config reports operator_mode so a client can grey those sections out before a write is ever attempted.

Method Path Description
GET /api/auth/config How to sign in: {"password_login", "oidc": {"enabled","name"}, "device_login", "operator_mode"}. No session needed
GET /api/version {"version"} — this build's version string. No session needed, the same as /healthz
POST /api/login {"username","password"} → sets the session cookie, returns {user, has_password}. 429 after too many failures; 403 when TERDUT_PASSWORD_LOGIN=false
GET /api/oidc/login Starts a single sign-on sign-in: redirects the browser to the provider. ?next=/path is where to land afterwards; only a path on this server is honoured. Only exists when SSO is configured
POST /api/oidc/device Starts a device login: returns {device_code, user_code, verification_url, interval, expires_in}. Only exists when SSO is configured
POST /api/oidc/device/token {"device_code"} → 202 {"status":"pending"}, then 200 with the session cookie once approved (once only). 410 with {"error":"expired"} or {"error":"denied"}; 429 {"error":"slow_down"} if polled faster than interval
POST /api/oidc/device/approve session — {"user_code"}. Approves a pending device login as the caller. 403 for an API key; 404 for an unknown, expired or already decided code
POST /api/oidc/device/deny session — {"user_code"}. Refuses it
GET /api/oidc/callback Where the provider sends the browser back. Sets the session cookie and redirects to /, or to /?sso_error=<code> — one of denied, expired, failed, unavailable, not_allowed, no_email, email_conflict, disabled, not_bootstrapped (no user exists on this install yet — sign in again once something has called /api/bootstrap)
POST /api/logout Ends the session and clears the cookie
GET /api/me The caller: {user, has_password}

Users

Method Path Description
admin marks an endpoint that requires the administrator flag; **self or
admin** marks one you may use on your own account and an administrator may use
on anybody's.
Method Path Who Description
GET /api/signup — Whether sign-up is open, and whether ?invite= is usable. No session needed: the caller has no account yet
POST /api/signup — Create an account {"username","email","password","invite"?,"team_name"?} and sign in. 403 without a usable invite when the mode is invite-only
POST /api/bootstrap — Create first user + API key {"username","email","password"?} (only works on empty DB). The user is an administrator
GET /api/users any List users. Open to everybody: the queue's assignment control and the schedule both have to name people
GET /api/users/{id}/teams self or admin The teams that user is in, each with their role. /api/teams is always about the caller; this one answers it about somebody else, for the admin page's per-user view. 404 for a user who does not exist, so "no teams" and "no such person" are distinguishable
POST /api/users admin Create user {"username","email"}. Not an administrator
DELETE /api/users/{id} admin Delete user (cascades to keys). 409 for yourself or the last administrator
PUT /api/users/{id}/admin admin Grant or revoke the administrator flag {"is_admin"}. 409 for yourself, the last administrator, or an administrator granted by single sign-on
PUT /api/users/{id}/disabled admin Take an account out of use, or put it back {"disabled"}. 409 for yourself or the last administrator
PUT /api/users/{id}/notify self or admin Set push notification target {"ntfy_topic"} — empty string clears it
PUT /api/users/{id}/password self or admin Set web UI password {"password","current_password"}. current_password is required only when changing your own existing password. Ends the user's other sessions
POST /api/users/{id}/api-keys self or admin Issue API key {"name"} — key shown once
DELETE /api/users/{id}/api-keys/{keyID} self or admin Revoke API key

Administration

Method Path Who Description
GET /api/admin/teams admin Every team on the server, with its member and open-incident counts. /api/teams answers "what am I in"; this answers "what is there"
GET /api/admin/teams/{teamID} admin One team and who is in it: {"team", "members"}. 404 for a team that does not exist. GET /api/teams/{teamID}/members is member-only and still 404s an administrator from outside the team — reading a team's shape and reading its work are different questions, so they are different endpoints
GET /api/admin/settings admin The editable settings with their bounds, plus the environment-configured ones, read-only. Never credentials
PUT /api/admin/settings admin Change one or more {"key": seconds}, or {"signup_mode": "open"|"invite_only"}. 400 for an unknown key or a value outside its bounds

Service accounts

A service account is a scoped, non-human credential for automation — not a users row, so it never signs in, is never a team member, and never carries the administrator flag. Two scopes:

  • instance — the same reach system administration has over teams: create one, and mint a team-scoped account against any of them. There is no cap on how many instance-scoped accounts exist, but ordinarily there is one, belonging to whatever is provisioning this install end to end.
  • team — owner-equivalent for that one team, and nothing else: every owner-gated endpoint under Teams, membership and invites included. Nothing narrower is enforced server-side; what actually keeps membership out of automation's hands is that no operator built against this scope should ever call those two endpoints — see operator mode and SERVICE-ACCOUNTS.md's note on this.

A key is shown once, at creation or rotation, and only its hash is stored — the same handling as a user's API key. Losing it means minting a new one; there is no way to recover a raw key from the server.

Method Path Who Description
GET /api/service-accounts admin Every service account. Pass ?name= instead to look one up by its exact name — open to any authenticated caller (human or service account), since it returns no key material and is how an account finds its own id
POST /api/service-accounts owner* Create one and mint its first key {"name","scope","team_id"?} (team_id required for scope:"team", absent for scope:"instance"). Returns {"service_account", "key"} — key.key shown once
POST /api/service-accounts/{id}/keys owner* Mint an additional key {"name"} — rotation without recreating the account. Shown once
DELETE /api/service-accounts/{id}/keys/{keyID} owner* Revoke one key

* For an instance-scoped account: a system administrator only. For a team-scoped account: a system administrator, that team's own human owner, an instance-scoped service account (minting a narrower credential for a team it just created), or — for the two key endpoints only — the account rotating or revoking its own key, which is not a privilege escalation, the same reasoning a user's own API keys rest on.

Alert ingestion

Alerts arrive on a team's integration key. The key is both the credential and the routing: it says that the sender may post, and which team the alerts belong to. Create one with POST /api/teams/{teamID}/integrations, which returns the key and the full URL once and stores only a SHA-256 hash.

Method Path Description
POST /api/integrations/{key}/alertmanager Alertmanager v4 webhook receiver for the key's team. 401 for an unknown key

This is the only way in. The pre-teams POST /api/alertmanager/webhook took no credential at all — anything able to reach the port could open an incident — and was removed in v0.13.0 once senders had moved onto keys.

Teams

owner below means an owner of that team, a system administrator (who passes every one of these without being a member), or that team's own team-scoped service account — including membership and invites, technically, though no automation this scope was designed for (a Kubernetes operator's CRDs, see SERVICE-ACCOUNTS.md) ever models team membership or would call those two. See Authentication. member means membership and nothing else: an administrator who is not in the team gets the same 404 as anybody else.

Method Path Who Description
GET /api/teams any The caller's own teams, each with their role
POST /api/teams any Create a team {"name"}; a human creator becomes its first owner. An instance-scoped service account may also create one, and it gets no owner at all — expected for a team an operator is about to hand a team-scoped credential to, not an orphaned team a human made
PUT /api/teams/{teamID} owner Rename it {"name"}. 409 if the name is taken
DELETE /api/teams/{teamID} owner Delete a team and everything under it. 409 while it has open incidents
GET /api/teams/{teamID}/members member Who is in the team, with status (oncall if the rota has them today, unpageable when a page to them would go nowhere — even if they are on call — else reachable), on_call, next_shift (first rota day after today), pageable and problem (has no ntfy topic / account is disabled; never the topic itself) and last_active_at (their newest session or API-key use). Every member sees the same list
POST /api/teams/{teamID}/members owner Add a member, or change their role {"user_id","role"}. 409 when it would demote the last owner, or the membership is managed by single sign-on
DELETE /api/teams/{teamID}/members/{userID} owner Remove a member. 409 for the last owner, or a membership managed by single sign-on
GET /api/teams/{teamID}/oidc-groups member Which groups control this team's membership: {"member_group","owner_group"}. An empty string means no group grants that role here
PUT /api/teams/{teamID}/oidc-groups owner Set them. An empty string clears a binding
GET /api/teams/{teamID}/integrations member List integrations. Never returns keys. Each carries status (active if its key posted within 24h, quiet if it has but not lately, never), last_used_at (last webhook, usable or not), last_alert_at (when an alert last arrived on it) and alerts_24h (distinct alerts it refreshed in the last day). Alerts delivered before the source was recorded (migration 010) have none, so the last two fill in as Alertmanager re-sends them
PATCH /api/teams/{teamID}/integrations/{integrationID} owner Rename {"name"}. The key does not change
POST /api/teams/{teamID}/integrations owner Mint an integration {"name","kind"} — key and URL shown once
DELETE /api/teams/{teamID}/integrations/{integrationID} owner Revoke an integration. Alerts it delivered stay, unattributed
GET /api/teams/{teamID}/invites owner The team's invite links, with their uses and expiry. Never the tokens
POST /api/teams/{teamID}/invites owner Mint one {"role","max_uses"} — the full URL is returned once
DELETE /api/teams/{teamID}/invites/{inviteID} owner Revoke a link before it expires
GET /api/teams/{teamID}/escalation member The team's escalation ladder {repeat_count, fallback_topic, levels[], last_escalated_at?, last_escalated_incident_id?}. Empty levels means the team has none. Each level also carries status (ready, escalating when an unanswered incident has climbed to it, unreachable when nobody on it could be woken), waiting (ids of the open incidents on it) and, per target, username (who it means today — the person on call, for a rota target), reachable and problem. The extra fields are output only; PUT takes the plain shape
PUT /api/teams/{teamID}/escalation owner Replace it wholesale. 400 for a level with no targets or no timeout — a rung that pages nobody is a silence with a number on it
GET /api/teams/{teamID}/deadman/switches member The team's dead man's switches, each {id, name, matcher, timeout_seconds, severity, status, last_heartbeat_at, last_triggered_at, open_incident_id, sources[]}. status is healthy, dead or dormant; sources has one entry per heartbeat fingerprint. Empty when the team watches nothing
POST /api/teams/{teamID}/deadman/switches owner Add one: {name?, matcher, timeout_seconds, severity?}. 400 when the matcher names no alertname or holds several, or the timeout is not positive — a switch that silently watches nothing is the failure this feature exists to prevent
PUT /api/teams/{teamID}/deadman/switches/{switchID} owner Replace one in place, same body and validation as create. Its id is unchanged — for an automated caller reconciling a spec change, unlike delete-and-recreate
DELETE /api/teams/{teamID}/deadman/switches/{switchID} owner Stop watching. An incident it opened stays open. 404 for a switch of another team

Notifications

Method Path Description
POST /api/notify/ack/{token} Acknowledge an incident from a push notification's Acknowledge button. No auth: the token in the path is the credential — one incident, one action, 24 hours, idempotent. Must stay publicly reachable

Incidents

Method Path Description
GET /api/incidents List incidents. Filters: ?status=triggered|acknowledged|resolved, ?severity=, ?assigned_to=<user id>, ?archived=true, ?snoozed=true, ?from=YYYY-MM-DD, ?to=YYYY-MM-DD, ?sort=severity, ?cluster=<value of the cluster group label>, ?limit= (default 50, max 500)
GET /api/incidents/clusters The distinct cluster values on the caller's incidents from the last 90 days, sorted (?team_id= narrows it). An empty array when nothing carries the label
GET /api/incidents/{id} Get single incident, with its alerts inline
GET /api/incidents/{id}/alerts Alerts under this incident
GET /api/incidents/{id}/timeline Full event history, chronological
POST /api/incidents/{id}/acknowledge Acknowledge (stamps authed user + time)
DELETE /api/incidents/{id}/acknowledge Clear acknowledgement, back to triggered
POST /api/incidents/{id}/resolve Close by hand — terminal, see above
POST /api/incidents/{id}/assign Reassign {"user_id"}
POST /api/incidents/{id}/snooze Hide until {"until": RFC3339} or {"duration": "2h"}
DELETE /api/incidents/{id}/snooze Un-snooze
POST /api/incidents/{id}/archive Archive (hides from the default list)
DELETE /api/incidents/{id}/archive Un-archive
POST /api/incidents/{id}/notes Add a note {"content"}
DELETE /api/incidents/{id}/notes/{eventID} Delete own note

With no ?status= filter, GET /api/incidents returns open incidents only — the queue an on-call person wants. Currently snoozed and archived incidents are excluded unless asked for. Actions that only make sense on an open incident return 409 once it is resolved.

Notes are ordinary timeline events of type note; only they are deletable, and only by their author. The rest of the timeline is a record of what happened.

The incident object

Field Type Notes
id integer Server-assigned
group_key string Alertmanager's groupKey — opaque, treat as an identifier
title string Rendered from groupLabels
group_labels object String→string, as sent by Alertmanager
status string "triggered", "acknowledged" or "resolved"
severity string optional — high-water mark across the incident's alerts; never lowered
triggered_at timestamp When the incident opened
acknowledged_by_id / acknowledged_by / acknowledged_at optional — user id, username, time
assigned_to_id / assigned_to optional — user id, username
snoozed_until timestamp optional — a value in the past reads as not snoozed
resolved_at timestamp optional
resolution_source string optional — "alerts", "manual" or "recovered"
archived_at timestamp optional
alerts array Only on GET /api/incidents/{id}

Treat resolution_source as an open set, as with the alert field of the same name: degrade unknown values to "resolved, reason unknown".

The timeline event object

Field Type Notes
id integer
incident_id integer
type string See below — treat as an open set
user_id / username optional — absent when the server acted rather than a person
alert_id integer optional — the alert an alert_added / alert_resolved event refers to
detail string optional — the note body, the snooze deadline, etc.
created_at timestamp

Types written today: triggered, alert_added, alert_resolved, acknowledged, unacknowledged, assigned, archived, unarchived, snoozed, unsnoozed, resolved, note, notified, notify_failed, deadman_silent. On an assigned event user_id is the assignee, not the actor; the actor is in actor_user_id/actor_username or actor_service_account_id/actor_service_account_name (absent on assignments made before they were recorded). New types may be added; render unknown ones generically rather than dropping them.

On notified and notify_failed, detail carries the notification kind (triggered | reminder | resolved), and on a failure the reason after it. user_id is who was paged — absent means the page went to the shared fallback topic and so belongs to nobody. The topic itself is never written to the timeline: it is a shared secret with the ntfy server, and every API key can read this.

Alerts

Alerts are read-only. Everything a person does happens on the incident.

Method Path Description
GET /api/alerts List alerts. Filters: ?status=firing|resolved, ?name=, ?incident_id=, ?archived=true, ?from=YYYY-MM-DD, ?to=YYYY-MM-DD, ?limit= (default 50, max 500)
GET /api/alerts/{id} Get single alert

Archived alerts are hidden from GET /api/alerts unless ?archived=true is passed; alert archiving is automatic housekeeping by the sweeper, not a user action. Resolved alerts carry resolution_source: "alertmanager" for a real resolved webhook, "expiry" when the sweeper inferred it (see Stale alert expiry), "deadman" for a heartbeat declared dead (see Dead man's switch).

The alert object

Returned by GET /api/alerts (as an array) and GET /api/alerts/{id}. Timestamps are RFC 3339 in UTC. Fields marked optional are omitted entirely when unset, so clients must treat them as nullable.

Field Type Notes
id integer Server-assigned; stable for the life of the row
fingerprint string Alertmanager's fingerprint — the upsert key
name string From the alertname label
status string "firing" or "resolved"
labels object String→string, as sent by Alertmanager
annotations object String→string, as sent by Alertmanager
starts_at timestamp When the alert instance began, per Prometheus
ends_at timestamp optional — absent while no end is known
generator_url string Link back to the originating Prometheus
received_at timestamp When the server last accepted a webhook for this alert — see below
incident_id integer optional — the most recent incident this alert belongs to
resolution_source string optional — "alertmanager", "expiry" or "deadman"
archived_at timestamp optional — set while archived

received_at is a liveness heartbeat

starts_at comes from Prometheus and never changes for the lifetime of an alert instance. It says when the problem began, not whether it is still happening — an alert that started twelve days ago looks identical whether Alertmanager refreshed it a minute ago or went silent a week ago.

received_at is the field that answers "is this still live". It is set to the server's clock on every accepted webhook for that fingerprint, including the unchanged firing notifications Alertmanager re-sends every repeat_interval. Clients may rely on this:

  • A firing alert whose received_at is advancing is still being refreshed. Stale-dating it against repeat_interval is a valid liveness check, and it is what the built-in sweeper does (see Stale alert expiry).
  • received_at tracks accepted payloads, not delivery attempts. A retry that describes an older instance than the stored one is discarded, and a discarded payload does not move received_at.
  • It stops advancing once the alert resolves, because Alertmanager stops re-sending. On an alert resolved by the sweeper ("resolution_source": "expiry") it therefore marks the last time Alertmanager was actually heard from, which is earlier than ends_at.

GET /api/alerts is ordered by received_at descending — most recently refreshed first — and the ?from= / ?to= filters on both the alert and stats endpoints select on received_at, not starts_at.

resolution_source says how much to trust ends_at

An alert can leave the firing state two ways, and resolution_source records which happened. Clients may rely on this:

  • Absent while firing. It is set only on resolve, and a re-fire under the same fingerprint clears it again, so its presence always agrees with "status": "resolved".

  • "alertmanager" — a real resolved webhook arrived. ends_at is the end time Alertmanager reported. It is an observed value and can be displayed as fact.

  • "expiry" — the sweeper inferred the resolve because Alertmanager stopped refreshing the alert (see Stale alert expiry). Nothing ever reported an end, so ends_at is approximate: it is either the stale endsAt watermark from the last notification, or — when that notification carried none — the time the sweep ran, which lags the last real contact by up to TERDUT_STALE_AFTER plus a sweep interval. Treat it as "no later than", not as when the problem stopped.

    On these alerts received_at is the more truthful signal: it marks the last time Alertmanager was actually heard from. Surfacing the distinction is worthwhile, since "expiry" can also mean the alert is still firing and the notification path broke.

  • "deadman" — a heartbeat was declared dead (see Dead man's switch). Like "expiry", an inference from silence rather than an observed end, so ends_at is approximate — but a much tighter one, bounded by the switch's timeout. It is also the one resolution a re-fire under the same starts_at can undo, since the switch coming back is exactly the evidence that the inference was wrong.

Treat the value as an open set and tolerate ones you do not recognise — new sources may be added, and unknown values should degrade to "resolved, reason unknown" rather than being rejected.

On-call schedule

Method Path Description
Each team keeps its own rota, so two teams can have two different people on call
on the same day. The person taking a shift has to be in the team — paging
somebody who cannot open the incident is worse than paging nobody.
Method Path Who Description
POST /api/teams/{teamID}/schedule owner Assign user to dates {"user_id", "dates":["YYYY-MM-DD",...], "replace"} — all-or-nothing
GET /api/teams/{teamID}/schedule member List entries. Filters: ?from=YYYY-MM-DD, ?to=YYYY-MM-DD
DELETE /api/teams/{teamID}/schedule/{id} owner Remove schedule entry
GET /api/schedule/current any Who is on call today (UTC) in every team the caller is in — one entry per team, [] when nobody anywhere

Statistics

Every figure counts the caller's own teams only: a report that counted other teams' incidents would leak their volume, and their alert names through the top-alerts list, and would not be a number about the reader's work anyway.

All stat endpoints accept optional ?from=YYYY-MM-DD and ?to=YYYY-MM-DD, and exclude archived rows to match the default list views. Alert stats filter on received_at; incident stats filter on triggered_at.

Method Path Description
GET /api/stats/incidents {total, triggered, acknowledged, resolved, mtta_seconds, mttr_seconds}
GET /api/stats/alerts {total, firing, resolved} counts
GET /api/stats/alerts/top Most frequent alert names. ?limit= (default 10, max 100)
GET /api/stats/alerts/by-hour Count per hour-of-day (UTC), all 24 slots returned
GET /api/stats/alerts/by-day Count per day-of-week, all 7 slots with names returned

mtta_seconds (time to acknowledge) and mttr_seconds (time to resolve) are averages over incidents that have actually been acknowledged or resolved, and are null until there are any — null means "no data", not zero.