Rewrite the README as highlights with screenshots; move the detail into docs/
The README was 1,240 lines of reference material and still described a SQLite quick start. It is now a short tour (highlights, screenshots of the web UI, an accurate quick start against Postgres), and each topic has its own page under docs/ with an index: deployment, configuration, Alertmanager, incidents, notifications, escalation, dead man's switches, single sign-on, web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it described. The "Upgrading to ..." sections for an unreleased product are dropped. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
@@ -12,7 +12,7 @@ Preconditions and the plan, without side effects:
|
||||
```
|
||||
|
||||
Config is `.release.conf` here plus `make release-vars`. The process itself lives in
|
||||
`~/.claude/skills/release/`; why it is shaped this way is in README.md §Releasing.
|
||||
`~/.claude/skills/release/`; why it is shaped this way is in docs/development.md (Releasing).
|
||||
|
||||
Three things about this repo specifically:
|
||||
|
||||
|
||||
@@ -1,263 +1,60 @@
|
||||
# Service accounts: a scoped, non-human credential type
|
||||
# Service accounts
|
||||
|
||||
This is a design note for a feature, not an implementation plan — it exists to
|
||||
propose the shape before writing code. It's raised directly by `terdut-operator`
|
||||
(a separate repo, no shared code — see its `DESIGN.md` §6, §9, §13), which needs
|
||||
a credential for unattended, repeatable API access and currently has no good one
|
||||
available. Anything automating terdut-server long-term (this operator, CI, future
|
||||
integrations) hits the same gap, so this is written as a general primitive, not
|
||||
operator-specific.
|
||||
A non-human credential for automation (terdut-operator, CI, scripts). It is not a
|
||||
`users` row: no password, no `is_admin`, no OIDC identity, so it can never be
|
||||
pulled into login or group sync, and it is never mistaken for a person in an
|
||||
audit trail. The bearer token has the same shape as an API key (SHA-256 hash
|
||||
stored, raw value shown once), prefixed `tdsa_`.
|
||||
|
||||
## The problem
|
||||
## Scopes
|
||||
|
||||
terdut-server has two credential types today, and neither fits "an unattended
|
||||
process that manages teams/schedules/policies on someone's behalf":
|
||||
- **instance** — acts as owner of every team's *configuration* (rename, OIDC
|
||||
groups, escalation, dead man's switches, integrations, members, delete) and may
|
||||
create teams. It is not a member of any team, so it reads no incidents or
|
||||
queue. It is never an administrator: user management and
|
||||
`/api/admin/settings` stay human-only.
|
||||
- **team** — acts as owner of exactly one team, through a single synthetic
|
||||
membership. It may also mint another service account for its own team.
|
||||
|
||||
- **User API keys** (`api_keys`, `internal/api/users.go`) are always tied to a
|
||||
real `users` row and carry that user's full rights — every team they're a
|
||||
member of, their admin flag if set. There's no `kind`/`service` marker
|
||||
distinguishing "a human's personal automation key" from "a login session," and
|
||||
no way to mint one scoped to less than the full user.
|
||||
- **Integration keys** (`integrations`, `internal/api/*teams*.go`) are team-scoped,
|
||||
but narrowly: they authenticate exactly one inbound Alertmanager webhook call
|
||||
(`POST /api/integrations/{key}/alertmanager`) and nothing else. They're not a
|
||||
general management-API credential and shouldn't become one — overloading a
|
||||
narrow, one-way ingestion credential with broad read/write access would weaken
|
||||
the one property that makes it safe to embed in an Alertmanager config today.
|
||||
An account has many keys, so rotating is "mint a new key, revoke the old one"
|
||||
without losing the account's identity or history.
|
||||
|
||||
The result: any automation that needs to create teams, set escalation policies,
|
||||
manage dead-man switches, or rotate integration keys has to hold a real human
|
||||
admin's or team owner's API key. That key is exactly as powerful as that person
|
||||
logging in — full team access, and full instance access if they're an admin.
|
||||
`terdut-operator`'s design ran directly into this (its DESIGN.md §6): its
|
||||
described bootstrap/rotation flow assumed a repeatable, identity-scoped way to
|
||||
get a credential, and `/api/bootstrap`'s actual behavior (single-shot per
|
||||
install, gated on `COUNT(*) FROM users`, confirmed via `internal/api/users.go`
|
||||
and `charts/terdut-server/templates/bootstrap-job.yaml`) doesn't provide one —
|
||||
it mints exactly one founding admin, once, ever.
|
||||
## Endpoints
|
||||
|
||||
## Goals
|
||||
- `POST /api/service-accounts` `{name, scope, team_id}` — returns the account and
|
||||
its first key. An instance-scoped account is granted by a human administrator;
|
||||
a team-scoped one by an administrator, that team's owner, or an instance-scoped
|
||||
account.
|
||||
- `GET /api/service-accounts?name=` — look one up by name.
|
||||
- `POST /api/service-accounts/{id}/keys`, `DELETE .../keys/{keyID}` — mint or
|
||||
revoke a key. An instance-scoped account may manage any team-scoped account's
|
||||
keys, and any account may manage its own.
|
||||
|
||||
- A credential type that isn't a human: doesn't touch OIDC group sync, login,
|
||||
session, or the `is_admin`/account-management semantics that come with a real
|
||||
`users` row.
|
||||
- Two scopes matching the two shapes automation actually needs: instance-wide
|
||||
(create/list teams — what a server-owning controller needs) and team-scoped
|
||||
(manage one team's escalation policy, dead-man switches, integrations,
|
||||
schedule, OIDC group bindings — what a per-team controller or integration
|
||||
needs).
|
||||
- Repeatable issuance and rotation — unlike `/api/bootstrap`, callable more than
|
||||
once, by anything that already holds admin rights, without destroying and
|
||||
recreating state to get a fresh credential.
|
||||
- Visibly distinct from a human in every place identity shows up (audit trails,
|
||||
timeline entries, UI attribution) — a service account acting on a team should
|
||||
never be indistinguishable from a person.
|
||||
## Seeding the operator's account
|
||||
|
||||
## Non-goals
|
||||
`TERDUT_OPERATOR_KEY` (at least 32 characters) creates the instance-scoped account
|
||||
`terdut-operator` if missing and replaces its `seed` key with this value at every
|
||||
start (`internal/api/operator_key.go`). The deployer generates the key and
|
||||
nothing has to call `/api/bootstrap` for it; rotating is a restart with a new
|
||||
value. With `TERDUT_OPERATOR_MODE` on, configuration writes by humans are refused
|
||||
and a service account of either scope passes.
|
||||
|
||||
- Not a general OAuth2/OIDC client-credentials flow — this is a bearer-token
|
||||
primitive matching the shape `api_keys` already uses (SHA-256 hash stored,
|
||||
raw key shown once at creation), not a new auth protocol.
|
||||
- Not replacing integration keys — those stay as the narrow, one-way webhook
|
||||
credential they are today.
|
||||
- Not modeling per-endpoint or per-verb permissions within a scope — `instance`
|
||||
and `team` are the only two scopes for now; finer-grained scoping is future
|
||||
work if a real need shows up.
|
||||
## How it is enforced
|
||||
|
||||
## Proposed shape
|
||||
Every request resolves to one `Caller` (`internal/api/caller.go`): a human
|
||||
(session or API key) or a service account.
|
||||
|
||||
### Schema
|
||||
- `Caller.IsAdmin()` is true only for a human administrator. `AdminOnly` and
|
||||
`requireSelfOrAdmin` key on it alone; do not widen them — each time a gap came up
|
||||
the fix was a narrower purpose-built capability instead.
|
||||
- `Caller.IsInstanceServiceAccount()` is true only for an instance-scoped account,
|
||||
never for a human. `requireTeamOwner` and `callerOwnsTeam` admit it for any team.
|
||||
- `Caller.Role(teamID)`/`TeamIDs()` are a human's memberships or a team-scoped
|
||||
account's single owner membership; instance scope has none.
|
||||
- `Caller.AsHuman()` is what a handler must call when it needs a real `user_id`;
|
||||
handlers meant for people answer 403 to a service account instead of writing a
|
||||
zero id.
|
||||
|
||||
```sql
|
||||
CREATE TABLE service_accounts (
|
||||
id BIGSERIAL PRIMARY KEY,
|
||||
name TEXT NOT NULL UNIQUE, -- e.g. "terdut-operator"
|
||||
scope TEXT NOT NULL CHECK (scope IN ('instance', 'team')),
|
||||
team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE,
|
||||
-- team_id required iff scope = 'team'; NULL iff scope = 'instance'
|
||||
created_by BIGINT REFERENCES users(id),
|
||||
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
|
||||
);
|
||||
|
||||
CREATE TABLE service_account_keys (
|
||||
id BIGSERIAL PRIMARY KEY,
|
||||
service_account_id BIGINT NOT NULL REFERENCES service_accounts(id) ON DELETE CASCADE,
|
||||
key_hash TEXT NOT NULL UNIQUE,
|
||||
name TEXT NOT NULL, -- e.g. "initial", "2026-Q4-rotation"
|
||||
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
|
||||
last_used_at TIMESTAMPTZ
|
||||
);
|
||||
```
|
||||
|
||||
Deliberately not a `users` row: no `password_hash`, no `is_admin`, no
|
||||
`user_identities` linkage, so it's structurally impossible for a service account
|
||||
to be pulled into OIDC group sync or password login. Multiple keys per account
|
||||
(mirroring `api_keys`' existing one-user-many-keys shape) so rotation is "mint a
|
||||
new key, revoke the old one," not "recreate the account."
|
||||
|
||||
### Endpoints
|
||||
|
||||
- `POST /api/service-accounts` — instance-scope/admin-only. Body:
|
||||
`{"name": ..., "scope": "instance"|"team", "teamID": ... }` (teamID required
|
||||
iff scope=team, and caller must be that team's owner or a system admin).
|
||||
Returns the account plus its first raw key (shown once, same pattern as
|
||||
`POST /api/users/{id}/api-keys`). Safe to call again with the same `name` —
|
||||
see "idempotent lookup" below — unlike `/api/bootstrap`, which is inherently
|
||||
one-shot by design (it's answering "does any user exist yet," a question with
|
||||
no analogue once one already does).
|
||||
- `POST /api/service-accounts/{id}/keys` — mint an additional key on an existing
|
||||
account (self-service-equivalent: instance admin for `instance` scope, team
|
||||
owner or system admin for `team` scope). Enables rotation without recreating
|
||||
the account or losing its identity/audit history.
|
||||
- `DELETE /api/service-accounts/{id}/keys/{keyID}` — revoke one key, mirroring
|
||||
`DELETE /api/users/{id}/api-keys/{keyID}`.
|
||||
- `GET /api/service-accounts?name=` — look up an existing account by name.
|
||||
This is what turns "I tried to create my account and got a conflict" into a
|
||||
normal flow instead of an error: a controller that expects to have already
|
||||
registered itself calls this first, and only falls through to `POST` if
|
||||
nothing comes back.
|
||||
|
||||
### Auth middleware
|
||||
|
||||
**Revised** (this section originally described an aspiration that didn't
|
||||
match what shipped — `TEAM-LOOKUP.md` already caught one instance of that,
|
||||
and a fuller audit found three more; this is the corrected, as-built
|
||||
description, not the original proposal).
|
||||
|
||||
`internal/api/middleware.go`'s dual resolution (`Authorization: Bearer` →
|
||||
`apiKeyUser()`, or session cookie → `sessionUser()`) and the service-account
|
||||
path (`serviceAccountFor()`) both resolve into one `Caller` type
|
||||
(`internal/api/caller.go`), not two parallel, un-unified context
|
||||
representations the way an earlier version of this server kept them. Every
|
||||
authorization predicate reads `Caller`'s methods:
|
||||
|
||||
- `Caller.IsAdmin()` — true **only** for a human system administrator, never
|
||||
for a service account of either scope, under any circumstance. `AdminOnly`
|
||||
and `requireSelfOrAdmin` key on this alone — user management
|
||||
(`POST /api/users`, `PUT /api/users/{id}/admin`, etc.) and
|
||||
`GET/PUT /api/admin/settings` stay human-only, forever. The original text
|
||||
here claimed an instance-scoped service account satisfies `AdminOnly` "for
|
||||
team-creation/listing purposes" — that was never true of the shipped code
|
||||
(`TEAM-LOOKUP.md` caught the listing half; the creation half was always a
|
||||
separate, bespoke check in `handleCreateTeam`, not `AdminOnly` itself) and
|
||||
is not being made true now. Don't widen `AdminOnly`: every time this has
|
||||
come up, the fix has been a narrower, purpose-built capability instead
|
||||
(`?name=` lookups for teams and service accounts; now
|
||||
`terdut-operator`'s own invite-minting feature for the one real gap this
|
||||
boundary left — how a human ever gets a first login on a no-OIDC,
|
||||
operator-managed install. See the bottom of "What this unblocks.")
|
||||
- `Caller.IsInstanceServiceAccount()` — true only for an instance-scoped
|
||||
service account, never for a human (including a human admin).
|
||||
`handleCreateTeam` uses exactly this: a human creates a team by being a
|
||||
human (and becomes its owner); an instance-scoped service account creates
|
||||
one with no human owner at all. The two paths are not interchangeable, so
|
||||
this predicate deliberately does not also admit a human admin.
|
||||
- `Caller.Role(teamID)`/`TeamIDs()` — a human's real `team_members` rows, or
|
||||
a team-scoped service account's single synthetic owner membership
|
||||
(`serveAsServiceAccount`). This is what makes `requireTeamMember`/
|
||||
`requireTeamOwner` treat a team-scoped service account as owner-equivalent
|
||||
for that one team, with no separate branch needed in either function.
|
||||
- `Caller.ServiceAccountID()` — used by `OperatorModeBlock` ("any service
|
||||
account passes") and by `callerMayManageServiceAccount`'s self-rotation
|
||||
check.
|
||||
- `Caller.AsHuman()` — the accessor every handler that needs a real
|
||||
`user_id` to act on behalf of must call and check, instead of reading a
|
||||
user off context unconditionally. Before the `Caller` type existed, four
|
||||
handlers did the latter and silently misbehaved for a service-account
|
||||
caller: `handleMe` and `handleTestNotification` 500'd (a zero-value user id
|
||||
that matches no row), `handleDismissOnboarding` silently no-op'd (`UPDATE
|
||||
... WHERE id = 0` affects nothing, still returns 204), and
|
||||
`handleCreateInvite` wrote that same zero value into `invites.created_by`
|
||||
— a real foreign-key violation, not just a wrong answer, since that column
|
||||
is nullable but was never passed as `nil`. All four now call `AsHuman()`
|
||||
and return an explicit 403 ("this endpoint is for human accounts only")
|
||||
or, for the invite case, leave `created_by` `NULL` the same way
|
||||
`handleCreateServiceAccount` already did for the analogous situation.
|
||||
|
||||
**Team scope is owner-equivalent for every `requireTeamOwner` endpoint,
|
||||
membership and invites included — by design, not by an unclosed gap.** An
|
||||
earlier version of this document flagged this as "acknowledged rather than
|
||||
closed," kept in check only by the social convention that nobody *builds*
|
||||
automation against those two routes. That convention is retired:
|
||||
`terdut-operator`'s `TerdutTeam` controller now mints and revokes its own
|
||||
team's invite link through exactly this capability (its existing
|
||||
team-scoped credential, `POST`/`DELETE /api/teams/{teamID}/invites`), which
|
||||
is the real fix for the human-onboarding gap below — not a narrower
|
||||
carve-out of this capability. `service_accounts_test.go`'s
|
||||
`TestServiceAccount_TeamScopeManagesItsOwnInvites` pins it.
|
||||
|
||||
**A team-scoped account can also mint another service account scoped to its
|
||||
own team** (`handleCreateServiceAccount`'s `callerOwnsTeam` branch, which a
|
||||
team-scoped caller already satisfies for its own team via the synthetic
|
||||
membership above). Kept, not restricted, for the same reason: a team-scoped
|
||||
credential is that team's owner's reach, full stop — carving this one
|
||||
capability out while leaving membership/invites alone would be an arbitrary
|
||||
asymmetry. Pinned by
|
||||
`TestServiceAccount_TeamScopeCanMintAnotherAccountForItsOwnTeam`.
|
||||
|
||||
**`callerMayManageServiceAccount` gained the one load-bearing fix this
|
||||
redesign exists for:** an instance-scoped service account may manage
|
||||
(mint/revoke a key on) *any* team-scoped account, not only one admin, that
|
||||
team's human owner, or the account itself. `handleCreateServiceAccount`
|
||||
already let an instance-scoped caller *create* a team-scoped account for
|
||||
any team; this closes the gap where adopting or rotating one it didn't just
|
||||
create in the same call — exactly `terdut-operator`'s documented
|
||||
adopt-on-409 crash-window recovery (its own `DESIGN.md` §5) — 403'd forever
|
||||
instead of succeeding (`terdut-operator#3`). Pinned by
|
||||
`TestServiceAccount_InstanceScopeAdoptsAnExistingTeamScopedAccountsKey`.
|
||||
|
||||
Anywhere identity is recorded for a human (incident timeline
|
||||
`acknowledged_by`/`assigned_to`, audit-relevant fields), a service-account
|
||||
caller is still coerced into a bare `user_id` of `0` today — `Caller`'s new
|
||||
`Identity()` accessor exists for exactly this follow-up, but wiring it in
|
||||
needs a schema migration (an actor-attribution column distinct from
|
||||
`user_id`) and is deliberately out of scope here. Tracked separately, not by
|
||||
this document.
|
||||
|
||||
## What this unblocks
|
||||
|
||||
Directly resolves `terdut-operator` DESIGN.md §6's two broken assumptions:
|
||||
1. **Bootstrap becomes single-purpose again.** `/api/bootstrap` mints exactly
|
||||
the founding human admin, once. The operator's actual first-reconcile flow:
|
||||
call `/api/bootstrap` only on a genuinely empty install; otherwise (or
|
||||
immediately after, if it won the bootstrap race) call
|
||||
`GET /api/service-accounts?name=terdut-operator`, and `POST` one if it
|
||||
doesn't exist yet. From then on the operator never touches `/api/bootstrap`
|
||||
again.
|
||||
2. **Rotation becomes real.** `POST /api/service-accounts/{id}/keys` + revoke the
|
||||
old one — no destructive DB-level workaround, no re-triggering a single-shot
|
||||
endpoint that can't fire twice.
|
||||
3. **Cross-namespace credential mirroring is no longer needed at all.**
|
||||
`terdut-operator`'s current design holds every credential — instance- and
|
||||
team-scoped alike — privately in the operator's own namespace, never in
|
||||
the namespace of the CR each one authenticates for; reconciliation happens
|
||||
entirely inside the operator's controller loop, so no CR owner ever needs
|
||||
read access to a terdut-server credential regardless of same- or
|
||||
cross-namespace `serverRef`. Team scoping is still what bounds the blast
|
||||
radius of any individual credential: a leaked team-scoped key exposes
|
||||
exactly one team's resources, never the whole server, which is what makes
|
||||
holding many credentials in one place (the operator's namespace) an
|
||||
acceptable trade rather than reintroducing the mirrored design's
|
||||
server-admin-equivalent-everywhere problem.
|
||||
4. **A human can get a first login on a no-OIDC, operator-managed install —
|
||||
without ever touching `AdminOnly` or `/api/admin/settings`.** This was
|
||||
filed as `terdut-server#23` ("no API path to create a human login after
|
||||
bootstrap") and diagnosed, at the time, as this server needing to let a
|
||||
service account through `AdminOnly`. It doesn't: the fix lives entirely
|
||||
in `terdut-operator`, because a team-scoped credential was *already*
|
||||
owner-equivalent for `POST /api/teams/{teamID}/invites`, and invite
|
||||
redemption (`POST /api/signup` with an `invite` token) bypasses
|
||||
`signup_mode` entirely — `terdut-operator` just never grew a feature to
|
||||
use either fact. Its `TerdutTeam` controller now mints and surfaces one
|
||||
via its own existing team-scoped credential (`spec.invite`,
|
||||
`status.inviteSecretRef`, see that repo's own docs), so a human joins a
|
||||
CRD-managed team by a real invite link, the same way anyone else would.
|
||||
`terdut-server#23` is closed with this note once that feature ships — its
|
||||
named routes stay human-only, correctly, not a gap.
|
||||
|
||||
## Suggested sequencing
|
||||
|
||||
Land this before `terdut-operator` implements any bootstrap/credential-handling
|
||||
code — that code would otherwise be written against the current one-shot,
|
||||
user-only credential model as a known-temporary workaround, which is wasted
|
||||
effort on a repo that currently has zero implementation to begin with.
|
||||
Where a service account acts on an incident (acknowledge, resolve), the
|
||||
timeline and `acknowledged_by` record it through parallel `*_service_account_id`
|
||||
columns, never as a user.
|
||||
|
||||
@@ -1,73 +0,0 @@
|
||||
# Team lookup for service accounts: closing terdut-operator's create-path crash window
|
||||
|
||||
This is a design note for a feature, not an implementation plan — same posture as
|
||||
`SERVICE-ACCOUNTS.md`, and raised for the same reason: `terdut-operator`'s `TerdutTeam`
|
||||
controller (ROADMAP.md Stage 2, a separate repo, no shared code) hit a gap this server has
|
||||
no answer for yet.
|
||||
|
||||
## The problem
|
||||
|
||||
`POST /api/teams` (`handleCreateTeam`, confirmed against `internal/api/teams.go`) lets an
|
||||
instance-scoped service account create a team — it has its own explicit
|
||||
`isInstanceServiceAccount(...)` branch alongside the human-user path, not gated by
|
||||
`AdminOnly`. If that call succeeds server-side but the caller (`TerdutTeam`'s controller)
|
||||
crashes before persisting the resulting team ID locally, a retry's `POST` 409s on the name's
|
||||
unique constraint (confirmed: the `isUniqueViolation` branch in the same handler).
|
||||
|
||||
Recovering from that 409 means looking the team up by name, and nothing today permits that
|
||||
for a service account:
|
||||
|
||||
- `GET /api/teams` (`handleListTeams`) answers "what teams does the *caller* belong to", via
|
||||
a `team_members` join keyed on `userFromContext`'s `caller.ID` — confirmed against source.
|
||||
A service account is never a member of anything, so this always returns empty for one,
|
||||
regardless of what exists.
|
||||
- `GET /api/admin/teams` (`handleAdminListTeams`) is gated by `AdminOnly`, and `AdminOnly`'s
|
||||
actual code (`internal/api/middleware.go`) checks only `userFromContext(...).IsAdmin` — no
|
||||
branch for a service account at all, confirmed against source. This contradicts
|
||||
`SERVICE-ACCOUNTS.md`'s own text, which claims "an instance-scoped [service account
|
||||
satisfies] `AdminOnly` for team-creation/listing purposes" — that claim doesn't match this
|
||||
endpoint's actual, shipped code. (Team *creation* is fine: `handleCreateTeam` isn't behind
|
||||
`AdminOnly` at all, it has its own check. Only the listing half of that sentence is wrong.)
|
||||
|
||||
This is exactly the shape of gap `SERVICE-ACCOUNTS.md`'s own `GET /api/service-accounts?name=`
|
||||
closed for service accounts themselves (confirmed: that endpoint's own comment —
|
||||
"the name lookup is open to any authenticated caller... what lets a service account find its
|
||||
own account on the 403 that follows a second POST"). Teams never got the equivalent, because
|
||||
nothing needed it until an operator started creating them unattended.
|
||||
|
||||
## Goals
|
||||
|
||||
- A service-account-accessible way to look up one team by exact name, mirroring
|
||||
`GET /api/service-accounts?name=` as closely as possible — same shape, same reasoning,
|
||||
same low sensitivity of what it discloses.
|
||||
- No change to today's behavior for an empty/no-name request.
|
||||
|
||||
## Proposed shape
|
||||
|
||||
Extend `GET /api/teams` itself, the same way `handleListServiceAccounts` already branches on
|
||||
a `?name=` query param, rather than adding a new route:
|
||||
|
||||
- `name` unset (today's behavior, unchanged): the caller's own teams, via `team_members`.
|
||||
- `name=<value>` set: look up that one team by exact name — a one-or-zero-length array, not
|
||||
an error on no match, mirroring `GET /api/service-accounts?name=`'s own response shape and
|
||||
status codes exactly. Deliberately **not** gated by `isInstanceServiceAccount` or
|
||||
`AdminOnly`: a human caller who's already a member sees this same information in their own
|
||||
team list regardless, and a non-member learning only that a name is taken — not who's in
|
||||
the team, not any of its data — is the same low-sensitivity disclosure
|
||||
`GET /api/service-accounts?name=` already accepts for service-account names.
|
||||
|
||||
## What this unblocks
|
||||
|
||||
Directly resolves the crash-window gap in `terdut-operator`'s `TerdutTeam` controller: on a
|
||||
409 from `POST /api/teams`, `GET /api/teams?name=<the same name>` — authenticated with the
|
||||
same instance-scoped credential that just got the 409 — finds the id, and the controller
|
||||
proceeds as if its own create had returned it directly. The same adopt-on-409 pattern already
|
||||
proven for service accounts (that repo's `DESIGN.md` §6 point 1, §5's general rule), not a
|
||||
new one.
|
||||
|
||||
## Suggested sequencing
|
||||
|
||||
Land this before `TerdutTeam`'s create path is implemented — the same reasoning
|
||||
`SERVICE-ACCOUNTS.md` gave for its own sequencing: writing that code against today's gap as a
|
||||
"known-temporary workaround" is wasted effort when the fix is this small and this
|
||||
well-precedented.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Terminal Duty documentation
|
||||
|
||||
The [README](../README.md) is the short tour. These pages hold the detail.
|
||||
|
||||
**Running it**
|
||||
- [Deployment](./deployment.md): Docker, the Helm chart, the database, backups, and the operator.
|
||||
- [Configuration](./configuration.md): environment variables and settings.
|
||||
- [Single sign-on](./single-sign-on.md): OIDC, group mapping, the terminal device flow.
|
||||
|
||||
**Using it**
|
||||
- [The web UI](./web-ui.md): sessions, the Team and Admin tabs.
|
||||
- [Alertmanager configuration](./alertmanager.md): routes, integration keys and webhooks.
|
||||
- [Alerts and incidents](./incidents.md): correlation, lifecycle, on-call assignment, stale-alert expiry.
|
||||
- [Push notifications](./notifications.md): ntfy pages and acknowledging from them.
|
||||
- [Escalation](./escalation.md): ladders, repeats and the fallback topic.
|
||||
- [Dead man's switches](./dead-mans-switch.md): noticing that alerts stopped arriving.
|
||||
|
||||
**Integrating and contributing**
|
||||
- [API reference](./api.md): every endpoint, authentication and error shape.
|
||||
- [Service accounts](../SERVICE-ACCOUNTS.md): non-human credentials for automation.
|
||||
- [Development and releasing](./development.md): tests, the CI gate, the release pipeline.
|
||||
@@ -0,0 +1,62 @@
|
||||
# Alertmanager configuration
|
||||
|
||||
_Pointing Alertmanager at the server._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
Alerts arrive on a team's **integration key**, which says both that the sender
|
||||
may post and which team the alerts belong to. Mint one as an owner of the team:
|
||||
|
||||
```bash
|
||||
curl -X POST https://terdut.example.com/api/teams/1/integrations \
|
||||
-H "Authorization: Bearer $TERDUT_API_KEY" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"name":"prod alertmanager"}'
|
||||
```
|
||||
|
||||
The response carries the key and the full URL **once**; only a SHA-256 hash is
|
||||
stored. Put it in your `alertmanager.yml`:
|
||||
|
||||
```yaml
|
||||
receivers:
|
||||
- name: terdut
|
||||
webhook_configs:
|
||||
- url: http://terdut-server:8080/api/integrations/<key>/alertmanager
|
||||
send_resolved: true
|
||||
|
||||
route:
|
||||
receiver: terdut
|
||||
```
|
||||
|
||||
The whole URL is a credential, so treat it like one. Alertmanager 0.26 and
|
||||
later can read it from a file with `url_file:` instead, which keeps it out of
|
||||
your configuration repository:
|
||||
|
||||
```yaml
|
||||
- url_file: /etc/alertmanager/secrets/terdut-webhook-url/url
|
||||
send_resolved: true
|
||||
```
|
||||
|
||||
The webhook endpoint requires no authentication.
|
||||
|
||||
If you use the [dead man's switch](./dead-mans-switch.md) — and the default configuration does — give
|
||||
the heartbeat a route of its own, because the deadline is only as tight as the interval feeding it:
|
||||
|
||||
```yaml
|
||||
route:
|
||||
receiver: terdut
|
||||
repeat_interval: 4h
|
||||
routes:
|
||||
- matchers: [ 'alertname = "Watchdog"' ]
|
||||
receiver: terdut
|
||||
group_wait: 0s
|
||||
group_interval: 1m
|
||||
repeat_interval: 1m
|
||||
```
|
||||
|
||||
That delivers a heartbeat every **2 minutes**, not every minute. Alertmanager only reconsiders a
|
||||
group every `group_interval`, and at exactly one elapsed interval `repeat_interval` has not *quite*
|
||||
passed, so the send slips to the next tick — equal values give 2×. Two minutes against the 15 minute
|
||||
default is seven heartbeats per window, which is the point; use `group_interval: 30s` if you want
|
||||
the numbers to mean what they say.
|
||||
|
||||
kube-prometheus-stack users get the `Watchdog` alert (`expr: vector(1)`) for free; it just needs
|
||||
routing to terdut rather than to `null`.
|
||||
@@ -0,0 +1,436 @@
|
||||
# API reference
|
||||
|
||||
_The REST API._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
## Authentication
|
||||
|
||||
All endpoints except `/api/bootstrap`, `/api/integrations/{key}/alertmanager`,
|
||||
`/api/notify/ack/{token}`, `/api/login`, `/api/logout`, `/api/auth/config`,
|
||||
`/api/version`, `/api/oidc/login`, `/api/oidc/callback`, `/api/oidc/device` and
|
||||
`/api/oidc/device/token` require either an API key:
|
||||
|
||||
```
|
||||
Authorization: Bearer <api-key>
|
||||
```
|
||||
|
||||
or the web UI's session cookie. A request that carries an `Authorization` header
|
||||
is judged on that header alone.
|
||||
|
||||
Two kinds of user exist. An **administrator** manages accounts: creating and
|
||||
deleting users, setting anybody's password, minting keys for anybody, and
|
||||
granting the flag itself. Everybody else works incidents — acknowledging,
|
||||
assigning, snoozing, resolving, noting — and manages their own account and
|
||||
nobody else's. An API key carries exactly the rights of the user it belongs to.
|
||||
|
||||
A third principal, the **service account**, exists for automation (a
|
||||
Kubernetes operator, most likely) that needs to manage teams, escalation
|
||||
policies, dead man's switches and integrations without impersonating a human.
|
||||
It is not a user — it never signs in, never appears in a team's member list,
|
||||
and never holds the administrator flag — and its key is prefixed `tdsa_` so it
|
||||
reads as one at a glance in a log line. See [Service accounts](#service-accounts).
|
||||
|
||||
**Getting an account.** The first one comes from `/api/bootstrap`. After that
|
||||
it depends on `signup_mode`, an administrator setting:
|
||||
|
||||
- `invite_only` (the default) — a team owner mints a link with
|
||||
`POST /api/teams/{teamID}/invites`, and the person who opens it picks a
|
||||
username and password and lands in that team with the role the link carries.
|
||||
Links are single-use unless told otherwise, expire after seven days, and can
|
||||
be revoked before that.
|
||||
- `open` — anybody who can reach the server can create an account, and must
|
||||
name a team, which they then own.
|
||||
|
||||
Invites are **links, not email**: this server has no SMTP, and adding it to send
|
||||
one message would be a subsystem to run, secure and monitor. Send the link
|
||||
however you already talk to the person.
|
||||
|
||||
A domain-restricted third mode was considered and dropped: with no email there
|
||||
is nothing to verify an address against, so it would only check the domain of a
|
||||
string somebody typed.
|
||||
|
||||
The first user, from `/api/bootstrap`, is an administrator. Users created
|
||||
afterwards are not, until an administrator says so. An install always keeps at
|
||||
least one: the last administrator can be neither deleted nor demoted, and
|
||||
nobody can delete or demote themselves.
|
||||
|
||||
Endpoints that require the flag answer `403` with
|
||||
`{"error":"administrator access required"}`.
|
||||
|
||||
**Teams** are the unit of tenancy, and are a separate axis from the administrator
|
||||
flag. A team owns its incidents, alerts, schedule and integrations, and a user
|
||||
sees exactly the teams they belong to. Within a team an **owner** configures it
|
||||
(schedule, integrations, membership) and a **member** works its incidents.
|
||||
|
||||
An administrator crosses that line in one direction only. They **configure any
|
||||
team** without being in it — every owner-only endpoint accepts the flag, because
|
||||
otherwise a team whose last owner left could never be repaired. They do **not
|
||||
read any team**: the queue, the alerts and the incidents are filtered by real
|
||||
membership, so an administrator sees a team's work only by joining it, which is
|
||||
a membership change and shows up as one. Administration is about accounts and
|
||||
the shape of a team, not about reading other people's incidents.
|
||||
|
||||
Anything belonging to a team you are not in answers `404`, not `403`: whether an
|
||||
incident exists is itself something only its team should learn.
|
||||
|
||||
**Operator mode** (`TERDUT_OPERATOR_MODE`, see [Configuration](./configuration.md#configuration))
|
||||
declares this install gitops-managed. When it is on, a session or a user's own
|
||||
API key gets `403 {"error": "...", "reason": "operator_managed"}` on every
|
||||
write this page marks **owner**-gated under Teams below (creating, renaming
|
||||
or deleting a team; its OIDC group binding; its escalation ladder; its dead
|
||||
man's switches; its integrations) — a service account's writes are unaffected.
|
||||
Team membership and invites are deliberately excluded: they are never
|
||||
gitops-managed, in operator mode or out of it. `GET /api/auth/config` reports
|
||||
`operator_mode` so a client can grey those sections out before a write is ever
|
||||
attempted.
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `GET` | `/api/auth/config` | How to sign in: `{"password_login", "oidc": {"enabled","name"}, "device_login", "operator_mode"}`. No session needed |
|
||||
| `GET` | `/api/version` | `{"version"}` — this build's version string. No session needed, the same as `/healthz` |
|
||||
| `POST` | `/api/login` | `{"username","password"}` → sets the session cookie, returns `{user, has_password}`. `429` after too many failures; `403` when `TERDUT_PASSWORD_LOGIN=false` |
|
||||
| `GET` | `/api/oidc/login` | Starts a single sign-on sign-in: redirects the browser to the provider. `?next=/path` is where to land afterwards; only a path on this server is honoured. Only exists when SSO is configured |
|
||||
| `POST` | `/api/oidc/device` | Starts a device login: returns `{device_code, user_code, verification_url, interval, expires_in}`. Only exists when SSO is configured |
|
||||
| `POST` | `/api/oidc/device/token` | `{"device_code"}` → `202 {"status":"pending"}`, then `200` with the session cookie once approved (once only). `410` with `{"error":"expired"}` or `{"error":"denied"}`; `429 {"error":"slow_down"}` if polled faster than `interval` |
|
||||
| `POST` | `/api/oidc/device/approve` | **session** — `{"user_code"}`. Approves a pending device login as the caller. `403` for an API key; `404` for an unknown, expired or already decided code |
|
||||
| `POST` | `/api/oidc/device/deny` | **session** — `{"user_code"}`. Refuses it |
|
||||
| `GET` | `/api/oidc/callback` | Where the provider sends the browser back. Sets the session cookie and redirects to `/`, or to `/?sso_error=<code>` — one of `denied`, `expired`, `failed`, `unavailable`, `not_allowed`, `no_email`, `email_conflict`, `disabled`, `not_bootstrapped` (no user exists on this install yet — sign in again once something has called `/api/bootstrap`) |
|
||||
| `POST` | `/api/logout` | Ends the session and clears the cookie |
|
||||
| `GET` | `/api/me` | The caller: `{user, has_password}` |
|
||||
|
||||
## Users
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
**admin** marks an endpoint that requires the administrator flag; **self or
|
||||
admin** marks one you may use on your own account and an administrator may use
|
||||
on anybody's.
|
||||
|
||||
| Method | Path | Who | Description |
|
||||
|---|---|---|---|
|
||||
| `GET` | `/api/signup` | — | Whether sign-up is open, and whether `?invite=` is usable. No session needed: the caller has no account yet |
|
||||
| `POST` | `/api/signup` | — | Create an account `{"username","email","password","invite"?,"team_name"?}` and sign in. `403` without a usable invite when the mode is invite-only |
|
||||
| `POST` | `/api/bootstrap` | — | Create first user + API key `{"username","email","password"?}` (only works on empty DB). The user is an administrator |
|
||||
| `GET` | `/api/users` | any | List users. Open to everybody: the queue's assignment control and the schedule both have to name people |
|
||||
| `GET` | `/api/users/{id}/teams` | self or admin | The teams that user is in, each with their role. `/api/teams` is always about the caller; this one answers it about somebody else, for the admin page's per-user view. `404` for a user who does not exist, so "no teams" and "no such person" are distinguishable |
|
||||
| `POST` | `/api/users` | **admin** | Create user `{"username","email"}`. Not an administrator |
|
||||
| `DELETE` | `/api/users/{id}` | **admin** | Delete user (cascades to keys). `409` for yourself or the last administrator |
|
||||
| `PUT` | `/api/users/{id}/admin` | **admin** | Grant or revoke the administrator flag `{"is_admin"}`. `409` for yourself, the last administrator, or an administrator granted by single sign-on |
|
||||
| `PUT` | `/api/users/{id}/disabled` | **admin** | Take an account out of use, or put it back `{"disabled"}`. `409` for yourself or the last administrator |
|
||||
| `PUT` | `/api/users/{id}/notify` | self or admin | Set push notification target `{"ntfy_topic"}` — empty string clears it |
|
||||
| `PUT` | `/api/users/{id}/password` | self or admin | Set web UI password `{"password","current_password"}`. `current_password` is required only when changing your own existing password. Ends the user's other sessions |
|
||||
| `POST` | `/api/users/{id}/api-keys` | self or admin | Issue API key `{"name"}` — key shown once |
|
||||
| `DELETE` | `/api/users/{id}/api-keys/{keyID}` | self or admin | Revoke API key |
|
||||
|
||||
## Administration
|
||||
|
||||
| Method | Path | Who | Description |
|
||||
|---|---|---|---|
|
||||
| `GET` | `/api/admin/teams` | **admin** | Every team on the server, with its member and open-incident counts. `/api/teams` answers "what am I in"; this answers "what is there" |
|
||||
| `GET` | `/api/admin/teams/{teamID}` | **admin** | One team and who is in it: `{"team", "members"}`. `404` for a team that does not exist. `GET /api/teams/{teamID}/members` is **member**-only and still `404`s an administrator from outside the team — reading a team's shape and reading its work are different questions, so they are different endpoints |
|
||||
| `GET` | `/api/admin/settings` | **admin** | The editable settings with their bounds, plus the environment-configured ones, read-only. Never credentials |
|
||||
| `PUT` | `/api/admin/settings` | **admin** | Change one or more `{"key": seconds}`, or `{"signup_mode": "open"\|"invite_only"}`. `400` for an unknown key or a value outside its bounds |
|
||||
|
||||
## Service accounts
|
||||
|
||||
A service account is a scoped, non-human credential for automation — not a
|
||||
`users` row, so it never signs in, is never a team member, and never carries
|
||||
the administrator flag. Two scopes:
|
||||
|
||||
- **instance** — the same reach system administration has over teams: create
|
||||
one, and mint a **team**-scoped account against any of them. There is no
|
||||
cap on how many instance-scoped accounts exist, but ordinarily there is one,
|
||||
belonging to whatever is provisioning this install end to end.
|
||||
- **team** — owner-equivalent for that one team, and nothing else: every
|
||||
**owner**-gated endpoint under [Teams](#teams), membership and invites
|
||||
included. Nothing narrower is enforced server-side; what actually keeps
|
||||
membership out of automation's hands is that no operator built against this
|
||||
scope should ever call those two endpoints — see
|
||||
[operator mode](#authentication) and [`SERVICE-ACCOUNTS.md`](../SERVICE-ACCOUNTS.md)'s note on this.
|
||||
|
||||
A key is shown once, at creation or rotation, and only its hash is stored —
|
||||
the same handling as a user's API key. Losing it means minting a new one;
|
||||
there is no way to recover a raw key from the server.
|
||||
|
||||
| Method | Path | Who | Description |
|
||||
|---|---|---|---|
|
||||
| `GET` | `/api/service-accounts` | **admin** | Every service account. Pass `?name=` instead to look one up by its exact name — open to **any** authenticated caller (human or service account), since it returns no key material and is how an account finds its own id |
|
||||
| `POST` | `/api/service-accounts` | owner\* | Create one and mint its first key `{"name","scope","team_id"?}` (`team_id` required for `scope:"team"`, absent for `scope:"instance"`). Returns `{"service_account", "key"}` — `key.key` shown once |
|
||||
| `POST` | `/api/service-accounts/{id}/keys` | owner\* | Mint an additional key `{"name"}` — rotation without recreating the account. Shown once |
|
||||
| `DELETE` | `/api/service-accounts/{id}/keys/{keyID}` | owner\* | Revoke one key |
|
||||
|
||||
\* For an **instance**-scoped account: a system administrator only. For a
|
||||
**team**-scoped account: a system administrator, that team's own human owner,
|
||||
an instance-scoped service account (minting a narrower credential for a team
|
||||
it just created), or — for the two key endpoints only — the account rotating
|
||||
or revoking its own key, which is not a privilege escalation, the same
|
||||
reasoning a user's own API keys rest on.
|
||||
|
||||
## Alert ingestion
|
||||
|
||||
Alerts arrive on a team's integration key. The key is both the credential and the
|
||||
routing: it says that the sender may post, and which team the alerts belong to.
|
||||
Create one with `POST /api/teams/{teamID}/integrations`, which returns the key
|
||||
and the full URL once and stores only a SHA-256 hash.
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `POST` | `/api/integrations/{key}/alertmanager` | Alertmanager v4 webhook receiver for the key's team. `401` for an unknown key |
|
||||
|
||||
This is the only way in. The pre-teams `POST /api/alertmanager/webhook` took no
|
||||
credential at all — anything able to reach the port could open an incident —
|
||||
and was removed in v0.13.0 once senders had moved onto keys.
|
||||
|
||||
## Teams
|
||||
|
||||
**owner** below means an owner of that team, a system administrator (who
|
||||
passes every one of these without being a member), or that team's own
|
||||
team-scoped [service account](#service-accounts) — including membership and
|
||||
invites, technically, though no automation this scope was designed for
|
||||
(a Kubernetes operator's CRDs, see [`SERVICE-ACCOUNTS.md`](../SERVICE-ACCOUNTS.md)) ever models team
|
||||
membership or would call those two. See [Authentication](#authentication).
|
||||
**member** means membership and nothing else: an administrator who is not in
|
||||
the team gets the same `404` as anybody else.
|
||||
|
||||
| Method | Path | Who | Description |
|
||||
|---|---|---|---|
|
||||
| `GET` | `/api/teams` | any | The caller's own teams, each with their role |
|
||||
| `POST` | `/api/teams` | any | Create a team `{"name"}`; a human creator becomes its first owner. An instance-scoped [service account](#service-accounts) may also create one, and it gets no owner at all — expected for a team an operator is about to hand a team-scoped credential to, not an orphaned team a human made |
|
||||
| `PUT` | `/api/teams/{teamID}` | **owner** | Rename it `{"name"}`. `409` if the name is taken |
|
||||
| `DELETE` | `/api/teams/{teamID}` | **owner** | Delete a team and everything under it. `409` while it has open incidents |
|
||||
| `GET` | `/api/teams/{teamID}/members` | member | Who is in the team, with `status` (`oncall` if the rota has them today, `unpageable` when a page to them would go nowhere — even if they are on call — else `reachable`), `on_call`, `next_shift` (first rota day after today), `pageable` and `problem` (`has no ntfy topic` / `account is disabled`; never the topic itself) and `last_active_at` (their newest session or API-key use). Every member sees the same list |
|
||||
| `POST` | `/api/teams/{teamID}/members` | **owner** | Add a member, or change their role `{"user_id","role"}`. `409` when it would demote the last owner, or the membership is managed by single sign-on |
|
||||
| `DELETE` | `/api/teams/{teamID}/members/{userID}` | **owner** | Remove a member. `409` for the last owner, or a membership managed by single sign-on |
|
||||
| `GET` | `/api/teams/{teamID}/oidc-groups` | member | Which groups control this team's membership: `{"member_group","owner_group"}`. An empty string means no group grants that role here |
|
||||
| `PUT` | `/api/teams/{teamID}/oidc-groups` | **owner** | Set them. An empty string clears a binding |
|
||||
| `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys. Each carries `status` (`active` if its key posted within 24h, `quiet` if it has but not lately, `never`), `last_used_at` (last webhook, usable or not), `last_alert_at` (when an alert last arrived on it) and `alerts_24h` (distinct alerts it refreshed in the last day). Alerts delivered before the source was recorded (migration 010) have none, so the last two fill in as Alertmanager re-sends them |
|
||||
| `PATCH` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Rename `{"name"}`. The key does not change |
|
||||
| `POST` | `/api/teams/{teamID}/integrations` | **owner** | Mint an integration `{"name","kind"}` — key and URL shown once |
|
||||
| `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration. Alerts it delivered stay, unattributed |
|
||||
| `GET` | `/api/teams/{teamID}/invites` | **owner** | The team's invite links, with their uses and expiry. Never the tokens |
|
||||
| `POST` | `/api/teams/{teamID}/invites` | **owner** | Mint one `{"role","max_uses"}` — the full URL is returned once |
|
||||
| `DELETE` | `/api/teams/{teamID}/invites/{inviteID}` | **owner** | Revoke a link before it expires |
|
||||
| `GET` | `/api/teams/{teamID}/escalation` | member | The team's [escalation ladder](./escalation.md#escalation) `{repeat_count, fallback_topic, levels[], last_escalated_at?, last_escalated_incident_id?}`. Empty levels means the team has none. Each level also carries `status` (`ready`, `escalating` when an unanswered incident has climbed to it, `unreachable` when nobody on it could be woken), `waiting` (ids of the open incidents on it) and, per target, `username` (who it means today — the person on call, for a rota target), `reachable` and `problem`. The extra fields are output only; `PUT` takes the plain shape |
|
||||
| `PUT` | `/api/teams/{teamID}/escalation` | **owner** | Replace it wholesale. `400` for a level with no targets or no timeout — a rung that pages nobody is a silence with a number on it |
|
||||
| `GET` | `/api/teams/{teamID}/deadman/switches` | member | The team's [dead man's switches](./dead-mans-switch.md), each `{id, name, matcher, timeout_seconds, severity, status, last_heartbeat_at, last_triggered_at, open_incident_id, sources[]}`. `status` is `healthy`, `dead` or `dormant`; `sources` has one entry per heartbeat fingerprint. Empty when the team watches nothing |
|
||||
| `POST` | `/api/teams/{teamID}/deadman/switches` | **owner** | Add one: `{name?, matcher, timeout_seconds, severity?}`. `400` when the matcher names no `alertname` or holds several, or the timeout is not positive — a switch that silently watches nothing is the failure this feature exists to prevent |
|
||||
| `PUT` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Replace one in place, same body and validation as create. Its id is unchanged — for an automated caller reconciling a spec change, unlike delete-and-recreate |
|
||||
| `DELETE` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Stop watching. An incident it opened stays open. `404` for a switch of another team |
|
||||
|
||||
## Notifications
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `POST` | `/api/notify/ack/{token}` | Acknowledge an incident from a push notification's Acknowledge button. No auth: the token in the path is the credential — one incident, one action, 24 hours, idempotent. Must stay publicly reachable |
|
||||
|
||||
## Incidents
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `GET` | `/api/incidents` | List incidents. Filters: `?status=triggered\|acknowledged\|resolved`, `?severity=`, `?assigned_to=<user id>`, `?archived=true`, `?snoozed=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?sort=severity`, `?cluster=<value of the cluster group label>`, `?limit=` (default 50, max 500) |
|
||||
| `GET` | `/api/incidents/clusters` | The distinct `cluster` values on the caller's incidents from the last 90 days, sorted (`?team_id=` narrows it). An empty array when nothing carries the label |
|
||||
| `GET` | `/api/incidents/{id}` | Get single incident, with its alerts inline |
|
||||
| `GET` | `/api/incidents/{id}/alerts` | Alerts under this incident |
|
||||
| `GET` | `/api/incidents/{id}/timeline` | Full event history, chronological |
|
||||
| `POST` | `/api/incidents/{id}/acknowledge` | Acknowledge (stamps authed user + time) |
|
||||
| `DELETE` | `/api/incidents/{id}/acknowledge` | Clear acknowledgement, back to `triggered` |
|
||||
| `POST` | `/api/incidents/{id}/resolve` | Close by hand — **terminal**, see above |
|
||||
| `POST` | `/api/incidents/{id}/assign` | Reassign `{"user_id"}` |
|
||||
| `POST` | `/api/incidents/{id}/snooze` | Hide until `{"until": RFC3339}` or `{"duration": "2h"}` |
|
||||
| `DELETE` | `/api/incidents/{id}/snooze` | Un-snooze |
|
||||
| `POST` | `/api/incidents/{id}/archive` | Archive (hides from the default list) |
|
||||
| `DELETE` | `/api/incidents/{id}/archive` | Un-archive |
|
||||
| `POST` | `/api/incidents/{id}/notes` | Add a note `{"content"}` |
|
||||
| `DELETE` | `/api/incidents/{id}/notes/{eventID}` | Delete own note |
|
||||
|
||||
With no `?status=` filter, `GET /api/incidents` returns **open** incidents only —
|
||||
the queue an on-call person wants. Currently snoozed and archived incidents are
|
||||
excluded unless asked for. Actions that only make sense on an open incident
|
||||
return `409` once it is resolved.
|
||||
|
||||
Notes are ordinary timeline events of type `note`; only they are deletable, and
|
||||
only by their author. The rest of the timeline is a record of what happened.
|
||||
|
||||
### The incident object
|
||||
|
||||
| Field | Type | Notes |
|
||||
|---|---|---|
|
||||
| `id` | integer | Server-assigned |
|
||||
| `group_key` | string | Alertmanager's `groupKey` — opaque, treat as an identifier |
|
||||
| `title` | string | Rendered from `groupLabels` |
|
||||
| `group_labels` | object | String→string, as sent by Alertmanager |
|
||||
| `status` | string | `"triggered"`, `"acknowledged"` or `"resolved"` |
|
||||
| `severity` | string | *optional* — high-water mark across the incident's alerts; never lowered |
|
||||
| `triggered_at` | timestamp | When the incident opened |
|
||||
| `acknowledged_by_id` / `acknowledged_by` / `acknowledged_at` | | *optional* — user id, username, time |
|
||||
| `assigned_to_id` / `assigned_to` | | *optional* — user id, username |
|
||||
| `snoozed_until` | timestamp | *optional* — a value in the past reads as not snoozed |
|
||||
| `resolved_at` | timestamp | *optional* |
|
||||
| `resolution_source` | string | *optional* — `"alerts"`, `"manual"` or `"recovered"` |
|
||||
| `archived_at` | timestamp | *optional* |
|
||||
| `alerts` | array | Only on `GET /api/incidents/{id}` |
|
||||
|
||||
Treat `resolution_source` as an open set, as with the alert field of the same
|
||||
name: degrade unknown values to "resolved, reason unknown".
|
||||
|
||||
### The timeline event object
|
||||
|
||||
| Field | Type | Notes |
|
||||
|---|---|---|
|
||||
| `id` | integer | |
|
||||
| `incident_id` | integer | |
|
||||
| `type` | string | See below — treat as an open set |
|
||||
| `user_id` / `username` | | *optional* — absent when the server acted rather than a person |
|
||||
| `alert_id` | integer | *optional* — the alert an `alert_added` / `alert_resolved` event refers to |
|
||||
| `detail` | string | *optional* — the note body, the snooze deadline, etc. |
|
||||
| `created_at` | timestamp | |
|
||||
|
||||
Types written today: `triggered`, `alert_added`, `alert_resolved`,
|
||||
`acknowledged`, `unacknowledged`, `assigned`, `archived`, `unarchived`, `snoozed`,
|
||||
`unsnoozed`, `resolved`, `note`, `notified`, `notify_failed`, `deadman_silent`. On an
|
||||
`assigned` event `user_id` is the **assignee**, not the actor; the actor is in
|
||||
`actor_user_id`/`actor_username` or `actor_service_account_id`/`actor_service_account_name`
|
||||
(absent on assignments made before they were recorded). New types may be added; render
|
||||
unknown ones generically rather than dropping them.
|
||||
|
||||
On `notified` and `notify_failed`, `detail` carries the notification kind
|
||||
(`triggered` | `reminder` | `resolved`), and on a failure the reason after it.
|
||||
`user_id` is who was paged — absent means the page went to the shared fallback
|
||||
topic and so belongs to nobody. The topic itself is never written to the
|
||||
timeline: it is a shared secret with the ntfy server, and every API key can read
|
||||
this.
|
||||
|
||||
## Alerts
|
||||
|
||||
Alerts are read-only. Everything a person does happens on the incident.
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `GET` | `/api/alerts` | List alerts. Filters: `?status=firing\|resolved`, `?name=`, `?incident_id=`, `?archived=true`, `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD`, `?limit=` (default 50, max 500) |
|
||||
| `GET` | `/api/alerts/{id}` | Get single alert |
|
||||
|
||||
Archived alerts are hidden from `GET /api/alerts` unless `?archived=true` is
|
||||
passed; alert archiving is automatic housekeeping by the sweeper, not a user
|
||||
action. Resolved alerts carry `resolution_source`: `"alertmanager"` for a real
|
||||
resolved webhook, `"expiry"` when the sweeper inferred it (see
|
||||
[Stale alert expiry](./incidents.md#stale-alert-expiry)), `"deadman"` for a heartbeat declared
|
||||
dead (see [Dead man's switch](./dead-mans-switch.md)).
|
||||
|
||||
### The alert object
|
||||
|
||||
Returned by `GET /api/alerts` (as an array) and `GET /api/alerts/{id}`.
|
||||
Timestamps are RFC 3339 in UTC. Fields marked *optional* are omitted entirely
|
||||
when unset, so clients must treat them as nullable.
|
||||
|
||||
| Field | Type | Notes |
|
||||
|---|---|---|
|
||||
| `id` | integer | Server-assigned; stable for the life of the row |
|
||||
| `fingerprint` | string | Alertmanager's fingerprint — the upsert key |
|
||||
| `name` | string | From the `alertname` label |
|
||||
| `status` | string | `"firing"` or `"resolved"` |
|
||||
| `labels` | object | String→string, as sent by Alertmanager |
|
||||
| `annotations` | object | String→string, as sent by Alertmanager |
|
||||
| `starts_at` | timestamp | When the alert instance began, **per Prometheus** |
|
||||
| `ends_at` | timestamp | *optional* — absent while no end is known |
|
||||
| `generator_url` | string | Link back to the originating Prometheus |
|
||||
| `received_at` | timestamp | When the server last accepted a webhook for this alert — see below |
|
||||
| `incident_id` | integer | *optional* — the most recent incident this alert belongs to |
|
||||
| `resolution_source` | string | *optional* — `"alertmanager"`, `"expiry"` or `"deadman"` |
|
||||
| `archived_at` | timestamp | *optional* — set while archived |
|
||||
|
||||
#### `received_at` is a liveness heartbeat
|
||||
|
||||
`starts_at` comes from Prometheus and **never changes** for the lifetime of an
|
||||
alert instance. It says when the problem began, not whether it is still
|
||||
happening — an alert that started twelve days ago looks identical whether
|
||||
Alertmanager refreshed it a minute ago or went silent a week ago.
|
||||
|
||||
`received_at` is the field that answers "is this still live". It is set to the
|
||||
server's clock on **every accepted webhook** for that fingerprint, including the
|
||||
unchanged firing notifications Alertmanager re-sends every `repeat_interval`.
|
||||
Clients may rely on this:
|
||||
|
||||
- **A firing alert whose `received_at` is advancing is still being refreshed.**
|
||||
Stale-dating it against `repeat_interval` is a valid liveness check, and it is
|
||||
what the built-in sweeper does (see
|
||||
[Stale alert expiry](./incidents.md#stale-alert-expiry)).
|
||||
- **`received_at` tracks accepted payloads, not delivery attempts.** A retry
|
||||
that describes an older instance than the stored one is discarded, and a
|
||||
discarded payload does not move `received_at`.
|
||||
- **It stops advancing once the alert resolves,** because Alertmanager stops
|
||||
re-sending. On an alert resolved by the sweeper
|
||||
(`"resolution_source": "expiry"`) it therefore marks the last time
|
||||
Alertmanager was actually heard from, which is earlier than `ends_at`.
|
||||
|
||||
`GET /api/alerts` is ordered by `received_at` descending — most recently
|
||||
refreshed first — and the `?from=` / `?to=` filters on both the alert and stats
|
||||
endpoints select on `received_at`, not `starts_at`.
|
||||
|
||||
#### `resolution_source` says how much to trust `ends_at`
|
||||
|
||||
An alert can leave the firing state two ways, and `resolution_source` records
|
||||
which happened. Clients may rely on this:
|
||||
|
||||
- **Absent while firing.** It is set only on resolve, and a re-fire under the
|
||||
same fingerprint clears it again, so its presence always agrees with
|
||||
`"status": "resolved"`.
|
||||
- **`"alertmanager"` — a real resolved webhook arrived.** `ends_at` is the end
|
||||
time Alertmanager reported. It is an observed value and can be displayed as
|
||||
fact.
|
||||
- **`"expiry"` — the sweeper inferred the resolve** because Alertmanager stopped
|
||||
refreshing the alert (see [Stale alert expiry](./incidents.md#stale-alert-expiry)). Nothing
|
||||
ever reported an end, so **`ends_at` is approximate**: it is either the stale
|
||||
`endsAt` watermark from the last notification, or — when that notification
|
||||
carried none — the time the sweep ran, which lags the last real contact by up
|
||||
to `TERDUT_STALE_AFTER` plus a sweep interval. Treat it as "no later than",
|
||||
not as when the problem stopped.
|
||||
|
||||
On these alerts `received_at` is the more truthful signal: it marks the last
|
||||
time Alertmanager was actually heard from. Surfacing the distinction is
|
||||
worthwhile, since `"expiry"` can also mean the alert is still firing and the
|
||||
notification path broke.
|
||||
|
||||
- **`"deadman"` — a heartbeat was declared dead** (see
|
||||
[Dead man's switch](./dead-mans-switch.md)). Like `"expiry"`, an inference from
|
||||
silence rather than an observed end, so `ends_at` is approximate — but a much
|
||||
tighter one, bounded by the switch's timeout. It is also the one resolution
|
||||
a re-fire under the same `starts_at` can undo, since the switch coming back is
|
||||
exactly the evidence that the inference was wrong.
|
||||
|
||||
Treat the value as an open set and tolerate ones you do not recognise — new
|
||||
sources may be added, and unknown values should degrade to "resolved, reason
|
||||
unknown" rather than being rejected.
|
||||
|
||||
## On-call schedule
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
Each team keeps its own rota, so two teams can have two different people on call
|
||||
on the same day. The person taking a shift has to be in the team — paging
|
||||
somebody who cannot open the incident is worse than paging nobody.
|
||||
|
||||
| Method | Path | Who | Description |
|
||||
|---|---|---|---|
|
||||
| `POST` | `/api/teams/{teamID}/schedule` | **owner** | Assign user to dates `{"user_id", "dates":["YYYY-MM-DD",...], "replace"}` — all-or-nothing |
|
||||
| `GET` | `/api/teams/{teamID}/schedule` | member | List entries. Filters: `?from=YYYY-MM-DD`, `?to=YYYY-MM-DD` |
|
||||
| `DELETE` | `/api/teams/{teamID}/schedule/{id}` | **owner** | Remove schedule entry |
|
||||
| `GET` | `/api/schedule/current` | any | Who is on call today (UTC) in **every** team the caller is in — one entry per team, `[]` when nobody anywhere |
|
||||
|
||||
## Statistics
|
||||
|
||||
Every figure counts the caller's own teams only: a report that counted other
|
||||
teams' incidents would leak their volume, and their alert names through the
|
||||
top-alerts list, and would not be a number about the reader's work anyway.
|
||||
|
||||
All stat endpoints accept optional `?from=YYYY-MM-DD` and `?to=YYYY-MM-DD`, and exclude archived rows to match the default list views. Alert stats filter on `received_at`; incident stats filter on `triggered_at`.
|
||||
|
||||
| Method | Path | Description |
|
||||
|---|---|---|
|
||||
| `GET` | `/api/stats/incidents` | `{total, triggered, acknowledged, resolved, mtta_seconds, mttr_seconds}` |
|
||||
| `GET` | `/api/stats/alerts` | `{total, firing, resolved}` counts |
|
||||
| `GET` | `/api/stats/alerts/top` | Most frequent alert names. `?limit=` (default 10, max 100) |
|
||||
| `GET` | `/api/stats/alerts/by-hour` | Count per hour-of-day (UTC), all 24 slots returned |
|
||||
| `GET` | `/api/stats/alerts/by-day` | Count per day-of-week, all 7 slots with names returned |
|
||||
|
||||
`mtta_seconds` (time to acknowledge) and `mttr_seconds` (time to resolve) are
|
||||
averages over incidents that have actually been acknowledged or resolved, and are
|
||||
**null** until there are any — null means "no data", not zero.
|
||||
@@ -0,0 +1,49 @@
|
||||
# Configuration
|
||||
|
||||
_Environment variables and settings._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
Two kinds of setting, split by who changes them and how often.
|
||||
|
||||
**Where the server is plugged in** stays in the environment: the listen address,
|
||||
the database DSN, the ntfy URL and token, the public URL. They are needed before
|
||||
the database is open, and two of them are credentials.
|
||||
|
||||
**How the server behaves** lives in the database and is edited by an
|
||||
administrator in the web UI or through `PUT /api/admin/settings`, taking effect
|
||||
on the next sweep rather than at the next restart. The variables below marked
|
||||
**seed** are the value each of those starts from: written once, on first start,
|
||||
and never overwritten afterwards — a redeploy cannot put a chart's default back
|
||||
over an administrator's edit.
|
||||
|
||||
| Variable | Default | Description |
|
||||
|---|---|---|
|
||||
| `TERDUT_ADDR` | `:8080` | TCP address to listen on |
|
||||
| `TERDUT_DB_DSN` | — | **Required.** Postgres connection string, e.g. `postgres://terdut:secret@localhost:5432/terdut?sslmode=require` |
|
||||
| `TERDUT_ARCHIVE_AFTER` | `168h` (7d) | **seed.** How long a resolved alert or incident stays in the default list before being auto-archived |
|
||||
| `TERDUT_STALE_AFTER` | `6h` | **seed.** How long a firing alert may go without a refreshing webhook before it is treated as resolved — **must exceed your Alertmanager `repeat_interval`** |
|
||||
| `TERDUT_NTFY_URL` | — | ntfy server to publish push notifications to. Empty disables notifications entirely |
|
||||
| `TERDUT_NTFY_TOKEN` | — | Bearer token for an access-controlled ntfy |
|
||||
| `TERDUT_NTFY_FALLBACK_TOPIC` | — | Topic used when nobody is on call |
|
||||
| `TERDUT_PUBLIC_URL` | — | Base URL a phone uses to reach this server: the notification's link into the web UI, its Acknowledge button, and whether the session cookie is `Secure` |
|
||||
| `TERDUT_NOTIFY_REPEAT` | `15m` | **seed.** How long an incident may sit unacknowledged before it is paged again. `0` notifies once and never repeats |
|
||||
| `TERDUT_PASSWORD_LOGIN` | `true` | `false` refuses password login and password sign-up (`403`), leaving single sign-on the only way in. Refused at startup unless SSO is configured |
|
||||
| `TERDUT_TRUSTED_PROXIES` | `1` | How many reverse proxies in front of the server append to `X-Forwarded-For`; the per-address rate limits use the entry that many hops from the right. `0` ignores the header |
|
||||
| `TERDUT_OPERATOR_KEY` | — | At least 32 characters. When set, the instance-scoped service account `terdut-operator` is created if missing and its `seed` key replaced with this value at every start — how terdut-operator authenticates without a bootstrap handshake. An instance-scoped account acts as owner of every team (team configuration) but is not a member of any, so it reads no incidents |
|
||||
| `TERDUT_OPERATOR_MODE` | `false` | Declares this install gitops-managed: a session's or a user's own API key's writes to teams, escalation policies, dead man's switches and integrations are refused (`403 reason:"operator_managed"`); a [service account](./api.md#service-accounts)'s are not. Team membership and the schedule stay editable regardless |
|
||||
| `TERDUT_OIDC_ISSUER` | — | Turns single sign-on on. The provider's issuer URL; discovery is read from `<issuer>/.well-known/openid-configuration`. See [Single sign-on](./single-sign-on.md#single-sign-on-oidc) |
|
||||
| `TERDUT_OIDC_CLIENT_ID` / `TERDUT_OIDC_CLIENT_SECRET` | — | **Required with an issuer.** The confidential client registered at the provider. Keep the secret in a Secret, not in values |
|
||||
| `TERDUT_OIDC_NAME` | `SSO` | What the sign-in button calls the provider |
|
||||
| `TERDUT_OIDC_SCOPES` | `openid profile email` | Scopes requested, comma or space separated. Authentik puts `groups` behind `profile` |
|
||||
| `TERDUT_OIDC_USERNAME_CLAIM` / `_EMAIL_CLAIM` / `_GROUPS_CLAIM` | `preferred_username` / `email` / `groups` | ID token claims read for the username, email and groups |
|
||||
| `TERDUT_OIDC_TRUST_EMAIL` | `false` | Link a first sign-in to an existing local user by email even if the provider does not mark the address verified |
|
||||
| `TERDUT_OIDC_ALLOWED_GROUPS` | — | Comma-separated. Only people in one of these may sign in. Empty admits everybody the provider authenticates |
|
||||
| `TERDUT_OIDC_ADMIN_GROUP` | — | Members are system administrators |
|
||||
| `TERDUT_OIDC_SESSION_MAX_AGE` | `12h` | Hard ceiling on a session made by an SSO sign-in |
|
||||
|
||||
Durations use Go syntax (`30m`, `12h`, `168h`). An unparseable value falls back to the default.
|
||||
|
||||
Note that `TERDUT_STALE_AFTER` and a dead man's switch timeout point in opposite directions. Staleness
|
||||
is a generous grace period around a `repeat_interval` you do not control; a dead man's switch is a
|
||||
deadline you set deliberately, and the heartbeat's route is configured to beat faster than it.
|
||||
|
||||
In the Helm chart the two sweeper durations are set via `sweeper.staleAfter` and `sweeper.archiveAfter`, notifications via the `notify.*` values, single sign-on via `oidc.*` and `passwordLogin`, and operator mode via `operatorMode`.
|
||||
@@ -0,0 +1,81 @@
|
||||
# Dead man's switches
|
||||
|
||||
_Detecting that alerts have stopped arriving._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
Everything above assumes alerts arrive. If Prometheus stops evaluating, or
|
||||
Alertmanager cannot reach this server, nothing arrives — and silence looks
|
||||
exactly like everything being fine. A dead man's switch inverts the handling for
|
||||
one designated alert so that silence is the signal:
|
||||
|
||||
- **receiving** it opens no incident, and
|
||||
- the **absence** of it does.
|
||||
|
||||
kube-prometheus-stack already ships the alert for this. `Watchdog` is
|
||||
`expr: vector(1)`, so it fires permanently and is re-sent forever; it is worth
|
||||
nothing unless something downstream notices it stop. It is the usual first switch.
|
||||
|
||||
**Switches belong to a team**, which decides which of its own alerts are
|
||||
heartbeats and how long a silence has to last. Each **switch** is a row of its
|
||||
own — a name, one matcher, a timeout and a severity — so switches in one team
|
||||
can have different deadlines. An owner adds and removes them on **Team →
|
||||
Switches**, which lists each with a status (**healthy**, **dead**, or
|
||||
**dormant** until its first heartbeat), when it was last heard from, and when it
|
||||
last opened an incident; a matcher that several clusters satisfy is broken down
|
||||
per cluster. The API is `POST`/`DELETE /api/teams/{teamID}/deadman/switches`. A
|
||||
missed heartbeat opens an incident in the team whose integration received it.
|
||||
Removing a switch stops the watching; an incident it already opened stays open
|
||||
until somebody resolves it.
|
||||
|
||||
A new team watches nothing until its owner (or terdut-operator, from a
|
||||
`TerdutTeam`) adds a switch: inheriting an install-wide heartbeat would page a
|
||||
new team about a source it has never heard of.
|
||||
|
||||
A matcher is a set of exact label conditions, one of which must be the
|
||||
`alertname`, , one matcher per switch, `,` between the label conditions:
|
||||
|
||||
```
|
||||
alertname=Watchdog,cluster=prod
|
||||
```
|
||||
|
||||
**The unit of monitoring is the fingerprint, not the alert name.** Two clusters
|
||||
sending the same `Watchdog` are two independent switches, so a healthy one can
|
||||
never mask a dead one.
|
||||
|
||||
## The lifecycle
|
||||
|
||||
A switch is **dormant** until its first heartbeat arrives. A configured matcher
|
||||
that has never been heard from opens nothing, so a fresh deploy or a restored
|
||||
database does not page. It also means a matcher that never matches anything is
|
||||
silently inert.
|
||||
|
||||
Once armed, the sweeper declares it **dead** when either the heartbeat has not
|
||||
been refreshed within the switch's `timeout_seconds`, or Alertmanager explicitly
|
||||
resolved it — the sender saying the heartbeat stopped needs no further waiting.
|
||||
That opens an incident at the switch's `severity`, assigned and paged like any
|
||||
other, and marks the heartbeat alert `"resolution_source": "deadman"` so the
|
||||
alert list stops claiming a dead switch is firing.
|
||||
|
||||
It **recovers** when the heartbeat starts arriving again: the incident resolves
|
||||
with `"resolution_source": "recovered"` and the all-clear goes to whoever was
|
||||
paged.
|
||||
|
||||
Resolving the incident by hand sticks, the same way it does for an alert-backed
|
||||
one. While the switch stays silent nothing new opens — so a decommissioned
|
||||
source is a one-time page rather than a nag. The switch **re-arms** on the next
|
||||
heartbeat: come back and die again, and that is a new incident.
|
||||
|
||||
## Two things to know
|
||||
|
||||
A switch's timeout must be **shorter** than the `repeat_interval` of the
|
||||
route carrying the heartbeat, which is the exact opposite of
|
||||
`TERDUT_STALE_AFTER`. Inheriting a default `repeat_interval` of 4h gives you a
|
||||
switch that takes four hours to notice anything, so give the heartbeat
|
||||
[its own route](./alertmanager.md#alertmanager-configuration). Matched alerts are exempt from
|
||||
stale-alert expiry — a heartbeat answers to its own timeout and nothing else.
|
||||
|
||||
A dead man's switch incident has **no member alerts**:
|
||||
`GET /api/incidents/{id}/alerts` returns an empty list. There is no alert
|
||||
describing the problem, because the problem is that no alert arrived. What
|
||||
happened is on the timeline instead, as a `deadman_silent` event carrying the age
|
||||
of the last heartbeat, and the heartbeat's labels are on the incident's
|
||||
`group_labels`.
|
||||
@@ -0,0 +1,74 @@
|
||||
# Deployment
|
||||
|
||||
_Running the server in a container and on Kubernetes with the Helm chart._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
## Docker
|
||||
|
||||
```bash
|
||||
docker build -t terdut-server .
|
||||
docker run -p 8080:8080 \
|
||||
-e TERDUT_DB_DSN='postgres://terdut:secret@host.docker.internal:5432/terdut?sslmode=disable' \
|
||||
terdut-server
|
||||
```
|
||||
|
||||
The server creates its own schema on startup and needs a reachable Postgres; it stores nothing on
|
||||
disk, so there is no volume to mount.
|
||||
|
||||
## Kubernetes
|
||||
|
||||
A Helm chart is published from this repository as an OCI artifact, versioned in lockstep
|
||||
with the app — chart `x.y.z` is always app `vx.y.z`:
|
||||
|
||||
```bash
|
||||
helm upgrade --install terdut-server oci://git.ryuvia.com/niklas/terdut-server \
|
||||
--version 0.9.2 \
|
||||
--namespace terdut-server --create-namespace \
|
||||
--set networking.hostname=terdut.example.com
|
||||
```
|
||||
|
||||
The chart expects a [Gateway API](https://gateway-api.sigs.k8s.io/) Gateway named `envoy-main` in
|
||||
the `envoy-gateway-system` namespace to already exist — it renders an `HTTPRoute` against it rather
|
||||
than an `Ingress`. TLS is terminated at the gateway, so the server itself never sees a certificate.
|
||||
|
||||
| Value | Default | Description |
|
||||
|---|---|---|
|
||||
| `networking.hostname` | `terdut.example.com` | Hostname the `HTTPRoute` serves |
|
||||
| `networking.listener` | `""` | Gateway listener (`sectionName`) to bind to. Empty attaches to every matching listener, **including plaintext HTTP** — set it to the HTTPS listener's name to serve TLS only |
|
||||
| `networking.servicePort` | `8080` | Port the route forwards to; keep in sync with `service.port` |
|
||||
| `bootstrap.enabled` | `true` | Runs a post-install hook that creates the first user and stores its API key in the `<release>-admin-key` Secret. Already-bootstrapped servers are left alone |
|
||||
| `database.dsn` | `""` | **Required.** Postgres DSN, with no password in it. The chart provisions no database |
|
||||
| `database.passwordSecret.name` | `""` | Secret supplying `PGPASSWORD`. With the Zalando postgres operator, the Secret it generates for the role |
|
||||
| `database.passwordSecret.key` | `password` | Key within that Secret |
|
||||
|
||||
The API key travels in an `Authorization: Bearer` header, so set `networking.listener` whenever the
|
||||
hostname is reachable outside a trusted network.
|
||||
|
||||
### The database
|
||||
|
||||
The chart provisions no database: it takes a DSN and expects a Postgres that already exists. In this
|
||||
cluster the wrapper chart declares an `acid.zalan.do/v1 postgresql` CR; anywhere else, any reachable
|
||||
Postgres 14+ will do.
|
||||
|
||||
The DSN carries no password. pgx falls back to libpq's environment variables for whatever the DSN
|
||||
leaves out, so the password arrives as `PGPASSWORD` from a Secret and never appears in values, in
|
||||
the rendered manifest or in `kubectl describe pod`. With the postgres operator that Secret is the
|
||||
one it generates for the role, so a rebuild mints a new password with nothing to keep in sync —
|
||||
the same wiring miniflux uses.
|
||||
|
||||
The server migrates its own schema on startup, so a new database only has to exist and be writable.
|
||||
|
||||
### Backups
|
||||
|
||||
Postgres is backed up where it runs, not from here. The database pod carries a
|
||||
[k8up](https://k8up.io/) `k8up.io/backupcommand` annotation that streams a `pg_dump`, the same way
|
||||
gitea and immich do in this cluster.
|
||||
|
||||
## On Kubernetes with the operator
|
||||
|
||||
[terdut-operator](https://git.ryuvia.com/niklas/terdut-operator) runs a server for you from a
|
||||
`TerdutServer` object and manages its teams, escalation ladders, dead man's switches and alert
|
||||
sources as Kubernetes objects. It hands the server a generated key through `TERDUT_OPERATOR_KEY`
|
||||
(see [Configuration](./configuration.md) and [`SERVICE-ACCOUNTS.md`](../SERVICE-ACCOUNTS.md)), and
|
||||
the server then treats configuration as operator-managed (`TERDUT_OPERATOR_MODE`), refusing edits
|
||||
made by hand in the web UI. Use the Helm chart above for a plain install, the operator when you
|
||||
want that configuration in gitops.
|
||||
@@ -0,0 +1,93 @@
|
||||
# Development and releasing
|
||||
|
||||
_Building, testing and releasing the server._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
## Upgrading
|
||||
|
||||
The schema is a single baseline (`internal/db/migrations/001_schema.sql`) and no
|
||||
release has shipped yet, so there is no upgrade path from earlier development
|
||||
databases: start from an empty one. Changes after the first release arrive as
|
||||
new numbered migrations.
|
||||
|
||||
## Development
|
||||
|
||||
```bash
|
||||
make test-db # start a local Postgres for the tests (podman or docker)
|
||||
make test # run all tests
|
||||
go build ./... # compile all packages
|
||||
go run ./cmd/terdut # run locally (needs TERDUT_DB_DSN)
|
||||
```
|
||||
|
||||
The tests need a real Postgres, because the server does — there is no in-memory Postgres.
|
||||
`TERDUT_TEST_DSN` says where it is, `make test-db` starts
|
||||
one on port 5433 and prints the DSN, and `make test-db-stop` removes it. Each test gets its
|
||||
own schema on that server, so tests cannot see each other's rows. An unset `TERDUT_TEST_DSN`
|
||||
fails the suite rather than skipping it: a run that quietly tests nothing is worse than one
|
||||
that does not run.
|
||||
|
||||
`make fmt lint test helm-lint` is the gate. It mirrors `.gitea/workflows/ci.yaml` step for
|
||||
step, so a green run here means a green pipeline — with one deliberate exception: `make test`
|
||||
adds `-race`, which CI does not. The sweeper, the notifier goroutine and the dead man's switch
|
||||
sweep all run concurrently against the same database, and a race between them would surface as
|
||||
a flaky incident in production rather than as a red build.
|
||||
|
||||
The web UI lives in `internal/web/static/` as plain HTML, CSS and ES modules,
|
||||
embedded into the binary with `go:embed`. It has no build step and no npm, so
|
||||
editing a file and restarting the server is the whole loop.
|
||||
|
||||
## Releasing
|
||||
|
||||
```
|
||||
push or PR → ci.yaml gofmt, go vet, go test -race
|
||||
govulncheck, gitleaks
|
||||
helm lint + render
|
||||
push tag vX.Y.Z → release.yaml the same gate, then publish:
|
||||
git.ryuvia.com/niklas/terdut-server:vX.Y.Z
|
||||
oci://git.ryuvia.com/niklas/terdut-server X.Y.Z
|
||||
then trivy-scan the pushed image
|
||||
PR to Ryuvia/charts → bump the wrapper chart to X.Y.Z; on merge
|
||||
Flux reconciles and the release rolls out
|
||||
```
|
||||
|
||||
Both artifacts go to the **personal** Gitea namespace rather than `ryuvia`, because Gitea
|
||||
scopes package visibility to the owner with no per-package override — so `ryuvia/*` is private
|
||||
because the org is. Publishing to `niklas` keeps them anonymously pullable, which is why no
|
||||
pull secret is needed in the cluster. Same reasoning, and the same choice, as riksdata and
|
||||
rd-web.
|
||||
|
||||
Saying **"Release"** runs all three rows: the `release` skill commits, pushes, tags, waits for
|
||||
the pipeline, and opens the `Ryuvia/charts` PR, stopping before the merge. See
|
||||
`~/.claude/skills/release/`, or `.release.conf` here for this repo's part of it.
|
||||
|
||||
The chart is published **only** from the tag, by the `chart` job. There used to be a second
|
||||
publisher on every `charts/**` push to main, and the two raced for the same chart version with
|
||||
different answers — chart 0.9.0 went out reading `appVersion: "latest"` that way. One
|
||||
publisher, triggered by the tag (`766f439`). The cost is that a chart-only change has no
|
||||
version of its own and rides the next app tag.
|
||||
|
||||
Both workflows are thin drivers over the Makefile: `ci.yaml` runs `make fmt lint test` and
|
||||
`make helm-lint`, `release.yaml` adds `make binaries`, `make push`, `make helm-package` and
|
||||
`make helm-push`. That is deliberate — it is what makes a green local gate and a green
|
||||
pipeline the same code rather than two descriptions of it, and it is how riksdata and rd-web
|
||||
have always worked.
|
||||
|
||||
`make push` builds and pushes in one step, unlike those two, because the image is
|
||||
`linux/amd64,linux/arm64` and buildx cannot load a multi-platform result into the local image
|
||||
store. `make build` stays single-platform and local-only. Both refuse `VERSION=dev`:
|
||||
publishing is one command, so it is also one command to run by accident. Publishing happens
|
||||
by pushing a tag.
|
||||
|
||||
Two things the release process needs to know about this repo:
|
||||
|
||||
- **The image scan runs after publishing**, like riksdata's and rd-web's: trivy cannot read
|
||||
a locally built image on this runner, so it pulls the pushed one. A red `scan-image` means
|
||||
do not bump the wrapper chart to that version — it cannot unpublish anything. The image is
|
||||
`FROM scratch`, so trivy sees exactly one target, the Go binary and its module graph.
|
||||
- **The wrapper chart's `values.yaml` has two `tag:` lines** — the app image and the python
|
||||
backup sidecar — so `chart-bump` is given `--image` to say which one moves. Once the wrapper
|
||||
chart drops the sidecar and declares a `postgresql` CR instead, there is one `tag:` line
|
||||
again, and `--image` becomes belt and braces.
|
||||
|
||||
The wrapper chart must have **its own `version:` bumped in the same commit**. Flux reconciles
|
||||
with `reconcileStrategy: ChartVersion`, so a chart whose version did not change produces no
|
||||
new artifact and the change is never deployed — with no error anywhere.
|
||||
@@ -0,0 +1,45 @@
|
||||
# Escalation
|
||||
|
||||
_Escalation ladders: who is paged next when nobody acknowledges._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
Without a ladder, an unacknowledged incident re-pages the same topic every
|
||||
`notify_repeat` forever. That is a louder version of the same silence: if the
|
||||
person on call is asleep, out of signal, or has left, nothing else happens.
|
||||
|
||||
A team can configure an ordered ladder instead. Each level has a timeout and a
|
||||
set of targets, and a target is either a named person or **whoever the team's
|
||||
rota says is on call today** — the target that keeps working when the rota
|
||||
changes and nobody remembers to edit the policy.
|
||||
|
||||
```
|
||||
level 1 5m oncall the rota gets first refusal
|
||||
level 2 5m user:bob then a named second
|
||||
then repeat_count more rounds
|
||||
then the team's fallback topic, once
|
||||
```
|
||||
|
||||
When a level's timeout passes with the incident still `triggered`, the next
|
||||
level is paged. Off the end of the ladder the whole thing runs again
|
||||
`repeat_count` times, and after that the team's `fallback_topic` is paged once
|
||||
as the end of the line. The incident stays open throughout: running out of
|
||||
people to wake is not the same as somebody answering.
|
||||
|
||||
**Acknowledging or resolving stops it**, which is the point — continuing to wake
|
||||
people after somebody has said "I have this" is how a tool teaches people to
|
||||
mute it. **Snoozing pauses it**: a deliberate "not now" holds the ladder where
|
||||
it is, and it resumes when the snooze runs out.
|
||||
|
||||
Every step is on the incident's timeline with the level and the names it woke,
|
||||
so somebody reading it afterwards can tell why their phone rang at 04:00. A
|
||||
level whose targets are all unreachable — no ntfy topic, a disabled account, an
|
||||
empty rota — is recorded as `nobody reachable` and the ladder moves on rather
|
||||
than stalling on a rung that cannot ring.
|
||||
|
||||
**Reminders and escalation never both run.** A team with a ladder gets
|
||||
escalation; a team without keeps the reminder behaviour exactly as it was. Two
|
||||
pages for one silence is the surest way to get a tool muted.
|
||||
|
||||
The ladder's `fallback_topic` is per team, unlike `TERDUT_NTFY_FALLBACK_TOPIC`,
|
||||
which is the install-wide topic used when an incident opens with nobody on call.
|
||||
They answer different questions: one is "nobody was scheduled", the other is
|
||||
"everybody scheduled has been tried".
|
||||
|
After Width: | Height: | Size: 23 KiB |
|
After Width: | Height: | Size: 55 KiB |
|
After Width: | Height: | Size: 88 KiB |
|
After Width: | Height: | Size: 52 KiB |
|
After Width: | Height: | Size: 53 KiB |
|
After Width: | Height: | Size: 100 KiB |
|
After Width: | Height: | Size: 93 KiB |
|
After Width: | Height: | Size: 94 KiB |
|
After Width: | Height: | Size: 42 KiB |
|
After Width: | Height: | Size: 35 KiB |
|
After Width: | Height: | Size: 44 KiB |
@@ -0,0 +1,113 @@
|
||||
# Alerts and incidents
|
||||
|
||||
_How alerts become incidents and how incidents are worked, notified, escalated and expired._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
There are two objects, and the difference between them is the whole design.
|
||||
|
||||
**An alert is Alertmanager's record.** It has two states, `firing` and
|
||||
`resolved`, one row per fingerprint, and no human ever writes to it. The API
|
||||
exposes alerts read-only.
|
||||
|
||||
**An incident is the work item.** It goes `triggered → acknowledged → resolved`,
|
||||
carries an assignee, a snooze, notes and a timeline, and is the only thing people
|
||||
act on. Many alerts belong to one incident.
|
||||
|
||||
## Correlation uses Alertmanager's `groupKey`
|
||||
|
||||
Alertmanager has already grouped alerts according to the `group_by` routing tree
|
||||
you configured, and it sends the resulting `groupKey` and `groupLabels` on every
|
||||
webhook. Incidents adopt that answer rather than re-grouping alerts a second
|
||||
time — if you want different correlation, change `group_by` in
|
||||
`alertmanager.yml` and terdut follows.
|
||||
|
||||
At most one incident is open per `groupKey` at a time. Alerts firing in a group
|
||||
that already has an open incident join it. The incident's `severity` is a
|
||||
high-water mark — the highest `severity` label any of its alerts has carried — so
|
||||
an incident that hit `critical` still reads as critical after the critical alert
|
||||
clears.
|
||||
|
||||
## Several clusters, one team
|
||||
|
||||
A team with one Alertmanager per Kubernetes cluster, each posting to its own
|
||||
source, needs two settings or the clusters run together.
|
||||
|
||||
1. Give every alert a `cluster` label at the source. In Prometheus that is
|
||||
`externalLabels: {cluster: prod-eu}` (kube-prometheus-stack:
|
||||
`prometheus.prometheusSpec.externalLabels`).
|
||||
2. Add `cluster` to `group_by` in `alertmanager.yml`.
|
||||
|
||||
The second one is the one that matters. Incidents are matched on the team and
|
||||
Alertmanager's `groupKey`, and the `groupKey` does not include external labels:
|
||||
without `cluster` in `group_by`, the same alert in two clusters has the same
|
||||
key and joins one incident. With it, each cluster gets its own, `cluster` is in
|
||||
the incident's `group_labels`, and the web UI shows it as a coloured chip on the
|
||||
queue, the incident and the alert list, instead of leaving it in the title.
|
||||
An alert that is not grouped by `cluster` still shows the chip on the alert
|
||||
list, which reads the label from the alert itself.
|
||||
|
||||
The queue has a cluster dropdown once there are two or more values to choose
|
||||
between. It filters on the incident's `cluster` group label
|
||||
(`GET /api/incidents?cluster=...`), so it only sees incidents grouped by it.
|
||||
|
||||
## An incident opens only on a new occurrence
|
||||
|
||||
An incident opens when an alert **transitions into firing**: a fingerprint that
|
||||
was never seen, an alert with a newer `startsAt`, or a resolved alert that
|
||||
started again. The unchanged firing notifications Alertmanager re-sends every
|
||||
`repeat_interval` are none of those, and open nothing.
|
||||
|
||||
This is what makes closing an incident by hand mean something. Without the rule,
|
||||
`POST /api/incidents/{id}/resolve` would be undone by the next re-send of an
|
||||
alert that never stopped firing.
|
||||
|
||||
## Leaving the open state
|
||||
|
||||
- **Automatically**, once every alert under the incident has stopped firing —
|
||||
whether by a resolved webhook or by the sweeper's
|
||||
[stale-alert expiry](#stale-alert-expiry). The incident gets
|
||||
`"resolution_source": "alerts"`.
|
||||
- **By hand**, via `POST /api/incidents/{id}/resolve`
|
||||
(`"resolution_source": "manual"`). This is **terminal**: a later occurrence in
|
||||
that group opens a *new* incident rather than reopening this one. If the alert
|
||||
underneath never stops firing, the incident stays closed — that is what
|
||||
resolving by hand asserts.
|
||||
- **On recovery**, for a [dead man's switch](./dead-mans-switch.md) incident whose
|
||||
heartbeat started arriving again (`"resolution_source": "recovered"`). These
|
||||
incidents have no member alerts, so the automatic cascade above cannot reach
|
||||
them.
|
||||
|
||||
To quieten an incident you expect to come back, snooze it instead
|
||||
(`POST /api/incidents/{id}/snooze`). A snooze hides the incident from the default
|
||||
list without closing it, and expires by simply falling into the past.
|
||||
|
||||
## On-call assignment
|
||||
|
||||
A new incident is assigned to whoever holds today's schedule entry at the moment
|
||||
it opens (`GET /api/schedule/current`). If nobody is scheduled it opens
|
||||
unassigned. Reassign with `POST /api/incidents/{id}/assign`.
|
||||
|
||||
One person holds a given day, so `POST /api/schedule` refuses a date somebody
|
||||
already has: taking a shift off the person expecting to be paged for it should
|
||||
not be something a plain call does by accident. Pass `"replace": true` to take
|
||||
them anyway. Either way the whole request is one transaction — a week where some
|
||||
days are free and some are taken moves as a unit, and a failure leaves the rota
|
||||
exactly as it was rather than with a hole in it.
|
||||
|
||||
## Stale alert expiry
|
||||
|
||||
A resolved webhook is the only signal that an alert has stopped firing, so a
|
||||
notification that is dropped, silenced, or lost to a restart would otherwise pin
|
||||
that alert as firing forever. A background sweeper resolves firing alerts that
|
||||
Alertmanager has stopped refreshing, using either signal:
|
||||
|
||||
- the `endsAt` watermark on the last notification has passed, or
|
||||
- no webhook has refreshed the alert within `TERDUT_STALE_AFTER`.
|
||||
|
||||
Alertmanager re-sends firing notifications every `repeat_interval`, which is what
|
||||
keeps a live alert fresh — so `TERDUT_STALE_AFTER` must be comfortably larger
|
||||
than your `repeat_interval` (default 4h), or live alerts will be resolved
|
||||
prematurely. Alerts resolved this way are marked `"resolution_source": "expiry"`
|
||||
to distinguish them from a real Alertmanager resolve (`"alertmanager"`).
|
||||
|
||||
An expiry cascades: once it leaves an incident with nothing firing under it, the
|
||||
incident resolves too, in the same sweep.
|
||||
@@ -0,0 +1,62 @@
|
||||
# Push notifications
|
||||
|
||||
_Pages through ntfy, who gets them and how to acknowledge from the notification._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
With `TERDUT_NTFY_URL` set, an incident that opens is pushed to the on-call
|
||||
person's phone through [ntfy](https://ntfy.sh). Everybody sets their own topic
|
||||
under *Account* in the web UI, where a **Send a test push** button proves it
|
||||
before an incident has to; `PUT /api/users/{id}/notify` is the same thing over
|
||||
the API, and an administrator may set somebody else's. A user with no topic
|
||||
falls back to `TERDUT_NTFY_FALLBACK_TOPIC`, as does an incident that opens with
|
||||
nobody on call. If neither yields a topic, nothing is queued.
|
||||
|
||||
The **server** is the install's one ntfy, from `TERDUT_NTFY_URL`, and is not
|
||||
something a user picks. Only the topic is per-person.
|
||||
|
||||
A topic is a shared secret with the ntfy server: anyone who knows it can both
|
||||
read the pages and publish to it, so an unguessable one is worth the trouble.
|
||||
That is also why the topic never appears in an incident's timeline, which every
|
||||
API key can read.
|
||||
|
||||
Three things get pushed:
|
||||
|
||||
- **triggered** — an incident opened. Priority follows severity (`critical` maps
|
||||
to ntfy's max priority, the one that overrides the phone's quiet settings).
|
||||
- **reminder** — the incident is still `triggered` after `TERDUT_NOTIFY_REPEAT`.
|
||||
Repeats until somebody acts. Acknowledging, snoozing, resolving or archiving
|
||||
all stop it — snooze is the mute button.
|
||||
- **resolved** — every alert under the incident stopped firing. Only sent to
|
||||
whoever was paged in the first place, and only for the automatic cascade:
|
||||
resolving by hand pushes nothing, since the person who did it already knows.
|
||||
|
||||
Notifications carry an **Acknowledge** button that acknowledges the incident
|
||||
without opening anything. It POSTs to `/api/notify/ack/{token}`, an
|
||||
unauthenticated route authorised by the 256-bit token in its path — minted fresh
|
||||
per notification, scoped to one incident and one action, and valid for 24 hours.
|
||||
A real API key is never put in a notification, because the message is stored on
|
||||
the ntfy server and cached on the device.
|
||||
|
||||
The token is **not** consumed by use. Acknowledging is idempotent, so a token
|
||||
stays valid for its full 24 hours and a second tap is a no-op that reports the
|
||||
incident's current state rather than an error — which is what you want when a
|
||||
tap is retried on a flaky mobile connection. What bounds it is scope, not a use
|
||||
count: one incident, one action, one day. Expired tokens are purged by the
|
||||
sweeper.
|
||||
|
||||
Two consequences worth planning for:
|
||||
|
||||
- `/api/notify/ack/{token}` **must stay publicly reachable**, or the button will
|
||||
not work when the responder is off your network.
|
||||
- Notifications sent to the fallback topic carry **no** Acknowledge button. The
|
||||
topic is shared, and a button on it would let any subscriber acknowledge as
|
||||
somebody else.
|
||||
|
||||
Delivery is a queue, not an inline call: the webhook writes a row and a
|
||||
background notifier sends it within 30 seconds, retrying with exponential
|
||||
backoff up to 8 attempts. Nothing about ingestion blocks on ntfy being reachable.
|
||||
|
||||
Every delivery is recorded on the incident's timeline: a `notified` event once
|
||||
ntfy accepts the publish, and a `notify_failed` event when a notification
|
||||
exhausts its retries. Written from the result rather than at enqueue, so the
|
||||
timeline says what actually happened — and a page that never landed is visible
|
||||
instead of looking the same as one that did.
|
||||
@@ -0,0 +1,105 @@
|
||||
# Single sign-on (OIDC)
|
||||
|
||||
_Signing in through an OpenID Connect provider, and mapping its groups to teams and administrators._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
terdut can sign people in through any OpenID Connect provider; the examples use
|
||||
[Authentik](https://goauthentik.io/). Groups at the provider decide who may sign
|
||||
in, which teams they belong to and whether they administer the install, much as
|
||||
Grafana's OAuth role and org mapping does. Password login keeps working alongside
|
||||
it unless you turn it off.
|
||||
|
||||
**At the provider**, create an OAuth2/OpenID provider and an application for it:
|
||||
a *confidential* client, redirect URI `<TERDUT_PUBLIC_URL>/api/oidc/callback`, and
|
||||
the `openid`, `profile` and `email` scopes. The issuer is the application's, e.g.
|
||||
`https://auth.example.com/application/o/terdut/`. Then set:
|
||||
|
||||
```sh
|
||||
TERDUT_PUBLIC_URL=https://terdut.example.com
|
||||
TERDUT_OIDC_ISSUER=https://auth.example.com/application/o/terdut/
|
||||
TERDUT_OIDC_CLIENT_ID=terdut
|
||||
TERDUT_OIDC_CLIENT_SECRET=...
|
||||
TERDUT_OIDC_ALLOWED_GROUPS=terdut-users,terdut-admins
|
||||
TERDUT_OIDC_ADMIN_GROUP=terdut-admins
|
||||
```
|
||||
|
||||
Which team a group grants is not server-wide config: each team names its own
|
||||
group(s), set by that team's own owner (or an administrator) from its Members
|
||||
tab, or `PUT /api/teams/{teamID}/oidc-groups {"member_group":"sre","owner_group":"sre-leads"}`.
|
||||
A team must already exist before a group can grant access to it — the sync
|
||||
never creates one.
|
||||
|
||||
The web UI's sign-in page shows a "Sign in with <name>" button (a plain link to
|
||||
`/api/oidc/login`) above the password form, or instead of it when
|
||||
`TERDUT_PASSWORD_LOGIN=false`; it asks `GET /api/auth/config` what the server offers
|
||||
(`password_login`, `oidc.enabled`, `oidc.name`). A refused sign-in comes back to that
|
||||
page with the reason spelled out. Access the groups grant is badged **SSO** on the
|
||||
Team, Admin and per-user pages, with its edit and remove controls disabled, and the
|
||||
Account page does not offer to set a password nobody could use.
|
||||
|
||||
**What a sign-in does**
|
||||
|
||||
1. *Who.* The provider's `(issuer, subject)` is the identity. The first time, a
|
||||
user is found by email — only when the provider marks it verified, or
|
||||
`TERDUT_OIDC_TRUST_EMAIL` is set — or created with no password. A username taken
|
||||
by somebody else gets a numeric suffix (`alice-2`). Username and email follow the
|
||||
provider at each sign-in. Authentik reports `email_verified` as false unless
|
||||
configured otherwise, so linking existing users usually needs
|
||||
`TERDUT_OIDC_TRUST_EMAIL=true`.
|
||||
2. *Whether.* With `TERDUT_OIDC_ALLOWED_GROUPS` set, somebody in none of them is
|
||||
refused and nothing is created.
|
||||
3. *What.* The administrator flag follows `TERDUT_OIDC_ADMIN_GROUP`. Team roles
|
||||
follow each team's own `oidc_member_group`/`oidc_owner_group`; where both of a
|
||||
team's groups match, the owner group wins.
|
||||
|
||||
**Managed access.** What the sync grants is marked as managed by single sign-on,
|
||||
and only that is ever changed by it. It is added at sign-in, and removed at the
|
||||
next sign-in after the group is gone, even if that leaves a team without an owner
|
||||
(an administrator can always repair a team) — the provider is the source of truth
|
||||
for what it grants, so the last-owner and last-administrator guards do not apply.
|
||||
Memberships and administrators added by hand are left alone; the exception is a
|
||||
hand-added member whose team's own group grants a *higher* role, who is raised and
|
||||
from then on managed. Editing managed access by hand (`POST` or `DELETE` on a
|
||||
team's members, revoking an SSO-granted administrator) is refused with `409`, since
|
||||
the next sign-in would undo it.
|
||||
|
||||
> **Upgrading past migration 013: reconfigure every team's groups.**
|
||||
> `TERDUT_OIDC_GROUP_MAPPINGS` is gone, and the sync no longer creates a team by
|
||||
> name. Group-to-team-role mapping is now each team's own setting — an owner sets
|
||||
> it from the Members tab, or `PUT /api/teams/{teamID}/oidc-groups`. Until a team's
|
||||
> owner does that, an OIDC-sourced membership in it is dropped at that user's next
|
||||
> SSO sign-in, the same as any other loss of group access. Set every team's groups
|
||||
> before affected users next sign in, to avoid a visible gap in access.
|
||||
|
||||
**How fast changes arrive.** Groups are read only at sign-in. A session made by an
|
||||
SSO sign-in has a hard ceiling (`TERDUT_OIDC_SESSION_MAX_AGE`, default 12h) that
|
||||
sliding never extends, so a change at the provider reaches terdut within that time.
|
||||
Password sessions are unaffected.
|
||||
|
||||
> **API keys are not revoked when somebody is removed at the provider.** terdut
|
||||
> holds no refresh token and never asks the provider again, so a person removed
|
||||
> from every allowed group loses their sessions within `TERDUT_OIDC_SESSION_MAX_AGE`
|
||||
> and cannot sign in again, but keeps any API key they made (the TUI and scripts use
|
||||
> them) until an administrator disables the user in terdut.
|
||||
|
||||
**Signing in from a terminal.** A client with no browser of its own, such as the
|
||||
TUI over SSH, signs in with a device code, run by terdut itself so the terminal
|
||||
never talks to the provider:
|
||||
|
||||
1. The terminal calls `POST /api/oidc/device` and shows the person a link
|
||||
(`<TERDUT_PUBLIC_URL>/device?code=XXXX-XXXX`) and the code.
|
||||
2. On any device the person opens the link, signs in (by the provider or by
|
||||
password, whatever the login page offers), sees the code and the account, and
|
||||
presses **Approve**. Only a browser session can approve; an API key cannot.
|
||||
3. The terminal polls `POST /api/oidc/device/token` every 5 seconds and is given the
|
||||
ordinary `terdut_session` cookie once. A person who signs in through the provider
|
||||
gets the same `TERDUT_OIDC_SESSION_MAX_AGE` ceiling on the terminal's session as
|
||||
on their browser's.
|
||||
|
||||
A login expires after 10 minutes. `GET /api/auth/config` reports `device_login`.
|
||||
|
||||
**If the provider is down**, terdut still starts (discovery is fetched on first
|
||||
use) and password login is the way in. With `TERDUT_PASSWORD_LOGIN=false` that way
|
||||
is closed: set it back to `true`. The first administrator comes from the bootstrap
|
||||
endpoint, and stays a manual administrator that no group can revoke; on an SSO-only
|
||||
install set `bootstrap.enabled: false` in the chart if you don't want that account,
|
||||
or keep it and never give it a password.
|
||||
@@ -0,0 +1,90 @@
|
||||
# The web UI
|
||||
|
||||
_What the web UI offers, how sign-in and sessions work, and the Team and Admin tabs._ Back to the [README](../README.md) and the [documentation index](./README.md).
|
||||
|
||||
The server serves a web UI at `/`: the incident queue, each incident's alerts
|
||||
and timeline with every action (acknowledge, assign, snooze, note, resolve,
|
||||
archive), who is on call, the alert feed, and an *Account* tab for your own
|
||||
password and the ntfy topic your pages go to. It is built for a phone first. On a phone
|
||||
it navigates through a hamburger menu and has a sticky action bar, it follows the
|
||||
system's dark mode, and it can be added to the home screen. From 900px wide it switches
|
||||
to a sidebar with the queue and the incident side by side. The Stats page shows
|
||||
incident counts, MTTA and MTTR, and alert frequency by name, hour and day over a
|
||||
chosen range.
|
||||
|
||||
You sign in with a username and password. Users have no password until one is
|
||||
set, and a user without one can only use API keys:
|
||||
|
||||
```bash
|
||||
# an admin sets someone's first password with their API key
|
||||
curl -X PUT http://localhost:8080/api/users/2/password \
|
||||
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
|
||||
-d '{"password": "<at least 10 characters>"}'
|
||||
```
|
||||
|
||||
After that, users change it themselves under *Account*. Changing your own
|
||||
password requires the current one.
|
||||
|
||||
How a browser stays signed in:
|
||||
|
||||
- A successful login sets an `HttpOnly`, `SameSite=Lax` session cookie. It lasts
|
||||
30 days and slides forward while it is used, so an on-call phone stays signed
|
||||
in.
|
||||
- The cookie is marked `Secure` when `TERDUT_PUBLIC_URL` starts with `https://`,
|
||||
so set it to the HTTPS address. TLS terminates at the gateway and the server
|
||||
itself only ever sees plain HTTP.
|
||||
- Requests authenticated by the cookie are checked for cross-origin use (Go's
|
||||
`http.CrossOriginProtection`). That is the CSRF guard. Bearer-key clients are
|
||||
not affected.
|
||||
- Setting a password signs that user out everywhere else.
|
||||
- Ten failed logins for one username within 15 minutes lock that username for
|
||||
the rest of the window.
|
||||
|
||||
With `TERDUT_PUBLIC_URL` set, tapping a push notification opens the incident in
|
||||
the web UI (`/incidents/{id}`).
|
||||
|
||||
A **Team** tab holds everything a team owns, in five sub-sections with a URL
|
||||
each and a strip across the top to move between them: the on-call rota
|
||||
(`/team/rota`), the membership (`/team/members`), the escalation ladder
|
||||
(`/team/escalation`), the alert sources with their keys (`/team/sources`) and
|
||||
the dead man's switches (`/team/deadman`). `/team` itself is an overview — who
|
||||
is on call today, how many members and owners, how many ladder levels, how many
|
||||
keys and how many switches — so a page fetches only what it shows. An owner
|
||||
edits it; a member sees the same pages read-only, because the server refuses
|
||||
their writes anyway. Somebody in more than one team picks between them above
|
||||
the strip, since the choice changes the subject of all five.
|
||||
|
||||
The rota is a month at a time, one coloured initial per day with a legend
|
||||
underneath, and it says how many days are left uncovered — the question a rota
|
||||
is read for is who holds which stretch, and a run of one colour answers it
|
||||
where a list of dates does not. An owner taps a day to hand it to somebody or
|
||||
empty it, and fills a whole shift from the range form folded in below.
|
||||
|
||||
The **Admin** tab appears only for a system administrator, and holds what
|
||||
belongs to the whole server rather than to one team. It has three sub-sections,
|
||||
each with a URL of its own and a strip across the top to move between them:
|
||||
every team (`/admin/teams`), every user (`/admin/users`), and the settings that
|
||||
used to be environment variables (`/admin/settings`). `/admin` itself is an
|
||||
overview — how many of each, and what each section is for. Adding somebody is
|
||||
minting them an invite link into a team, rather than creating a bare account:
|
||||
the person who accepts it picks their own password, so one never passes through
|
||||
an administrator, and the link carries the team, so they land somewhere with a
|
||||
queue in it. That happens on the team's own page, since an invite is a fact
|
||||
about a team; the user list points there rather than asking which team beside a
|
||||
form.
|
||||
|
||||
A name in the team list opens **that team's page**, at `/admin/teams/{id}`: when it
|
||||
was created, how many are in it and how much is open, a field to rename it, the
|
||||
members with their roles, the invites into it, and deletion. The member list is the
|
||||
one thing there that needed a new endpoint — `GET /api/teams/{id}/members` is
|
||||
member-only and answers `404` to an administrator who is not in the team, which is
|
||||
the rule and not an oversight, so the page reads `GET /api/admin/teams/{id}` instead.
|
||||
An administrator still sees none of that team's incidents, alerts or rota.
|
||||
|
||||
A name in the user list opens **that person's page**, at `/admin/users/{id}`: their
|
||||
email and when they joined, where their notifications go, whether they are an
|
||||
administrator, whether the account is disabled, the teams they are in with their
|
||||
role in each, a password field for a first or forgotten one, and deletion. It is
|
||||
the one place membership is edited from the person's side — the Team tab answers
|
||||
"who is in this team", and answering "which teams is this person in" there means
|
||||
visiting each team in turn.
|
||||
@@ -1,182 +0,0 @@
|
||||
-- The Postgres baseline: the schema as it stood at the end of the SQLite line,
|
||||
-- in one file rather than ten.
|
||||
--
|
||||
-- The ten SQLite migrations are in git history up to the commit that introduced
|
||||
-- this one, and they replay against nothing here: their shape was incremental
|
||||
-- (columns added, then dropped again in 008) and 008's backfill rewrote data
|
||||
-- that a Postgres install never had. An existing SQLite database is carried over
|
||||
-- by scripts/sqlite-to-postgres.go, which copies rows into this schema.
|
||||
--
|
||||
-- Two conventions inherited deliberately:
|
||||
--
|
||||
-- * Timestamps are BIGINT unix seconds, not timestamptz. Everything in Go
|
||||
-- already speaks epochs, and converting was a second change riding along
|
||||
-- with the port. Worth revisiting on its own.
|
||||
--
|
||||
-- * Ids are GENERATED BY DEFAULT, not ALWAYS, so the migration script can
|
||||
-- insert rows with their original ids and keep every foreign key intact.
|
||||
-- setval at the end of the copy puts the sequences past them.
|
||||
|
||||
CREATE TABLE users (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
username TEXT NOT NULL UNIQUE,
|
||||
email TEXT NOT NULL UNIQUE,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
-- Where this user's notifications go. NULL means they get none; incidents
|
||||
-- assigned to them fall back to the configured fallback topic.
|
||||
ntfy_topic TEXT,
|
||||
-- NULL means the user has no password and can only use API keys.
|
||||
password_hash TEXT
|
||||
);
|
||||
|
||||
CREATE TABLE api_keys (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
user_id BIGINT NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
||||
key_hash TEXT NOT NULL UNIQUE,
|
||||
name TEXT NOT NULL,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
last_used_at BIGINT
|
||||
);
|
||||
|
||||
-- A session is a browser's credential, the cookie counterpart of an API key:
|
||||
-- only the hash of the token is stored. expires_at slides forward while the
|
||||
-- session is in use, so an on-call phone stays signed in.
|
||||
CREATE TABLE sessions (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
token_hash TEXT NOT NULL UNIQUE,
|
||||
user_id BIGINT NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
||||
created_at BIGINT NOT NULL,
|
||||
last_seen_at BIGINT NOT NULL,
|
||||
expires_at BIGINT NOT NULL,
|
||||
user_agent TEXT
|
||||
);
|
||||
|
||||
CREATE INDEX idx_sessions_user ON sessions(user_id);
|
||||
|
||||
-- The machine-owned signal record: what Alertmanager says is true right now.
|
||||
-- Workflow state lives on incidents, never here, because the webhook upsert owns
|
||||
-- these rows and would overwrite it.
|
||||
CREATE TABLE alerts (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
fingerprint TEXT NOT NULL UNIQUE,
|
||||
name TEXT NOT NULL,
|
||||
status TEXT NOT NULL CHECK (status IN ('firing', 'resolved')),
|
||||
labels JSONB NOT NULL DEFAULT '{}'::jsonb,
|
||||
annotations JSONB NOT NULL DEFAULT '{}'::jsonb,
|
||||
starts_at BIGINT NOT NULL,
|
||||
ends_at BIGINT,
|
||||
generator_url TEXT NOT NULL DEFAULT '',
|
||||
received_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
archived_at BIGINT,
|
||||
-- Why the alert left the firing state: 'alertmanager' when a resolved
|
||||
-- webhook set it, 'expiry' when the sweeper inferred it from staleness.
|
||||
resolution_source TEXT
|
||||
);
|
||||
|
||||
CREATE INDEX alerts_status_idx ON alerts(status);
|
||||
CREATE INDEX alerts_name_idx ON alerts(name);
|
||||
CREATE INDEX alerts_received_at_idx ON alerts(received_at DESC);
|
||||
CREATE INDEX alerts_archived_at_idx ON alerts(archived_at);
|
||||
|
||||
CREATE TABLE schedule_entries (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
user_id BIGINT NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
||||
date TEXT NOT NULL UNIQUE, -- YYYY-MM-DD; one person per day
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
|
||||
CREATE INDEX schedule_entries_date_idx ON schedule_entries(date);
|
||||
|
||||
-- The human work item: what people acknowledge, assign, snooze, discuss and
|
||||
-- resolve. Correlation uses Alertmanager's own groupKey, so incidents follow the
|
||||
-- group_by routing tree the operator already tuned.
|
||||
CREATE TABLE incidents (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
group_key TEXT NOT NULL, -- Alertmanager groupKey, opaque
|
||||
title TEXT NOT NULL, -- rendered from group_labels
|
||||
group_labels JSONB NOT NULL DEFAULT '{}'::jsonb,
|
||||
status TEXT NOT NULL CHECK (status IN ('triggered', 'acknowledged', 'resolved')),
|
||||
severity TEXT, -- highest `severity` label across firing members
|
||||
triggered_at BIGINT NOT NULL,
|
||||
acknowledged_by BIGINT REFERENCES users(id) ON DELETE SET NULL,
|
||||
acknowledged_at BIGINT,
|
||||
assigned_to BIGINT REFERENCES users(id) ON DELETE SET NULL,
|
||||
snoozed_until BIGINT,
|
||||
resolved_at BIGINT,
|
||||
resolution_source TEXT, -- 'alerts' | 'manual'
|
||||
archived_at BIGINT
|
||||
);
|
||||
|
||||
-- Load-bearing: at most one OPEN incident per group_key. This is what makes
|
||||
-- "resolved incident + a new alert occurrence = a new incident" work, and it is
|
||||
-- the constraint the webhook's find-or-open lookup relies on.
|
||||
CREATE UNIQUE INDEX incidents_open_group_key_idx ON incidents(group_key) WHERE resolved_at IS NULL;
|
||||
CREATE INDEX incidents_status_idx ON incidents(status);
|
||||
CREATE INDEX incidents_triggered_at_idx ON incidents(triggered_at DESC);
|
||||
CREATE INDEX incidents_archived_at_idx ON incidents(archived_at);
|
||||
|
||||
-- Membership is historical, not a pointer on alerts: one alert row (one
|
||||
-- fingerprint) resolves and re-fires over time and belongs to a different
|
||||
-- incident each occurrence.
|
||||
CREATE TABLE incident_alerts (
|
||||
incident_id BIGINT NOT NULL REFERENCES incidents(id) ON DELETE CASCADE,
|
||||
alert_id BIGINT NOT NULL REFERENCES alerts(id) ON DELETE CASCADE,
|
||||
added_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
PRIMARY KEY (incident_id, alert_id)
|
||||
);
|
||||
|
||||
CREATE INDEX incident_alerts_alert_id_idx ON incident_alerts(alert_id);
|
||||
|
||||
-- The timeline. Append-only, and the only history this server keeps: alert rows
|
||||
-- are mutated in place, so without this there is no record that anything
|
||||
-- happened. Notes are events too, so one query renders the whole story.
|
||||
CREATE TABLE incident_events (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
incident_id BIGINT NOT NULL REFERENCES incidents(id) ON DELETE CASCADE,
|
||||
-- triggered | alert_added | alert_resolved | acknowledged | unacknowledged
|
||||
-- | assigned | snoozed | unsnoozed | resolved | note | notified | notify_failed
|
||||
type TEXT NOT NULL,
|
||||
user_id BIGINT REFERENCES users(id) ON DELETE SET NULL, -- NULL = the server acted
|
||||
alert_id BIGINT REFERENCES alerts(id) ON DELETE SET NULL,
|
||||
detail TEXT,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
|
||||
CREATE INDEX incident_events_incident_idx ON incident_events(incident_id, created_at);
|
||||
|
||||
-- Delivery is an outbox rather than an inline HTTP call: a POST made while
|
||||
-- holding the webhook's transaction would hold a connection open across a
|
||||
-- network round trip. The webhook inserts a row; the notifier goroutine
|
||||
-- delivers it.
|
||||
CREATE TABLE notifications (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
incident_id BIGINT NOT NULL REFERENCES incidents(id) ON DELETE CASCADE,
|
||||
-- Nullable: a notification sent to the fallback topic belongs to nobody,
|
||||
-- because nobody was on call when the incident opened.
|
||||
user_id BIGINT REFERENCES users(id) ON DELETE SET NULL,
|
||||
topic TEXT NOT NULL, -- resolved at enqueue: who was on call then
|
||||
kind TEXT NOT NULL CHECK (kind IN ('triggered', 'reminder', 'resolved')),
|
||||
created_at BIGINT NOT NULL,
|
||||
send_after BIGINT NOT NULL, -- retry backoff watermark
|
||||
attempts BIGINT NOT NULL DEFAULT 0,
|
||||
sent_at BIGINT,
|
||||
last_error TEXT -- kept after the last attempt, for debugging
|
||||
);
|
||||
|
||||
-- The delivery loop's only query: what is due and still unsent.
|
||||
CREATE INDEX notifications_pending_idx ON notifications(send_after) WHERE sent_at IS NULL;
|
||||
-- Reminders and resolved notices both look up an incident's newest row.
|
||||
CREATE INDEX notifications_incident_idx ON notifications(incident_id, id DESC);
|
||||
|
||||
-- A notification body is stored on the ntfy server and cached on the device, so
|
||||
-- a real API key must never appear in one. Each delivery mints its own token
|
||||
-- instead: one incident, one action, one day.
|
||||
CREATE TABLE incident_ack_tokens (
|
||||
token_hash TEXT PRIMARY KEY, -- SHA-256 of the raw token, as with api_keys
|
||||
incident_id BIGINT NOT NULL REFERENCES incidents(id) ON DELETE CASCADE,
|
||||
user_id BIGINT NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
||||
created_at BIGINT NOT NULL,
|
||||
expires_at BIGINT NOT NULL
|
||||
);
|
||||
|
||||
CREATE INDEX incident_ack_tokens_expires_idx ON incident_ack_tokens(expires_at);
|
||||
@@ -1,25 +0,0 @@
|
||||
-- A system administrator role, and the first thing in this server that one user
|
||||
-- can do and another cannot.
|
||||
--
|
||||
-- Until now every authenticated caller could create and delete users, set
|
||||
-- anybody's password and mint API keys for anybody — auth.go said so in a
|
||||
-- comment. That was defensible with one operator and a hand-made account; it is
|
||||
-- not once people sign themselves up (see #7).
|
||||
--
|
||||
-- EVERY EXISTING USER BECOMES AN ADMIN. They already hold these powers, so
|
||||
-- this migration changes nobody's access: it names what is already true, and
|
||||
-- leaves demotion as a deliberate act somebody performs afterwards. The
|
||||
-- alternative — promoting only user 1 — would silently strip the others, and
|
||||
-- could leave an install whose only admin is an account nobody has a password
|
||||
-- for.
|
||||
--
|
||||
-- New users are not admins: the column defaults to false, and the only ways to
|
||||
-- become one are this backfill, the bootstrap endpoint, or an existing admin
|
||||
-- granting it.
|
||||
ALTER TABLE users ADD COLUMN is_admin BOOLEAN NOT NULL DEFAULT false;
|
||||
|
||||
UPDATE users SET is_admin = true;
|
||||
|
||||
-- The queue's assignment dropdown and the on-call schedule read every user, and
|
||||
-- the admin screens in #5 will filter on this.
|
||||
CREATE INDEX users_is_admin_idx ON users(is_admin) WHERE is_admin;
|
||||
@@ -1,103 +0,0 @@
|
||||
-- Teams: the unit of tenancy. Everything a person works on now belongs to one.
|
||||
--
|
||||
-- Until this migration the install was one shared space — every user saw every
|
||||
-- alert and every incident, and the Alertmanager webhook was unauthenticated, so
|
||||
-- anything that could reach the port could open an incident for everybody.
|
||||
--
|
||||
-- The shape, in one paragraph: a team owns its incidents, alerts, schedule and
|
||||
-- integrations. A user belongs to as many teams as they like, with a role in
|
||||
-- each: an `owner` configures the team, a `member` works its incidents. An
|
||||
-- integration key is what an alert arrives on, and the key is what says which
|
||||
-- team the alert belongs to.
|
||||
--
|
||||
-- EVERYTHING EXISTING MOVES INTO ONE DEFAULT TEAM, and every existing user
|
||||
-- becomes an owner of it. That keeps an upgrade a no-op for the people using it:
|
||||
-- the same queue, the same schedule, the same incidents, with a name on them.
|
||||
|
||||
CREATE TABLE teams (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
name TEXT NOT NULL UNIQUE,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
|
||||
-- role is free text with a CHECK rather than an enum, so adding a third role
|
||||
-- later is a migration and not a type rewrite.
|
||||
CREATE TABLE team_members (
|
||||
team_id BIGINT NOT NULL REFERENCES teams(id) ON DELETE CASCADE,
|
||||
user_id BIGINT NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
||||
role TEXT NOT NULL CHECK (role IN ('owner', 'member')),
|
||||
joined_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
PRIMARY KEY (team_id, user_id)
|
||||
);
|
||||
|
||||
CREATE INDEX team_members_user_idx ON team_members(user_id);
|
||||
|
||||
-- How alerts get in, and the only thing that says which team they belong to.
|
||||
-- The key is stored as a SHA-256 hash, like api_keys and the ack tokens: a
|
||||
-- leaked database gives nobody the ability to post alerts.
|
||||
CREATE TABLE integrations (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
team_id BIGINT NOT NULL REFERENCES teams(id) ON DELETE CASCADE,
|
||||
kind TEXT NOT NULL CHECK (kind IN ('alertmanager')),
|
||||
name TEXT NOT NULL,
|
||||
key_hash TEXT NOT NULL UNIQUE,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
last_used_at BIGINT
|
||||
);
|
||||
|
||||
CREATE INDEX integrations_team_idx ON integrations(team_id);
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- The default team, and everything that already exists moving into it.
|
||||
--
|
||||
-- Created unconditionally, even on an empty install, so there is always a team
|
||||
-- for the bootstrap user to land in and for the first integration to hang off.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
INSERT INTO teams (name) VALUES ('Default');
|
||||
|
||||
INSERT INTO team_members (team_id, user_id, role)
|
||||
SELECT (SELECT id FROM teams WHERE name = 'Default'), id, 'owner' FROM users;
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- team_id on everything a team owns.
|
||||
--
|
||||
-- Added nullable, backfilled, then made NOT NULL: adding a NOT NULL column with
|
||||
-- no default to a table with rows is rejected, and a DEFAULT pointing at the
|
||||
-- default team would quietly keep working after the default team is gone.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
ALTER TABLE alerts ADD COLUMN team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE;
|
||||
ALTER TABLE incidents ADD COLUMN team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE;
|
||||
ALTER TABLE schedule_entries ADD COLUMN team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE;
|
||||
|
||||
UPDATE alerts SET team_id = (SELECT id FROM teams WHERE name = 'Default');
|
||||
UPDATE incidents SET team_id = (SELECT id FROM teams WHERE name = 'Default');
|
||||
UPDATE schedule_entries SET team_id = (SELECT id FROM teams WHERE name = 'Default');
|
||||
|
||||
ALTER TABLE alerts ALTER COLUMN team_id SET NOT NULL;
|
||||
ALTER TABLE incidents ALTER COLUMN team_id SET NOT NULL;
|
||||
ALTER TABLE schedule_entries ALTER COLUMN team_id SET NOT NULL;
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- The uniqueness rules were all written for one tenant, and every one of them
|
||||
-- is wrong now: two teams monitoring two clusters legitimately see the same
|
||||
-- fingerprint, the same groupKey, and want somebody on call on the same day.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
ALTER TABLE alerts DROP CONSTRAINT alerts_fingerprint_key;
|
||||
CREATE UNIQUE INDEX alerts_team_fingerprint_idx ON alerts(team_id, fingerprint);
|
||||
|
||||
DROP INDEX incidents_open_group_key_idx;
|
||||
-- Still load-bearing, now per team: at most one OPEN incident per group_key
|
||||
-- within a team. This is what makes "resolved incident + a new alert occurrence
|
||||
-- = a new incident" work, and what the webhook's find-or-open lookup relies on.
|
||||
CREATE UNIQUE INDEX incidents_open_group_key_idx
|
||||
ON incidents(team_id, group_key) WHERE resolved_at IS NULL;
|
||||
|
||||
ALTER TABLE schedule_entries DROP CONSTRAINT schedule_entries_date_key;
|
||||
CREATE UNIQUE INDEX schedule_entries_team_date_idx ON schedule_entries(team_id, date);
|
||||
|
||||
-- The list views all filter by team first.
|
||||
CREATE INDEX alerts_team_received_idx ON alerts(team_id, received_at DESC);
|
||||
CREATE INDEX incidents_team_triggered_idx ON incidents(team_id, triggered_at DESC);
|
||||
@@ -1,39 +0,0 @@
|
||||
-- Dead man's switches become a team's own configuration.
|
||||
--
|
||||
-- They were three environment variables — TERDUT_DEADMAN_MATCHERS, _TIMEOUT and
|
||||
-- _SEVERITY — which made them one setting for the whole install. That was the
|
||||
-- last piece of the alerting path a team could not control: a team could take
|
||||
-- its own alerts on its own key and still not say which of them were
|
||||
-- heartbeats, or how long a silence had to last before somebody was paged.
|
||||
--
|
||||
-- One row per team rather than one row per switch. The unit of monitoring is
|
||||
-- still the fingerprint, as it always was — two clusters sending the same
|
||||
-- heartbeat alertname are two independent switches — and the matcher string
|
||||
-- keeps the format the environment variable used, so a value can be moved from
|
||||
-- one to the other unchanged.
|
||||
--
|
||||
-- No rows are seeded here: a migration cannot read the environment. The server
|
||||
-- inserts a row per team at startup from its own configuration, and the same
|
||||
-- values therefore carry forward into the first team's row without anybody
|
||||
-- retyping them. See seedDeadmanConfigs.
|
||||
CREATE TABLE deadman_configs (
|
||||
team_id BIGINT PRIMARY KEY REFERENCES teams(id) ON DELETE CASCADE,
|
||||
|
||||
-- ";" separates matchers, "," the label conditions within one, "=" is exact
|
||||
-- equality: `alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat`.
|
||||
-- Every matcher must name an alertname. Empty watches nothing.
|
||||
matchers TEXT NOT NULL DEFAULT '',
|
||||
|
||||
-- Seconds rather than a Go duration string: the column is compared and
|
||||
-- arithmetic is done on it, and a value that has to be parsed before it can
|
||||
-- be believed is a value that can be stored unparseable. Zero disables the
|
||||
-- team's switches entirely.
|
||||
timeout_seconds BIGINT NOT NULL DEFAULT 0,
|
||||
|
||||
-- The severity these incidents open at. They have no member alerts to
|
||||
-- derive one from, and a heartbeat's own severity label is meaningless —
|
||||
-- Watchdog ships as "none".
|
||||
severity TEXT NOT NULL DEFAULT 'critical',
|
||||
|
||||
updated_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
@@ -1,35 +0,0 @@
|
||||
-- Settings that an administrator can change without a redeploy, and the flag
|
||||
-- that takes an account out of use without deleting it.
|
||||
--
|
||||
-- Three of the server's tunables were environment variables, which meant
|
||||
-- changing how long an incident waits before it is paged again required editing
|
||||
-- a chart, merging it, and waiting for a reconcile. They are behaviour, not
|
||||
-- infrastructure, and the difference is who needs to change them and how often.
|
||||
--
|
||||
-- What stays in the environment: the ntfy URL and token, the database DSN, the
|
||||
-- listen address and the public URL. Those are where the server is plugged in
|
||||
-- rather than how it behaves, they are needed before the database is open, and
|
||||
-- two of them are credentials.
|
||||
--
|
||||
-- Key/value rather than a column per setting. A settings table with one row and
|
||||
-- a column per knob needs a migration for every new knob, and #6 and #7 will
|
||||
-- both add some. The cost is that values are text and the accessor has to say
|
||||
-- what type it wanted; settings.go does that in one place.
|
||||
--
|
||||
-- No rows are seeded here: a migration cannot read the environment. The server
|
||||
-- inserts each key from its own configuration at startup, once, so an install
|
||||
-- that upgrades keeps exactly the behaviour it had. See SeedSettings.
|
||||
CREATE TABLE settings (
|
||||
key TEXT PRIMARY KEY,
|
||||
value TEXT NOT NULL,
|
||||
updated_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
|
||||
-- Disabling an account rather than deleting it: the person has left, or the
|
||||
-- credential is suspect, and their incidents, acknowledgements and timeline
|
||||
-- entries must stay exactly where they are. Deleting a user nulls their
|
||||
-- acknowledged_by and assigned_to, which quietly rewrites history.
|
||||
--
|
||||
-- A disabled user cannot sign in and their API keys stop working, but they are
|
||||
-- still a name the timeline can show and still a member of their teams.
|
||||
ALTER TABLE users ADD COLUMN disabled_at BIGINT;
|
||||
@@ -1,95 +0,0 @@
|
||||
-- Escalation: page somebody else when the first person does not answer.
|
||||
--
|
||||
-- This is the gap the whole multi-tenancy line of work was opened to close.
|
||||
-- Until now an unacknowledged incident re-paged the same topic every
|
||||
-- notify_repeat forever, which is a louder version of the same silence: if the
|
||||
-- person on call is asleep, has no signal, or has left, nothing else happens.
|
||||
--
|
||||
-- Shape: one policy per team, an ordered list of levels, each level with a
|
||||
-- timeout and a set of targets. When a level's timeout passes and the incident
|
||||
-- is still triggered, the next level is paged. When the last level passes, the
|
||||
-- chain repeats repeat_count times, and then the team's fallback topic is paged
|
||||
-- once as the end of the line.
|
||||
--
|
||||
-- A team WITHOUT a policy keeps exactly today's behaviour: page the assignee,
|
||||
-- then remind on the same topic. Escalation is opt-in per team, and the two
|
||||
-- never both run for one incident -- see enqueueReminders.
|
||||
CREATE TABLE escalation_policies (
|
||||
-- One per team for now, hence the team as the key rather than an id with a
|
||||
-- unique index: routing different alerts to different chains needs the
|
||||
-- alert to carry something to route ON, which is a separate question.
|
||||
team_id BIGINT PRIMARY KEY REFERENCES teams(id) ON DELETE CASCADE,
|
||||
|
||||
-- How many extra times to run the whole chain after it has been walked
|
||||
-- once. 0 means walk it once and stop at the fallback.
|
||||
repeat_count BIGINT NOT NULL DEFAULT 0 CHECK (repeat_count >= 0 AND repeat_count <= 10),
|
||||
|
||||
-- Where the last page goes when every level has been tried. Per team now:
|
||||
-- TERDUT_NTFY_FALLBACK_TOPIC was one topic for the whole install, which in
|
||||
-- a multi-team server pages the wrong people. Empty means the chain simply
|
||||
-- ends.
|
||||
fallback_topic TEXT NOT NULL DEFAULT '',
|
||||
|
||||
updated_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
|
||||
CREATE TABLE escalation_levels (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
team_id BIGINT NOT NULL REFERENCES escalation_policies(team_id) ON DELETE CASCADE,
|
||||
-- 1-based, dense. The API rewrites the whole ladder on every edit rather
|
||||
-- than patching one rung, so there is no way to leave a gap.
|
||||
position BIGINT NOT NULL,
|
||||
-- How long this level has to produce an acknowledgement before the next one
|
||||
-- is paged. Seconds, like every other duration in this schema.
|
||||
timeout_seconds BIGINT NOT NULL CHECK (timeout_seconds > 0),
|
||||
|
||||
UNIQUE (team_id, position)
|
||||
);
|
||||
|
||||
-- Who a level pages. Either a named person, or whoever the team's rota says is
|
||||
-- on call today -- which is the target that keeps working when the rota
|
||||
-- changes and nobody remembers to edit the policy.
|
||||
CREATE TABLE escalation_targets (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
level_id BIGINT NOT NULL REFERENCES escalation_levels(id) ON DELETE CASCADE,
|
||||
kind TEXT NOT NULL CHECK (kind IN ('user', 'oncall')),
|
||||
-- Set for kind='user', NULL for kind='oncall'.
|
||||
user_id BIGINT REFERENCES users(id) ON DELETE CASCADE,
|
||||
|
||||
CHECK ((kind = 'user' AND user_id IS NOT NULL) OR (kind = 'oncall' AND user_id IS NULL))
|
||||
);
|
||||
|
||||
CREATE INDEX escalation_targets_level_idx ON escalation_targets(level_id);
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- Where an incident is in its chain.
|
||||
--
|
||||
-- On the incident rather than in a side table: it is read on every notifier
|
||||
-- tick alongside the incident's status, and one row per incident is exactly
|
||||
-- what the state is.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
-- 0 means no level has been paged yet, which is the state of every incident
|
||||
-- that existed before escalation and of every incident in a team with no
|
||||
-- policy. 1 is the first level.
|
||||
ALTER TABLE incidents ADD COLUMN escalation_level BIGINT NOT NULL DEFAULT 0;
|
||||
|
||||
-- When the current level was entered, and therefore what its timeout is
|
||||
-- measured from. NULL while escalation_level is 0.
|
||||
ALTER TABLE incidents ADD COLUMN escalation_level_at BIGINT;
|
||||
|
||||
-- How many times the chain has been walked in full. Compared against the
|
||||
-- policy's repeat_count.
|
||||
ALTER TABLE incidents ADD COLUMN escalation_round BIGINT NOT NULL DEFAULT 0;
|
||||
|
||||
-- The notifier's escalation query: incidents still waiting, oldest level first.
|
||||
CREATE INDEX incidents_escalation_idx
|
||||
ON incidents(escalation_level_at)
|
||||
WHERE resolved_at IS NULL AND status = 'triggered';
|
||||
|
||||
-- 'escalated' joins the outbox kinds: a page that went out because nobody
|
||||
-- answered the last one, which is worth telling apart from the first page and
|
||||
-- from a reminder when reading the timeline or debugging a delivery.
|
||||
ALTER TABLE notifications DROP CONSTRAINT notifications_kind_check;
|
||||
ALTER TABLE notifications ADD CONSTRAINT notifications_kind_check
|
||||
CHECK (kind IN ('triggered', 'reminder', 'resolved', 'escalated'));
|
||||
@@ -1,49 +0,0 @@
|
||||
-- Self-service sign-up, and the invite links that make it useful.
|
||||
--
|
||||
-- Until now the only way to get an account was for somebody who already had one
|
||||
-- to create it, and the login page told people to "ask an admin". That is a
|
||||
-- workable arrangement for one operator and an impossible one for a team.
|
||||
--
|
||||
-- An invite is a link, not an email: this server has no SMTP and adding it to
|
||||
-- send one message would be a new subsystem to run, secure and monitor. The
|
||||
-- person inviting sends the link however they already talk to the person they
|
||||
-- are inviting.
|
||||
CREATE TABLE invites (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
|
||||
-- SHA-256 of the raw token, like api_keys, the integration keys and the
|
||||
-- acknowledgement tokens. A leaked database hands nobody an account.
|
||||
token_hash TEXT NOT NULL UNIQUE,
|
||||
|
||||
-- Which team the invitee lands in, and as what. An invite always names a
|
||||
-- team: an account in no team sees an empty queue and can be paged by
|
||||
-- nobody, which is not a state to invite somebody into.
|
||||
team_id BIGINT NOT NULL REFERENCES teams(id) ON DELETE CASCADE,
|
||||
role TEXT NOT NULL CHECK (role IN ('owner', 'member')),
|
||||
|
||||
created_by BIGINT REFERENCES users(id) ON DELETE SET NULL,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
|
||||
-- Invites expire. A link that works forever is a credential nobody
|
||||
-- remembers issuing, sitting in a chat log.
|
||||
expires_at BIGINT NOT NULL,
|
||||
|
||||
-- Single-use by default: max_uses 1. A team onboarding six people at once
|
||||
-- can raise it rather than minting six links.
|
||||
max_uses BIGINT NOT NULL DEFAULT 1 CHECK (max_uses > 0 AND max_uses <= 100),
|
||||
uses BIGINT NOT NULL DEFAULT 0,
|
||||
|
||||
-- Revoked by hand, separately from expiry, so "this link is no longer
|
||||
-- wanted" and "this link timed out" stay distinguishable in the listing.
|
||||
revoked_at BIGINT
|
||||
);
|
||||
|
||||
CREATE INDEX invites_team_idx ON invites(team_id);
|
||||
|
||||
-- Who redeemed which invite. Kept after the invite is gone — the answer to "how
|
||||
-- did this account get here" should outlive the link that made it.
|
||||
ALTER TABLE users ADD COLUMN invited_via BIGINT REFERENCES invites(id) ON DELETE SET NULL;
|
||||
|
||||
-- Where a person is in the first-run checklist, so it can be resumed and
|
||||
-- dismissed rather than nagging forever. One row per user, created on demand.
|
||||
ALTER TABLE users ADD COLUMN onboarding_dismissed_at BIGINT;
|
||||
@@ -1,23 +0,0 @@
|
||||
-- Similar incidents: a signature per incident, so "has this happened before"
|
||||
-- is an indexed equality instead of a search.
|
||||
--
|
||||
-- The signature is the alert name plus the group labels that identify WHAT is
|
||||
-- broken, minus the ones that only say WHERE it happened to run this time
|
||||
-- (instance, pod, ...). Two incidents with the same signature in the same team
|
||||
-- are the same problem for a responder's purposes.
|
||||
--
|
||||
-- Computed in Go for new incidents (incidentSignature in incident_store.go).
|
||||
-- The backfill below MUST produce the same string; keep the volatile list in
|
||||
-- both places in step.
|
||||
ALTER TABLE incidents ADD COLUMN signature TEXT NOT NULL DEFAULT '';
|
||||
|
||||
UPDATE incidents SET signature =
|
||||
COALESCE(NULLIF(group_labels->>'alertname', ''), title) || '|' ||
|
||||
COALESCE((
|
||||
SELECT string_agg(e.k || '=' || e.v, ',' ORDER BY e.k)
|
||||
FROM jsonb_each_text(incidents.group_labels) AS e(k, v)
|
||||
WHERE e.k <> 'alertname'
|
||||
AND e.k NOT IN ('instance', 'pod', 'pod_name', 'pod_ip', 'container', 'container_name', 'endpoint')
|
||||
), '');
|
||||
|
||||
CREATE INDEX incidents_signature_idx ON incidents(team_id, signature, triggered_at DESC);
|
||||
@@ -1,54 +0,0 @@
|
||||
-- Dead man's switches become rows of their own.
|
||||
--
|
||||
-- 004 kept a team's switches in one string with one timeout and one severity,
|
||||
-- which was enough to configure them and not enough to show them: there was no
|
||||
-- thing to list, nothing to hang a status on, and every switch in a team had to
|
||||
-- share a deadline. A row per switch gives each its own name, matcher, timeout
|
||||
-- and severity, and gives the Team → Switches page something to be a list of.
|
||||
--
|
||||
-- The matcher keeps the syntax the string used, one matcher per row:
|
||||
-- `alertname=Watchdog,cluster=prod`. The unit of monitoring is still the
|
||||
-- fingerprint, so a matcher that many clusters satisfy is still one switch row
|
||||
-- watching several independent heartbeats.
|
||||
CREATE TABLE deadman_switches (
|
||||
id BIGSERIAL PRIMARY KEY,
|
||||
team_id BIGINT NOT NULL REFERENCES teams(id) ON DELETE CASCADE,
|
||||
|
||||
-- What the owner calls it. Defaults to the matcher when they do not say.
|
||||
name TEXT NOT NULL,
|
||||
|
||||
-- "," separates the label conditions, "=" is exact equality, and alertname is
|
||||
-- mandatory: it is what keeps the sweeper's candidate query on an index.
|
||||
matcher TEXT NOT NULL,
|
||||
|
||||
-- Seconds of silence before the switch is declared dead. Never zero: a switch
|
||||
-- that cannot fire is deleted, not disabled.
|
||||
timeout_seconds BIGINT NOT NULL CHECK (timeout_seconds > 0),
|
||||
|
||||
-- The severity its incidents open at. See 004 for why they carry their own.
|
||||
severity TEXT NOT NULL DEFAULT 'critical',
|
||||
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
|
||||
CREATE INDEX deadman_switches_team_idx ON deadman_switches (team_id);
|
||||
|
||||
-- Carry every team's configuration over, one row per matcher. A team whose
|
||||
-- timeout was zero had switches turned off, which is now "no rows".
|
||||
INSERT INTO deadman_switches (team_id, name, matcher, timeout_seconds, severity)
|
||||
SELECT c.team_id, btrim(m), btrim(m), c.timeout_seconds, c.severity
|
||||
FROM deadman_configs c,
|
||||
LATERAL regexp_split_to_table(c.matchers, ';') AS m
|
||||
WHERE c.timeout_seconds > 0
|
||||
AND btrim(m) <> ''
|
||||
ORDER BY c.team_id;
|
||||
|
||||
-- The server seeds environment defaults into teams once, and remembers that it
|
||||
-- did. An install that had a row per team was already seeded; without this
|
||||
-- marker the first start after upgrading would seed teams that had switched
|
||||
-- theirs off.
|
||||
INSERT INTO settings (key, value)
|
||||
SELECT 'deadman_seeded', '1'
|
||||
WHERE EXISTS (SELECT 1 FROM deadman_configs);
|
||||
|
||||
DROP TABLE deadman_configs;
|
||||
@@ -1,21 +0,0 @@
|
||||
-- Which alert source an alert last arrived on.
|
||||
--
|
||||
-- Team -> Sources shows when each source last posted, which integrations
|
||||
-- already knew (last_used_at, stamped on every webhook). What it could not say
|
||||
-- was what a source delivered: an alert never recorded the key it came in on, so
|
||||
-- "prod alertmanager" and "staging alertmanager" were indistinguishable once
|
||||
-- inside. This column is that link, and lets the page show each source's last
|
||||
-- alert and how many alerts it has kept fresh over the past day.
|
||||
--
|
||||
-- Last sender wins: every accepted payload restamps it, the way it advances
|
||||
-- received_at. Two sources posting the same fingerprint into one team is
|
||||
-- already one alert, and it is attributed to whichever spoke last.
|
||||
--
|
||||
-- Nullable, and not backfilled. Alerts that arrived before this migration have
|
||||
-- no source, and NULL says so honestly rather than guessing. It heals by itself:
|
||||
-- Alertmanager re-sends every alert each repeat_interval, and each re-send is an
|
||||
-- accepted payload. Deleting a source keeps its alerts, unattributed.
|
||||
ALTER TABLE alerts ADD COLUMN integration_id BIGINT REFERENCES integrations(id) ON DELETE SET NULL;
|
||||
|
||||
CREATE INDEX alerts_integration_idx ON alerts (integration_id, received_at)
|
||||
WHERE integration_id IS NOT NULL;
|
||||
@@ -1,60 +0,0 @@
|
||||
-- Single sign-on through an OpenID Connect provider (Authentik, and anything
|
||||
-- else that speaks OIDC).
|
||||
--
|
||||
-- Four things change, and none of them touches a password user: every new column
|
||||
-- has a default that says "this is how it has always worked".
|
||||
--
|
||||
-- 1. user_identities says which provider account a user is. It is keyed on
|
||||
-- (issuer, subject), never on email or username: those are mutable at the
|
||||
-- provider, and a recycled address must not inherit somebody's account. A
|
||||
-- user can have several identities (a second provider later), and none at all
|
||||
-- (a local, password-only user), which is why this is a table and not two
|
||||
-- columns on users.
|
||||
--
|
||||
-- 2. team_members.source and users.admin_source record who granted a role. 'oidc'
|
||||
-- rows are owned by the group sync: it adds them when a group grants access
|
||||
-- and removes them when it stops, and nothing else may edit them. 'manual' rows
|
||||
-- are everything that existed before this migration, and are never touched by
|
||||
-- the sync. Without the marker the sync could not tell a membership it created
|
||||
-- from one an owner added by hand, and would have to either leave stale access
|
||||
-- behind or delete people it had no business deleting.
|
||||
--
|
||||
-- 3. sessions.max_expires_at is a hard ceiling on a session's life. Ordinary
|
||||
-- sessions slide for as long as they are used; a session made by an SSO login
|
||||
-- must not, because the login is the only moment the groups are re-read.
|
||||
-- Capping the session is what makes "removed from the group in the provider"
|
||||
-- take effect within a bounded time. NULL means no ceiling.
|
||||
--
|
||||
-- 4. oidc_logins holds a login that has been started and not yet finished: the
|
||||
-- state, nonce and PKCE verifier the callback must see again. A row rather
|
||||
-- than a signed cookie, so it survives a restart and needs no signing key.
|
||||
-- Only the hash of the state is stored, like every other token here; the
|
||||
-- nonce and verifier are useless without the state that names the row.
|
||||
CREATE TABLE user_identities (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
user_id BIGINT NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
||||
issuer TEXT NOT NULL,
|
||||
subject TEXT NOT NULL,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
last_login_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
UNIQUE (issuer, subject)
|
||||
);
|
||||
|
||||
CREATE INDEX user_identities_user_idx ON user_identities (user_id);
|
||||
|
||||
ALTER TABLE team_members
|
||||
ADD COLUMN source TEXT NOT NULL DEFAULT 'manual' CHECK (source IN ('manual', 'oidc'));
|
||||
|
||||
ALTER TABLE users
|
||||
ADD COLUMN admin_source TEXT NOT NULL DEFAULT 'manual' CHECK (admin_source IN ('manual', 'oidc'));
|
||||
|
||||
ALTER TABLE sessions ADD COLUMN max_expires_at BIGINT;
|
||||
|
||||
CREATE TABLE oidc_logins (
|
||||
state_hash TEXT PRIMARY KEY,
|
||||
nonce TEXT NOT NULL,
|
||||
pkce_verifier TEXT NOT NULL,
|
||||
expires_at BIGINT NOT NULL
|
||||
);
|
||||
|
||||
CREATE INDEX oidc_logins_expires_idx ON oidc_logins (expires_at);
|
||||
@@ -1,40 +0,0 @@
|
||||
-- Signing in from a terminal, for clients that cannot open a browser on the
|
||||
-- machine they run on (the TUI over SSH is the reason).
|
||||
--
|
||||
-- The flow is the OAuth device authorization grant, run by terdut itself rather
|
||||
-- than the identity provider, so the terminal never talks to the provider and
|
||||
-- the server issues its ordinary session at the end:
|
||||
--
|
||||
-- 1. The terminal asks for a login and gets two secrets: a device code it
|
||||
-- keeps and polls with, and a short user code it shows the person.
|
||||
-- 2. The person opens the verification URL on any device, signs in by whatever
|
||||
-- means the server offers, sees the user code, and approves it.
|
||||
-- 3. The terminal's next poll finds the row approved and is given a session.
|
||||
--
|
||||
-- Only the hash of the device code is stored, like every other token here: the
|
||||
-- device code is what earns a session, so a database read must not yield one.
|
||||
-- The user code is shown on screens and typed by people, so it is stored as is;
|
||||
-- on its own it can only be approved, never redeemed.
|
||||
--
|
||||
-- user_id is the person who approved. It is empty until then, and the session
|
||||
-- is minted at redemption, not at approval: an approval nobody collects must not
|
||||
-- leave a live session lying about.
|
||||
--
|
||||
-- last_polled_at lets the server refuse a client that polls faster than the
|
||||
-- interval it was told.
|
||||
CREATE TABLE device_logins (
|
||||
device_hash TEXT PRIMARY KEY,
|
||||
user_code TEXT NOT NULL UNIQUE,
|
||||
status TEXT NOT NULL DEFAULT 'pending' CHECK (status IN ('pending', 'approved', 'denied')),
|
||||
user_id BIGINT REFERENCES users(id) ON DELETE CASCADE,
|
||||
expires_at BIGINT NOT NULL,
|
||||
last_polled_at BIGINT NOT NULL DEFAULT 0
|
||||
);
|
||||
|
||||
CREATE INDEX device_logins_expires_idx ON device_logins (expires_at);
|
||||
|
||||
-- Where to send the browser once a single sign-on login completes. A person who
|
||||
-- opens /device?code=... without a session has to sign in first and then come
|
||||
-- back to it, and the same is true of any other deep link. Validated when it is
|
||||
-- stored: only a path on this server is ever kept.
|
||||
ALTER TABLE oidc_logins ADD COLUMN next TEXT NOT NULL DEFAULT '/';
|
||||
@@ -1,26 +0,0 @@
|
||||
-- Per-team OIDC group configuration, replacing the global
|
||||
-- TERDUT_OIDC_GROUP_MAPPINGS env var.
|
||||
--
|
||||
-- Group -> team -> role used to be one global list an operator set for the
|
||||
-- whole install, matched against a team by name, and the sync would create
|
||||
-- the team if no team by that name existed yet. That put the decision of
|
||||
-- which group controls a team in the server's environment rather than the
|
||||
-- team's own hands, meant changing it needed an env var edit and a restart,
|
||||
-- and let a typo in a team name silently create a stray team.
|
||||
--
|
||||
-- Each team now names, itself, which group grants membership and which
|
||||
-- grants ownership. Nullable: most teams need neither. No uniqueness
|
||||
-- constraint on either column — two teams may legitimately watch the same
|
||||
-- provider group (a broad team and a narrower one both keyed off overlapping
|
||||
-- groups is a choice for their owners to make, not one the schema should
|
||||
-- refuse).
|
||||
--
|
||||
-- BREAKING CHANGE, deliberately not auto-migrated: TERDUT_OIDC_GROUP_MAPPINGS
|
||||
-- stops being read as of this version, and the sync no longer creates a team
|
||||
-- by name. Every team's group binding must be set again through
|
||||
-- PUT /api/teams/{teamID}/oidc-groups. Until an owner does that, an
|
||||
-- OIDC-sourced membership in that team is dropped at that user's next SSO
|
||||
-- sign-in, the same way any other loss of group access is handled. See the
|
||||
-- README's OIDC section.
|
||||
ALTER TABLE teams ADD COLUMN oidc_member_group TEXT;
|
||||
ALTER TABLE teams ADD COLUMN oidc_owner_group TEXT;
|
||||
@@ -1,43 +0,0 @@
|
||||
-- Service accounts: a scoped, non-human credential for automation (e.g.
|
||||
-- terdut-operator) that needs to manage teams, escalation policies, dead
|
||||
-- man's switches, integrations and OIDC group bindings without impersonating
|
||||
-- a human user. See SERVICE-ACCOUNTS.md for the design this implements.
|
||||
--
|
||||
-- Deliberately not a users row: no password_hash, no is_admin, no
|
||||
-- user_identities linkage, so a service account can never be pulled into
|
||||
-- OIDC group sync or password login, and is never mistaken for a human in an
|
||||
-- audit trail.
|
||||
--
|
||||
-- scope is 'instance' (acts with the same reach system administration has
|
||||
-- over teams: create one, list them, mint a 'team'-scoped account against
|
||||
-- any of them) or 'team' (acts as that one team's owner, and nothing else).
|
||||
-- The CHECK ties team_id's presence to scope directly, rather than leaving it
|
||||
-- to application code to keep the two consistent.
|
||||
CREATE TABLE service_accounts (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
name TEXT NOT NULL UNIQUE,
|
||||
scope TEXT NOT NULL CHECK (scope IN ('instance', 'team')),
|
||||
team_id BIGINT REFERENCES teams(id) ON DELETE CASCADE,
|
||||
created_by BIGINT REFERENCES users(id) ON DELETE SET NULL,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
CONSTRAINT service_accounts_scope_team_id_chk CHECK (
|
||||
(scope = 'team' AND team_id IS NOT NULL) OR
|
||||
(scope = 'instance' AND team_id IS NULL)
|
||||
)
|
||||
);
|
||||
|
||||
CREATE INDEX service_accounts_team_id_idx ON service_accounts(team_id);
|
||||
|
||||
-- One account, many keys: rotation is minting a new one and revoking the
|
||||
-- old, the same shape api_keys already has, so an account's identity and
|
||||
-- audit history survive a rotation instead of being recreated by it.
|
||||
CREATE TABLE service_account_keys (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
service_account_id BIGINT NOT NULL REFERENCES service_accounts(id) ON DELETE CASCADE,
|
||||
key_hash TEXT NOT NULL UNIQUE,
|
||||
name TEXT NOT NULL,
|
||||
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint,
|
||||
last_used_at BIGINT
|
||||
);
|
||||
|
||||
CREATE INDEX service_account_keys_service_account_id_idx ON service_account_keys(service_account_id);
|
||||
@@ -1,39 +0,0 @@
|
||||
-- Service-account actors on incident mutations (terdut-server#25). A
|
||||
-- team-scoped service account acknowledging/resolving/snoozing/noting an
|
||||
-- incident is not a users row, so it cannot be written into
|
||||
-- acknowledged_by/incident_events.user_id — doing so either violates the
|
||||
-- users(id) FK (new rows) or, for incident_events.user_id, silently matches
|
||||
-- zero rows on delete. These columns are the service-account-shaped parallel
|
||||
-- to the existing human ones: nullable, mutually exclusive with their human
|
||||
-- counterpart, ON DELETE SET NULL so a deleted service account doesn't take
|
||||
-- the incident history with it.
|
||||
ALTER TABLE incidents
|
||||
ADD COLUMN acknowledged_by_service_account_id BIGINT
|
||||
REFERENCES service_accounts(id) ON DELETE SET NULL;
|
||||
|
||||
ALTER TABLE incident_events
|
||||
ADD COLUMN service_account_id BIGINT
|
||||
REFERENCES service_accounts(id) ON DELETE SET NULL;
|
||||
|
||||
-- At most one actor kind per row: both NULL ("the server acted") is valid,
|
||||
-- exactly one set is valid, both set is a bug this constraint refuses to
|
||||
-- store rather than silently accepting.
|
||||
ALTER TABLE incidents
|
||||
ADD CONSTRAINT incidents_ack_actor_xor_chk CHECK (
|
||||
acknowledged_by IS NULL OR acknowledged_by_service_account_id IS NULL
|
||||
);
|
||||
|
||||
ALTER TABLE incident_events
|
||||
ADD CONSTRAINT incident_events_actor_xor_chk CHECK (
|
||||
user_id IS NULL OR service_account_id IS NULL
|
||||
);
|
||||
|
||||
CREATE INDEX incidents_acknowledged_by_service_account_id_idx
|
||||
ON incidents(acknowledged_by_service_account_id);
|
||||
CREATE INDEX incident_events_service_account_id_idx
|
||||
ON incident_events(service_account_id);
|
||||
|
||||
-- assigned_to_service_account_id is deliberately not added here: it would sit
|
||||
-- unpopulated until handleIncidentAssign itself tracks an actor, which is a
|
||||
-- separate, pre-existing gap (it records the assignee today, never the
|
||||
-- actor, for humans either) tracked in its own follow-up issue.
|
||||
@@ -1,16 +0,0 @@
|
||||
-- Backs the rate limiters (failed logins, sign-ups, OIDC/device start) with
|
||||
-- Postgres instead of an in-memory map, now that the server runs more than
|
||||
-- one replica in production (v0.37.0): a counter that only ever sees its own
|
||||
-- pod's traffic quietly let every one of these limits through multiplied by
|
||||
-- the replica count.
|
||||
--
|
||||
-- window_start is the start of the current fixed window for key, in the same
|
||||
-- "unix seconds" shape every other timestamp in this schema uses. The window
|
||||
-- resets rather than slides, matching the in-memory limiter it replaces:
|
||||
-- once a key's window is older than the limiter's window length, the next
|
||||
-- failure starts a fresh one instead of extending the stale one.
|
||||
CREATE TABLE rate_limit_counters (
|
||||
key TEXT PRIMARY KEY,
|
||||
window_start BIGINT NOT NULL,
|
||||
count INT NOT NULL
|
||||
);
|
||||
@@ -1,7 +0,0 @@
|
||||
-- Optional expiry on a user's own API keys. NULL (the existing default for
|
||||
-- every row already in this table) means "never expires" -- the same
|
||||
-- behavior these keys have always had, so no existing integration breaks.
|
||||
-- Service account keys are deliberately NOT touched: they are a different
|
||||
-- table, managed by automation, and already distinguished by their own
|
||||
-- "tdsa_" prefix.
|
||||
ALTER TABLE api_keys ADD COLUMN expires_at BIGINT;
|
||||
@@ -1,21 +0,0 @@
|
||||
-- Who performed an assignment (terdut-server#35). On an 'assigned' event
|
||||
-- incident_events.user_id is the assignee, so the actor needs columns of its
|
||||
-- own. Only populated for 'assigned' events; every other event type keeps
|
||||
-- using user_id/service_account_id for the actor. Older 'assigned' rows stay
|
||||
-- NULL (the actor was never recorded). Same shape as migration 015: nullable,
|
||||
-- mutually exclusive, ON DELETE SET NULL.
|
||||
--
|
||||
-- assigned_to_service_account_id is still deliberately not added: making
|
||||
-- service accounts assignable is a separate change (request body, assignee
|
||||
-- picker, notifier, filters).
|
||||
ALTER TABLE incident_events
|
||||
ADD COLUMN actor_user_id BIGINT REFERENCES users(id) ON DELETE SET NULL,
|
||||
ADD COLUMN actor_service_account_id BIGINT REFERENCES service_accounts(id) ON DELETE SET NULL;
|
||||
|
||||
ALTER TABLE incident_events
|
||||
ADD CONSTRAINT incident_events_assign_actor_xor_chk CHECK (
|
||||
actor_user_id IS NULL OR actor_service_account_id IS NULL
|
||||
);
|
||||
|
||||
CREATE INDEX incident_events_actor_user_id_idx ON incident_events(actor_user_id);
|
||||
CREATE INDEX incident_events_actor_service_account_id_idx ON incident_events(actor_service_account_id);
|
||||