Scope everything to a team, and route alerts by integration key #11

Merged
niklas merged 1 commits from teams into admin-role 2026-09-20 11:39:34 +00:00
Owner

The core of #4. Stacked on #10 — based on admin-role, so merge that first and this retargets to main.

terdut stops being one shared space. A team owns its incidents, alerts, schedule and integrations; a user sees exactly the teams they are in. Everything that exists moves into one Default team with every current user as an owner, so the upgrade is a no-op for the people using it.

Ingestion is the load-bearing half

An alert arrives on a team's integration key, and the key is both the credential and the routing. That also closes the unauthenticated webhook — the old path stays for one release, deprecated and routed to the oldest team, so an upgrade doesn't stop delivering while the Alertmanager config is edited.

Scoping is enforced in few places, because the failure mode is silent

  • serveAs loads the caller's memberships once per request.
  • List queries carry team_id = ANY(...).
  • Every incident route goes through incidentIDParam, which now parses the id and checks the team in the same call — so a new handler can't remember the first half and forget the second.
  • Anything in another team is 404, never 403: whether an incident exists is that team's business.

Two bugs this surfaced, both silent

  • upsertAlerts decided "is this a new occurrence" from the fingerprint alone. Across teams that made team B's first alert look like a re-send of team A's, so it opened no incident at all. Now keyed on (team_id, fingerprint), as the index is.
  • Every uniqueness rule was written for one tenant. Two teams watching two clusters legitimately see the same fingerprint, the same groupKey, and want somebody on call the same day. All three constraints move to include team_id.

Roles

Team roles are a separate axis from the system administrator flag. An owner configures the team, a member works its incidents, and an admin is not implicitly in every team — administration is about accounts, not about reading other people's incidents. An admin can still repair a team whose owner left, which is why requireTeamOwner lets them through.

A shift can only be given to somebody in the team: paging a person who cannot open the incident is worse than paging nobody.

Verified

make fmt lint test helm-lint green with -race against Postgres 17, plus seven new tests in teams_test.go that build two teams and check the boundary from both sides: incidents, alerts and stats scoped; the other team's incident unreadable and unactionable by id; the same fingerprint and the same date legal in both; an unknown key rejected with 401 and nothing ingested; a member refused team configuration but allowed to read it; an outsider seeing nothing at all.

Also exercised against a live server: bootstrap → default team → mint an integration → post an alert on its key → the incident appears with its team_id and team_name.

Breaking for API clients

The schedule endpoints moved under /api/teams/{teamID}/schedule, and GET /api/schedule/current now returns an array (one entry per team with somebody on call) rather than an object or a 404. terdut-tui will need a version for that.

Deliberately not here

  • Per-team dead-man configuration. A heartbeat's incident already opens in the team whose key received it, which is the part that matters for isolation; moving the matchers out of env into per-team rows changes how deadman.go is configured, not who sees what.
  • The UI beyond keeping it working. It loads the viewer's teams with the session and uses the first, since nobody has a second yet; "On call now" shows every team, named only when there is more than one. The switcher, badges and per-team settings pages are the next step.

https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7

The core of #4. **Stacked on #10** — based on `admin-role`, so merge that first and this retargets to `main`. terdut stops being one shared space. A team owns its incidents, alerts, schedule and integrations; a user sees exactly the teams they are in. Everything that exists moves into one **Default** team with every current user as an owner, so the upgrade is a no-op for the people using it. ### Ingestion is the load-bearing half An alert arrives on a team's integration key, and the key is both the credential and the routing. That also closes the unauthenticated webhook — the old path stays for one release, deprecated and routed to the oldest team, so an upgrade doesn't stop delivering while the Alertmanager config is edited. ### Scoping is enforced in few places, because the failure mode is silent - `serveAs` loads the caller's memberships once per request. - List queries carry `team_id = ANY(...)`. - **Every incident route goes through `incidentIDParam`, which now parses the id *and* checks the team in the same call** — so a new handler can't remember the first half and forget the second. - Anything in another team is `404`, never `403`: whether an incident exists is that team's business. ### Two bugs this surfaced, both silent - **`upsertAlerts` decided "is this a new occurrence" from the fingerprint alone.** Across teams that made team B's first alert look like a re-send of team A's, so it opened **no incident at all**. Now keyed on `(team_id, fingerprint)`, as the index is. - **Every uniqueness rule was written for one tenant.** Two teams watching two clusters legitimately see the same fingerprint, the same groupKey, and want somebody on call the same day. All three constraints move to include `team_id`. ### Roles Team roles are a separate axis from the system administrator flag. An owner configures the team, a member works its incidents, and **an admin is not implicitly in every team** — administration is about accounts, not about reading other people's incidents. An admin can still repair a team whose owner left, which is why `requireTeamOwner` lets them through. A shift can only be given to somebody in the team: paging a person who cannot open the incident is worse than paging nobody. ### Verified `make fmt lint test helm-lint` green with `-race` against Postgres 17, plus seven new tests in `teams_test.go` that build two teams and check the boundary from both sides: incidents, alerts and stats scoped; the other team's incident unreadable and unactionable by id; the same fingerprint and the same date legal in both; an unknown key rejected with `401` and nothing ingested; a member refused team configuration but allowed to read it; an outsider seeing nothing at all. Also exercised against a live server: bootstrap → default team → mint an integration → post an alert on its key → the incident appears with its `team_id` and `team_name`. ### Breaking for API clients The schedule endpoints moved under `/api/teams/{teamID}/schedule`, and `GET /api/schedule/current` now returns an **array** (one entry per team with somebody on call) rather than an object or a `404`. terdut-tui will need a version for that. ### Deliberately not here - **Per-team dead-man configuration.** A heartbeat's incident already opens in the team whose key received it, which is the part that matters for isolation; moving the matchers out of env into per-team rows changes how `deadman.go` is configured, not who sees what. - **The UI beyond keeping it working.** It loads the viewer's teams with the session and uses the first, since nobody has a second yet; "On call now" shows every team, named only when there is more than one. The switcher, badges and per-team settings pages are the next step. https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
niklas added 1 commit 2026-09-20 11:36:47 +00:00
Scope everything to a team, and route alerts by integration key
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 13s
CI / test (pull_request) Successful in 1m49s
a4fbd60441
The core of #4, and what #1 is for: terdut stops being one shared space.
A team owns its incidents, alerts, schedule and integrations; a user sees
exactly the teams they are in. Everything that existed moves into one
Default team and every existing user becomes an owner of it, so the
upgrade is a no-op for the people using it.

Ingestion is the load-bearing half. An alert arrives on a team's
integration key, and the key is both the credential and the routing: it
says that the sender may post, and which team the alerts belong to. That
also closes the unauthenticated webhook -- the old path stays for one
release, deprecated and routed to the oldest team, so an upgrade does not
stop delivering while somebody edits the Alertmanager config.

Scoping is enforced in as few places as possible, because the failure
mode is silent. serveAs loads the caller's memberships once; list queries
carry `team_id = ANY(...)`; and every incident route goes through
incidentIDParam, which now parses the id AND checks the team in the same
call, so a new handler cannot remember the first half and forget the
second. Anything in another team is 404, never 403: whether an incident
exists is that team's business.

Two bugs this found, both of which would have been silent:

  * upsertAlerts decided "is this a new occurrence" by looking up the
    fingerprint alone. Across teams that made team B's first alert look
    like a re-send of team A's, so it opened no incident at all. The
    lookups are keyed on (team_id, fingerprint) now, as the index is.

  * Every uniqueness rule was written for one tenant. Two teams watching
    two clusters legitimately see the same fingerprint, the same
    groupKey, and want somebody on call on the same day; all three
    constraints move to include team_id.

Roles inside a team are separate from the system administrator flag: an
owner configures the team, a member works its incidents, and an admin is
NOT implicitly in every team -- administration is about accounts, not
about reading other people's incidents. An admin can still repair a team
whose owner has left, which is why requireTeamOwner lets them through.

A shift can only be given to somebody in the team. Paging a person who
cannot open the incident is worse than paging nobody.

The UI is updated only as far as keeping it working: it loads the
viewer's teams with the session and uses the first one, since nobody has
a second yet. "On call now" shows every team the viewer is in, named only
when there is more than one, so the common case reads exactly as before.
The team switcher, badges and per-team settings pages are the next step.

Breaking for API clients: the schedule endpoints moved under the team,
and /api/schedule/current returns an array rather than an object or a
404. terdut-tui will need a version for that.

Per-team dead-man configuration is deliberately not here. A heartbeat's
incident already opens in the team whose key received it, which is the
part that matters for isolation; moving the matchers out of env into
per-team rows is a change to how deadman.go is configured rather than to
who sees what.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
niklas merged commit 2de5c8412d into admin-role 2026-09-20 11:39:34 +00:00
Sign in to join this conversation.