c9f8494b00b5c7e9e3d9e02cd24b8badcdf3d72c
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9029d48584 |
Let an operator authenticate with a seeded key, and reset the schema
- TERDUT_OPERATOR_KEY creates or re-keys the instance-scoped service account
"terdut-operator" at every start, so terdut-operator needs no bootstrap
handshake. An instance-scoped account now acts as owner of every team's
configuration, but is not a member of any team.
- POST /api/teams takes an external_id (instance service accounts only) and
is idempotent on it, so automation finds its own team again after a crash
instead of adopting by display name. GET /api/teams?name= is removed.
- Integration and dead man's switch names are unique per team (409). The
escalation PUT accepts usernames and resolves them itself.
- The 18 migrations are squashed into 001_schema.sql, with no Default team.
TERDUT_DEADMAN_* and the env seeding of switches are removed: teams carry
their own. Existing development databases must be recreated.
Security and robustness:
- GET /api/users no longer returns other people's email or ntfy topic to
non-admins.
- The access log records the route pattern, so integration keys and ack
tokens in the path are not written to the log. Server errors are logged.
- Rate limits take the client address TERDUT_TRUSTED_PROXIES hops from the
right of X-Forwarded-For instead of trusting the first, forgeable entry.
- /api/bootstrap runs in a transaction under an advisory lock, so two
concurrent calls cannot both create an administrator.
- API key last_used_at is written at most every five minutes.
Cleanup: remove GET /api/incidents/{id}/alerts, unused exports, SQLite
remnants in comments and config.
Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
|
||
|
|
bc9f793f1f |
Add GET /api/teams?name= (TEAM-LOOKUP.md)
Resolves the gap TEAM-LOOKUP.md raised: an instance-scoped service account had no way to recover a team's id after a 409 on POST /api/teams, unlike the already-solved equivalent for service accounts themselves (GET /api/service-accounts?name=). Same route, extended the same way handleListServiceAccounts already branches on ?name=: unset behaves exactly as before (the caller's own teams via team_members); set looks up one team by exact name, open to any authenticated caller -- not gated by isInstanceServiceAccount or AdminOnly, since what it discloses (a name is taken, nothing about who's in it) is the same low sensitivity that lookup already accepts for service-account names. Tests cover the exact motivating scenario (create, 409 on a retry, recover the id via ?name=), the empty-array-not-an-error case, that no role/source is reported for a non-member match, and that a caller who isn't a member of the matched team still gets it. |
||
|
|
5b4683febf |
Let each team name its own OIDC group, not a global mapping
Team membership from single sign-on used to come from one env var,
TERDUT_OIDC_GROUP_MAPPINGS, matched against a team by name and creating
the team if none existed. That put the decision in the server's
environment rather than the team's own hands, needed a restart to
change, and let a typo in a team name silently create a stray team.
Each team now carries its own oidc_member_group and oidc_owner_group,
set by its owner (or an administrator) from the Members tab, or PUT
/api/teams/{teamID}/oidc-groups. The "highest role wins" rule
TERDUT_OIDC_GROUP_MAPPINGS used to apply across mappings now applies
across one team's own two fields: being in both makes somebody an
owner. The sync no longer creates a team by name; a group only ever
grants into a team that already exists.
This is a breaking change for anyone already using
TERDUT_OIDC_GROUP_MAPPINGS, deliberately not auto-migrated: an
OIDC-sourced membership is dropped at a user's next sign-in until its
team's owner re-sets the group. The README's OIDC section spells out
the migration and the risk of a visible access gap during it.
TERDUT_OIDC_ADMIN_GROUP and TERDUT_OIDC_ALLOWED_GROUPS are untouched --
only team membership moved. terdut-tui needs no change: it only reads
GET /api/teams and GET /api/teams/{id}/members, and neither response
shape moved.
|
||
|
|
a4fbd60441 |
Scope everything to a team, and route alerts by integration key
The core of #4, and what #1 is for: terdut stops being one shared space. A team owns its incidents, alerts, schedule and integrations; a user sees exactly the teams they are in. Everything that existed moves into one Default team and every existing user becomes an owner of it, so the upgrade is a no-op for the people using it. Ingestion is the load-bearing half. An alert arrives on a team's integration key, and the key is both the credential and the routing: it says that the sender may post, and which team the alerts belong to. That also closes the unauthenticated webhook -- the old path stays for one release, deprecated and routed to the oldest team, so an upgrade does not stop delivering while somebody edits the Alertmanager config. Scoping is enforced in as few places as possible, because the failure mode is silent. serveAs loads the caller's memberships once; list queries carry `team_id = ANY(...)`; and every incident route goes through incidentIDParam, which now parses the id AND checks the team in the same call, so a new handler cannot remember the first half and forget the second. Anything in another team is 404, never 403: whether an incident exists is that team's business. Two bugs this found, both of which would have been silent: * upsertAlerts decided "is this a new occurrence" by looking up the fingerprint alone. Across teams that made team B's first alert look like a re-send of team A's, so it opened no incident at all. The lookups are keyed on (team_id, fingerprint) now, as the index is. * Every uniqueness rule was written for one tenant. Two teams watching two clusters legitimately see the same fingerprint, the same groupKey, and want somebody on call on the same day; all three constraints move to include team_id. Roles inside a team are separate from the system administrator flag: an owner configures the team, a member works its incidents, and an admin is NOT implicitly in every team -- administration is about accounts, not about reading other people's incidents. An admin can still repair a team whose owner has left, which is why requireTeamOwner lets them through. A shift can only be given to somebody in the team. Paging a person who cannot open the incident is worse than paging nobody. The UI is updated only as far as keeping it working: it loads the viewer's teams with the session and uses the first one, since nobody has a second yet. "On call now" shows every team the viewer is in, named only when there is more than one, so the common case reads exactly as before. The team switcher, badges and per-team settings pages are the next step. Breaking for API clients: the schedule endpoints moved under the team, and /api/schedule/current returns an array rather than an object or a 404. terdut-tui will need a version for that. Per-team dead-man configuration is deliberately not here. A heartbeat's incident already opens in the team whose key received it, which is the part that matters for isolation; moving the matchers out of env into per-team rows is a change to how deadman.go is configured rather than to who sees what. Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7 |