01922291f600d302fe61ed3654fb28061f2bb1f3
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9029d48584 |
Let an operator authenticate with a seeded key, and reset the schema
- TERDUT_OPERATOR_KEY creates or re-keys the instance-scoped service account
"terdut-operator" at every start, so terdut-operator needs no bootstrap
handshake. An instance-scoped account now acts as owner of every team's
configuration, but is not a member of any team.
- POST /api/teams takes an external_id (instance service accounts only) and
is idempotent on it, so automation finds its own team again after a crash
instead of adopting by display name. GET /api/teams?name= is removed.
- Integration and dead man's switch names are unique per team (409). The
escalation PUT accepts usernames and resolves them itself.
- The 18 migrations are squashed into 001_schema.sql, with no Default team.
TERDUT_DEADMAN_* and the env seeding of switches are removed: teams carry
their own. Existing development databases must be recreated.
Security and robustness:
- GET /api/users no longer returns other people's email or ntfy topic to
non-admins.
- The access log records the route pattern, so integration keys and ack
tokens in the path are not written to the log. Server errors are logged.
- Rate limits take the client address TERDUT_TRUSTED_PROXIES hops from the
right of X-Forwarded-For instead of trusting the first, forgeable entry.
- /api/bootstrap runs in a transaction under an advisory lock, so two
concurrent calls cannot both create an administrator.
- API key last_used_at is written at most every five minutes.
Cleanup: remove GET /api/incidents/{id}/alerts, unused exports, SQLite
remnants in comments and config.
Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
|
||
|
|
1f1faa437c |
Show the escalation ladder as a list, with who it would page and where it is
Team -> Escalation was the draft form on the page, which showed the ladder only as inputs. It is now a table in the style of Switches and Sources: a row per level with a status badge, who it pages, the wait before the next level, and the open incidents currently waiting on it. Below it, the repeat count, the fallback topic and when the ladder last escalated (linking the incident). The editor moved into an "Edit ladder" sheet, so a poll of the page underneath can no longer throw away half an edit, and the page-level draft state went with it. Targets are resolved to who they mean today, and the badge says what would actually happen: Ready, Escalating (an unanswered incident has climbed to level 2 or higher), or Pages nobody. The last is the one worth seeing before an incident finds it: an empty rota, a person with no ntfy topic or a disabled account each make a rung a silence with a number on it, and the target says which. The rules are pageLevel's own, so the page cannot promise a page the notifier would skip. "Last escalated" comes from the escalated timeline events that already exist, so there is no migration. Acknowledging or resolving takes an incident off the ladder, so Escalating clears then while the history stays. API: GET /escalation gains status and waiting per level, username, reachable and problem per target, and last_escalated_at and last_escalated_incident_id. Output only and additive; PUT is unchanged and terdut-tui needs nothing. |
||
|
|
3183e7e5c5 |
Page the next person when nobody answers
Closes #6, and closes the thing this whole line of work was opened for. Until now an unacknowledged incident re-paged the same topic every notify_repeat forever, which is a louder version of the same silence: if the person on call is asleep, out of signal or has left the company, nothing else happened. A team can now configure an ordered ladder. Each level has a timeout and a set of targets; a target is a named person or whoever the team's rota says is on call today. That second kind is the one that keeps working when the rota changes and nobody remembers to edit the policy. When a level's timeout passes with the incident still triggered, the next level is paged; off the end the chain repeats repeat_count times and then the team's fallback topic is paged once. The incident stays open throughout, because running out of people to wake is not somebody answering. Escalation rides the notifier's existing 30-second tick and its outbox rather than adding a second scheduler, and runs before delivery so a level that comes due on a tick is paged on that tick. Each target gets its own outbox row and therefore its own Acknowledge token: the button in a notification must acknowledge as the person holding the phone, not as whoever was paged first. Acknowledging or resolving takes the incident off the ladder. Snoozing pauses it -- a deliberate "not now" holds the ladder where it is and it resumes when the snooze runs out, rather than carrying on without the person who asked for quiet. Reminders and escalation never both run. A team with a ladder gets escalation; a team without keeps today's behaviour exactly. Both would mean two pages for one silence, which is how a tool gets muted. A level whose targets cannot be reached -- no topic, a disabled account, an empty rota -- is entered anyway, recorded as "nobody reachable", and the ladder moves on. Stalling on a rung that cannot ring would be the failure this feature exists to prevent, wearing the feature's clothes. A policy with such a level cannot be created, but an older row could hold one. The API replaces the ladder wholesale rather than patching a rung, because the levels are an order: editing one has to answer what happens to the numbering of the others, and a whole-ladder PUT makes that the client's decision and the edit atomic. Verified against a live server as well as in tests: alice paged, nobody answers, bob paged, nobody answers, the fallback topic paged once and the timeline reading "level 2: bob" then "escalation exhausted: paged terdut-oncall-all" -- and a second incident acknowledged before its timeout, which woke nobody else. No UI yet. The team-settings screens for escalation, integrations and dead man's switches are all still missing, and they are one piece of work rather than three. Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7 |