Files
terdut-server/docs/escalation.md
T
Niklas Ye 44b2eb2cc3 Rewrite the README as highlights with screenshots; move the detail into docs/
The README was 1,240 lines of reference material and still described a
SQLite quick start. It is now a short tour (highlights, screenshots of the
web UI, an accurate quick start against Postgres), and each topic has its
own page under docs/ with an index: deployment, configuration, Alertmanager,
incidents, notifications, escalation, dead man's switches, single sign-on,
web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a
proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it
described. The "Upgrading to ..." sections for an unreleased product are
dropped.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:13 +02:00

2.4 KiB

Escalation

Escalation ladders: who is paged next when nobody acknowledges. Back to the README and the documentation index.

Without a ladder, an unacknowledged incident re-pages the same topic every notify_repeat forever. That is a louder version of the same silence: if the person on call is asleep, out of signal, or has left, nothing else happens.

A team can configure an ordered ladder instead. Each level has a timeout and a set of targets, and a target is either a named person or whoever the team's rota says is on call today — the target that keeps working when the rota changes and nobody remembers to edit the policy.

level 1   5m    oncall            the rota gets first refusal
level 2   5m    user:bob          then a named second
                                  then repeat_count more rounds
                                  then the team's fallback topic, once

When a level's timeout passes with the incident still triggered, the next level is paged. Off the end of the ladder the whole thing runs again repeat_count times, and after that the team's fallback_topic is paged once as the end of the line. The incident stays open throughout: running out of people to wake is not the same as somebody answering.

Acknowledging or resolving stops it, which is the point — continuing to wake people after somebody has said "I have this" is how a tool teaches people to mute it. Snoozing pauses it: a deliberate "not now" holds the ladder where it is, and it resumes when the snooze runs out.

Every step is on the incident's timeline with the level and the names it woke, so somebody reading it afterwards can tell why their phone rang at 04:00. A level whose targets are all unreachable — no ntfy topic, a disabled account, an empty rota — is recorded as nobody reachable and the ladder moves on rather than stalling on a rung that cannot ring.

Reminders and escalation never both run. A team with a ladder gets escalation; a team without keeps the reminder behaviour exactly as it was. Two pages for one silence is the surest way to get a tool muted.

The ladder's fallback_topic is per team, unlike TERDUT_NTFY_FALLBACK_TOPIC, which is the install-wide topic used when an incident opens with nobody on call. They answer different questions: one is "nobody was scheduled", the other is "everybody scheduled has been tried".