The README was 1,240 lines of reference material and still described a SQLite quick start. It is now a short tour (highlights, screenshots of the web UI, an accurate quick start against Postgres), and each topic has its own page under docs/ with an index: deployment, configuration, Alertmanager, incidents, notifications, escalation, dead man's switches, single sign-on, web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it described. The "Upgrading to ..." sections for an unreleased product are dropped. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2.4 KiB
Escalation
Escalation ladders: who is paged next when nobody acknowledges. Back to the README and the documentation index.
Without a ladder, an unacknowledged incident re-pages the same topic every
notify_repeat forever. That is a louder version of the same silence: if the
person on call is asleep, out of signal, or has left, nothing else happens.
A team can configure an ordered ladder instead. Each level has a timeout and a set of targets, and a target is either a named person or whoever the team's rota says is on call today — the target that keeps working when the rota changes and nobody remembers to edit the policy.
level 1 5m oncall the rota gets first refusal
level 2 5m user:bob then a named second
then repeat_count more rounds
then the team's fallback topic, once
When a level's timeout passes with the incident still triggered, the next
level is paged. Off the end of the ladder the whole thing runs again
repeat_count times, and after that the team's fallback_topic is paged once
as the end of the line. The incident stays open throughout: running out of
people to wake is not the same as somebody answering.
Acknowledging or resolving stops it, which is the point — continuing to wake people after somebody has said "I have this" is how a tool teaches people to mute it. Snoozing pauses it: a deliberate "not now" holds the ladder where it is, and it resumes when the snooze runs out.
Every step is on the incident's timeline with the level and the names it woke,
so somebody reading it afterwards can tell why their phone rang at 04:00. A
level whose targets are all unreachable — no ntfy topic, a disabled account, an
empty rota — is recorded as nobody reachable and the ladder moves on rather
than stalling on a rung that cannot ring.
Reminders and escalation never both run. A team with a ladder gets escalation; a team without keeps the reminder behaviour exactly as it was. Two pages for one silence is the surest way to get a tool muted.
The ladder's fallback_topic is per team, unlike TERDUT_NTFY_FALLBACK_TOPIC,
which is the install-wide topic used when an incident opens with nobody on call.
They answer different questions: one is "nobody was scheduled", the other is
"everybody scheduled has been tried".