The README was 1,240 lines of reference material and still described a SQLite quick start. It is now a short tour (highlights, screenshots of the web UI, an accurate quick start against Postgres), and each topic has its own page under docs/ with an index: deployment, configuration, Alertmanager, incidents, notifications, escalation, dead man's switches, single sign-on, web UI, API and development. SERVICE-ACCOUNTS.md is rewritten from a proposal into a reference, and TEAM-LOOKUP.md is gone with the endpoint it described. The "Upgrading to ..." sections for an unreleased product are dropped. Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
3.4 KiB
Push notifications
Pages through ntfy, who gets them and how to acknowledge from the notification. Back to the README and the documentation index.
With TERDUT_NTFY_URL set, an incident that opens is pushed to the on-call
person's phone through ntfy. Everybody sets their own topic
under Account in the web UI, where a Send a test push button proves it
before an incident has to; PUT /api/users/{id}/notify is the same thing over
the API, and an administrator may set somebody else's. A user with no topic
falls back to TERDUT_NTFY_FALLBACK_TOPIC, as does an incident that opens with
nobody on call. If neither yields a topic, nothing is queued.
The server is the install's one ntfy, from TERDUT_NTFY_URL, and is not
something a user picks. Only the topic is per-person.
A topic is a shared secret with the ntfy server: anyone who knows it can both read the pages and publish to it, so an unguessable one is worth the trouble. That is also why the topic never appears in an incident's timeline, which every API key can read.
Three things get pushed:
- triggered — an incident opened. Priority follows severity (
criticalmaps to ntfy's max priority, the one that overrides the phone's quiet settings). - reminder — the incident is still
triggeredafterTERDUT_NOTIFY_REPEAT. Repeats until somebody acts. Acknowledging, snoozing, resolving or archiving all stop it — snooze is the mute button. - resolved — every alert under the incident stopped firing. Only sent to whoever was paged in the first place, and only for the automatic cascade: resolving by hand pushes nothing, since the person who did it already knows.
Notifications carry an Acknowledge button that acknowledges the incident
without opening anything. It POSTs to /api/notify/ack/{token}, an
unauthenticated route authorised by the 256-bit token in its path — minted fresh
per notification, scoped to one incident and one action, and valid for 24 hours.
A real API key is never put in a notification, because the message is stored on
the ntfy server and cached on the device.
The token is not consumed by use. Acknowledging is idempotent, so a token stays valid for its full 24 hours and a second tap is a no-op that reports the incident's current state rather than an error — which is what you want when a tap is retried on a flaky mobile connection. What bounds it is scope, not a use count: one incident, one action, one day. Expired tokens are purged by the sweeper.
Two consequences worth planning for:
/api/notify/ack/{token}must stay publicly reachable, or the button will not work when the responder is off your network.- Notifications sent to the fallback topic carry no Acknowledge button. The topic is shared, and a button on it would let any subscriber acknowledge as somebody else.
Delivery is a queue, not an inline call: the webhook writes a row and a background notifier sends it within 30 seconds, retrying with exponential backoff up to 8 attempts. Nothing about ingestion blocks on ntfy being reachable.
Every delivery is recorded on the incident's timeline: a notified event once
ntfy accepts the publish, and a notify_failed event when a notification
exhausts its retries. Written from the result rather than at enqueue, so the
timeline says what actually happened — and a page that never landed is visible
instead of looking the same as one that did.