Compare commits

...

6 Commits

Author SHA1 Message Date
Niklas Ye 4e8c52c28c Set the chart's placeholder version to 0.14.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 12s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 24s
Release / image (push) Successful in 53s
Release / scan-image (push) Successful in 2s
Cosmetic, as in 4c85e76 and 041e159. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.14.0 that still says 0.13.0
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
2026-09-20 21:29:05 +02:00
niklas fb927aa67b Merge pull request 'Put a team's own settings in the web UI' (#18) from team-settings-ui into main
CI / test (push) Successful in 5s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Reviewed-on: #18
2026-09-20 19:26:22 +00:00
Niklas Ye d728af53b1 Put a team's own settings in the web UI
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 13s
CI / test (pull_request) Successful in 2m19s
Closes #17. Everything a team owner configures was API-only: escalation,
integrations, dead man's switches, membership, and the rota -- which the
on-call view still described as the TUI's job, and the TUI has been broken
against this server since teams landed. Setting up the feature this whole
line of work exists for meant using curl.

A Team tab now holds all of it, one team at a time, with a picker for
somebody in more than one. An owner edits; a member sees the same page
without the controls, because the server refuses their writes anyway --
hiding a button is a courtesy to the reader, not the thing enforcing
anything.

The escalation editor holds a draft and sends the whole ladder, because
the API replaces it wholesale: the levels are an order, and patching one
rung leaves the numbering of the others undecided. Adding a level
defaults to five minutes and the rota, which is the shape almost every
ladder starts as.

An integration key is returned exactly once, so creating one opens a
panel that says so, shows the URL large with a copy button, and renders
the Alertmanager receiver snippet with the URL already in it -- the next
thing anybody does with that key is paste it into a config. The panel
stays until it is dismissed rather than disappearing on the next
re-render.

The incident view gains where an incident is on the ladder and when the
next page is due, which is the question somebody looking at an
unacknowledged incident actually has. The API carries it: the incident
payload now includes escalation_level and escalation_due_at, the latter
computed in the incident SELECT by joining the level's timeout, so a list
costs no extra queries.

Verified against a live server by making every call the page makes,
including the writes: the six reads the Team tab issues, a two-level
ladder saved and read back, an integration created and its key returned
once, three days of rota assigned, switches set, and an incident showing
level 1 with a due time five minutes out.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-20 21:22:24 +02:00
Niklas Ye 53e5e03f4e Set the chart's placeholder version to 0.13.0
CI / test (push) Successful in 4s
CI / chart (push) Successful in 1s
CI / security (push) Successful in 11s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 18s
Release / image (push) Successful in 58s
Release / scan-image (push) Successful in 7s
Cosmetic, as in 4c85e76 and 041e159. `make helm-package` passes --version
and --app-version from the tag, so neither field decides anything about
what release.yaml publishes.

Done anyway because a tree heading for v0.13.0 that still says 0.12.0
tells a reader something false. appVersion keeps the v, per
APPVERSION_PREFIX in .release.conf.
2026-09-20 20:47:52 +02:00
niklas 4d62c1130b Merge pull request 'Page the next person when nobody answers' (#16) from escalation into main
CI / chart (push) Successful in 1s
CI / test (push) Successful in 6s
CI / security (push) Successful in 12s
Reviewed-on: #16
2026-09-20 16:43:03 +00:00
Niklas Ye 3183e7e5c5 Page the next person when nobody answers
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 16s
CI / test (pull_request) Successful in 2m21s
Closes #6, and closes the thing this whole line of work was opened for.
Until now an unacknowledged incident re-paged the same topic every
notify_repeat forever, which is a louder version of the same silence: if
the person on call is asleep, out of signal or has left the company,
nothing else happened.

A team can now configure an ordered ladder. Each level has a timeout and
a set of targets; a target is a named person or whoever the team's rota
says is on call today. That second kind is the one that keeps working
when the rota changes and nobody remembers to edit the policy. When a
level's timeout passes with the incident still triggered, the next level
is paged; off the end the chain repeats repeat_count times and then the
team's fallback topic is paged once. The incident stays open throughout,
because running out of people to wake is not somebody answering.

Escalation rides the notifier's existing 30-second tick and its outbox
rather than adding a second scheduler, and runs before delivery so a
level that comes due on a tick is paged on that tick. Each target gets
its own outbox row and therefore its own Acknowledge token: the button in
a notification must acknowledge as the person holding the phone, not as
whoever was paged first.

Acknowledging or resolving takes the incident off the ladder. Snoozing
pauses it -- a deliberate "not now" holds the ladder where it is and it
resumes when the snooze runs out, rather than carrying on without the
person who asked for quiet.

Reminders and escalation never both run. A team with a ladder gets
escalation; a team without keeps today's behaviour exactly. Both would
mean two pages for one silence, which is how a tool gets muted.

A level whose targets cannot be reached -- no topic, a disabled account,
an empty rota -- is entered anyway, recorded as "nobody reachable", and
the ladder moves on. Stalling on a rung that cannot ring would be the
failure this feature exists to prevent, wearing the feature's clothes. A
policy with such a level cannot be created, but an older row could hold
one.

The API replaces the ladder wholesale rather than patching a rung,
because the levels are an order: editing one has to answer what happens
to the numbering of the others, and a whole-ladder PUT makes that the
client's decision and the edit atomic.

Verified against a live server as well as in tests: alice paged, nobody
answers, bob paged, nobody answers, the fallback topic paged once and the
timeline reading "level 2: bob" then "escalation exhausted: paged
terdut-oncall-all" -- and a second incident acknowledged before its
timeout, which woke nobody else.

No UI yet. The team-settings screens for escalation, integrations and
dead man's switches are all still missing, and they are one piece of work
rather than three.

Claude-Session: https://claude.ai/code/session_01RHPj4ggeFdEjKKfm4SHbD7
2026-09-20 18:37:54 +02:00
18 changed files with 1648 additions and 8 deletions
+56
View File
@@ -85,6 +85,16 @@ How a browser stays signed in:
With `TERDUT_PUBLIC_URL` set, tapping a push notification opens the incident in With `TERDUT_PUBLIC_URL` set, tapping a push notification opens the incident in
the web UI (`/incidents/{id}`). the web UI (`/incidents/{id}`).
A **Team** tab holds everything a team owns: the on-call rota, the escalation
ladder, the alert sources with their keys, the dead man's switches and the
membership. An owner edits it; a member sees the same page read-only, because
the server refuses their writes anyway. Somebody in more than one team picks
between them at the top.
The **Admin** tab appears only for a system administrator, and holds what
belongs to the whole server rather than to one team: every team, every user, and
the settings that used to be environment variables.
### Docker ### Docker
```bash ```bash
@@ -377,6 +387,50 @@ exhausts its retries. Written from the result rather than at enqueue, so the
timeline says what actually happened — and a page that never landed is visible timeline says what actually happened — and a page that never landed is visible
instead of looking the same as one that did. instead of looking the same as one that did.
### Escalation
Without a ladder, an unacknowledged incident re-pages the same topic every
`notify_repeat` forever. That is a louder version of the same silence: if the
person on call is asleep, out of signal, or has left, nothing else happens.
A team can configure an ordered ladder instead. Each level has a timeout and a
set of targets, and a target is either a named person or **whoever the team's
rota says is on call today** — the target that keeps working when the rota
changes and nobody remembers to edit the policy.
```
level 1 5m oncall the rota gets first refusal
level 2 5m user:bob then a named second
then repeat_count more rounds
then the team's fallback topic, once
```
When a level's timeout passes with the incident still `triggered`, the next
level is paged. Off the end of the ladder the whole thing runs again
`repeat_count` times, and after that the team's `fallback_topic` is paged once
as the end of the line. The incident stays open throughout: running out of
people to wake is not the same as somebody answering.
**Acknowledging or resolving stops it**, which is the point — continuing to wake
people after somebody has said "I have this" is how a tool teaches people to
mute it. **Snoozing pauses it**: a deliberate "not now" holds the ladder where
it is, and it resumes when the snooze runs out.
Every step is on the incident's timeline with the level and the names it woke,
so somebody reading it afterwards can tell why their phone rang at 04:00. A
level whose targets are all unreachable — no ntfy topic, a disabled account, an
empty rota — is recorded as `nobody reachable` and the ladder moves on rather
than stalling on a rung that cannot ring.
**Reminders and escalation never both run.** A team with a ladder gets
escalation; a team without keeps the reminder behaviour exactly as it was. Two
pages for one silence is the surest way to get a tool muted.
The ladder's `fallback_topic` is per team, unlike `TERDUT_NTFY_FALLBACK_TOPIC`,
which is the install-wide topic used when an incident opens with nobody on call.
They answer different questions: one is "nobody was scheduled", the other is
"everybody scheduled has been tried".
### Stale alert expiry ### Stale alert expiry
A resolved webhook is the only signal that an alert has stopped firing, so a A resolved webhook is the only signal that an alert has stopped firing, so a
@@ -578,6 +632,8 @@ and was removed in v0.13.0 once senders had moved onto keys.
| `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys | | `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys |
| `POST` | `/api/teams/{teamID}/integrations` | **owner** | Mint an integration `{"name","kind"}` — key and URL shown once | | `POST` | `/api/teams/{teamID}/integrations` | **owner** | Mint an integration `{"name","kind"}` — key and URL shown once |
| `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration | | `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration |
| `GET` | `/api/teams/{teamID}/escalation` | member | The team's [escalation ladder](#escalation) `{repeat_count, fallback_topic, levels[]}`. Empty levels means the team has none |
| `PUT` | `/api/teams/{teamID}/escalation` | **owner** | Replace it wholesale. `400` for a level with no targets or no timeout — a rung that pages nobody is a silence with a number on it |
| `GET` | `/api/teams/{teamID}/deadman` | member | The team's [dead man's switch](#dead-mans-switch) configuration `{matchers, timeout_seconds, severity}` | | `GET` | `/api/teams/{teamID}/deadman` | member | The team's [dead man's switch](#dead-mans-switch) configuration `{matchers, timeout_seconds, severity}` |
| `PUT` | `/api/teams/{teamID}/deadman` | **owner** | Replace it. `400` when no matcher names an `alertname`, because a switch that silently watches nothing is the failure this feature exists to prevent | | `PUT` | `/api/teams/{teamID}/deadman` | **owner** | Replace it. `400` when no matcher names an `alertname`, because a switch that silently watches nothing is the failure this feature exists to prevent |
+2 -2
View File
@@ -15,5 +15,5 @@ type: application
# appVersion and image.tag in values.yaml no longer agree, and that is not an oversight: # appVersion and image.tag in values.yaml no longer agree, and that is not an oversight:
# image.tag stays "latest", which is what a local install actually pulls. appVersion is # image.tag stays "latest", which is what a local install actually pulls. appVersion is
# metadata and drives nothing. # metadata and drives nothing.
version: 0.12.0 version: 0.14.0
appVersion: "v0.12.0" appVersion: "v0.14.0"
+9 -3
View File
@@ -403,12 +403,18 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID in
} }
} }
// Queue the page, but do not send it here: this runs inside a transaction on // Queue the page, but do not send it here: this runs inside the webhook's
// a single-connection pool, so an HTTP call would hold up every other // transaction, and an HTTP call would hold a connection open across a
// request. The notifier picks the row up within a tick. // network round trip. The notifier picks the row up within a tick.
if err := enqueueOpened(ctx, q, notify, id, onCall); err != nil { if err := enqueueOpened(ctx, q, notify, id, onCall); err != nil {
return 0, err return 0, err
} }
// And start the escalation clock, if the team keeps one. In the same
// transaction, so an incident is never briefly open with nobody counting.
if err := startEscalation(ctx, q, id, teamID); err != nil {
return 0, err
}
return id, nil return id, nil
} }
+506
View File
@@ -0,0 +1,506 @@
package api
import (
"context"
"database/sql"
"log"
"net/http"
"strconv"
"strings"
"time"
)
// evEscalated records a rung of the ladder on the incident's timeline: which
// level, and who it woke.
const evEscalated = "escalated"
// escalationPolicy is a team's ladder, loaded whole. It is small — a handful of
// levels with a few targets each — and every use needs all of it, so there is
// no point reading it a level at a time.
type escalationPolicy struct {
teamID int64
repeatCount int64
fallbackTopic string
levels []escalationLevel
}
type escalationLevel struct {
id int64
position int64
timeout time.Duration
targets []escalationTarget
}
type escalationTarget struct {
kind string // "user" or "oncall"
userID *int64
}
// configured reports whether this team has anything to escalate through. A
// policy row with no levels is the same as no policy: the team gets the
// pre-escalation behaviour, which is reminders on the assignee's topic.
func (p *escalationPolicy) configured() bool { return p != nil && len(p.levels) > 0 }
// level returns the level at a 1-based position.
func (p *escalationPolicy) level(pos int64) (escalationLevel, bool) {
for _, l := range p.levels {
if l.position == pos {
return l, true
}
}
return escalationLevel{}, false
}
// loadEscalationPolicy reads one team's ladder. A team with no policy row
// returns nil, which every caller treats as "not configured" rather than as an
// error: most teams will never set one up.
func loadEscalationPolicy(ctx context.Context, q querier, teamID int64) (*escalationPolicy, error) {
p := &escalationPolicy{teamID: teamID}
err := q.QueryRowContext(ctx,
"SELECT repeat_count, fallback_topic FROM escalation_policies WHERE team_id = $1",
teamID).Scan(&p.repeatCount, &p.fallbackTopic)
if err == sql.ErrNoRows {
return nil, nil
}
if err != nil {
return nil, err
}
rows, err := q.QueryContext(ctx, `
SELECT l.id, l.position, l.timeout_seconds, t.kind, t.user_id
FROM escalation_levels l
LEFT JOIN escalation_targets t ON t.level_id = l.id
WHERE l.team_id = $1
ORDER BY l.position, t.id`, teamID)
if err != nil {
return nil, err
}
defer rows.Close()
byPosition := map[int64]int{} // position -> index in p.levels
for rows.Next() {
var id, position, timeout int64
var kind *string
var userID *int64
if err := rows.Scan(&id, &position, &timeout, &kind, &userID); err != nil {
return nil, err
}
idx, seen := byPosition[position]
if !seen {
p.levels = append(p.levels, escalationLevel{
id: id,
position: position,
timeout: time.Duration(timeout) * time.Second,
})
idx = len(p.levels) - 1
byPosition[position] = idx
}
// LEFT JOIN: a level with no targets yet still produces a row, with a
// NULL kind. It is a rung that pages nobody, which the API refuses to
// store but an older row could still hold.
if kind != nil {
p.levels[idx].targets = append(p.levels[idx].targets,
escalationTarget{kind: *kind, userID: userID})
}
}
return p, rows.Err()
}
// escalate advances every incident whose current level has run out of time.
//
// Runs on the notifier's tick, beside the reminder pass, because it is the same
// question asked differently: reminders ask "has this been ignored long
// enough to say it again", escalation asks "long enough to say it to somebody
// else". Sharing the tick means one query cadence and one outbox.
func escalate(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
rows, err := db.QueryContext(ctx, `
SELECT i.id, i.team_id, i.escalation_level, i.escalation_level_at, i.escalation_round
FROM incidents i
JOIN escalation_policies p ON p.team_id = i.team_id
WHERE i.resolved_at IS NULL
AND i.archived_at IS NULL
AND i.status = 'triggered'
AND (i.snoozed_until IS NULL OR i.snoozed_until <= $1)
AND i.escalation_level > 0`, time.Now().Unix())
if err != nil {
log.Printf("escalation: find due: %v", err)
return
}
type pending struct {
incidentID, teamID, level, round int64
levelAt int64
}
var due []pending
for rows.Next() {
var p pending
var levelAt *int64
if err := rows.Scan(&p.incidentID, &p.teamID, &p.level, &levelAt, &p.round); err != nil {
rows.Close()
log.Printf("escalation: scan: %v", err)
return
}
if levelAt == nil {
continue
}
p.levelAt = *levelAt
due = append(due, p)
}
rows.Close()
if err := rows.Err(); err != nil {
log.Printf("escalation: iterate: %v", err)
return
}
now := time.Now()
for _, d := range due {
policy, err := loadEscalationPolicy(ctx, db, d.teamID)
if err != nil {
log.Printf("escalation: load policy for team %d: %v", d.teamID, err)
continue
}
if !policy.configured() {
continue
}
current, ok := policy.level(d.level)
if !ok {
continue
}
if now.Sub(time.Unix(d.levelAt, 0)) < current.timeout {
continue
}
if err := advanceEscalation(ctx, db, cfg, policy, d.incidentID, d.level, d.round, now); err != nil {
log.Printf("escalation: advance incident %d: %v", d.incidentID, err)
}
}
}
// advanceEscalation moves one incident to its next rung, or off the end of the
// ladder.
//
// The whole move is one transaction: the level, the page and the timeline entry
// are one event, and an incident recorded as being at level 3 that nobody at
// level 3 was told about is the worst of the possible half-states.
func advanceEscalation(ctx context.Context, db *sql.DB, cfg NotifyConfig, policy *escalationPolicy, incidentID, level, round int64, now time.Time) error {
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return err
}
defer tx.Rollback() //nolint:errcheck
next := level + 1
nextRound := round
if _, ok := policy.level(next); !ok {
// Off the end. Either start the chain again, or make the last call.
if round < policy.repeatCount {
next, nextRound = 1, round+1
} else {
if err := escalationExhausted(ctx, tx, policy, incidentID, now); err != nil {
return err
}
return tx.Commit()
}
}
target, ok := policy.level(next)
if !ok {
return nil
}
paged, err := pageLevel(ctx, tx, cfg, policy, incidentID, target)
if err != nil {
return err
}
if _, err := tx.ExecContext(ctx, `
UPDATE incidents
SET escalation_level = $1, escalation_level_at = $2, escalation_round = $3
WHERE id = $4`, next, now.Unix(), nextRound, incidentID); err != nil {
return err
}
detail := "level " + strconv.FormatInt(next, 10)
if nextRound > round {
detail += " (round " + strconv.FormatInt(nextRound+1, 10) + ")"
}
if len(paged) > 0 {
detail += ": " + strings.Join(paged, ", ")
} else {
// Worth recording loudly: the rung exists, its turn came, and it woke
// nobody. That is a policy that looks configured and is not.
detail += ": nobody reachable"
}
if err := logEvent(ctx, tx, incidentID, evEscalated, nil, nil, &detail); err != nil {
return err
}
return tx.Commit()
}
// escalationExhausted is the end of the line: the fallback topic, once, and a
// timeline entry saying the ladder is finished. The incident stays triggered —
// escalation running out is not the same as somebody answering.
func escalationExhausted(ctx context.Context, tx *sql.Tx, policy *escalationPolicy, incidentID int64, now time.Time) error {
detail := "escalation exhausted"
if policy.fallbackTopic != "" {
if err := enqueueNotification(ctx, tx, incidentID, nil, policy.fallbackTopic, notifyEscalated); err != nil {
return err
}
detail += ": paged " + policy.fallbackTopic
} else {
detail += ": no fallback topic configured"
}
// Level 0 again, so the sweep stops considering it. The round counter is
// left where it is, as the record of how far it got.
if _, err := tx.ExecContext(ctx,
"UPDATE incidents SET escalation_level = 0, escalation_level_at = NULL WHERE id = $1",
incidentID); err != nil {
return err
}
return logEvent(ctx, tx, incidentID, evEscalated, nil, nil, &detail)
}
// pageLevel notifies every target of one level and reports who was woken.
//
// Each target gets its own outbox row, so each gets its own Acknowledge token:
// the button in a notification must acknowledge as the person holding the
// phone, not as whoever was paged first.
func pageLevel(ctx context.Context, tx *sql.Tx, cfg NotifyConfig, policy *escalationPolicy, incidentID int64, level escalationLevel) ([]string, error) {
var paged []string
seen := map[int64]bool{}
for _, t := range level.targets {
userID := t.userID
if t.kind == "oncall" {
onCall, err := currentOnCall(ctx, tx, policy.teamID)
if err != nil {
return nil, err
}
if onCall == nil {
continue
}
userID = onCall
}
if userID == nil || seen[*userID] {
continue
}
seen[*userID] = true
var topic *string
var username string
if err := tx.QueryRowContext(ctx,
"SELECT ntfy_topic, username FROM users WHERE id = $1 AND disabled_at IS NULL",
*userID).Scan(&topic, &username); err != nil {
// A disabled or deleted account is not an error in the middle of an
// escalation: it is a target that cannot be woken, and the next
// level is the answer to that.
continue
}
if topic == nil || *topic == "" {
continue
}
if err := enqueueNotification(ctx, tx, incidentID, userID, *topic, notifyEscalated); err != nil {
return nil, err
}
paged = append(paged, username)
}
return paged, nil
}
// startEscalation puts a newly opened incident on the first rung, when its team
// has a ladder. Called from openIncident, inside the same transaction, so an
// incident is never briefly open with no escalation clock running.
func startEscalation(ctx context.Context, q querier, incidentID, teamID int64) error {
policy, err := loadEscalationPolicy(ctx, q, teamID)
if err != nil || !policy.configured() {
return err
}
_, err = q.ExecContext(ctx,
"UPDATE incidents SET escalation_level = 1, escalation_level_at = $1 WHERE id = $2",
time.Now().Unix(), incidentID)
return err
}
// stopEscalation takes an incident off the ladder. Acknowledging or resolving
// is somebody saying "I have this", and continuing to wake people after that is
// the behaviour that teaches people to ignore the tool.
func stopEscalation(ctx context.Context, q querier, incidentID int64) error {
_, err := q.ExecContext(ctx,
"UPDATE incidents SET escalation_level = 0, escalation_level_at = NULL WHERE id = $1",
incidentID)
return err
}
// handleGetEscalation returns a team's ladder.
func handleGetEscalation(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
if !ok {
return
}
if !requireTeamMember(w, r, teamID) {
return
}
policy, err := loadEscalationPolicy(r.Context(), db, teamID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, escalationResponse(policy, teamID))
}
}
type escalationLevelJSON struct {
Position int64 `json:"position"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Targets []escalationTargetJSON `json:"targets"`
}
type escalationTargetJSON struct {
Kind string `json:"kind"`
UserID *int64 `json:"user_id,omitempty"`
}
type escalationJSON struct {
TeamID int64 `json:"team_id"`
RepeatCount int64 `json:"repeat_count"`
FallbackTopic string `json:"fallback_topic"`
Levels []escalationLevelJSON `json:"levels"`
}
func escalationResponse(p *escalationPolicy, teamID int64) escalationJSON {
out := escalationJSON{TeamID: teamID, Levels: []escalationLevelJSON{}}
if p == nil {
return out
}
out.RepeatCount = p.repeatCount
out.FallbackTopic = p.fallbackTopic
for _, l := range p.levels {
level := escalationLevelJSON{
Position: l.position,
TimeoutSeconds: int64(l.timeout.Seconds()),
Targets: []escalationTargetJSON{},
}
for _, t := range l.targets {
level.Targets = append(level.Targets, escalationTargetJSON{Kind: t.kind, UserID: t.userID})
}
out.Levels = append(out.Levels, level)
}
return out
}
// handleSetEscalation replaces a team's ladder wholesale.
//
// Replace rather than patch: the levels are an order, and an API that edits one
// rung has to answer what happens to the numbering of the others. Sending the
// whole ladder makes the order the client's to decide and the server's to
// store, and makes an edit atomic — there is no moment where level 2 exists
// twice.
func handleSetEscalation(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
if !ok {
return
}
if !requireTeamOwner(w, r, teamID) {
return
}
var req escalationJSON
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
if req.RepeatCount < 0 || req.RepeatCount > 10 {
respond(w, http.StatusBadRequest, errResp("repeat_count must be between 0 and 10"))
return
}
for i, l := range req.Levels {
if l.TimeoutSeconds <= 0 {
respond(w, http.StatusBadRequest, errResp("every level needs a timeout"))
return
}
if len(l.Targets) == 0 {
// A rung that pages nobody is not a delay, it is a silence with
// a number on it.
respond(w, http.StatusBadRequest,
errResp("level "+strconv.FormatInt(int64(i+1), 10)+" has no targets"))
return
}
for _, t := range l.Targets {
switch t.Kind {
case "oncall":
if t.UserID != nil {
respond(w, http.StatusBadRequest, errResp("an oncall target takes no user_id"))
return
}
case "user":
if t.UserID == nil {
respond(w, http.StatusBadRequest, errResp("a user target needs a user_id"))
return
}
default:
respond(w, http.StatusBadRequest, errResp("target kind must be user or oncall"))
return
}
}
}
tx, err := db.BeginTx(r.Context(), nil)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer tx.Rollback() //nolint:errcheck
if _, err := tx.ExecContext(r.Context(), `
INSERT INTO escalation_policies (team_id, repeat_count, fallback_topic, updated_at)
VALUES ($1, $2, $3, `+nowEpoch+`)
ON CONFLICT (team_id) DO UPDATE SET
repeat_count = excluded.repeat_count,
fallback_topic = excluded.fallback_topic,
updated_at = excluded.updated_at`,
teamID, req.RepeatCount, req.FallbackTopic); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
// The levels are replaced, not merged; the cascade takes the targets.
if _, err := tx.ExecContext(r.Context(),
"DELETE FROM escalation_levels WHERE team_id = $1", teamID); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
for i, l := range req.Levels {
var levelID int64
if err := tx.QueryRowContext(r.Context(), `
INSERT INTO escalation_levels (team_id, position, timeout_seconds)
VALUES ($1, $2, $3) RETURNING id`,
teamID, int64(i+1), l.TimeoutSeconds).Scan(&levelID); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
for _, t := range l.Targets {
if _, err := tx.ExecContext(r.Context(), `
INSERT INTO escalation_targets (level_id, kind, user_id)
VALUES ($1, $2, $3)`, levelID, t.Kind, t.UserID); err != nil {
// The only foreign key here is the user.
respond(w, http.StatusBadRequest, errResp("unknown user in targets"))
return
}
}
}
if err := tx.Commit(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
policy, err := loadEscalationPolicy(r.Context(), db, teamID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, escalationResponse(policy, teamID))
}
}
+385
View File
@@ -0,0 +1,385 @@
package api_test
import (
"net/http"
"strings"
"testing"
"time"
"git.ryuvia.com/niklas/terdut-server/internal/api"
)
// teamUser creates a user in the default team with an ntfy topic, so they can
// actually be paged.
func teamUser(t *testing.T, s *ts, username, topic string) int64 {
t.Helper()
var user struct {
ID int64 `json:"id"`
}
decode(t, s.req(t, http.MethodPost, "/api/users",
map[string]string{"username": username, "email": username + "@test.com"}), &user)
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/members",
map[string]any{"user_id": user.ID, "role": "member"})
resp.Body.Close()
setTopic(t, s, int(user.ID), topic)
return user.ID
}
// Escalation is all timeouts, and there is no fake clock in this package. The
// tests back-date escalation_level_at instead, which is the same trick the dead
// man's switch tests use on received_at: the sweeper reads a stored timestamp,
// so moving the timestamp is moving the clock.
// ladder configures the default team with two levels: the rota first, then a
// named person, then the fallback topic.
func ladder(t *testing.T, s *ts, secondUserID int64, repeat int64, fallback string) {
t.Helper()
resp := s.req(t, http.MethodPut, "/api/teams/"+defaultTeam+"/escalation", map[string]any{
"repeat_count": repeat,
"fallback_topic": fallback,
"levels": []map[string]any{
{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "oncall"}}},
{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "user", "user_id": secondUserID}}},
},
})
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("configure the ladder: %d", resp.StatusCode)
}
}
// overdue back-dates an incident's current level so its timeout has passed.
func overdue(t *testing.T, s *ts, incidentID int64) {
t.Helper()
s.exec(t, "UPDATE incidents SET escalation_level_at = $1 WHERE id = $2",
time.Now().Add(-time.Hour).Unix(), incidentID)
}
func escalationLevel(t *testing.T, s *ts, incidentID int64) (level, round int64) {
t.Helper()
if err := s.db.QueryRow(
"SELECT escalation_level, escalation_round FROM incidents WHERE id = $1",
incidentID).Scan(&level, &round); err != nil {
t.Fatalf("read escalation state: %v", err)
}
return level, round
}
// The whole point: nobody answers, so somebody else is woken.
func TestEscalation_PagesTheNextLevel(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
second := teamUser(t, s, "second", "terdut-second")
ladder(t, s, second, 0, "terdut-fallback")
postWebhook(t, s, []map[string]any{
amAlert("fp-esc", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
// Level 1 is the rota, so the first page went to the admin.
if level, _ := escalationLevel(t, s, 1); level != 1 {
t.Fatalf("a new incident should start at level 1, got %d", level)
}
if got := f.topicsSince(t); len(got) == 0 || got[0] != "terdut-admin" {
t.Fatalf("the first page should go to the on-call user, went to %v", got)
}
// Time passes with no acknowledgement.
f.forget()
overdue(t, s, 1)
s.sweepNotify(t)
if level, _ := escalationLevel(t, s, 1); level != 2 {
t.Errorf("expected level 2, got %d", level)
}
if got := f.topicsSince(t); len(got) != 1 || got[0] != "terdut-second" {
t.Errorf("level 2 should page the named user, paged %v", got)
}
// And the timeline says so, which is what somebody reads afterwards to
// understand why their phone rang at 04:00.
timeline := list(t, s.req(t, http.MethodGet, "/api/incidents/1/timeline", nil))
found := ""
for _, e := range timeline {
if e["type"] == "escalated" {
found, _ = e["detail"].(string)
}
}
if found == "" {
t.Error("the timeline should record the escalation")
} else if !strings.HasPrefix(found, "level 2") || !strings.Contains(found, "second") {
t.Errorf("the escalation entry should say which level and who: %q", found)
}
}
// Acknowledging is somebody saying "I have this". Nobody else should be woken.
func TestEscalation_AcknowledgementStopsIt(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
second := teamUser(t, s, "second", "terdut-second")
ladder(t, s, second, 0, "terdut-fallback")
postWebhook(t, s, []map[string]any{
amAlert("fp-ack", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil).Body.Close()
if level, _ := escalationLevel(t, s, 1); level != 0 {
t.Errorf("acknowledging should take the incident off the ladder, level is %d", level)
}
f.forget()
overdue(t, s, 1) // no-op: level is 0, so there is nothing due
s.sweepNotify(t)
if got := f.topicsSince(t); len(got) != 0 {
t.Errorf("an acknowledged incident should page nobody, paged %v", got)
}
}
// Resolving stops it too, and by the same mechanism.
func TestEscalation_ResolutionStopsIt(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
second := teamUser(t, s, "second", "terdut-second")
ladder(t, s, second, 0, "terdut-fallback")
postWebhook(t, s, []map[string]any{
amAlert("fp-res", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
f.forget()
overdue(t, s, 1)
s.sweepNotify(t)
if level, _ := escalationLevel(t, s, 1); level != 0 {
t.Errorf("a resolved incident should be off the ladder, level is %d", level)
}
}
// Snoozing is a deliberate "not now", so the ladder waits rather than carrying
// on without the person who asked for quiet.
func TestEscalation_SnoozePausesIt(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
second := teamUser(t, s, "second", "terdut-second")
ladder(t, s, second, 0, "terdut-fallback")
postWebhook(t, s, []map[string]any{
amAlert("fp-snooze", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
resp := s.req(t, http.MethodPost, "/api/incidents/1/snooze", map[string]any{"duration": "1h"})
resp.Body.Close()
f.forget()
overdue(t, s, 1)
s.sweepNotify(t)
if level, _ := escalationLevel(t, s, 1); level != 1 {
t.Errorf("a snoozed incident should stay where it is, level is %d", level)
}
if got := f.topicsSince(t); len(got) != 0 {
t.Errorf("a snoozed incident should page nobody, paged %v", got)
}
// When the snooze ends, the ladder picks up where it left off.
s.exec(t, "UPDATE incidents SET snoozed_until = $1 WHERE id = 1", time.Now().Add(-time.Minute).Unix())
s.sweepNotify(t)
if level, _ := escalationLevel(t, s, 1); level != 2 {
t.Errorf("after the snooze the ladder should resume, level is %d", level)
}
}
// Running out of ladder pages the team's fallback topic once, and says so.
func TestEscalation_ExhaustionPagesTheFallback(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
second := teamUser(t, s, "second", "terdut-second")
ladder(t, s, second, 0, "terdut-fallback")
postWebhook(t, s, []map[string]any{
amAlert("fp-end", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
overdue(t, s, 1)
s.sweepNotify(t) // level 2
f.forget()
overdue(t, s, 1)
s.sweepNotify(t) // off the end
if got := f.topicsSince(t); len(got) != 1 || got[0] != "terdut-fallback" {
t.Errorf("exhaustion should page the fallback topic once, paged %v", got)
}
level, _ := escalationLevel(t, s, 1)
if level != 0 {
t.Errorf("an exhausted ladder should stop asking, level is %d", level)
}
// The incident is still open: running out of people is not an answer.
var status string
if err := s.db.QueryRow("SELECT status FROM incidents WHERE id = 1").Scan(&status); err != nil {
t.Fatal(err)
}
if status != "triggered" {
t.Errorf("exhaustion must not resolve the incident, status is %q", status)
}
}
// repeat_count walks the whole ladder again before giving up.
func TestEscalation_RepeatsTheChain(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
second := teamUser(t, s, "second", "terdut-second")
ladder(t, s, second, 1, "terdut-fallback") // one extra round
postWebhook(t, s, []map[string]any{
amAlert("fp-repeat", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
overdue(t, s, 1)
s.sweepNotify(t) // level 2
f.forget()
overdue(t, s, 1)
s.sweepNotify(t) // back to level 1, round 2
level, round := escalationLevel(t, s, 1)
if level != 1 || round != 1 {
t.Errorf("expected level 1 round 1, got level %d round %d", level, round)
}
if got := f.topicsSince(t); len(got) != 1 || got[0] != "terdut-admin" {
t.Errorf("the second round should start at the top again, paged %v", got)
}
}
// A team without a ladder keeps exactly the behaviour it had, and never gets
// both a reminder and an escalation for the same silence.
func TestEscalation_WithoutAPolicyRemindersStillRun(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
postWebhook(t, s, []map[string]any{
amAlert("fp-noesc", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
// Age the first notification past the repeat interval.
f.forget()
s.exec(t, "UPDATE notifications SET created_at = $1, sent_at = $1",
time.Now().Add(-time.Hour).Unix())
s.sweepNotify(t)
if got := f.topicsSince(t); len(got) != 1 || got[0] != "terdut-admin" {
t.Errorf("without a ladder the reminder should still fire, paged %v", got)
}
if level, _ := escalationLevel(t, s, 1); level != 0 {
t.Errorf("an incident in a team with no ladder should not be on one, level is %d", level)
}
}
// With a ladder, reminders stop: two pages for one silence is how people learn
// to mute the tool.
func TestEscalation_WithAPolicyRemindersDoNotAlsoFire(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
second := teamUser(t, s, "second", "terdut-second")
ladder(t, s, second, 0, "terdut-fallback")
postWebhook(t, s, []map[string]any{
amAlert("fp-both", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
f.forget()
// Old enough for a reminder, but not yet due for escalation.
s.exec(t, "UPDATE notifications SET created_at = $1, sent_at = $1",
time.Now().Add(-time.Hour).Unix())
s.sweepNotify(t)
if got := f.topicsSince(t); len(got) != 0 {
t.Errorf("a team with a ladder should not also get reminders, paged %v", got)
}
}
// The API refuses a ladder that cannot page anybody.
func TestEscalation_RejectsAnUnusablePolicy(t *testing.T) {
s := newTS(t)
for _, c := range []struct {
name string
body map[string]any
}{
{"a level with no targets", map[string]any{
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{}}},
}},
{"a level with no timeout", map[string]any{
"levels": []map[string]any{{"timeout_seconds": 0, "targets": []map[string]any{{"kind": "oncall"}}}},
}},
{"a user target with no user", map[string]any{
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "user"}}}},
}},
{"an unknown target kind", map[string]any{
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "everybody"}}}},
}},
{"an absurd repeat count", map[string]any{
"repeat_count": 99,
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "oncall"}}}},
}},
} {
resp := s.req(t, http.MethodPut, "/api/teams/"+defaultTeam+"/escalation", c.body)
resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("%s: expected 400, got %d", c.name, resp.StatusCode)
}
}
}
// Editing the ladder is an owner's job; reading it is any member's.
func TestEscalation_OwnerOnlyToEdit(t *testing.T) {
s := newTS(t)
_, call := member(t, s, "plain")
resp := call(http.MethodPut, "/api/teams/"+defaultTeam+"/escalation", map[string]any{
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "oncall"}}}},
})
resp.Body.Close()
if resp.StatusCode != http.StatusForbidden {
t.Errorf("a member editing the ladder: expected 403, got %d", resp.StatusCode)
}
resp = call(http.MethodGet, "/api/teams/"+defaultTeam+"/escalation", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Errorf("a member reading the ladder: expected 200, got %d", resp.StatusCode)
}
}
// A target who cannot be woken is not a reason to stop: the next level is the
// answer to an unreachable one.
func TestEscalation_SkipsUnreachableTargets(t *testing.T) {
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
// Second user has no ntfy topic at all.
var user struct {
ID int64 `json:"id"`
}
decode(t, s.req(t, http.MethodPost, "/api/users",
map[string]string{"username": "silent", "email": "silent@test.com"}), &user)
s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/members",
map[string]any{"user_id": user.ID, "role": "member"}).Body.Close()
ladder(t, s, user.ID, 0, "terdut-fallback")
postWebhook(t, s, []map[string]any{
amAlert("fp-silent", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
f.forget()
overdue(t, s, 1)
s.sweepNotify(t)
// Level 2 was entered even though it woke nobody, so the ladder keeps
// moving toward the fallback rather than stalling on a silent rung.
if level, _ := escalationLevel(t, s, 1); level != 2 {
t.Errorf("expected the ladder to advance past an unreachable target, level is %d", level)
}
if got := f.topicsSince(t); len(got) != 0 {
t.Errorf("a target with no topic should page nothing, paged %v", got)
}
}
+17 -1
View File
@@ -52,6 +52,13 @@ type querier interface {
const incidentSelectFrom = ` const incidentSelectFrom = `
SELECT i.id, i.team_id, t.name, i.group_key, i.title, i.group_labels, i.status, i.severity, SELECT i.id, i.team_id, t.name, i.group_key, i.title, i.group_labels, i.status, i.severity,
i.escalation_level,
-- When this level runs out. Computed here rather than in Go because
-- the timeout lives beside the level in the policy, and one join is
-- cheaper than a second query per incident in a list.
(SELECT i.escalation_level_at + el.timeout_seconds
FROM escalation_levels el
WHERE el.team_id = i.team_id AND el.position = i.escalation_level),
i.triggered_at, i.triggered_at,
i.acknowledged_by, i.acknowledged_at, ack.username, i.acknowledged_by, i.acknowledged_at, ack.username,
i.assigned_to, asg.username, i.snoozed_until, i.assigned_to, asg.username, i.snoozed_until,
@@ -65,10 +72,11 @@ func scanIncident(s scanner) (models.Incident, error) {
var i models.Incident var i models.Incident
var groupLabelsJSON string var groupLabelsJSON string
var triggeredAt int64 var triggeredAt int64
var ackAt, snoozedUntil, resolvedAt, archivedAt *int64 var ackAt, snoozedUntil, resolvedAt, archivedAt, escalationDue *int64
if err := s.Scan( if err := s.Scan(
&i.ID, &i.TeamID, &i.TeamName, &i.GroupKey, &i.Title, &groupLabelsJSON, &i.Status, &i.Severity, &i.ID, &i.TeamID, &i.TeamName, &i.GroupKey, &i.Title, &groupLabelsJSON, &i.Status, &i.Severity,
&i.EscalationLevel, &escalationDue,
&triggeredAt, &triggeredAt,
&i.AcknowledgedByID, &ackAt, &i.AcknowledgedByUser, &i.AcknowledgedByID, &ackAt, &i.AcknowledgedByUser,
&i.AssignedToID, &i.AssignedToUser, &snoozedUntil, &i.AssignedToID, &i.AssignedToUser, &snoozedUntil,
@@ -83,6 +91,7 @@ func scanIncident(s scanner) (models.Incident, error) {
i.SnoozedUntil = unixPtr(snoozedUntil) i.SnoozedUntil = unixPtr(snoozedUntil)
i.ResolvedAt = unixPtr(resolvedAt) i.ResolvedAt = unixPtr(resolvedAt)
i.ArchivedAt = unixPtr(archivedAt) i.ArchivedAt = unixPtr(archivedAt)
i.EscalationDueAt = unixPtr(escalationDue)
return i, nil return i, nil
} }
@@ -235,6 +244,9 @@ func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, e
if n == 0 { if n == 0 {
return false, nil return false, nil
} }
if err := stopEscalation(ctx, q, incidentID); err != nil {
return false, err
}
if err := logEvent(ctx, q, incidentID, evResolved, nil, nil, nil); err != nil { if err := logEvent(ctx, q, incidentID, evResolved, nil, nil, nil); err != nil {
return false, err return false, err
} }
@@ -260,6 +272,10 @@ func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int6
if n, _ := res.RowsAffected(); n == 0 { if n, _ := res.RowsAffected(); n == 0 {
return false, nil return false, nil
} }
// Somebody has it: stop waking anybody else.
if err := stopEscalation(ctx, q, incidentID); err != nil {
return false, err
}
return true, logEvent(ctx, q, incidentID, evAcknowledged, &userID, nil, nil) return true, logEvent(ctx, q, incidentID, evAcknowledged, &userID, nil, nil)
} }
+5
View File
@@ -243,6 +243,11 @@ func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
time.Now().Unix(), incidentResolutionManual, id) { time.Now().Unix(), incidentResolutionManual, id) {
return return
} }
// A person closing an incident is the clearest possible "I have this".
if err := stopEscalation(r.Context(), db, id); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if err := logEvent(r.Context(), db, id, evResolved, &user.ID, nil, nil); err != nil { if err := logEvent(r.Context(), db, id, evResolved, &user.ID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error")) respond(w, http.StatusInternalServerError, errResp("internal error"))
return return
+13 -1
View File
@@ -43,6 +43,10 @@ const (
notifyTriggered = "triggered" notifyTriggered = "triggered"
notifyReminder = "reminder" notifyReminder = "reminder"
notifyResolved = "resolved" notifyResolved = "resolved"
// notifyEscalated is a page that went out because nobody answered the last
// one. Told apart from a reminder because it goes to somebody else.
notifyEscalated = "escalated"
) )
// Timeline event types the notifier writes, so an incident's history says who // Timeline event types the notifier writes, so an incident's history says who
@@ -120,6 +124,9 @@ func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
// Exported so tests can drive a pass without waiting on the ticker. // Exported so tests can drive a pass without waiting on the ticker.
func NotifySweep(ctx context.Context, db *sql.DB, cfg NotifyConfig) { func NotifySweep(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
enqueueReminders(ctx, db, cfg) enqueueReminders(ctx, db, cfg)
// Escalation before delivery, so a level that comes due on this tick is
// paged on this tick rather than waiting for the next one.
escalate(ctx, db, cfg)
deliverPending(ctx, db, cfg) deliverPending(ctx, db, cfg)
} }
@@ -159,7 +166,12 @@ func enqueueReminders(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
AND i.resolved_at IS NULL AND i.resolved_at IS NULL
AND i.archived_at IS NULL AND i.archived_at IS NULL
AND i.status = 'triggered' AND i.status = 'triggered'
AND (i.snoozed_until IS NULL OR i.snoozed_until <= $2)`, AND (i.snoozed_until IS NULL OR i.snoozed_until <= $2)
-- A team with an escalation ladder gets escalation instead. Both
-- would mean two pages for one silence, which is how people learn to
-- mute a tool.
AND NOT EXISTS (
SELECT 1 FROM escalation_levels el WHERE el.team_id = i.team_id)`,
now.Add(-repeat).Unix(), now.Unix()) now.Add(-repeat).Unix(), now.Unix())
if err != nil { if err != nil {
log.Printf("notifier: find reminders: %v", err) log.Printf("notifier: find reminders: %v", err)
+21
View File
@@ -69,6 +69,27 @@ func (f *fakeNtfy) messages() []pushed {
return append([]pushed(nil), f.got...) return append([]pushed(nil), f.got...)
} }
// topicsSince lists the topics published to since the last forget, which is how
// the escalation tests ask "who did this tick wake".
func (f *fakeNtfy) topicsSince(t *testing.T) []string {
t.Helper()
f.mu.Lock()
defer f.mu.Unlock()
out := make([]string, 0, len(f.got))
for _, m := range f.got {
out = append(out, m.Topic)
}
return out
}
// forget drops what has been published so far, so the next assertion is about
// this tick rather than the whole test.
func (f *fakeNtfy) forget() {
f.mu.Lock()
defer f.mu.Unlock()
f.got = nil
}
func (f *fakeNtfy) failWith(status int) { func (f *fakeNtfy) failWith(status int) {
f.mu.Lock() f.mu.Lock()
defer f.mu.Unlock() defer f.mu.Unlock()
+4
View File
@@ -109,6 +109,10 @@ func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config) http.Handler
r.Post("/api/teams/{teamID}/members", handleAddTeamMember(db)) r.Post("/api/teams/{teamID}/members", handleAddTeamMember(db))
r.Delete("/api/teams/{teamID}/members/{userID}", handleRemoveTeamMember(db)) r.Delete("/api/teams/{teamID}/members/{userID}", handleRemoveTeamMember(db))
// A team's escalation ladder: who is paged when nobody answers.
r.Get("/api/teams/{teamID}/escalation", handleGetEscalation(db))
r.Put("/api/teams/{teamID}/escalation", handleSetEscalation(db))
// A team's own dead man's switches: which of its alerts are heartbeats, // A team's own dead man's switches: which of its alerts are heartbeats,
// and how long a silence has to last before somebody is paged. // and how long a silence has to last before somebody is paged.
r.Get("/api/teams/{teamID}/deadman", handleGetTeamDeadman(db)) r.Get("/api/teams/{teamID}/deadman", handleGetTeamDeadman(db))
+95
View File
@@ -0,0 +1,95 @@
-- Escalation: page somebody else when the first person does not answer.
--
-- This is the gap the whole multi-tenancy line of work was opened to close.
-- Until now an unacknowledged incident re-paged the same topic every
-- notify_repeat forever, which is a louder version of the same silence: if the
-- person on call is asleep, has no signal, or has left, nothing else happens.
--
-- Shape: one policy per team, an ordered list of levels, each level with a
-- timeout and a set of targets. When a level's timeout passes and the incident
-- is still triggered, the next level is paged. When the last level passes, the
-- chain repeats repeat_count times, and then the team's fallback topic is paged
-- once as the end of the line.
--
-- A team WITHOUT a policy keeps exactly today's behaviour: page the assignee,
-- then remind on the same topic. Escalation is opt-in per team, and the two
-- never both run for one incident -- see enqueueReminders.
CREATE TABLE escalation_policies (
-- One per team for now, hence the team as the key rather than an id with a
-- unique index: routing different alerts to different chains needs the
-- alert to carry something to route ON, which is a separate question.
team_id BIGINT PRIMARY KEY REFERENCES teams(id) ON DELETE CASCADE,
-- How many extra times to run the whole chain after it has been walked
-- once. 0 means walk it once and stop at the fallback.
repeat_count BIGINT NOT NULL DEFAULT 0 CHECK (repeat_count >= 0 AND repeat_count <= 10),
-- Where the last page goes when every level has been tried. Per team now:
-- TERDUT_NTFY_FALLBACK_TOPIC was one topic for the whole install, which in
-- a multi-team server pages the wrong people. Empty means the chain simply
-- ends.
fallback_topic TEXT NOT NULL DEFAULT '',
updated_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
);
CREATE TABLE escalation_levels (
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
team_id BIGINT NOT NULL REFERENCES escalation_policies(team_id) ON DELETE CASCADE,
-- 1-based, dense. The API rewrites the whole ladder on every edit rather
-- than patching one rung, so there is no way to leave a gap.
position BIGINT NOT NULL,
-- How long this level has to produce an acknowledgement before the next one
-- is paged. Seconds, like every other duration in this schema.
timeout_seconds BIGINT NOT NULL CHECK (timeout_seconds > 0),
UNIQUE (team_id, position)
);
-- Who a level pages. Either a named person, or whoever the team's rota says is
-- on call today -- which is the target that keeps working when the rota
-- changes and nobody remembers to edit the policy.
CREATE TABLE escalation_targets (
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
level_id BIGINT NOT NULL REFERENCES escalation_levels(id) ON DELETE CASCADE,
kind TEXT NOT NULL CHECK (kind IN ('user', 'oncall')),
-- Set for kind='user', NULL for kind='oncall'.
user_id BIGINT REFERENCES users(id) ON DELETE CASCADE,
CHECK ((kind = 'user' AND user_id IS NOT NULL) OR (kind = 'oncall' AND user_id IS NULL))
);
CREATE INDEX escalation_targets_level_idx ON escalation_targets(level_id);
-- ---------------------------------------------------------------------------
-- Where an incident is in its chain.
--
-- On the incident rather than in a side table: it is read on every notifier
-- tick alongside the incident's status, and one row per incident is exactly
-- what the state is.
-- ---------------------------------------------------------------------------
-- 0 means no level has been paged yet, which is the state of every incident
-- that existed before escalation and of every incident in a team with no
-- policy. 1 is the first level.
ALTER TABLE incidents ADD COLUMN escalation_level BIGINT NOT NULL DEFAULT 0;
-- When the current level was entered, and therefore what its timeout is
-- measured from. NULL while escalation_level is 0.
ALTER TABLE incidents ADD COLUMN escalation_level_at BIGINT;
-- How many times the chain has been walked in full. Compared against the
-- policy's repeat_count.
ALTER TABLE incidents ADD COLUMN escalation_round BIGINT NOT NULL DEFAULT 0;
-- The notifier's escalation query: incidents still waiting, oldest level first.
CREATE INDEX incidents_escalation_idx
ON incidents(escalation_level_at)
WHERE resolved_at IS NULL AND status = 'triggered';
-- 'escalated' joins the outbox kinds: a page that went out because nobody
-- answered the last one, which is worth telling apart from the first page and
-- from a reminder when reading the timeline or debugging a delivery.
ALTER TABLE notifications DROP CONSTRAINT notifications_kind_check;
ALTER TABLE notifications ADD CONSTRAINT notifications_kind_check
CHECK (kind IN ('triggered', 'reminder', 'resolved', 'escalated'));
+7
View File
@@ -11,6 +11,13 @@ import "time"
// the webhook and the sweeper may flip to "resolved" once every member alert has // the webhook and the sweeper may flip to "resolved" once every member alert has
// stopped firing. // stopped firing.
type Incident struct { type Incident struct {
// EscalationLevel is which rung of its team's ladder this incident is on,
// 0 for none — either the team has no ladder, or somebody has answered.
// EscalationDueAt is when the current level runs out, so a client can say
// how long is left rather than only what already happened.
EscalationLevel int64 `json:"escalation_level"`
EscalationDueAt *time.Time `json:"escalation_due_at,omitempty"`
// TeamID is the team that owns this incident, fixed when it opens: an // TeamID is the team that owns this incident, fixed when it opens: an
// incident never moves between teams. TeamName rides along so the combined // incident never moves between teams. TeamName rides along so the combined
// queue can badge each row without a second request. // queue can badge each row without a second request.
+30
View File
@@ -657,3 +657,33 @@ kbd {
.admin-settings .setting-unit { max-width: 8em; } .admin-settings .setting-unit { max-width: 8em; }
.admin-settings button[type="submit"] { margin-top: 12px; } .admin-settings button[type="submit"] { margin-top: 12px; }
.small { font-size: 13px; } .small { font-size: 13px; }
/* --- team settings -------------------------------------------------------
Forms with a label above each control, rather than the queue's rows of
links. The escalation ladder is the only nested structure in the app, so it
gets a little indentation to make the levels read as an order. */
.stacked-form { display: flex; flex-direction: column; gap: 10px; margin-top: 12px; align-items: flex-start; }
.stacked-form label { display: flex; align-items: center; gap: 6px; flex-wrap: wrap; font-size: 14px; }
.stacked-form label.checkbox { gap: 8px; }
.stacked-form input.wide { min-width: min(420px, 100%); }
.team-picker { margin-top: 8px; max-width: 100%; }
.ladder-level {
border-left: 3px solid var(--border-strong);
padding: 8px 0 8px 12px; margin: 12px 0;
}
.ladder-head { display: flex; align-items: center; gap: 10px; margin-bottom: 6px; }
.ladder-targets { display: flex; flex-direction: column; gap: 6px; margin-top: 8px; }
.target-row { display: flex; gap: 6px; align-items: center; flex-wrap: wrap; }
/* An integration key is shown exactly once, so it should look like something
to act on rather than another row of text. */
.key-panel {
margin-top: 12px; padding: 12px;
border: 1px solid var(--accent); border-radius: 8px; background: var(--accent-soft);
}
.key-panel pre {
overflow-x: auto; background: var(--surface); border: 1px solid var(--border);
border-radius: 6px; padding: 8px; font-size: 12px;
}
.key-url code { word-break: break-all; }
+5
View File
@@ -59,6 +59,10 @@
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M6 16V11a6 6 0 0 1 12 0v5l1.5 2h-15z"/><path d="M10 20.5a2 2 0 0 0 4 0"/></svg> <svg viewBox="0 0 24 24" aria-hidden="true"><path d="M6 16V11a6 6 0 0 1 12 0v5l1.5 2h-15z"/><path d="M10 20.5a2 2 0 0 0 4 0"/></svg>
<span class="nav-label">Alerts</span> <span class="nav-label">Alerts</span>
</a> </a>
<a class="nav-link" href="/team" data-section="team">
<svg viewBox="0 0 24 24" aria-hidden="true"><circle cx="9" cy="8" r="3"/><circle cx="17" cy="9" r="2.5"/><path d="M3 19a6 6 0 0 1 12 0M15 19a5 5 0 0 1 6-4"/></svg>
<span class="nav-label">Team</span>
</a>
<!-- Hidden unless the signed-in user is a system administrator; app.js <!-- Hidden unless the signed-in user is a system administrator; app.js
unhides it once /api/me says so. The server refuses every admin unhides it once /api/me says so. The server refuses every admin
endpoint regardless, so this is a courtesy and not a gate. --> endpoint regardless, so this is a courtesy and not a gate. -->
@@ -87,6 +91,7 @@
<section id="view-oncall" class="view view-page" data-view="oncall" hidden></section> <section id="view-oncall" class="view view-page" data-view="oncall" hidden></section>
<section id="view-alerts" class="view view-page" data-view="alerts" hidden></section> <section id="view-alerts" class="view view-page" data-view="alerts" hidden></section>
<section id="view-team" class="view view-page" data-view="team" hidden></section>
<section id="view-admin" class="view view-page" data-view="admin" hidden></section> <section id="view-admin" class="view view-page" data-view="admin" hidden></section>
<section id="view-more" class="view view-page" data-view="more" hidden></section> <section id="view-more" class="view view-page" data-view="more" hidden></section>
</div> </div>
+24
View File
@@ -92,6 +92,30 @@ export const createTeam = (name) => call('POST', '/teams', { body: { name } });
export const renameTeam = (id, name) => call('PUT', `/teams/${id}`, { body: { name } }); export const renameTeam = (id, name) => call('PUT', `/teams/${id}`, { body: { name } });
export const deleteTeam = (id) => call('DELETE', `/teams/${id}`); export const deleteTeam = (id) => call('DELETE', `/teams/${id}`);
// A team's own settings. Every write is owner-only and every read is
// member-only; the server answers 403 and 404 respectively, so the UI shows
// what the role allows rather than guarding it.
export const teamMembers = (id) => call('GET', `/teams/${id}/members`);
export const addTeamMember = (id, userID, role) =>
call('POST', `/teams/${id}/members`, { body: { user_id: userID, role } });
export const removeTeamMember = (id, userID) => call('DELETE', `/teams/${id}/members/${userID}`);
export const integrations = (id) => call('GET', `/teams/${id}/integrations`);
export const createIntegration = (id, name) =>
call('POST', `/teams/${id}/integrations`, { body: { name } });
export const deleteIntegration = (id, integrationID) =>
call('DELETE', `/teams/${id}/integrations/${integrationID}`);
export const deadman = (id) => call('GET', `/teams/${id}/deadman`);
export const setDeadman = (id, body) => call('PUT', `/teams/${id}/deadman`, { body });
export const escalation = (id) => call('GET', `/teams/${id}/escalation`);
export const setEscalation = (id, body) => call('PUT', `/teams/${id}/escalation`, { body });
export const assignSchedule = (id, userID, dates, replace = false) =>
call('POST', `/teams/${id}/schedule`, { body: { user_id: userID, dates, replace } });
export const unassignSchedule = (id, entryID) => call('DELETE', `/teams/${id}/schedule/${entryID}`);
// Administration. Every one of these is refused with 403 for anybody without // Administration. Every one of these is refused with 403 for anybody without
// the flag, so the UI hides the section rather than guarding it. // the flag, so the UI hides the section rather than guarding it.
export const adminTeams = () => call('GET', '/admin/teams'); export const adminTeams = () => call('GET', '/admin/teams');
+3 -1
View File
@@ -9,6 +9,7 @@ import * as incident from './incident.js';
import * as oncall from './oncall.js'; import * as oncall from './oncall.js';
import * as alerts from './alerts.js'; import * as alerts from './alerts.js';
import * as account from './account.js'; import * as account from './account.js';
import * as team from './team.js';
import * as admin from './admin.js'; import * as admin from './admin.js';
const $ = (id) => document.getElementById(id); const $ = (id) => document.getElementById(id);
@@ -18,6 +19,7 @@ const SECTIONS = {
queue: { title: 'Queue', view: queue }, queue: { title: 'Queue', view: queue },
oncall: { title: 'On-call', view: oncall }, oncall: { title: 'On-call', view: oncall },
alerts: { title: 'Alerts', view: alerts }, alerts: { title: 'Alerts', view: alerts },
team: { title: 'Team', view: team },
admin: { title: 'Admin', view: admin }, admin: { title: 'Admin', view: admin },
more: { title: 'Account', view: account }, more: { title: 'Account', view: account },
}; };
@@ -26,7 +28,7 @@ function parseRoute(pathname) {
const m = pathname.match(/^\/incidents\/(\d+)\/?$/); const m = pathname.match(/^\/incidents\/(\d+)\/?$/);
if (m) return { section: 'queue', incident: Number(m[1]) }; if (m) return { section: 'queue', incident: Number(m[1]) };
const name = pathname.replace(/^\/|\/$/g, ''); const name = pathname.replace(/^\/|\/$/g, '');
if (name === 'oncall' || name === 'alerts' || name === 'admin' || name === 'more') return { section: name }; if (name === 'oncall' || name === 'alerts' || name === 'team' || name === 'admin' || name === 'more') return { section: name };
return { section: 'queue', incident: null }; return { section: 'queue', incident: null };
} }
+13
View File
@@ -95,6 +95,15 @@ function statusBadges() {
if (inc.status !== 'resolved' && isFuture(inc.snoozed_until)) { if (inc.status !== 'resolved' && isFuture(inc.snoozed_until)) {
out.push(badge(`Snoozed · ${until(inc.snoozed_until)} left`, 'st-snoozed')); out.push(badge(`Snoozed · ${until(inc.snoozed_until)} left`, 'st-snoozed'));
} }
// Where it is on the ladder, while it is still climbing. The queue shows
// what happened; this says what happens next, which is the question somebody
// looking at an unacknowledged incident actually has.
if (inc.escalation_level > 0) {
const left = inc.escalation_due_at && isFuture(inc.escalation_due_at)
? ` · next in ${until(inc.escalation_due_at)}`
: ' · next page due';
out.push(badge(`Escalating · level ${inc.escalation_level}${left}`, 'st-triggered'));
}
if (inc.archived_at) out.push(badge('Archived', 'plain')); if (inc.archived_at) out.push(badge('Archived', 'plain'));
return out; return out;
} }
@@ -115,6 +124,10 @@ function facts() {
if (inc.status !== 'resolved' && isFuture(inc.snoozed_until)) { if (inc.status !== 'resolved' && isFuture(inc.snoozed_until)) {
add('Snoozed until', when(inc.snoozed_until)); add('Snoozed until', when(inc.snoozed_until));
} }
if (inc.escalation_level > 0 && inc.escalation_due_at) {
add('Escalates next', when(inc.escalation_due_at),
h('span', { class: 'sub', text: ` · level ${inc.escalation_level}` }));
}
if (inc.resolved_at) { if (inc.resolved_at) {
const how = inc.resolution_source === 'manual' ? 'by hand' : 'alerts stopped firing'; const how = inc.resolution_source === 'manual' ? 'by hand' : 'alerts stopped firing';
add('Resolved', when(inc.resolved_at), h('span', { class: 'sub', text: ` · ${how}` })); add('Resolved', when(inc.resolved_at), h('span', { class: 'sub', text: ` · ${how}` }));
+453
View File
@@ -0,0 +1,453 @@
// Team settings: the rota, who is in the team, where its alerts come from,
// what it escalates through, and which of its alerts are heartbeats.
//
// Everything here was API-only until now, which meant a team owner had to use
// curl to set up escalation — the feature this whole line of work exists for.
//
// The server decides what a role may do: an owner's edits succeed, a member's
// are refused with 403, and a non-member gets 404 for the lot. This view hides
// the controls a member cannot use, because a form that always fails is worse
// than no form, but it is not the thing enforcing anything.
import * as api from './api.js';
import { h, clear, spinner, confirm } from './ui.js';
import { state, currentTeam, users as allUsers } from './state.js';
import { isoDate, addDays } from './format.js';
const view = () => document.getElementById('view-team');
let teamID = null;
let data = null; // { team, members, integrations, escalation, deadman, schedule, users }
let error = null;
let freshKey = null; // an integration key, shown once, until the view is left
export function show() {
if (!data) clear(view(), spinner());
refresh();
}
function selectedTeam() {
const teams = state.teams || [];
return teams.find((t) => t.id === teamID) || currentTeam();
}
export async function refresh() {
const team = selectedTeam();
if (!team) {
data = null;
render();
return;
}
teamID = team.id;
try {
// A member may read all of this; only the writes are owner-only.
const [members, integrations, escalation, deadman, schedule, users] = await Promise.all([
api.teamMembers(team.id),
api.integrations(team.id),
api.escalation(team.id),
api.deadman(team.id),
api.schedule(team.id, isoDate(new Date()), isoDate(addDays(new Date(), 30))),
allUsers(),
]);
data = { team, members, integrations, escalation, deadman, schedule, users };
error = null;
} catch (err) {
error = err.message;
}
render();
}
function isOwner() {
return data?.team?.role === 'owner' || state.me?.user?.is_admin;
}
function render() {
if (!data) {
clear(view(), error
? h('div', { class: 'load-error', text: error })
: h('div', { class: 'card' }, h('p', { class: 'muted', text: 'You are not in a team yet.' })));
return;
}
clear(view(),
error && h('div', { class: 'load-error', text: `Showing older data: ${error}` }),
teamPicker(),
!isOwner() && h('div', { class: 'card' },
h('p', { class: 'muted small', text: 'You are a member of this team. Only an owner can change its settings.' })),
scheduleCard(),
escalationCard(),
integrationsCard(),
deadmanCard(),
membersCard(),
);
}
// Only shown to somebody in more than one team, like the queue's filter chips.
function teamPicker() {
if ((state.teams || []).length < 2) {
return h('div', { class: 'card' }, h('h2', { text: data.team.name }));
}
const select = h('select', { class: 'team-picker' },
...state.teams.map((t) => h('option', {
value: String(t.id), text: t.name, selected: t.id === teamID,
})));
select.addEventListener('change', () => {
teamID = Number(select.value);
data = null;
freshKey = null;
show();
});
return h('div', { class: 'card' }, h('h2', { text: 'Team' }), select);
}
// --- schedule --------------------------------------------------------------
// The rota is one person per UTC day. The on-call page shows it; this is where
// it is set, which until now was the TUI's job and the TUI cannot do it any
// more.
function scheduleCard() {
const rows = (data.schedule || []).map((e) =>
h('tr', {},
h('td', { text: e.date }),
h('td', {}, h('strong', { text: e.username })),
h('td', {}, isOwner() && h('button', {
class: 'btn-sm danger', type: 'button', text: 'Clear',
onclick: () => act(() => api.unassignSchedule(teamID, e.id)),
})),
));
return h('div', { class: 'card' },
h('h2', { text: 'On-call rota' }),
h('p', { class: 'muted small', text: 'One person per UTC day, for the next 30 days.' }),
rows.length
? h('table', { class: 'admin-table' }, h('tbody', {}, rows))
: h('p', { class: 'muted', text: 'Nobody is scheduled.' }),
isOwner() && assignForm(),
);
}
function assignForm() {
const who = memberSelect();
const from = h('input', { type: 'date', required: true, value: isoDate(new Date()) });
const days = h('input', { type: 'number', min: '1', max: '31', value: '1', class: 'setting-value' });
const replace = h('input', { type: 'checkbox' });
const form = h('form', { class: 'stacked-form' },
h('label', {}, 'Who ', who),
h('label', {}, 'From ', from),
h('label', {}, 'Days ', days),
// Taking a day somebody else holds has to be asked for, the same rule the
// API enforces: a plain assignment that silently moved a shift would move
// who gets paged without telling either of them.
h('label', { class: 'checkbox' }, replace, ' Take days somebody else holds'),
h('button', { class: 'btn', type: 'submit', text: 'Assign' }));
form.addEventListener('submit', (e) => {
e.preventDefault();
const start = new Date(from.value + 'T00:00:00Z');
const dates = [];
for (let i = 0; i < Number(days.value || 1); i++) dates.push(isoDate(addDays(start, i)));
act(() => api.assignSchedule(teamID, Number(who.value), dates, replace.checked));
});
return form;
}
function memberSelect(selected) {
return h('select', {},
...(data.members || []).map((m) => h('option', {
value: String(m.user_id), text: m.username, selected: m.user_id === selected,
})));
}
// --- escalation ------------------------------------------------------------
// The ladder is edited as a whole and sent as a whole, because the API replaces
// it wholesale: the levels are an order, and patching one rung would leave the
// numbering of the others undecided.
let draft = null;
function escalationCard() {
const esc = data.escalation;
if (!draft) {
draft = {
repeat_count: esc.repeat_count || 0,
fallback_topic: esc.fallback_topic || '',
levels: (esc.levels || []).map((l) => ({
timeout_seconds: l.timeout_seconds,
targets: (l.targets || []).map((t) => ({ kind: t.kind, user_id: t.user_id })),
})),
};
}
const body = [];
if (!draft.levels.length) {
body.push(h('p', { class: 'muted' },
'No ladder. An unacknowledged incident re-pages the same person every ',
'reminder interval and nobody else is woken.'));
}
draft.levels.forEach((level, i) => {
body.push(h('div', { class: 'ladder-level' },
h('div', { class: 'ladder-head' },
h('strong', { text: `Level ${i + 1}` }),
isOwner() && h('button', {
class: 'btn-sm danger', type: 'button', text: 'Remove',
onclick: () => { draft.levels.splice(i, 1); render(); },
})),
h('label', {}, 'Wait ', minutesInput(level.timeout_seconds, (secs) => {
level.timeout_seconds = secs;
}), ' before the next level'),
h('div', { class: 'ladder-targets' },
...level.targets.map((t, ti) => targetRow(level, t, ti)),
isOwner() && h('button', {
class: 'btn-sm', type: 'button', text: '+ target',
onclick: () => { level.targets.push({ kind: 'oncall' }); render(); },
})),
));
});
if (isOwner()) {
body.push(h('button', {
class: 'btn-sm', type: 'button', text: '+ level',
onclick: () => {
draft.levels.push({ timeout_seconds: 300, targets: [{ kind: 'oncall' }] });
render();
},
}));
const repeat = h('input', {
type: 'number', min: '0', max: '10', class: 'setting-value',
value: String(draft.repeat_count),
oninput: (e) => { draft.repeat_count = Number(e.target.value); },
});
const fallback = h('input', {
type: 'text', value: draft.fallback_topic, placeholder: 'terdut-oncall-all',
oninput: (e) => { draft.fallback_topic = e.target.value; },
});
body.push(h('label', {}, 'Repeat the whole ladder ', repeat, ' more times'));
body.push(h('label', {}, 'Then page this ntfy topic once ', fallback));
body.push(h('button', {
class: 'btn', type: 'button', text: 'Save ladder',
onclick: () => act(() => api.setEscalation(teamID, draft), { resetDraft: true }),
}));
}
return h('div', { class: 'card' },
h('h2', { text: 'Escalation' }),
h('p', { class: 'muted small' },
'When a level’s wait passes and nobody has acknowledged, the next level is ',
'paged. Acknowledging or resolving stops it; snoozing pauses it.'),
...body,
);
}
function targetRow(level, target, index) {
const kind = h('select', {},
h('option', { value: 'oncall', text: 'Whoever is on call', selected: target.kind === 'oncall' }),
h('option', { value: 'user', text: 'A specific person', selected: target.kind === 'user' }));
kind.addEventListener('change', () => {
target.kind = kind.value;
target.user_id = kind.value === 'user' ? (data.members[0] || {}).user_id : undefined;
render();
});
const who = target.kind === 'user'
? memberSelect(target.user_id)
: null;
if (who) {
who.addEventListener('change', () => { target.user_id = Number(who.value); });
}
return h('div', { class: 'target-row' }, kind, who,
isOwner() && h('button', {
class: 'btn-sm danger', type: 'button', text: '×',
title: 'Remove this target',
onclick: () => { level.targets.splice(index, 1); render(); },
}));
}
function minutesInput(seconds, onChange) {
const input = h('input', {
type: 'number', min: '1', class: 'setting-value',
value: String(Math.max(1, Math.round(seconds / 60))),
oninput: (e) => onChange(Number(e.target.value) * 60),
});
return h('span', {}, input, ' minutes');
}
// --- integrations ----------------------------------------------------------
function integrationsCard() {
const rows = (data.integrations || []).map((i) =>
h('tr', {},
h('td', {}, h('strong', { text: i.name })),
h('td', { class: 'muted small', text: i.kind }),
h('td', { class: 'muted small', text: i.last_used_at ? 'in use' : 'never used' }),
h('td', {}, isOwner() && h('button', {
class: 'btn-sm danger', type: 'button', text: 'Revoke',
onclick: async () => {
if (!(await confirm({
title: `Revoke ${i.name}?`,
text: 'Anything posting with this key stops delivering immediately.',
confirmLabel: 'Revoke',
danger: true,
}))) return;
act(() => api.deleteIntegration(teamID, i.id));
},
})),
));
return h('div', { class: 'card' },
h('h2', { text: 'Alert sources' }),
h('p', { class: 'muted small' },
'Alerts arrive on an integration key, which says both that the sender may ',
'post and which team the alerts belong to.'),
rows.length
? h('table', { class: 'admin-table' }, h('tbody', {}, rows))
: h('p', { class: 'muted', text: 'No alert source yet, so nothing can reach this team.' }),
freshKey && newKeyPanel(),
isOwner() && !freshKey && newIntegrationForm(),
);
}
// The key is returned exactly once. Say so, show it large, and give the
// Alertmanager snippet with it already in place — the next thing anybody does
// with it is paste it into a config.
function newKeyPanel() {
const url = freshKey.url || `${location.origin}/api/integrations/${freshKey.key}/alertmanager`;
const snippet = `receivers:
- name: terdut
webhook_configs:
- url: ${url}
send_resolved: true`;
return h('div', { class: 'key-panel' },
h('strong', { text: 'Copy this now — it is not shown again.' }),
h('pre', { class: 'key-url' }, h('code', { text: url })),
h('button', {
class: 'btn-sm', type: 'button', text: 'Copy URL',
onclick: () => navigator.clipboard?.writeText(url),
}),
h('p', { class: 'muted small', text: 'Alertmanager receiver:' }),
h('pre', {}, h('code', { text: snippet })),
h('button', {
class: 'btn-sm', type: 'button', text: 'Done',
onclick: () => { freshKey = null; render(); },
}),
);
}
function newIntegrationForm() {
const name = h('input', { type: 'text', placeholder: 'prod alertmanager', required: true });
const form = h('form', { class: 'inline-form' }, name,
h('button', { class: 'btn', type: 'submit', text: 'Add' }));
form.addEventListener('submit', async (e) => {
e.preventDefault();
try {
freshKey = await api.createIntegration(teamID, name.value.trim());
await refresh();
} catch (err) {
error = err.message;
render();
}
});
return form;
}
// --- dead man's switches ---------------------------------------------------
function deadmanCard() {
const d = data.deadman || {};
const matchers = h('input', {
type: 'text', value: d.matchers || '', placeholder: 'alertname=Watchdog',
class: 'wide',
});
const timeout = h('input', {
type: 'number', min: '0', class: 'setting-value',
value: String(Math.round((d.timeout_seconds || 0) / 60)),
});
const severity = h('select', {},
...['critical', 'error', 'warning', 'info'].map((s) =>
h('option', { value: s, text: s, selected: (d.severity || 'critical') === s })));
const form = h('form', { class: 'stacked-form' },
h('label', {}, 'Heartbeat alerts ', matchers),
h('label', {}, 'Declare dead after ', timeout, ' minutes of silence'),
h('label', {}, 'Open the incident at severity ', severity),
h('button', { class: 'btn', type: 'submit', text: 'Save switches' }));
form.addEventListener('submit', (e) => {
e.preventDefault();
act(() => api.setDeadman(teamID, {
matchers: matchers.value.trim(),
timeout_seconds: Number(timeout.value) * 60,
severity: severity.value,
}));
});
return h('div', { class: 'card' },
h('h2', { text: 'Dead man’s switches' }),
h('p', { class: 'muted small' },
'Alerts whose ABSENCE is the signal. Receiving one opens nothing; going ',
'quiet for longer than the timeout opens an incident. ',
h('code', { text: 'alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat' }),
' — semicolons separate switches, commas separate conditions, and every ',
'switch must name an alertname. Leave empty to watch nothing.'),
isOwner() ? form : h('p', { class: 'muted', text: d.matchers || 'Nothing watched.' }),
);
}
// --- members ---------------------------------------------------------------
function membersCard() {
const rows = (data.members || []).map((m) =>
h('tr', {},
h('td', {}, h('strong', { text: m.username })),
h('td', { class: 'muted small', text: m.role }),
h('td', {}, isOwner() && h('button', {
class: 'btn-sm', type: 'button',
text: m.role === 'owner' ? 'Make member' : 'Make owner',
onclick: () => act(() =>
api.addTeamMember(teamID, m.user_id, m.role === 'owner' ? 'member' : 'owner')),
}), isOwner() && h('button', {
class: 'btn-sm danger', type: 'button', text: 'Remove',
onclick: () => act(() => api.removeTeamMember(teamID, m.user_id)),
})),
));
const inTeam = new Set((data.members || []).map((m) => m.user_id));
const candidates = (data.users || []).filter((u) => !inTeam.has(u.id) && !u.disabled_at);
const pick = h('select', {},
...candidates.map((u) => h('option', { value: String(u.id), text: u.username })));
const role = h('select', {},
h('option', { value: 'member', text: 'member' }),
h('option', { value: 'owner', text: 'owner' }));
const form = h('form', { class: 'inline-form' }, pick, role,
h('button', { class: 'btn', type: 'submit', text: 'Add' }));
form.addEventListener('submit', (e) => {
e.preventDefault();
act(() => api.addTeamMember(teamID, Number(pick.value), role.value));
});
return h('div', { class: 'card' },
h('h2', { text: 'Members' }),
h('table', { class: 'admin-table' }, h('tbody', {}, rows)),
isOwner() && candidates.length > 0 && form,
);
}
// --- plumbing --------------------------------------------------------------
// act runs a write and reloads. Errors are shown rather than thrown away: a
// 409 from the last-owner guard or the schedule's conflict rule is the server
// explaining itself, and the reader needs to see it.
async function act(fn, { resetDraft = false } = {}) {
try {
await fn();
error = null;
if (resetDraft) draft = null;
} catch (err) {
error = err.message;
}
if (!resetDraft) draft = null;
await refresh();
}