Compare commits
5 Commits
303e7a3365
...
v0.13.0
| Author | SHA1 | Date | |
|---|---|---|---|
| 53e5e03f4e | |||
| 4d62c1130b | |||
| 3183e7e5c5 | |||
| 94d23a593c | |||
| b0a02c010b |
@@ -155,12 +155,25 @@ somewhere to exec. The sidecar, the PVC and the `backupSidecar` values are all g
|
||||
|
||||
## Configuration
|
||||
|
||||
Two kinds of setting, split by who changes them and how often.
|
||||
|
||||
**Where the server is plugged in** stays in the environment: the listen address,
|
||||
the database DSN, the ntfy URL and token, the public URL. They are needed before
|
||||
the database is open, and two of them are credentials.
|
||||
|
||||
**How the server behaves** lives in the database and is edited by an
|
||||
administrator in the web UI or through `PUT /api/admin/settings`, taking effect
|
||||
on the next sweep rather than at the next restart. The variables below marked
|
||||
**seed** are the value each of those starts from: written once, on first start,
|
||||
and never overwritten afterwards — a redeploy cannot put a chart's default back
|
||||
over an administrator's edit.
|
||||
|
||||
| Variable | Default | Description |
|
||||
|---|---|---|
|
||||
| `TERDUT_ADDR` | `:8080` | TCP address to listen on |
|
||||
| `TERDUT_DB_DSN` | — | **Required.** Postgres connection string, e.g. `postgres://terdut:secret@localhost:5432/terdut?sslmode=require` |
|
||||
| `TERDUT_ARCHIVE_AFTER` | `168h` (7d) | How long a resolved alert or incident stays in the default list before being auto-archived |
|
||||
| `TERDUT_STALE_AFTER` | `6h` | How long a firing alert may go without a refreshing webhook before it is treated as resolved — **must exceed your Alertmanager `repeat_interval`** |
|
||||
| `TERDUT_ARCHIVE_AFTER` | `168h` (7d) | **seed.** How long a resolved alert or incident stays in the default list before being auto-archived |
|
||||
| `TERDUT_STALE_AFTER` | `6h` | **seed.** How long a firing alert may go without a refreshing webhook before it is treated as resolved — **must exceed your Alertmanager `repeat_interval`** |
|
||||
| `TERDUT_DEADMAN_MATCHERS` | `alertname=Watchdog` | The **default** matchers a team starts with — switches are per team now, and this seeds teams that have no configuration of their own. `;` separates matchers, `,` the label conditions within one, `=` is exact equality. Every matcher must name an `alertname` |
|
||||
| `TERDUT_DEADMAN_TIMEOUT` | `15m` | How long a heartbeat may go unheard before its switch is declared dead — **must be shorter than the `repeat_interval` of the route carrying it**. `0` disables dead man's switch handling |
|
||||
| `TERDUT_DEADMAN_SEVERITY` | `critical` | Severity a dead man's switch incident opens at |
|
||||
@@ -168,7 +181,7 @@ somewhere to exec. The sidecar, the PVC and the `backupSidecar` values are all g
|
||||
| `TERDUT_NTFY_TOKEN` | — | Bearer token for an access-controlled ntfy |
|
||||
| `TERDUT_NTFY_FALLBACK_TOPIC` | — | Topic used when nobody is on call |
|
||||
| `TERDUT_PUBLIC_URL` | — | Base URL a phone uses to reach this server: the notification's link into the web UI, its Acknowledge button, and whether the session cookie is `Secure` |
|
||||
| `TERDUT_NOTIFY_REPEAT` | `15m` | How long an incident may sit unacknowledged before it is paged again. `0` notifies once and never repeats |
|
||||
| `TERDUT_NOTIFY_REPEAT` | `15m` | **seed.** How long an incident may sit unacknowledged before it is paged again. `0` notifies once and never repeats |
|
||||
|
||||
Durations use Go syntax (`30m`, `12h`, `168h`). An unparseable value falls back to the default.
|
||||
|
||||
@@ -364,6 +377,50 @@ exhausts its retries. Written from the result rather than at enqueue, so the
|
||||
timeline says what actually happened — and a page that never landed is visible
|
||||
instead of looking the same as one that did.
|
||||
|
||||
### Escalation
|
||||
|
||||
Without a ladder, an unacknowledged incident re-pages the same topic every
|
||||
`notify_repeat` forever. That is a louder version of the same silence: if the
|
||||
person on call is asleep, out of signal, or has left, nothing else happens.
|
||||
|
||||
A team can configure an ordered ladder instead. Each level has a timeout and a
|
||||
set of targets, and a target is either a named person or **whoever the team's
|
||||
rota says is on call today** — the target that keeps working when the rota
|
||||
changes and nobody remembers to edit the policy.
|
||||
|
||||
```
|
||||
level 1 5m oncall the rota gets first refusal
|
||||
level 2 5m user:bob then a named second
|
||||
then repeat_count more rounds
|
||||
then the team's fallback topic, once
|
||||
```
|
||||
|
||||
When a level's timeout passes with the incident still `triggered`, the next
|
||||
level is paged. Off the end of the ladder the whole thing runs again
|
||||
`repeat_count` times, and after that the team's `fallback_topic` is paged once
|
||||
as the end of the line. The incident stays open throughout: running out of
|
||||
people to wake is not the same as somebody answering.
|
||||
|
||||
**Acknowledging or resolving stops it**, which is the point — continuing to wake
|
||||
people after somebody has said "I have this" is how a tool teaches people to
|
||||
mute it. **Snoozing pauses it**: a deliberate "not now" holds the ladder where
|
||||
it is, and it resumes when the snooze runs out.
|
||||
|
||||
Every step is on the incident's timeline with the level and the names it woke,
|
||||
so somebody reading it afterwards can tell why their phone rang at 04:00. A
|
||||
level whose targets are all unreachable — no ntfy topic, a disabled account, an
|
||||
empty rota — is recorded as `nobody reachable` and the ladder moves on rather
|
||||
than stalling on a rung that cannot ring.
|
||||
|
||||
**Reminders and escalation never both run.** A team with a ladder gets
|
||||
escalation; a team without keeps the reminder behaviour exactly as it was. Two
|
||||
pages for one silence is the surest way to get a tool muted.
|
||||
|
||||
The ladder's `fallback_topic` is per team, unlike `TERDUT_NTFY_FALLBACK_TOPIC`,
|
||||
which is the install-wide topic used when an incident opens with nobody on call.
|
||||
They answer different questions: one is "nobody was scheduled", the other is
|
||||
"everybody scheduled has been tried".
|
||||
|
||||
### Stale alert expiry
|
||||
|
||||
A resolved webhook is the only signal that an alert has stopped firing, so a
|
||||
@@ -522,11 +579,20 @@ on anybody's.
|
||||
| `POST` | `/api/users` | **admin** | Create user `{"username","email"}`. Not an administrator |
|
||||
| `DELETE` | `/api/users/{id}` | **admin** | Delete user (cascades to keys). `409` for yourself or the last administrator |
|
||||
| `PUT` | `/api/users/{id}/admin` | **admin** | Grant or revoke the administrator flag `{"is_admin"}`. `409` for yourself or the last administrator |
|
||||
| `PUT` | `/api/users/{id}/disabled` | **admin** | Take an account out of use, or put it back `{"disabled"}`. `409` for yourself or the last administrator |
|
||||
| `PUT` | `/api/users/{id}/notify` | self or admin | Set push notification target `{"ntfy_topic"}` — empty string clears it |
|
||||
| `PUT` | `/api/users/{id}/password` | self or admin | Set web UI password `{"password","current_password"}`. `current_password` is required only when changing your own existing password. Ends the user's other sessions |
|
||||
| `POST` | `/api/users/{id}/api-keys` | self or admin | Issue API key `{"name"}` — key shown once |
|
||||
| `DELETE` | `/api/users/{id}/api-keys/{keyID}` | self or admin | Revoke API key |
|
||||
|
||||
### Administration
|
||||
|
||||
| Method | Path | Who | Description |
|
||||
|---|---|---|---|
|
||||
| `GET` | `/api/admin/teams` | **admin** | Every team on the server, with its member and open-incident counts. `/api/teams` answers "what am I in"; this answers "what is there" |
|
||||
| `GET` | `/api/admin/settings` | **admin** | The editable settings with their bounds, plus the environment-configured ones, read-only. Never credentials |
|
||||
| `PUT` | `/api/admin/settings` | **admin** | Change one or more `{"key": seconds}`. `400` for an unknown key or a value outside its bounds |
|
||||
|
||||
### Alert ingestion
|
||||
|
||||
Alerts arrive on a team's integration key. The key is both the credential and the
|
||||
@@ -548,6 +614,7 @@ and was removed in v0.13.0 once senders had moved onto keys.
|
||||
|---|---|---|---|
|
||||
| `GET` | `/api/teams` | any | The caller's own teams, each with their role |
|
||||
| `POST` | `/api/teams` | any | Create a team `{"name"}`; the creator becomes its first owner |
|
||||
| `PUT` | `/api/teams/{teamID}` | **owner** | Rename it `{"name"}`. `409` if the name is taken |
|
||||
| `DELETE` | `/api/teams/{teamID}` | **owner** | Delete a team and everything under it. `409` while it has open incidents |
|
||||
| `GET` | `/api/teams/{teamID}/members` | member | Who is in the team |
|
||||
| `POST` | `/api/teams/{teamID}/members` | **owner** | Add a member, or change their role `{"user_id","role"}` |
|
||||
@@ -555,6 +622,8 @@ and was removed in v0.13.0 once senders had moved onto keys.
|
||||
| `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys |
|
||||
| `POST` | `/api/teams/{teamID}/integrations` | **owner** | Mint an integration `{"name","kind"}` — key and URL shown once |
|
||||
| `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration |
|
||||
| `GET` | `/api/teams/{teamID}/escalation` | member | The team's [escalation ladder](#escalation) `{repeat_count, fallback_topic, levels[]}`. Empty levels means the team has none |
|
||||
| `PUT` | `/api/teams/{teamID}/escalation` | **owner** | Replace it wholesale. `400` for a level with no targets or no timeout — a rung that pages nobody is a silence with a number on it |
|
||||
| `GET` | `/api/teams/{teamID}/deadman` | member | The team's [dead man's switch](#dead-mans-switch) configuration `{matchers, timeout_seconds, severity}` |
|
||||
| `PUT` | `/api/teams/{teamID}/deadman` | **owner** | Replace it. `400` when no matcher names an `alertname`, because a switch that silently watches nothing is the failure this feature exists to prevent |
|
||||
|
||||
|
||||
@@ -15,5 +15,5 @@ type: application
|
||||
# appVersion and image.tag in values.yaml no longer agree, and that is not an oversight:
|
||||
# image.tag stays "latest", which is what a local install actually pulls. appVersion is
|
||||
# metadata and drives nothing.
|
||||
version: 0.12.0
|
||||
appVersion: "v0.12.0"
|
||||
version: 0.13.0
|
||||
appVersion: "v0.13.0"
|
||||
|
||||
+7
-1
@@ -45,7 +45,13 @@ func main() {
|
||||
log.Fatalf("seed dead man's switch defaults: %v", err)
|
||||
}
|
||||
|
||||
router := api.NewRouter(database, notify)
|
||||
// The behaviour knobs move into the database on first start, after which an
|
||||
// administrator owns them and a redeploy leaves them alone.
|
||||
if err := api.SeedSettings(context.Background(), database, cfg); err != nil {
|
||||
log.Fatalf("seed settings: %v", err)
|
||||
}
|
||||
|
||||
router := api.NewRouter(database, notify, cfg)
|
||||
|
||||
srv := &http.Server{
|
||||
Addr: cfg.Addr,
|
||||
|
||||
@@ -403,12 +403,18 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID in
|
||||
}
|
||||
}
|
||||
|
||||
// Queue the page, but do not send it here: this runs inside a transaction on
|
||||
// a single-connection pool, so an HTTP call would hold up every other
|
||||
// request. The notifier picks the row up within a tick.
|
||||
// Queue the page, but do not send it here: this runs inside the webhook's
|
||||
// transaction, and an HTTP call would hold a connection open across a
|
||||
// network round trip. The notifier picks the row up within a tick.
|
||||
if err := enqueueOpened(ctx, q, notify, id, onCall); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
|
||||
// And start the escalation clock, if the team keeps one. In the same
|
||||
// transaction, so an incident is never briefly open with nobody counting.
|
||||
if err := startEscalation(ctx, q, id, teamID); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
return id, nil
|
||||
}
|
||||
|
||||
|
||||
@@ -52,7 +52,7 @@ func newDeadmanTS(t *testing.T, deadman api.DeadmanConfig, notify ...api.NotifyC
|
||||
}
|
||||
|
||||
database := newTestDB(t)
|
||||
srv := httptest.NewServer(api.NewRouter(database, cfg))
|
||||
srv := httptest.NewServer(api.NewRouter(database, cfg, testConfig()))
|
||||
t.Cleanup(srv.Close)
|
||||
|
||||
body, _ := json.Marshal(map[string]string{"username": "admin", "email": "admin@test.com"})
|
||||
|
||||
@@ -19,6 +19,10 @@ const (
|
||||
|
||||
// StartArchiver runs the alert sweeper until ctx is cancelled, starting with an
|
||||
// immediate pass so a restart reconciles state right away.
|
||||
// archiveAfter and staleAfter are the values the server started with. They are
|
||||
// the fallback, not the setting: each pass reads the current value from the
|
||||
// settings table, so an administrator's change takes effect on the next tick
|
||||
// instead of at the next restart.
|
||||
func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, notify NotifyConfig) {
|
||||
ticker := time.NewTicker(sweepInterval)
|
||||
defer ticker.Stop()
|
||||
@@ -45,6 +49,10 @@ func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter tim
|
||||
// staleness rules would otherwise resolve it as 'expiry' long before that.
|
||||
// Exported so tests can drive a pass without waiting on the ticker.
|
||||
func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, notify NotifyConfig) {
|
||||
settings := NewSettings(db)
|
||||
staleAfter = settings.Duration(ctx, SettingStaleAfter, staleAfter)
|
||||
archiveAfter = settings.Duration(ctx, SettingArchiveAfter, archiveAfter)
|
||||
|
||||
heartbeats := sweepDeadman(ctx, db, notify)
|
||||
expireStale(ctx, db, staleAfter, heartbeats)
|
||||
resolveSettledIncidents(ctx, db)
|
||||
|
||||
@@ -301,7 +301,7 @@ func TestSetPassword_EndsOtherSessionsButNotThisOne(t *testing.T) {
|
||||
|
||||
func TestBootstrap_WithPassword(t *testing.T) {
|
||||
database := newTestDB(t)
|
||||
srv := httptest.NewServer(api.NewRouter(database, api.NotifyConfig{}))
|
||||
srv := httptest.NewServer(api.NewRouter(database, api.NotifyConfig{}, testConfig()))
|
||||
t.Cleanup(srv.Close)
|
||||
|
||||
body := `{"username":"admin","email":"a@test.com","password":"` + adminPassword + `"}`
|
||||
|
||||
@@ -0,0 +1,506 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"log"
|
||||
"net/http"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// evEscalated records a rung of the ladder on the incident's timeline: which
|
||||
// level, and who it woke.
|
||||
const evEscalated = "escalated"
|
||||
|
||||
// escalationPolicy is a team's ladder, loaded whole. It is small — a handful of
|
||||
// levels with a few targets each — and every use needs all of it, so there is
|
||||
// no point reading it a level at a time.
|
||||
type escalationPolicy struct {
|
||||
teamID int64
|
||||
repeatCount int64
|
||||
fallbackTopic string
|
||||
levels []escalationLevel
|
||||
}
|
||||
|
||||
type escalationLevel struct {
|
||||
id int64
|
||||
position int64
|
||||
timeout time.Duration
|
||||
targets []escalationTarget
|
||||
}
|
||||
|
||||
type escalationTarget struct {
|
||||
kind string // "user" or "oncall"
|
||||
userID *int64
|
||||
}
|
||||
|
||||
// configured reports whether this team has anything to escalate through. A
|
||||
// policy row with no levels is the same as no policy: the team gets the
|
||||
// pre-escalation behaviour, which is reminders on the assignee's topic.
|
||||
func (p *escalationPolicy) configured() bool { return p != nil && len(p.levels) > 0 }
|
||||
|
||||
// level returns the level at a 1-based position.
|
||||
func (p *escalationPolicy) level(pos int64) (escalationLevel, bool) {
|
||||
for _, l := range p.levels {
|
||||
if l.position == pos {
|
||||
return l, true
|
||||
}
|
||||
}
|
||||
return escalationLevel{}, false
|
||||
}
|
||||
|
||||
// loadEscalationPolicy reads one team's ladder. A team with no policy row
|
||||
// returns nil, which every caller treats as "not configured" rather than as an
|
||||
// error: most teams will never set one up.
|
||||
func loadEscalationPolicy(ctx context.Context, q querier, teamID int64) (*escalationPolicy, error) {
|
||||
p := &escalationPolicy{teamID: teamID}
|
||||
err := q.QueryRowContext(ctx,
|
||||
"SELECT repeat_count, fallback_topic FROM escalation_policies WHERE team_id = $1",
|
||||
teamID).Scan(&p.repeatCount, &p.fallbackTopic)
|
||||
if err == sql.ErrNoRows {
|
||||
return nil, nil
|
||||
}
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
rows, err := q.QueryContext(ctx, `
|
||||
SELECT l.id, l.position, l.timeout_seconds, t.kind, t.user_id
|
||||
FROM escalation_levels l
|
||||
LEFT JOIN escalation_targets t ON t.level_id = l.id
|
||||
WHERE l.team_id = $1
|
||||
ORDER BY l.position, t.id`, teamID)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
byPosition := map[int64]int{} // position -> index in p.levels
|
||||
for rows.Next() {
|
||||
var id, position, timeout int64
|
||||
var kind *string
|
||||
var userID *int64
|
||||
if err := rows.Scan(&id, &position, &timeout, &kind, &userID); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
idx, seen := byPosition[position]
|
||||
if !seen {
|
||||
p.levels = append(p.levels, escalationLevel{
|
||||
id: id,
|
||||
position: position,
|
||||
timeout: time.Duration(timeout) * time.Second,
|
||||
})
|
||||
idx = len(p.levels) - 1
|
||||
byPosition[position] = idx
|
||||
}
|
||||
// LEFT JOIN: a level with no targets yet still produces a row, with a
|
||||
// NULL kind. It is a rung that pages nobody, which the API refuses to
|
||||
// store but an older row could still hold.
|
||||
if kind != nil {
|
||||
p.levels[idx].targets = append(p.levels[idx].targets,
|
||||
escalationTarget{kind: *kind, userID: userID})
|
||||
}
|
||||
}
|
||||
return p, rows.Err()
|
||||
}
|
||||
|
||||
// escalate advances every incident whose current level has run out of time.
|
||||
//
|
||||
// Runs on the notifier's tick, beside the reminder pass, because it is the same
|
||||
// question asked differently: reminders ask "has this been ignored long
|
||||
// enough to say it again", escalation asks "long enough to say it to somebody
|
||||
// else". Sharing the tick means one query cadence and one outbox.
|
||||
func escalate(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
|
||||
rows, err := db.QueryContext(ctx, `
|
||||
SELECT i.id, i.team_id, i.escalation_level, i.escalation_level_at, i.escalation_round
|
||||
FROM incidents i
|
||||
JOIN escalation_policies p ON p.team_id = i.team_id
|
||||
WHERE i.resolved_at IS NULL
|
||||
AND i.archived_at IS NULL
|
||||
AND i.status = 'triggered'
|
||||
AND (i.snoozed_until IS NULL OR i.snoozed_until <= $1)
|
||||
AND i.escalation_level > 0`, time.Now().Unix())
|
||||
if err != nil {
|
||||
log.Printf("escalation: find due: %v", err)
|
||||
return
|
||||
}
|
||||
|
||||
type pending struct {
|
||||
incidentID, teamID, level, round int64
|
||||
levelAt int64
|
||||
}
|
||||
var due []pending
|
||||
for rows.Next() {
|
||||
var p pending
|
||||
var levelAt *int64
|
||||
if err := rows.Scan(&p.incidentID, &p.teamID, &p.level, &levelAt, &p.round); err != nil {
|
||||
rows.Close()
|
||||
log.Printf("escalation: scan: %v", err)
|
||||
return
|
||||
}
|
||||
if levelAt == nil {
|
||||
continue
|
||||
}
|
||||
p.levelAt = *levelAt
|
||||
due = append(due, p)
|
||||
}
|
||||
rows.Close()
|
||||
if err := rows.Err(); err != nil {
|
||||
log.Printf("escalation: iterate: %v", err)
|
||||
return
|
||||
}
|
||||
|
||||
now := time.Now()
|
||||
for _, d := range due {
|
||||
policy, err := loadEscalationPolicy(ctx, db, d.teamID)
|
||||
if err != nil {
|
||||
log.Printf("escalation: load policy for team %d: %v", d.teamID, err)
|
||||
continue
|
||||
}
|
||||
if !policy.configured() {
|
||||
continue
|
||||
}
|
||||
current, ok := policy.level(d.level)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if now.Sub(time.Unix(d.levelAt, 0)) < current.timeout {
|
||||
continue
|
||||
}
|
||||
if err := advanceEscalation(ctx, db, cfg, policy, d.incidentID, d.level, d.round, now); err != nil {
|
||||
log.Printf("escalation: advance incident %d: %v", d.incidentID, err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// advanceEscalation moves one incident to its next rung, or off the end of the
|
||||
// ladder.
|
||||
//
|
||||
// The whole move is one transaction: the level, the page and the timeline entry
|
||||
// are one event, and an incident recorded as being at level 3 that nobody at
|
||||
// level 3 was told about is the worst of the possible half-states.
|
||||
func advanceEscalation(ctx context.Context, db *sql.DB, cfg NotifyConfig, policy *escalationPolicy, incidentID, level, round int64, now time.Time) error {
|
||||
tx, err := db.BeginTx(ctx, nil)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer tx.Rollback() //nolint:errcheck
|
||||
|
||||
next := level + 1
|
||||
nextRound := round
|
||||
if _, ok := policy.level(next); !ok {
|
||||
// Off the end. Either start the chain again, or make the last call.
|
||||
if round < policy.repeatCount {
|
||||
next, nextRound = 1, round+1
|
||||
} else {
|
||||
if err := escalationExhausted(ctx, tx, policy, incidentID, now); err != nil {
|
||||
return err
|
||||
}
|
||||
return tx.Commit()
|
||||
}
|
||||
}
|
||||
|
||||
target, ok := policy.level(next)
|
||||
if !ok {
|
||||
return nil
|
||||
}
|
||||
paged, err := pageLevel(ctx, tx, cfg, policy, incidentID, target)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
if _, err := tx.ExecContext(ctx, `
|
||||
UPDATE incidents
|
||||
SET escalation_level = $1, escalation_level_at = $2, escalation_round = $3
|
||||
WHERE id = $4`, next, now.Unix(), nextRound, incidentID); err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
detail := "level " + strconv.FormatInt(next, 10)
|
||||
if nextRound > round {
|
||||
detail += " (round " + strconv.FormatInt(nextRound+1, 10) + ")"
|
||||
}
|
||||
if len(paged) > 0 {
|
||||
detail += ": " + strings.Join(paged, ", ")
|
||||
} else {
|
||||
// Worth recording loudly: the rung exists, its turn came, and it woke
|
||||
// nobody. That is a policy that looks configured and is not.
|
||||
detail += ": nobody reachable"
|
||||
}
|
||||
if err := logEvent(ctx, tx, incidentID, evEscalated, nil, nil, &detail); err != nil {
|
||||
return err
|
||||
}
|
||||
return tx.Commit()
|
||||
}
|
||||
|
||||
// escalationExhausted is the end of the line: the fallback topic, once, and a
|
||||
// timeline entry saying the ladder is finished. The incident stays triggered —
|
||||
// escalation running out is not the same as somebody answering.
|
||||
func escalationExhausted(ctx context.Context, tx *sql.Tx, policy *escalationPolicy, incidentID int64, now time.Time) error {
|
||||
detail := "escalation exhausted"
|
||||
if policy.fallbackTopic != "" {
|
||||
if err := enqueueNotification(ctx, tx, incidentID, nil, policy.fallbackTopic, notifyEscalated); err != nil {
|
||||
return err
|
||||
}
|
||||
detail += ": paged " + policy.fallbackTopic
|
||||
} else {
|
||||
detail += ": no fallback topic configured"
|
||||
}
|
||||
|
||||
// Level 0 again, so the sweep stops considering it. The round counter is
|
||||
// left where it is, as the record of how far it got.
|
||||
if _, err := tx.ExecContext(ctx,
|
||||
"UPDATE incidents SET escalation_level = 0, escalation_level_at = NULL WHERE id = $1",
|
||||
incidentID); err != nil {
|
||||
return err
|
||||
}
|
||||
return logEvent(ctx, tx, incidentID, evEscalated, nil, nil, &detail)
|
||||
}
|
||||
|
||||
// pageLevel notifies every target of one level and reports who was woken.
|
||||
//
|
||||
// Each target gets its own outbox row, so each gets its own Acknowledge token:
|
||||
// the button in a notification must acknowledge as the person holding the
|
||||
// phone, not as whoever was paged first.
|
||||
func pageLevel(ctx context.Context, tx *sql.Tx, cfg NotifyConfig, policy *escalationPolicy, incidentID int64, level escalationLevel) ([]string, error) {
|
||||
var paged []string
|
||||
seen := map[int64]bool{}
|
||||
|
||||
for _, t := range level.targets {
|
||||
userID := t.userID
|
||||
if t.kind == "oncall" {
|
||||
onCall, err := currentOnCall(ctx, tx, policy.teamID)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if onCall == nil {
|
||||
continue
|
||||
}
|
||||
userID = onCall
|
||||
}
|
||||
if userID == nil || seen[*userID] {
|
||||
continue
|
||||
}
|
||||
seen[*userID] = true
|
||||
|
||||
var topic *string
|
||||
var username string
|
||||
if err := tx.QueryRowContext(ctx,
|
||||
"SELECT ntfy_topic, username FROM users WHERE id = $1 AND disabled_at IS NULL",
|
||||
*userID).Scan(&topic, &username); err != nil {
|
||||
// A disabled or deleted account is not an error in the middle of an
|
||||
// escalation: it is a target that cannot be woken, and the next
|
||||
// level is the answer to that.
|
||||
continue
|
||||
}
|
||||
if topic == nil || *topic == "" {
|
||||
continue
|
||||
}
|
||||
if err := enqueueNotification(ctx, tx, incidentID, userID, *topic, notifyEscalated); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
paged = append(paged, username)
|
||||
}
|
||||
return paged, nil
|
||||
}
|
||||
|
||||
// startEscalation puts a newly opened incident on the first rung, when its team
|
||||
// has a ladder. Called from openIncident, inside the same transaction, so an
|
||||
// incident is never briefly open with no escalation clock running.
|
||||
func startEscalation(ctx context.Context, q querier, incidentID, teamID int64) error {
|
||||
policy, err := loadEscalationPolicy(ctx, q, teamID)
|
||||
if err != nil || !policy.configured() {
|
||||
return err
|
||||
}
|
||||
_, err = q.ExecContext(ctx,
|
||||
"UPDATE incidents SET escalation_level = 1, escalation_level_at = $1 WHERE id = $2",
|
||||
time.Now().Unix(), incidentID)
|
||||
return err
|
||||
}
|
||||
|
||||
// stopEscalation takes an incident off the ladder. Acknowledging or resolving
|
||||
// is somebody saying "I have this", and continuing to wake people after that is
|
||||
// the behaviour that teaches people to ignore the tool.
|
||||
func stopEscalation(ctx context.Context, q querier, incidentID int64) error {
|
||||
_, err := q.ExecContext(ctx,
|
||||
"UPDATE incidents SET escalation_level = 0, escalation_level_at = NULL WHERE id = $1",
|
||||
incidentID)
|
||||
return err
|
||||
}
|
||||
|
||||
// handleGetEscalation returns a team's ladder.
|
||||
func handleGetEscalation(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamMember(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
policy, err := loadEscalationPolicy(r.Context(), db, teamID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, escalationResponse(policy, teamID))
|
||||
}
|
||||
}
|
||||
|
||||
type escalationLevelJSON struct {
|
||||
Position int64 `json:"position"`
|
||||
TimeoutSeconds int64 `json:"timeout_seconds"`
|
||||
Targets []escalationTargetJSON `json:"targets"`
|
||||
}
|
||||
|
||||
type escalationTargetJSON struct {
|
||||
Kind string `json:"kind"`
|
||||
UserID *int64 `json:"user_id,omitempty"`
|
||||
}
|
||||
|
||||
type escalationJSON struct {
|
||||
TeamID int64 `json:"team_id"`
|
||||
RepeatCount int64 `json:"repeat_count"`
|
||||
FallbackTopic string `json:"fallback_topic"`
|
||||
Levels []escalationLevelJSON `json:"levels"`
|
||||
}
|
||||
|
||||
func escalationResponse(p *escalationPolicy, teamID int64) escalationJSON {
|
||||
out := escalationJSON{TeamID: teamID, Levels: []escalationLevelJSON{}}
|
||||
if p == nil {
|
||||
return out
|
||||
}
|
||||
out.RepeatCount = p.repeatCount
|
||||
out.FallbackTopic = p.fallbackTopic
|
||||
for _, l := range p.levels {
|
||||
level := escalationLevelJSON{
|
||||
Position: l.position,
|
||||
TimeoutSeconds: int64(l.timeout.Seconds()),
|
||||
Targets: []escalationTargetJSON{},
|
||||
}
|
||||
for _, t := range l.targets {
|
||||
level.Targets = append(level.Targets, escalationTargetJSON{Kind: t.kind, UserID: t.userID})
|
||||
}
|
||||
out.Levels = append(out.Levels, level)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// handleSetEscalation replaces a team's ladder wholesale.
|
||||
//
|
||||
// Replace rather than patch: the levels are an order, and an API that edits one
|
||||
// rung has to answer what happens to the numbering of the others. Sending the
|
||||
// whole ladder makes the order the client's to decide and the server's to
|
||||
// store, and makes an edit atomic — there is no moment where level 2 exists
|
||||
// twice.
|
||||
func handleSetEscalation(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
var req escalationJSON
|
||||
if err := decodeJSON(r, &req); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid request body"))
|
||||
return
|
||||
}
|
||||
if req.RepeatCount < 0 || req.RepeatCount > 10 {
|
||||
respond(w, http.StatusBadRequest, errResp("repeat_count must be between 0 and 10"))
|
||||
return
|
||||
}
|
||||
for i, l := range req.Levels {
|
||||
if l.TimeoutSeconds <= 0 {
|
||||
respond(w, http.StatusBadRequest, errResp("every level needs a timeout"))
|
||||
return
|
||||
}
|
||||
if len(l.Targets) == 0 {
|
||||
// A rung that pages nobody is not a delay, it is a silence with
|
||||
// a number on it.
|
||||
respond(w, http.StatusBadRequest,
|
||||
errResp("level "+strconv.FormatInt(int64(i+1), 10)+" has no targets"))
|
||||
return
|
||||
}
|
||||
for _, t := range l.Targets {
|
||||
switch t.Kind {
|
||||
case "oncall":
|
||||
if t.UserID != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("an oncall target takes no user_id"))
|
||||
return
|
||||
}
|
||||
case "user":
|
||||
if t.UserID == nil {
|
||||
respond(w, http.StatusBadRequest, errResp("a user target needs a user_id"))
|
||||
return
|
||||
}
|
||||
default:
|
||||
respond(w, http.StatusBadRequest, errResp("target kind must be user or oncall"))
|
||||
return
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
tx, err := db.BeginTx(r.Context(), nil)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer tx.Rollback() //nolint:errcheck
|
||||
|
||||
if _, err := tx.ExecContext(r.Context(), `
|
||||
INSERT INTO escalation_policies (team_id, repeat_count, fallback_topic, updated_at)
|
||||
VALUES ($1, $2, $3, `+nowEpoch+`)
|
||||
ON CONFLICT (team_id) DO UPDATE SET
|
||||
repeat_count = excluded.repeat_count,
|
||||
fallback_topic = excluded.fallback_topic,
|
||||
updated_at = excluded.updated_at`,
|
||||
teamID, req.RepeatCount, req.FallbackTopic); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
// The levels are replaced, not merged; the cascade takes the targets.
|
||||
if _, err := tx.ExecContext(r.Context(),
|
||||
"DELETE FROM escalation_levels WHERE team_id = $1", teamID); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
|
||||
for i, l := range req.Levels {
|
||||
var levelID int64
|
||||
if err := tx.QueryRowContext(r.Context(), `
|
||||
INSERT INTO escalation_levels (team_id, position, timeout_seconds)
|
||||
VALUES ($1, $2, $3) RETURNING id`,
|
||||
teamID, int64(i+1), l.TimeoutSeconds).Scan(&levelID); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
for _, t := range l.Targets {
|
||||
if _, err := tx.ExecContext(r.Context(), `
|
||||
INSERT INTO escalation_targets (level_id, kind, user_id)
|
||||
VALUES ($1, $2, $3)`, levelID, t.Kind, t.UserID); err != nil {
|
||||
// The only foreign key here is the user.
|
||||
respond(w, http.StatusBadRequest, errResp("unknown user in targets"))
|
||||
return
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if err := tx.Commit(); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
|
||||
policy, err := loadEscalationPolicy(r.Context(), db, teamID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, escalationResponse(policy, teamID))
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,385 @@
|
||||
package api_test
|
||||
|
||||
import (
|
||||
"net/http"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"git.ryuvia.com/niklas/terdut-server/internal/api"
|
||||
)
|
||||
|
||||
// teamUser creates a user in the default team with an ntfy topic, so they can
|
||||
// actually be paged.
|
||||
func teamUser(t *testing.T, s *ts, username, topic string) int64 {
|
||||
t.Helper()
|
||||
var user struct {
|
||||
ID int64 `json:"id"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]string{"username": username, "email": username + "@test.com"}), &user)
|
||||
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/members",
|
||||
map[string]any{"user_id": user.ID, "role": "member"})
|
||||
resp.Body.Close()
|
||||
setTopic(t, s, int(user.ID), topic)
|
||||
return user.ID
|
||||
}
|
||||
|
||||
// Escalation is all timeouts, and there is no fake clock in this package. The
|
||||
// tests back-date escalation_level_at instead, which is the same trick the dead
|
||||
// man's switch tests use on received_at: the sweeper reads a stored timestamp,
|
||||
// so moving the timestamp is moving the clock.
|
||||
|
||||
// ladder configures the default team with two levels: the rota first, then a
|
||||
// named person, then the fallback topic.
|
||||
func ladder(t *testing.T, s *ts, secondUserID int64, repeat int64, fallback string) {
|
||||
t.Helper()
|
||||
resp := s.req(t, http.MethodPut, "/api/teams/"+defaultTeam+"/escalation", map[string]any{
|
||||
"repeat_count": repeat,
|
||||
"fallback_topic": fallback,
|
||||
"levels": []map[string]any{
|
||||
{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "oncall"}}},
|
||||
{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "user", "user_id": secondUserID}}},
|
||||
},
|
||||
})
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("configure the ladder: %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// overdue back-dates an incident's current level so its timeout has passed.
|
||||
func overdue(t *testing.T, s *ts, incidentID int64) {
|
||||
t.Helper()
|
||||
s.exec(t, "UPDATE incidents SET escalation_level_at = $1 WHERE id = $2",
|
||||
time.Now().Add(-time.Hour).Unix(), incidentID)
|
||||
}
|
||||
|
||||
func escalationLevel(t *testing.T, s *ts, incidentID int64) (level, round int64) {
|
||||
t.Helper()
|
||||
if err := s.db.QueryRow(
|
||||
"SELECT escalation_level, escalation_round FROM incidents WHERE id = $1",
|
||||
incidentID).Scan(&level, &round); err != nil {
|
||||
t.Fatalf("read escalation state: %v", err)
|
||||
}
|
||||
return level, round
|
||||
}
|
||||
|
||||
// The whole point: nobody answers, so somebody else is woken.
|
||||
func TestEscalation_PagesTheNextLevel(t *testing.T) {
|
||||
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
|
||||
second := teamUser(t, s, "second", "terdut-second")
|
||||
ladder(t, s, second, 0, "terdut-fallback")
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-esc", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.sweepNotify(t)
|
||||
|
||||
// Level 1 is the rota, so the first page went to the admin.
|
||||
if level, _ := escalationLevel(t, s, 1); level != 1 {
|
||||
t.Fatalf("a new incident should start at level 1, got %d", level)
|
||||
}
|
||||
if got := f.topicsSince(t); len(got) == 0 || got[0] != "terdut-admin" {
|
||||
t.Fatalf("the first page should go to the on-call user, went to %v", got)
|
||||
}
|
||||
|
||||
// Time passes with no acknowledgement.
|
||||
f.forget()
|
||||
overdue(t, s, 1)
|
||||
s.sweepNotify(t)
|
||||
|
||||
if level, _ := escalationLevel(t, s, 1); level != 2 {
|
||||
t.Errorf("expected level 2, got %d", level)
|
||||
}
|
||||
if got := f.topicsSince(t); len(got) != 1 || got[0] != "terdut-second" {
|
||||
t.Errorf("level 2 should page the named user, paged %v", got)
|
||||
}
|
||||
|
||||
// And the timeline says so, which is what somebody reads afterwards to
|
||||
// understand why their phone rang at 04:00.
|
||||
timeline := list(t, s.req(t, http.MethodGet, "/api/incidents/1/timeline", nil))
|
||||
found := ""
|
||||
for _, e := range timeline {
|
||||
if e["type"] == "escalated" {
|
||||
found, _ = e["detail"].(string)
|
||||
}
|
||||
}
|
||||
if found == "" {
|
||||
t.Error("the timeline should record the escalation")
|
||||
} else if !strings.HasPrefix(found, "level 2") || !strings.Contains(found, "second") {
|
||||
t.Errorf("the escalation entry should say which level and who: %q", found)
|
||||
}
|
||||
}
|
||||
|
||||
// Acknowledging is somebody saying "I have this". Nobody else should be woken.
|
||||
func TestEscalation_AcknowledgementStopsIt(t *testing.T) {
|
||||
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
|
||||
second := teamUser(t, s, "second", "terdut-second")
|
||||
ladder(t, s, second, 0, "terdut-fallback")
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-ack", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.sweepNotify(t)
|
||||
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil).Body.Close()
|
||||
if level, _ := escalationLevel(t, s, 1); level != 0 {
|
||||
t.Errorf("acknowledging should take the incident off the ladder, level is %d", level)
|
||||
}
|
||||
|
||||
f.forget()
|
||||
overdue(t, s, 1) // no-op: level is 0, so there is nothing due
|
||||
s.sweepNotify(t)
|
||||
if got := f.topicsSince(t); len(got) != 0 {
|
||||
t.Errorf("an acknowledged incident should page nobody, paged %v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// Resolving stops it too, and by the same mechanism.
|
||||
func TestEscalation_ResolutionStopsIt(t *testing.T) {
|
||||
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
|
||||
second := teamUser(t, s, "second", "terdut-second")
|
||||
ladder(t, s, second, 0, "terdut-fallback")
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-res", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.sweepNotify(t)
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
|
||||
|
||||
f.forget()
|
||||
overdue(t, s, 1)
|
||||
s.sweepNotify(t)
|
||||
if level, _ := escalationLevel(t, s, 1); level != 0 {
|
||||
t.Errorf("a resolved incident should be off the ladder, level is %d", level)
|
||||
}
|
||||
}
|
||||
|
||||
// Snoozing is a deliberate "not now", so the ladder waits rather than carrying
|
||||
// on without the person who asked for quiet.
|
||||
func TestEscalation_SnoozePausesIt(t *testing.T) {
|
||||
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
|
||||
second := teamUser(t, s, "second", "terdut-second")
|
||||
ladder(t, s, second, 0, "terdut-fallback")
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-snooze", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.sweepNotify(t)
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/incidents/1/snooze", map[string]any{"duration": "1h"})
|
||||
resp.Body.Close()
|
||||
|
||||
f.forget()
|
||||
overdue(t, s, 1)
|
||||
s.sweepNotify(t)
|
||||
|
||||
if level, _ := escalationLevel(t, s, 1); level != 1 {
|
||||
t.Errorf("a snoozed incident should stay where it is, level is %d", level)
|
||||
}
|
||||
if got := f.topicsSince(t); len(got) != 0 {
|
||||
t.Errorf("a snoozed incident should page nobody, paged %v", got)
|
||||
}
|
||||
|
||||
// When the snooze ends, the ladder picks up where it left off.
|
||||
s.exec(t, "UPDATE incidents SET snoozed_until = $1 WHERE id = 1", time.Now().Add(-time.Minute).Unix())
|
||||
s.sweepNotify(t)
|
||||
if level, _ := escalationLevel(t, s, 1); level != 2 {
|
||||
t.Errorf("after the snooze the ladder should resume, level is %d", level)
|
||||
}
|
||||
}
|
||||
|
||||
// Running out of ladder pages the team's fallback topic once, and says so.
|
||||
func TestEscalation_ExhaustionPagesTheFallback(t *testing.T) {
|
||||
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
|
||||
second := teamUser(t, s, "second", "terdut-second")
|
||||
ladder(t, s, second, 0, "terdut-fallback")
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-end", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.sweepNotify(t)
|
||||
|
||||
overdue(t, s, 1)
|
||||
s.sweepNotify(t) // level 2
|
||||
f.forget()
|
||||
overdue(t, s, 1)
|
||||
s.sweepNotify(t) // off the end
|
||||
|
||||
if got := f.topicsSince(t); len(got) != 1 || got[0] != "terdut-fallback" {
|
||||
t.Errorf("exhaustion should page the fallback topic once, paged %v", got)
|
||||
}
|
||||
level, _ := escalationLevel(t, s, 1)
|
||||
if level != 0 {
|
||||
t.Errorf("an exhausted ladder should stop asking, level is %d", level)
|
||||
}
|
||||
|
||||
// The incident is still open: running out of people is not an answer.
|
||||
var status string
|
||||
if err := s.db.QueryRow("SELECT status FROM incidents WHERE id = 1").Scan(&status); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if status != "triggered" {
|
||||
t.Errorf("exhaustion must not resolve the incident, status is %q", status)
|
||||
}
|
||||
}
|
||||
|
||||
// repeat_count walks the whole ladder again before giving up.
|
||||
func TestEscalation_RepeatsTheChain(t *testing.T) {
|
||||
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
|
||||
second := teamUser(t, s, "second", "terdut-second")
|
||||
ladder(t, s, second, 1, "terdut-fallback") // one extra round
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-repeat", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.sweepNotify(t)
|
||||
|
||||
overdue(t, s, 1)
|
||||
s.sweepNotify(t) // level 2
|
||||
f.forget()
|
||||
overdue(t, s, 1)
|
||||
s.sweepNotify(t) // back to level 1, round 2
|
||||
|
||||
level, round := escalationLevel(t, s, 1)
|
||||
if level != 1 || round != 1 {
|
||||
t.Errorf("expected level 1 round 1, got level %d round %d", level, round)
|
||||
}
|
||||
if got := f.topicsSince(t); len(got) != 1 || got[0] != "terdut-admin" {
|
||||
t.Errorf("the second round should start at the top again, paged %v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// A team without a ladder keeps exactly the behaviour it had, and never gets
|
||||
// both a reminder and an escalation for the same silence.
|
||||
func TestEscalation_WithoutAPolicyRemindersStillRun(t *testing.T) {
|
||||
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-noesc", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.sweepNotify(t)
|
||||
|
||||
// Age the first notification past the repeat interval.
|
||||
f.forget()
|
||||
s.exec(t, "UPDATE notifications SET created_at = $1, sent_at = $1",
|
||||
time.Now().Add(-time.Hour).Unix())
|
||||
s.sweepNotify(t)
|
||||
|
||||
if got := f.topicsSince(t); len(got) != 1 || got[0] != "terdut-admin" {
|
||||
t.Errorf("without a ladder the reminder should still fire, paged %v", got)
|
||||
}
|
||||
if level, _ := escalationLevel(t, s, 1); level != 0 {
|
||||
t.Errorf("an incident in a team with no ladder should not be on one, level is %d", level)
|
||||
}
|
||||
}
|
||||
|
||||
// With a ladder, reminders stop: two pages for one silence is how people learn
|
||||
// to mute the tool.
|
||||
func TestEscalation_WithAPolicyRemindersDoNotAlsoFire(t *testing.T) {
|
||||
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
|
||||
second := teamUser(t, s, "second", "terdut-second")
|
||||
ladder(t, s, second, 0, "terdut-fallback")
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-both", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.sweepNotify(t)
|
||||
|
||||
f.forget()
|
||||
// Old enough for a reminder, but not yet due for escalation.
|
||||
s.exec(t, "UPDATE notifications SET created_at = $1, sent_at = $1",
|
||||
time.Now().Add(-time.Hour).Unix())
|
||||
s.sweepNotify(t)
|
||||
|
||||
if got := f.topicsSince(t); len(got) != 0 {
|
||||
t.Errorf("a team with a ladder should not also get reminders, paged %v", got)
|
||||
}
|
||||
}
|
||||
|
||||
// The API refuses a ladder that cannot page anybody.
|
||||
func TestEscalation_RejectsAnUnusablePolicy(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
body map[string]any
|
||||
}{
|
||||
{"a level with no targets", map[string]any{
|
||||
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{}}},
|
||||
}},
|
||||
{"a level with no timeout", map[string]any{
|
||||
"levels": []map[string]any{{"timeout_seconds": 0, "targets": []map[string]any{{"kind": "oncall"}}}},
|
||||
}},
|
||||
{"a user target with no user", map[string]any{
|
||||
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "user"}}}},
|
||||
}},
|
||||
{"an unknown target kind", map[string]any{
|
||||
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "everybody"}}}},
|
||||
}},
|
||||
{"an absurd repeat count", map[string]any{
|
||||
"repeat_count": 99,
|
||||
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "oncall"}}}},
|
||||
}},
|
||||
} {
|
||||
resp := s.req(t, http.MethodPut, "/api/teams/"+defaultTeam+"/escalation", c.body)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusBadRequest {
|
||||
t.Errorf("%s: expected 400, got %d", c.name, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Editing the ladder is an owner's job; reading it is any member's.
|
||||
func TestEscalation_OwnerOnlyToEdit(t *testing.T) {
|
||||
s := newTS(t)
|
||||
_, call := member(t, s, "plain")
|
||||
|
||||
resp := call(http.MethodPut, "/api/teams/"+defaultTeam+"/escalation", map[string]any{
|
||||
"levels": []map[string]any{{"timeout_seconds": 300, "targets": []map[string]any{{"kind": "oncall"}}}},
|
||||
})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusForbidden {
|
||||
t.Errorf("a member editing the ladder: expected 403, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = call(http.MethodGet, "/api/teams/"+defaultTeam+"/escalation", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Errorf("a member reading the ladder: expected 200, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// A target who cannot be woken is not a reason to stop: the next level is the
|
||||
// answer to an unreachable one.
|
||||
func TestEscalation_SkipsUnreachableTargets(t *testing.T) {
|
||||
s, f := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
|
||||
// Second user has no ntfy topic at all.
|
||||
var user struct {
|
||||
ID int64 `json:"id"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]string{"username": "silent", "email": "silent@test.com"}), &user)
|
||||
s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/members",
|
||||
map[string]any{"user_id": user.ID, "role": "member"}).Body.Close()
|
||||
|
||||
ladder(t, s, user.ID, 0, "terdut-fallback")
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-silent", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.sweepNotify(t)
|
||||
|
||||
f.forget()
|
||||
overdue(t, s, 1)
|
||||
s.sweepNotify(t)
|
||||
|
||||
// Level 2 was entered even though it woke nobody, so the ladder keeps
|
||||
// moving toward the fallback rather than stalling on a silent rung.
|
||||
if level, _ := escalationLevel(t, s, 1); level != 2 {
|
||||
t.Errorf("expected the ladder to advance past an unreachable target, level is %d", level)
|
||||
}
|
||||
if got := f.topicsSince(t); len(got) != 0 {
|
||||
t.Errorf("a target with no topic should page nothing, paged %v", got)
|
||||
}
|
||||
}
|
||||
@@ -235,6 +235,9 @@ func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, e
|
||||
if n == 0 {
|
||||
return false, nil
|
||||
}
|
||||
if err := stopEscalation(ctx, q, incidentID); err != nil {
|
||||
return false, err
|
||||
}
|
||||
if err := logEvent(ctx, q, incidentID, evResolved, nil, nil, nil); err != nil {
|
||||
return false, err
|
||||
}
|
||||
@@ -260,6 +263,10 @@ func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int6
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
return false, nil
|
||||
}
|
||||
// Somebody has it: stop waking anybody else.
|
||||
if err := stopEscalation(ctx, q, incidentID); err != nil {
|
||||
return false, err
|
||||
}
|
||||
return true, logEvent(ctx, q, incidentID, evAcknowledged, &userID, nil, nil)
|
||||
}
|
||||
|
||||
|
||||
@@ -243,6 +243,11 @@ func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
|
||||
time.Now().Unix(), incidentResolutionManual, id) {
|
||||
return
|
||||
}
|
||||
// A person closing an incident is the clearest possible "I have this".
|
||||
if err := stopEscalation(r.Context(), db, id); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if err := logEvent(r.Context(), db, id, evResolved, &user.ID, nil, nil); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
|
||||
@@ -144,8 +144,12 @@ func sessionUser(ctx context.Context, db *sql.DB, token string) (sessionID, user
|
||||
func serveAs(w http.ResponseWriter, r *http.Request, next http.Handler, db *sql.DB, userID, sessionID int64) {
|
||||
var u models.User
|
||||
var createdUnix int64
|
||||
// disabled_at IS NULL is part of the lookup rather than a check afterwards:
|
||||
// a disabled account is one that cannot authenticate, by either credential,
|
||||
// and the way to be sure of that is for there to be no path where the row
|
||||
// is loaded and the flag is then forgotten.
|
||||
if err := db.QueryRowContext(r.Context(),
|
||||
"SELECT id, username, email, created_at, is_admin FROM users WHERE id = $1", userID,
|
||||
"SELECT id, username, email, created_at, is_admin FROM users WHERE id = $1 AND disabled_at IS NULL", userID,
|
||||
).Scan(&u.ID, &u.Username, &u.Email, &createdUnix, &u.IsAdmin); err != nil {
|
||||
respond(w, http.StatusUnauthorized, errResp("unauthorized"))
|
||||
return
|
||||
|
||||
@@ -43,6 +43,10 @@ const (
|
||||
notifyTriggered = "triggered"
|
||||
notifyReminder = "reminder"
|
||||
notifyResolved = "resolved"
|
||||
|
||||
// notifyEscalated is a page that went out because nobody answered the last
|
||||
// one. Told apart from a reminder because it goes to somebody else.
|
||||
notifyEscalated = "escalated"
|
||||
)
|
||||
|
||||
// Timeline event types the notifier writes, so an incident's history says who
|
||||
@@ -120,6 +124,9 @@ func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
|
||||
// Exported so tests can drive a pass without waiting on the ticker.
|
||||
func NotifySweep(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
|
||||
enqueueReminders(ctx, db, cfg)
|
||||
// Escalation before delivery, so a level that comes due on this tick is
|
||||
// paged on this tick rather than waiting for the next one.
|
||||
escalate(ctx, db, cfg)
|
||||
deliverPending(ctx, db, cfg)
|
||||
}
|
||||
|
||||
@@ -134,7 +141,11 @@ func NotifySweep(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
|
||||
// queued, so an ntfy outage produces a retry backlog rather than a reminder
|
||||
// backlog that all lands at once when it comes back.
|
||||
func enqueueReminders(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
|
||||
if cfg.RepeatEvery <= 0 {
|
||||
// cfg.RepeatEvery is what the server started with; the settings table is
|
||||
// what it runs on. Read per tick, so an administrator lengthening the
|
||||
// interval at 02:00 is obeyed at 02:00 and not at the next restart.
|
||||
repeat := NewSettings(db).Duration(ctx, SettingNotifyRepeat, cfg.RepeatEvery)
|
||||
if repeat <= 0 {
|
||||
return
|
||||
}
|
||||
now := time.Now()
|
||||
@@ -155,8 +166,13 @@ func enqueueReminders(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
|
||||
AND i.resolved_at IS NULL
|
||||
AND i.archived_at IS NULL
|
||||
AND i.status = 'triggered'
|
||||
AND (i.snoozed_until IS NULL OR i.snoozed_until <= $2)`,
|
||||
now.Add(-cfg.RepeatEvery).Unix(), now.Unix())
|
||||
AND (i.snoozed_until IS NULL OR i.snoozed_until <= $2)
|
||||
-- A team with an escalation ladder gets escalation instead. Both
|
||||
-- would mean two pages for one silence, which is how people learn to
|
||||
-- mute a tool.
|
||||
AND NOT EXISTS (
|
||||
SELECT 1 FROM escalation_levels el WHERE el.team_id = i.team_id)`,
|
||||
now.Add(-repeat).Unix(), now.Unix())
|
||||
if err != nil {
|
||||
log.Printf("notifier: find reminders: %v", err)
|
||||
return
|
||||
|
||||
@@ -69,6 +69,27 @@ func (f *fakeNtfy) messages() []pushed {
|
||||
return append([]pushed(nil), f.got...)
|
||||
}
|
||||
|
||||
// topicsSince lists the topics published to since the last forget, which is how
|
||||
// the escalation tests ask "who did this tick wake".
|
||||
func (f *fakeNtfy) topicsSince(t *testing.T) []string {
|
||||
t.Helper()
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
out := make([]string, 0, len(f.got))
|
||||
for _, m := range f.got {
|
||||
out = append(out, m.Topic)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// forget drops what has been published so far, so the next assertion is about
|
||||
// this tick rather than the whole test.
|
||||
func (f *fakeNtfy) forget() {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
f.got = nil
|
||||
}
|
||||
|
||||
func (f *fakeNtfy) failWith(status int) {
|
||||
f.mu.Lock()
|
||||
defer f.mu.Unlock()
|
||||
|
||||
+14
-1
@@ -4,6 +4,7 @@ import (
|
||||
"database/sql"
|
||||
"net/http"
|
||||
|
||||
"git.ryuvia.com/niklas/terdut-server/internal/config"
|
||||
"git.ryuvia.com/niklas/terdut-server/internal/web"
|
||||
"github.com/go-chi/chi/v5"
|
||||
"github.com/go-chi/chi/v5/middleware"
|
||||
@@ -13,7 +14,7 @@ import (
|
||||
// the only handler that has to decide where a new incident's page goes; a zero
|
||||
// notify disables notifications. Dead man's switches are per team and read from
|
||||
// the database, so nothing about them is wired in here.
|
||||
func NewRouter(db *sql.DB, notify NotifyConfig) http.Handler {
|
||||
func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config) http.Handler {
|
||||
r := chi.NewRouter()
|
||||
r.Use(middleware.Logger)
|
||||
r.Use(middleware.Recoverer)
|
||||
@@ -70,6 +71,13 @@ func NewRouter(db *sql.DB, notify NotifyConfig) http.Handler {
|
||||
r.Post("/api/users", handleCreateUser(db))
|
||||
r.Delete("/api/users/{id}", handleDeleteUser(db))
|
||||
r.Put("/api/users/{id}/admin", handleSetAdmin(db))
|
||||
r.Put("/api/users/{id}/disabled", handleSetUserDisabled(db))
|
||||
|
||||
// What exists on this server, and how it behaves. /api/teams
|
||||
// answers "what am I in"; this one answers "what is there".
|
||||
r.Get("/api/admin/teams", handleAdminListTeams(db))
|
||||
r.Get("/api/admin/settings", handleGetSettings(db, cfg))
|
||||
r.Put("/api/admin/settings", handleSetSettings(db))
|
||||
})
|
||||
|
||||
// Alerts are read-only: they are Alertmanager's record, not a work
|
||||
@@ -95,11 +103,16 @@ func NewRouter(db *sql.DB, notify NotifyConfig) http.Handler {
|
||||
// Teams. A user sees the teams they belong to; an owner configures one.
|
||||
r.Get("/api/teams", handleListTeams(db))
|
||||
r.Post("/api/teams", handleCreateTeam(db))
|
||||
r.Put("/api/teams/{teamID}", handleRenameTeam(db))
|
||||
r.Delete("/api/teams/{teamID}", handleDeleteTeam(db))
|
||||
r.Get("/api/teams/{teamID}/members", handleListTeamMembers(db))
|
||||
r.Post("/api/teams/{teamID}/members", handleAddTeamMember(db))
|
||||
r.Delete("/api/teams/{teamID}/members/{userID}", handleRemoveTeamMember(db))
|
||||
|
||||
// A team's escalation ladder: who is paged when nobody answers.
|
||||
r.Get("/api/teams/{teamID}/escalation", handleGetEscalation(db))
|
||||
r.Put("/api/teams/{teamID}/escalation", handleSetEscalation(db))
|
||||
|
||||
// A team's own dead man's switches: which of its alerts are heartbeats,
|
||||
// and how long a silence has to last before somebody is paged.
|
||||
r.Get("/api/teams/{teamID}/deadman", handleGetTeamDeadman(db))
|
||||
|
||||
@@ -0,0 +1,346 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"errors"
|
||||
"net/http"
|
||||
"strconv"
|
||||
"time"
|
||||
|
||||
"git.ryuvia.com/niklas/terdut-server/internal/config"
|
||||
"github.com/go-chi/chi/v5"
|
||||
)
|
||||
|
||||
// The settings an administrator can change at runtime. Each is behaviour rather
|
||||
// than infrastructure: what the server does, not where it is plugged in.
|
||||
//
|
||||
// The values are seconds, stored as text. A duration string would be friendlier
|
||||
// to read in psql and worse everywhere else — it can be stored unparseable, and
|
||||
// then the question is what a background loop should do at 02:00 with a
|
||||
// tuning knob it cannot understand.
|
||||
const (
|
||||
SettingNotifyRepeat = "notify_repeat_seconds"
|
||||
SettingStaleAfter = "stale_after_seconds"
|
||||
SettingArchiveAfter = "archive_after_seconds"
|
||||
)
|
||||
|
||||
// settingBounds keeps an edit from producing a server that cannot work. The
|
||||
// ceilings are loose — they exist to catch a slipped decimal point, not to have
|
||||
// an opinion about anybody's rota.
|
||||
var settingBounds = map[string]struct {
|
||||
min, max time.Duration
|
||||
label string
|
||||
}{
|
||||
SettingNotifyRepeat: {0, 24 * time.Hour, "how long an incident may sit unacknowledged before it is paged again; 0 disables reminders"},
|
||||
SettingStaleAfter: {5 * time.Minute, 30 * 24 * time.Hour, "how long a firing alert may go without a refreshing webhook before the sweeper resolves it"},
|
||||
SettingArchiveAfter: {time.Minute, 365 * 24 * time.Hour, "how long a resolved alert or incident stays in the default list"},
|
||||
}
|
||||
|
||||
// Settings reads the runtime configuration. It holds no cache: the readers are
|
||||
// two background loops that tick every 30 seconds and 15 minutes, and handlers
|
||||
// that run once per request, so a query each time costs nothing measurable and
|
||||
// means an administrator's change takes effect on the next tick rather than at
|
||||
// the next restart.
|
||||
type Settings struct{ db *sql.DB }
|
||||
|
||||
// NewSettings returns a reader over db.
|
||||
func NewSettings(db *sql.DB) *Settings { return &Settings{db: db} }
|
||||
|
||||
// Duration reads one setting, falling back to def when the row is missing or
|
||||
// unreadable. A tuning knob is never worth failing a sweep over: the fallback
|
||||
// is the value the server started with.
|
||||
func (s *Settings) Duration(ctx context.Context, key string, def time.Duration) time.Duration {
|
||||
var raw string
|
||||
err := s.db.QueryRowContext(ctx, "SELECT value FROM settings WHERE key = $1", key).Scan(&raw)
|
||||
if err != nil {
|
||||
return def
|
||||
}
|
||||
secs, err := strconv.ParseInt(raw, 10, 64)
|
||||
if err != nil {
|
||||
return def
|
||||
}
|
||||
return time.Duration(secs) * time.Second
|
||||
}
|
||||
|
||||
// SeedSettings writes each key from the server's environment configuration,
|
||||
// once. Never overwrites: after the first start the database owns these, and a
|
||||
// redeploy must not put a chart's default back over an administrator's edit —
|
||||
// the same rule as the per-team dead man's switches.
|
||||
func SeedSettings(ctx context.Context, db *sql.DB, cfg config.Config) error {
|
||||
seeds := map[string]time.Duration{
|
||||
SettingNotifyRepeat: cfg.NotifyRepeat,
|
||||
SettingStaleAfter: cfg.StaleAfter,
|
||||
SettingArchiveAfter: cfg.ArchiveAfter,
|
||||
}
|
||||
for key, d := range seeds {
|
||||
if _, err := db.ExecContext(ctx, `
|
||||
INSERT INTO settings (key, value) VALUES ($1, $2)
|
||||
ON CONFLICT (key) DO NOTHING`,
|
||||
key, strconv.FormatInt(int64(d.Seconds()), 10)); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// settingsResponse is what the admin page renders. The environment half is
|
||||
// included and marked read-only, so somebody looking for the ntfy URL finds out
|
||||
// where it lives rather than concluding the server does not have one.
|
||||
type settingsResponse struct {
|
||||
Editable map[string]settingValue `json:"editable"`
|
||||
FromEnv map[string]string `json:"from_env"`
|
||||
}
|
||||
|
||||
type settingValue struct {
|
||||
Seconds int64 `json:"seconds"`
|
||||
Description string `json:"description"`
|
||||
MinSeconds int64 `json:"min_seconds"`
|
||||
MaxSeconds int64 `json:"max_seconds"`
|
||||
}
|
||||
|
||||
func handleGetSettings(db *sql.DB, cfg config.Config) http.HandlerFunc {
|
||||
settings := NewSettings(db)
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
out := settingsResponse{
|
||||
Editable: map[string]settingValue{},
|
||||
FromEnv: map[string]string{
|
||||
// Never the ntfy token or the DSN: both are credentials, and an
|
||||
// admin page that renders them turns a browser tab into a place
|
||||
// they leak from.
|
||||
"ntfy_url": cfg.NtfyURL,
|
||||
"ntfy_configured": strconv.FormatBool(cfg.NtfyURL != ""),
|
||||
"ntfy_token_set": strconv.FormatBool(cfg.NtfyToken != ""),
|
||||
"public_url": cfg.PublicURL,
|
||||
"listen_address": cfg.Addr,
|
||||
},
|
||||
}
|
||||
for key, b := range settingBounds {
|
||||
def := map[string]time.Duration{
|
||||
SettingNotifyRepeat: cfg.NotifyRepeat,
|
||||
SettingStaleAfter: cfg.StaleAfter,
|
||||
SettingArchiveAfter: cfg.ArchiveAfter,
|
||||
}[key]
|
||||
out.Editable[key] = settingValue{
|
||||
Seconds: int64(settings.Duration(r.Context(), key, def).Seconds()),
|
||||
Description: b.label,
|
||||
MinSeconds: int64(b.min.Seconds()),
|
||||
MaxSeconds: int64(b.max.Seconds()),
|
||||
}
|
||||
}
|
||||
respond(w, http.StatusOK, out)
|
||||
}
|
||||
}
|
||||
|
||||
// handleSetSettings changes one or more settings. Unknown keys are refused
|
||||
// rather than stored: a typo that writes notify_repeat_second would otherwise
|
||||
// sit in the table looking like configuration and doing nothing.
|
||||
func handleSetSettings(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
var req map[string]int64
|
||||
if err := decodeJSON(r, &req); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid request body"))
|
||||
return
|
||||
}
|
||||
if len(req) == 0 {
|
||||
respond(w, http.StatusBadRequest, errResp("no settings given"))
|
||||
return
|
||||
}
|
||||
|
||||
for key, secs := range req {
|
||||
b, known := settingBounds[key]
|
||||
if !known {
|
||||
respond(w, http.StatusBadRequest, errResp("unknown setting: "+key))
|
||||
return
|
||||
}
|
||||
d := time.Duration(secs) * time.Second
|
||||
if d < b.min || d > b.max {
|
||||
respond(w, http.StatusBadRequest, errResp(
|
||||
key+" must be between "+b.min.String()+" and "+b.max.String()))
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
tx, err := db.BeginTx(r.Context(), nil)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer tx.Rollback() //nolint:errcheck
|
||||
|
||||
for key, secs := range req {
|
||||
if _, err := tx.ExecContext(r.Context(), `
|
||||
INSERT INTO settings (key, value, updated_at)
|
||||
VALUES ($1, $2, `+nowEpoch+`)
|
||||
ON CONFLICT (key) DO UPDATE SET
|
||||
value = excluded.value, updated_at = excluded.updated_at`,
|
||||
key, strconv.FormatInt(secs, 10)); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
}
|
||||
if err := tx.Commit(); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
// handleAdminListTeams lists every team on the server, with its size. The
|
||||
// ordinary /api/teams answers "what am I in"; this one answers "what exists",
|
||||
// which only an administrator may ask.
|
||||
func handleAdminListTeams(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
rows, err := db.QueryContext(r.Context(), `
|
||||
SELECT t.id, t.name, t.created_at,
|
||||
(SELECT COUNT(*) FROM team_members m WHERE m.team_id = t.id),
|
||||
(SELECT COUNT(*) FROM incidents i
|
||||
WHERE i.team_id = t.id AND i.resolved_at IS NULL)
|
||||
FROM teams t
|
||||
ORDER BY t.name`)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
type adminTeam struct {
|
||||
ID int64 `json:"id"`
|
||||
Name string `json:"name"`
|
||||
CreatedAt time.Time `json:"created_at"`
|
||||
Members int64 `json:"members"`
|
||||
OpenIncidents int64 `json:"open_incidents"`
|
||||
}
|
||||
teams := []adminTeam{}
|
||||
for rows.Next() {
|
||||
var t adminTeam
|
||||
var created int64
|
||||
if err := rows.Scan(&t.ID, &t.Name, &created, &t.Members, &t.OpenIncidents); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
t.CreatedAt = time.Unix(created, 0).UTC()
|
||||
teams = append(teams, t)
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, teams)
|
||||
}
|
||||
}
|
||||
|
||||
// handleRenameTeam renames a team. An owner's job, and an administrator's when
|
||||
// a team has nobody left to do it.
|
||||
func handleRenameTeam(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
teamID, ok := teamParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !requireTeamOwner(w, r, teamID) {
|
||||
return
|
||||
}
|
||||
|
||||
var req struct {
|
||||
Name string `json:"name"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil || req.Name == "" {
|
||||
respond(w, http.StatusBadRequest, errResp("name is required"))
|
||||
return
|
||||
}
|
||||
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"UPDATE teams SET name = $1 WHERE id = $2", req.Name, teamID)
|
||||
if err != nil {
|
||||
if isUniqueViolation(err) {
|
||||
respond(w, http.StatusConflict, errResp("a team with that name already exists"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("not found"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
// handleSetUserDisabled takes an account out of use, or puts it back.
|
||||
//
|
||||
// Not a delete: the person's acknowledgements, assignments and timeline entries
|
||||
// stay attached to them. Deleting a user nulls those columns, which rewrites
|
||||
// what happened during an incident months after the fact.
|
||||
func handleSetUserDisabled(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid user id"))
|
||||
return
|
||||
}
|
||||
var req struct {
|
||||
Disabled *bool `json:"disabled"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil || req.Disabled == nil {
|
||||
respond(w, http.StatusBadRequest, errResp("disabled is required"))
|
||||
return
|
||||
}
|
||||
|
||||
if *req.Disabled {
|
||||
caller, _ := userFromContext(r.Context())
|
||||
if caller.ID == id {
|
||||
respond(w, http.StatusConflict, errResp("cannot disable your own account"))
|
||||
return
|
||||
}
|
||||
last, err := isLastAdmin(r.Context(), db, id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if last {
|
||||
respond(w, http.StatusConflict, errResp("cannot disable the last administrator"))
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
var res sql.Result
|
||||
if *req.Disabled {
|
||||
res, err = db.ExecContext(r.Context(),
|
||||
"UPDATE users SET disabled_at = "+nowEpoch+" WHERE id = $1 AND disabled_at IS NULL", id)
|
||||
} else {
|
||||
res, err = db.ExecContext(r.Context(),
|
||||
"UPDATE users SET disabled_at = NULL WHERE id = $1", id)
|
||||
}
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
// Either no such user, or already in the state asked for. The
|
||||
// second is not a failure, so check which before answering.
|
||||
var exists int
|
||||
if err := db.QueryRowContext(r.Context(),
|
||||
"SELECT 1 FROM users WHERE id = $1", id).Scan(&exists); errors.Is(err, sql.ErrNoRows) {
|
||||
respond(w, http.StatusNotFound, errResp("user not found"))
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
// Signing back in is the only way to use a re-enabled account, and a
|
||||
// disabled one must not keep a live session.
|
||||
if *req.Disabled {
|
||||
db.ExecContext(r.Context(), "DELETE FROM sessions WHERE user_id = $1", id) //nolint:errcheck
|
||||
}
|
||||
|
||||
user, err := fetchUser(r.Context(), db, id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, user)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,275 @@
|
||||
package api_test
|
||||
|
||||
import (
|
||||
"net/http"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"git.ryuvia.com/niklas/terdut-server/internal/api"
|
||||
)
|
||||
|
||||
// The settings an administrator can change, and the ones they cannot.
|
||||
func TestSettings_EditableAndReadOnly(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
var got struct {
|
||||
Editable map[string]struct {
|
||||
Seconds int64 `json:"seconds"`
|
||||
Description string `json:"description"`
|
||||
MinSeconds int64 `json:"min_seconds"`
|
||||
MaxSeconds int64 `json:"max_seconds"`
|
||||
} `json:"editable"`
|
||||
FromEnv map[string]string `json:"from_env"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodGet, "/api/admin/settings", nil), &got)
|
||||
|
||||
// Seeded from the environment the server started with, not from zero.
|
||||
if v := got.Editable["notify_repeat_seconds"].Seconds; v != 900 {
|
||||
t.Errorf("notify_repeat_seconds seeded as %d, want 900", v)
|
||||
}
|
||||
if v := got.Editable["stale_after_seconds"].Seconds; v != 21600 {
|
||||
t.Errorf("stale_after_seconds seeded as %d, want 21600", v)
|
||||
}
|
||||
if got.Editable["archive_after_seconds"].Description == "" {
|
||||
t.Error("a setting without a description is a number nobody can act on")
|
||||
}
|
||||
|
||||
// The environment half is visible so somebody can see where it lives, but
|
||||
// never the credentials themselves.
|
||||
if _, ok := got.FromEnv["public_url"]; !ok {
|
||||
t.Error("public_url should be reported as environment-configured")
|
||||
}
|
||||
for _, leak := range []string{"ntfy_token", "dsn", "database_dsn", "password"} {
|
||||
if v, ok := got.FromEnv[leak]; ok {
|
||||
t.Errorf("%s must not be in the settings response (got %q)", leak, v)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Changing a setting takes effect on the next tick, without a restart. This is
|
||||
// the whole point of moving them out of the environment.
|
||||
func TestSettings_ChangeTakesEffectOnTheNextSweep(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
// An alert whose last webhook was two hours ago. Under the seeded
|
||||
// stale_after of six hours the sweeper leaves it alone.
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-settings", "Stale", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.exec(t, "UPDATE alerts SET received_at = $1 WHERE fingerprint = $2",
|
||||
time.Now().Add(-2*time.Hour).Unix(), "fp-settings")
|
||||
|
||||
sweep(t, s, noArchive)
|
||||
if status, _, _ := s.alertRow(t, "fp-settings"); status != "firing" {
|
||||
t.Fatalf("before the change the alert should still be firing, got %q", status)
|
||||
}
|
||||
|
||||
// Shorten it to an hour. Nothing restarts.
|
||||
resp := s.req(t, http.MethodPut, "/api/admin/settings",
|
||||
map[string]int64{"stale_after_seconds": 3600})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Fatalf("change setting: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
sweep(t, s, noArchive)
|
||||
status, source, _ := s.alertRow(t, "fp-settings")
|
||||
if status != "resolved" {
|
||||
t.Errorf("after the change the alert should have expired, got %q", status)
|
||||
}
|
||||
if source == nil || *source != "expiry" {
|
||||
t.Errorf("expected resolution_source expiry, got %v", source)
|
||||
}
|
||||
}
|
||||
|
||||
// A typo must not look like configuration, and a slipped decimal point must not
|
||||
// produce a server that sweeps every second.
|
||||
func TestSettings_RejectsUnknownKeysAndSillyValues(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
body map[string]int64
|
||||
}{
|
||||
{"unknown key", map[string]int64{"notify_repeat_second": 60}},
|
||||
{"below the floor", map[string]int64{"stale_after_seconds": 30}},
|
||||
{"above the ceiling", map[string]int64{"archive_after_seconds": 400 * 24 * 3600}},
|
||||
{"nothing at all", map[string]int64{}},
|
||||
} {
|
||||
resp := s.req(t, http.MethodPut, "/api/admin/settings", c.body)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusBadRequest {
|
||||
t.Errorf("%s: expected 400, got %d", c.name, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Settings are the server's behaviour, so only an administrator may change
|
||||
// them — or see where the rest of the configuration comes from.
|
||||
func TestSettings_AreAdminOnly(t *testing.T) {
|
||||
s := newTS(t)
|
||||
_, call := member(t, s, "member")
|
||||
|
||||
for _, c := range []struct {
|
||||
method string
|
||||
body any
|
||||
}{
|
||||
{http.MethodGet, nil},
|
||||
{http.MethodPut, map[string]int64{"notify_repeat_seconds": 60}},
|
||||
} {
|
||||
resp := call(c.method, "/api/admin/settings", c.body)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusForbidden {
|
||||
t.Errorf("%s /api/admin/settings: expected 403, got %d", c.method, resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
resp := call(http.MethodGet, "/api/admin/teams", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusForbidden {
|
||||
t.Errorf("GET /api/admin/teams: expected 403, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// An administrator sees every team, including ones they are not in — which is
|
||||
// exactly what /api/teams must not show them.
|
||||
func TestSettings_AdminSeesEveryTeam(t *testing.T) {
|
||||
s := newTS(t)
|
||||
newTeam(t, s, "red")
|
||||
newTeam(t, s, "blue")
|
||||
|
||||
all := list(t, s.req(t, http.MethodGet, "/api/admin/teams", nil))
|
||||
if len(all) != 3 { // Default, red, blue
|
||||
t.Fatalf("admin should see all 3 teams, saw %d", len(all))
|
||||
}
|
||||
for _, team := range all {
|
||||
if _, ok := team["members"]; !ok {
|
||||
t.Error("the admin listing should say how big each team is")
|
||||
}
|
||||
}
|
||||
|
||||
// The admin created them, so they own them — but they are not a member of
|
||||
// a team somebody else makes, and /api/teams still answers "what am I in".
|
||||
mine := list(t, s.req(t, http.MethodGet, "/api/teams", nil))
|
||||
if len(mine) != 3 {
|
||||
t.Errorf("the creator is an owner of what they created, saw %d", len(mine))
|
||||
}
|
||||
}
|
||||
|
||||
// Disabling is not deleting: the account stops working and the history stays.
|
||||
func TestSettings_DisablingAnAccountKeepsItsHistory(t *testing.T) {
|
||||
s := newTS(t)
|
||||
memberID, call := member(t, s, "leaver")
|
||||
|
||||
// They acknowledge an incident, so there is history to preserve.
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-leaver", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
resp := call(http.MethodPost, "/api/incidents/1/acknowledge", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("acknowledge: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodPut, "/api/users/"+id64(memberID)+"/disabled",
|
||||
map[string]bool{"disabled": true})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("disable: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
// Their API key stops working.
|
||||
resp = call(http.MethodGet, "/api/incidents", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusUnauthorized {
|
||||
t.Errorf("a disabled user's key: expected 401, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
// The acknowledgement still names them.
|
||||
var incident map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, "/api/incidents/1", nil), &incident)
|
||||
if incident["acknowledged_by"] != "leaver" {
|
||||
t.Errorf("the acknowledgement should still name leaver, got %v", incident["acknowledged_by"])
|
||||
}
|
||||
if incident["status"] != "acknowledged" {
|
||||
t.Errorf("the incident should still be acknowledged, got %v", incident["status"])
|
||||
}
|
||||
|
||||
// And re-enabling gives the account back.
|
||||
resp = s.req(t, http.MethodPut, "/api/users/"+id64(memberID)+"/disabled",
|
||||
map[string]bool{"disabled": false})
|
||||
resp.Body.Close()
|
||||
resp = call(http.MethodGet, "/api/incidents", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Errorf("after re-enabling: expected 200, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// The same two guards as deleting and demoting: an install must keep somebody
|
||||
// who can administer it.
|
||||
func TestSettings_CannotDisableYourselfOrTheLastAdmin(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
resp := s.req(t, http.MethodPut, "/api/users/1/disabled", map[string]bool{"disabled": true})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusConflict {
|
||||
t.Errorf("disabling yourself: expected 409, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// Renaming a team is an owner's job, and the name stays unique.
|
||||
func TestSettings_TeamRename(t *testing.T) {
|
||||
s := newTS(t)
|
||||
team := newTeam(t, s, "red")
|
||||
|
||||
resp := s.req(t, http.MethodPut, "/api/teams/"+id64(team.id), map[string]string{"name": "Platform"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Fatalf("rename: %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
teams := list(t, s.req(t, http.MethodGet, "/api/admin/teams", nil))
|
||||
found := false
|
||||
for _, x := range teams {
|
||||
if x["name"] == "Platform" {
|
||||
found = true
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
t.Error("the renamed team should be listed under its new name")
|
||||
}
|
||||
|
||||
// Taking a name that exists is a conflict, not a silent second team with
|
||||
// the same label.
|
||||
resp = s.req(t, http.MethodPut, "/api/teams/"+id64(team.id), map[string]string{"name": "Default"})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusConflict {
|
||||
t.Errorf("renaming onto an existing name: expected 409, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// The seed runs once. A redeploy must not put the chart's default back over an
|
||||
// administrator's edit — the rule the dead man's switches already follow.
|
||||
func TestSettings_SeedDoesNotOverwrite(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
resp := s.req(t, http.MethodPut, "/api/admin/settings",
|
||||
map[string]int64{"notify_repeat_seconds": 60})
|
||||
resp.Body.Close()
|
||||
|
||||
// A second start, with the environment still saying 15 minutes.
|
||||
if err := api.SeedSettings(t.Context(), s.db, testConfig()); err != nil {
|
||||
t.Fatalf("re-seed: %v", err)
|
||||
}
|
||||
|
||||
var got struct {
|
||||
Editable map[string]struct {
|
||||
Seconds int64 `json:"seconds"`
|
||||
} `json:"editable"`
|
||||
}
|
||||
decode(t, s.req(t, http.MethodGet, "/api/admin/settings", nil), &got)
|
||||
if v := got.Editable["notify_repeat_seconds"].Seconds; v != 60 {
|
||||
t.Errorf("the edit should survive a restart, got %d", v)
|
||||
}
|
||||
}
|
||||
@@ -7,7 +7,9 @@ import (
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"git.ryuvia.com/niklas/terdut-server/internal/config"
|
||||
"git.ryuvia.com/niklas/terdut-server/internal/db"
|
||||
)
|
||||
|
||||
@@ -30,6 +32,19 @@ import (
|
||||
// tests nothing is worse than one that does not run.
|
||||
const testDSNEnv = "TERDUT_TEST_DSN"
|
||||
|
||||
// testConfig is the environment half of the server's configuration, which the
|
||||
// admin settings page renders read-only and SeedSettings seeds the editable
|
||||
// half from. The durations match the defaults config.Load would produce, so a
|
||||
// test that never touches the settings table behaves as a fresh install does.
|
||||
func testConfig() config.Config {
|
||||
return config.Config{
|
||||
Addr: ":8080",
|
||||
ArchiveAfter: 7 * 24 * time.Hour,
|
||||
StaleAfter: 6 * time.Hour,
|
||||
NotifyRepeat: 15 * time.Minute,
|
||||
}
|
||||
}
|
||||
|
||||
// defaultTeam is the team migration 003 creates and the bootstrap user owns, as
|
||||
// a path segment. Every test that does not say otherwise works inside it.
|
||||
const defaultTeam = "1"
|
||||
|
||||
@@ -96,7 +96,7 @@ func handleBootstrap(db *sql.DB) http.HandlerFunc {
|
||||
func handleListUsers(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
rows, err := db.QueryContext(r.Context(),
|
||||
"SELECT id, username, email, created_at, ntfy_topic, is_admin FROM users ORDER BY id")
|
||||
"SELECT id, username, email, created_at, ntfy_topic, is_admin, disabled_at FROM users ORDER BY id")
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
@@ -107,11 +107,13 @@ func handleListUsers(db *sql.DB) http.HandlerFunc {
|
||||
for rows.Next() {
|
||||
var u models.User
|
||||
var ts int64
|
||||
if err := rows.Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic, &u.IsAdmin); err != nil {
|
||||
var disabled *int64
|
||||
if err := rows.Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic, &u.IsAdmin, &disabled); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
u.CreatedAt = time.Unix(ts, 0).UTC()
|
||||
u.DisabledAt = unixPtr(disabled)
|
||||
users = append(users, u)
|
||||
}
|
||||
respond(w, http.StatusOK, users)
|
||||
@@ -325,13 +327,15 @@ func randomToken() (raw, hash string, err error) {
|
||||
func fetchUser(ctx context.Context, db *sql.DB, id int64) (models.User, error) {
|
||||
var u models.User
|
||||
var ts int64
|
||||
var disabled *int64
|
||||
err := db.QueryRowContext(ctx,
|
||||
"SELECT id, username, email, created_at, ntfy_topic, is_admin FROM users WHERE id = $1", id).
|
||||
Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic, &u.IsAdmin)
|
||||
"SELECT id, username, email, created_at, ntfy_topic, is_admin, disabled_at FROM users WHERE id = $1", id).
|
||||
Scan(&u.ID, &u.Username, &u.Email, &ts, &u.NtfyTopic, &u.IsAdmin, &disabled)
|
||||
if err != nil {
|
||||
return u, err
|
||||
}
|
||||
u.CreatedAt = time.Unix(ts, 0).UTC()
|
||||
u.DisabledAt = unixPtr(disabled)
|
||||
return u, nil
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,35 @@
|
||||
-- Settings that an administrator can change without a redeploy, and the flag
|
||||
-- that takes an account out of use without deleting it.
|
||||
--
|
||||
-- Three of the server's tunables were environment variables, which meant
|
||||
-- changing how long an incident waits before it is paged again required editing
|
||||
-- a chart, merging it, and waiting for a reconcile. They are behaviour, not
|
||||
-- infrastructure, and the difference is who needs to change them and how often.
|
||||
--
|
||||
-- What stays in the environment: the ntfy URL and token, the database DSN, the
|
||||
-- listen address and the public URL. Those are where the server is plugged in
|
||||
-- rather than how it behaves, they are needed before the database is open, and
|
||||
-- two of them are credentials.
|
||||
--
|
||||
-- Key/value rather than a column per setting. A settings table with one row and
|
||||
-- a column per knob needs a migration for every new knob, and #6 and #7 will
|
||||
-- both add some. The cost is that values are text and the accessor has to say
|
||||
-- what type it wanted; settings.go does that in one place.
|
||||
--
|
||||
-- No rows are seeded here: a migration cannot read the environment. The server
|
||||
-- inserts each key from its own configuration at startup, once, so an install
|
||||
-- that upgrades keeps exactly the behaviour it had. See SeedSettings.
|
||||
CREATE TABLE settings (
|
||||
key TEXT PRIMARY KEY,
|
||||
value TEXT NOT NULL,
|
||||
updated_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
|
||||
-- Disabling an account rather than deleting it: the person has left, or the
|
||||
-- credential is suspect, and their incidents, acknowledgements and timeline
|
||||
-- entries must stay exactly where they are. Deleting a user nulls their
|
||||
-- acknowledged_by and assigned_to, which quietly rewrites history.
|
||||
--
|
||||
-- A disabled user cannot sign in and their API keys stop working, but they are
|
||||
-- still a name the timeline can show and still a member of their teams.
|
||||
ALTER TABLE users ADD COLUMN disabled_at BIGINT;
|
||||
@@ -0,0 +1,95 @@
|
||||
-- Escalation: page somebody else when the first person does not answer.
|
||||
--
|
||||
-- This is the gap the whole multi-tenancy line of work was opened to close.
|
||||
-- Until now an unacknowledged incident re-paged the same topic every
|
||||
-- notify_repeat forever, which is a louder version of the same silence: if the
|
||||
-- person on call is asleep, has no signal, or has left, nothing else happens.
|
||||
--
|
||||
-- Shape: one policy per team, an ordered list of levels, each level with a
|
||||
-- timeout and a set of targets. When a level's timeout passes and the incident
|
||||
-- is still triggered, the next level is paged. When the last level passes, the
|
||||
-- chain repeats repeat_count times, and then the team's fallback topic is paged
|
||||
-- once as the end of the line.
|
||||
--
|
||||
-- A team WITHOUT a policy keeps exactly today's behaviour: page the assignee,
|
||||
-- then remind on the same topic. Escalation is opt-in per team, and the two
|
||||
-- never both run for one incident -- see enqueueReminders.
|
||||
CREATE TABLE escalation_policies (
|
||||
-- One per team for now, hence the team as the key rather than an id with a
|
||||
-- unique index: routing different alerts to different chains needs the
|
||||
-- alert to carry something to route ON, which is a separate question.
|
||||
team_id BIGINT PRIMARY KEY REFERENCES teams(id) ON DELETE CASCADE,
|
||||
|
||||
-- How many extra times to run the whole chain after it has been walked
|
||||
-- once. 0 means walk it once and stop at the fallback.
|
||||
repeat_count BIGINT NOT NULL DEFAULT 0 CHECK (repeat_count >= 0 AND repeat_count <= 10),
|
||||
|
||||
-- Where the last page goes when every level has been tried. Per team now:
|
||||
-- TERDUT_NTFY_FALLBACK_TOPIC was one topic for the whole install, which in
|
||||
-- a multi-team server pages the wrong people. Empty means the chain simply
|
||||
-- ends.
|
||||
fallback_topic TEXT NOT NULL DEFAULT '',
|
||||
|
||||
updated_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
|
||||
);
|
||||
|
||||
CREATE TABLE escalation_levels (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
team_id BIGINT NOT NULL REFERENCES escalation_policies(team_id) ON DELETE CASCADE,
|
||||
-- 1-based, dense. The API rewrites the whole ladder on every edit rather
|
||||
-- than patching one rung, so there is no way to leave a gap.
|
||||
position BIGINT NOT NULL,
|
||||
-- How long this level has to produce an acknowledgement before the next one
|
||||
-- is paged. Seconds, like every other duration in this schema.
|
||||
timeout_seconds BIGINT NOT NULL CHECK (timeout_seconds > 0),
|
||||
|
||||
UNIQUE (team_id, position)
|
||||
);
|
||||
|
||||
-- Who a level pages. Either a named person, or whoever the team's rota says is
|
||||
-- on call today -- which is the target that keeps working when the rota
|
||||
-- changes and nobody remembers to edit the policy.
|
||||
CREATE TABLE escalation_targets (
|
||||
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
|
||||
level_id BIGINT NOT NULL REFERENCES escalation_levels(id) ON DELETE CASCADE,
|
||||
kind TEXT NOT NULL CHECK (kind IN ('user', 'oncall')),
|
||||
-- Set for kind='user', NULL for kind='oncall'.
|
||||
user_id BIGINT REFERENCES users(id) ON DELETE CASCADE,
|
||||
|
||||
CHECK ((kind = 'user' AND user_id IS NOT NULL) OR (kind = 'oncall' AND user_id IS NULL))
|
||||
);
|
||||
|
||||
CREATE INDEX escalation_targets_level_idx ON escalation_targets(level_id);
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- Where an incident is in its chain.
|
||||
--
|
||||
-- On the incident rather than in a side table: it is read on every notifier
|
||||
-- tick alongside the incident's status, and one row per incident is exactly
|
||||
-- what the state is.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
-- 0 means no level has been paged yet, which is the state of every incident
|
||||
-- that existed before escalation and of every incident in a team with no
|
||||
-- policy. 1 is the first level.
|
||||
ALTER TABLE incidents ADD COLUMN escalation_level BIGINT NOT NULL DEFAULT 0;
|
||||
|
||||
-- When the current level was entered, and therefore what its timeout is
|
||||
-- measured from. NULL while escalation_level is 0.
|
||||
ALTER TABLE incidents ADD COLUMN escalation_level_at BIGINT;
|
||||
|
||||
-- How many times the chain has been walked in full. Compared against the
|
||||
-- policy's repeat_count.
|
||||
ALTER TABLE incidents ADD COLUMN escalation_round BIGINT NOT NULL DEFAULT 0;
|
||||
|
||||
-- The notifier's escalation query: incidents still waiting, oldest level first.
|
||||
CREATE INDEX incidents_escalation_idx
|
||||
ON incidents(escalation_level_at)
|
||||
WHERE resolved_at IS NULL AND status = 'triggered';
|
||||
|
||||
-- 'escalated' joins the outbox kinds: a page that went out because nobody
|
||||
-- answered the last one, which is worth telling apart from the first page and
|
||||
-- from a reminder when reading the timeline or debugging a delivery.
|
||||
ALTER TABLE notifications DROP CONSTRAINT notifications_kind_check;
|
||||
ALTER TABLE notifications ADD CONSTRAINT notifications_kind_check
|
||||
CHECK (kind IN ('triggered', 'reminder', 'resolved', 'escalated'));
|
||||
@@ -13,6 +13,11 @@ type User struct {
|
||||
// fallback topic instead.
|
||||
NtfyTopic *string `json:"ntfy_topic,omitempty"`
|
||||
|
||||
// DisabledAt is when the account was taken out of use, or nil. A disabled
|
||||
// user cannot authenticate by either credential, and keeps their name on
|
||||
// every acknowledgement and timeline entry they made.
|
||||
DisabledAt *time.Time `json:"disabled_at,omitempty"`
|
||||
|
||||
// IsAdmin is the system administrator flag: managing users and API keys.
|
||||
// Not omitempty — a client has to be able to tell "false" from "this server
|
||||
// is too old to have the field", and the web UI decides what to show from
|
||||
|
||||
@@ -630,3 +630,30 @@ kbd {
|
||||
.toast, .app.detail-open ~ .toast { bottom: 24px; }
|
||||
.only-desktop { display: block; }
|
||||
}
|
||||
|
||||
/* --- admin ---------------------------------------------------------------
|
||||
The admin page is three tables of things you act on, so it needs table
|
||||
styling the rest of the app never did: the queue is a list of links and the
|
||||
account page is a form. */
|
||||
.admin-table { width: 100%; border-collapse: collapse; font-size: 14px; }
|
||||
.admin-table th {
|
||||
text-align: left; font-weight: 600; color: var(--muted); font-size: 12px;
|
||||
text-transform: uppercase; letter-spacing: 0.04em;
|
||||
padding: 4px 8px 4px 0; border-bottom: 1px solid var(--border);
|
||||
}
|
||||
.admin-table td { padding: 8px 8px 8px 0; border-bottom: 1px solid var(--border); vertical-align: middle; }
|
||||
.admin-table tr:last-child td { border-bottom: none; }
|
||||
.admin-table .num { text-align: right; font-variant-numeric: tabular-nums; }
|
||||
.admin-table td .btn-sm + .btn-sm { margin-left: 6px; }
|
||||
/* A disabled account stays readable — it is still the name on old
|
||||
acknowledgements — but should not look like a working one. */
|
||||
.disabled-row td { opacity: 0.55; }
|
||||
.btn-sm.danger { color: var(--crit); border-color: var(--crit-soft); }
|
||||
|
||||
.inline-form { display: flex; gap: 8px; margin-top: 12px; }
|
||||
.inline-form input { flex: 1; min-width: 0; }
|
||||
|
||||
.admin-settings .setting-value { width: 5.5em; margin-right: 6px; }
|
||||
.admin-settings .setting-unit { max-width: 8em; }
|
||||
.admin-settings button[type="submit"] { margin-top: 12px; }
|
||||
.small { font-size: 13px; }
|
||||
|
||||
@@ -59,6 +59,13 @@
|
||||
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M6 16V11a6 6 0 0 1 12 0v5l1.5 2h-15z"/><path d="M10 20.5a2 2 0 0 0 4 0"/></svg>
|
||||
<span class="nav-label">Alerts</span>
|
||||
</a>
|
||||
<!-- Hidden unless the signed-in user is a system administrator; app.js
|
||||
unhides it once /api/me says so. The server refuses every admin
|
||||
endpoint regardless, so this is a courtesy and not a gate. -->
|
||||
<a class="nav-link" href="/admin" data-section="admin" id="nav-admin" hidden>
|
||||
<svg viewBox="0 0 24 24" aria-hidden="true"><path d="M12 3l7 3v6c0 4-3 7-7 9-4-2-7-5-7-9V6z"/></svg>
|
||||
<span class="nav-label">Admin</span>
|
||||
</a>
|
||||
<a class="nav-link" href="/more" data-section="more">
|
||||
<svg viewBox="0 0 24 24" aria-hidden="true"><circle cx="12" cy="8" r="3.5"/><path d="M5 20a7 7 0 0 1 14 0"/></svg>
|
||||
<span class="nav-label">Account</span>
|
||||
@@ -80,6 +87,7 @@
|
||||
|
||||
<section id="view-oncall" class="view view-page" data-view="oncall" hidden></section>
|
||||
<section id="view-alerts" class="view view-page" data-view="alerts" hidden></section>
|
||||
<section id="view-admin" class="view view-page" data-view="admin" hidden></section>
|
||||
<section id="view-more" class="view view-page" data-view="more" hidden></section>
|
||||
</div>
|
||||
|
||||
|
||||
@@ -0,0 +1,308 @@
|
||||
// Administration: the teams on this server, the people who can sign in, and
|
||||
// the settings that change how the server behaves.
|
||||
//
|
||||
// Only rendered for a system administrator. The server enforces that on every
|
||||
// endpoint regardless — hiding a section is a courtesy to the reader, not a
|
||||
// permission — so this view simply says so rather than pretending to be a
|
||||
// gate.
|
||||
|
||||
import * as api from './api.js';
|
||||
import { h, clear, spinner, confirm } from './ui.js';
|
||||
import { state, myID } from './state.js';
|
||||
|
||||
const view = () => document.getElementById('view-admin');
|
||||
|
||||
let data = null; // { teams, users, settings }
|
||||
let error = null;
|
||||
let busy = false;
|
||||
|
||||
export function show() {
|
||||
if (!data) clear(view(), spinner());
|
||||
refresh();
|
||||
}
|
||||
|
||||
export async function refresh() {
|
||||
if (!state.me?.user?.is_admin) {
|
||||
data = null;
|
||||
render();
|
||||
return;
|
||||
}
|
||||
try {
|
||||
const [teams, users, settings] = await Promise.all([
|
||||
api.adminTeams(),
|
||||
api.users(),
|
||||
api.adminSettings(),
|
||||
]);
|
||||
data = { teams, users, settings };
|
||||
error = null;
|
||||
} catch (err) {
|
||||
error = err.message;
|
||||
}
|
||||
render();
|
||||
}
|
||||
|
||||
function render() {
|
||||
if (!state.me?.user?.is_admin) {
|
||||
clear(view(), h('div', { class: 'card' },
|
||||
h('p', { class: 'muted', text: 'Administration is for system administrators. Ask one for access.' })));
|
||||
return;
|
||||
}
|
||||
if (!data) {
|
||||
clear(view(), error ? h('div', { class: 'load-error', text: error }) : spinner());
|
||||
return;
|
||||
}
|
||||
clear(view(),
|
||||
error && h('div', { class: 'load-error', text: `Showing older data: ${error}` }),
|
||||
teamsCard(),
|
||||
usersCard(),
|
||||
settingsCard(),
|
||||
);
|
||||
}
|
||||
|
||||
// --- teams -----------------------------------------------------------------
|
||||
|
||||
function teamsCard() {
|
||||
const rows = data.teams.map((t) =>
|
||||
h('tr', {},
|
||||
h('td', {}, h('strong', { text: t.name })),
|
||||
h('td', { class: 'num', text: String(t.members) }),
|
||||
h('td', { class: 'num', text: String(t.open_incidents) }),
|
||||
h('td', {},
|
||||
h('button', {
|
||||
class: 'btn-sm',
|
||||
type: 'button',
|
||||
text: 'Rename',
|
||||
onclick: () => renameTeam(t),
|
||||
}),
|
||||
// A team with open incidents cannot be deleted, and saying so before
|
||||
// the click is kinder than a 409 afterwards.
|
||||
h('button', {
|
||||
class: 'btn-sm danger',
|
||||
type: 'button',
|
||||
text: 'Delete',
|
||||
disabled: t.open_incidents > 0,
|
||||
title: t.open_incidents > 0 ? 'Resolve its open incidents first' : '',
|
||||
onclick: () => deleteTeam(t),
|
||||
}),
|
||||
),
|
||||
));
|
||||
|
||||
return h('div', { class: 'card' },
|
||||
h('h2', { text: 'Teams' }),
|
||||
h('table', { class: 'admin-table' },
|
||||
h('thead', {}, h('tr', {},
|
||||
h('th', { text: 'Name' }),
|
||||
h('th', { class: 'num', text: 'Members' }),
|
||||
h('th', { class: 'num', text: 'Open' }),
|
||||
h('th', { text: '' }))),
|
||||
h('tbody', {}, rows)),
|
||||
newTeamForm(),
|
||||
);
|
||||
}
|
||||
|
||||
function newTeamForm() {
|
||||
const name = h('input', { name: 'name', type: 'text', placeholder: 'New team name', required: true });
|
||||
const form = h('form', { class: 'inline-form' }, name,
|
||||
h('button', { class: 'btn', type: 'submit', text: 'Create' }));
|
||||
form.addEventListener('submit', async (e) => {
|
||||
e.preventDefault();
|
||||
if (busy) return;
|
||||
busy = true;
|
||||
try {
|
||||
await api.createTeam(name.value.trim());
|
||||
name.value = '';
|
||||
await refresh();
|
||||
} catch (err) {
|
||||
error = err.message;
|
||||
render();
|
||||
} finally {
|
||||
busy = false;
|
||||
}
|
||||
});
|
||||
return form;
|
||||
}
|
||||
|
||||
async function renameTeam(team) {
|
||||
const next = window.prompt(`Rename ${team.name} to:`, team.name);
|
||||
if (!next || next === team.name) return;
|
||||
try {
|
||||
await api.renameTeam(team.id, next);
|
||||
} catch (err) {
|
||||
error = err.message;
|
||||
}
|
||||
refresh();
|
||||
}
|
||||
|
||||
async function deleteTeam(team) {
|
||||
if (!(await confirm({
|
||||
title: `Delete ${team.name}?`,
|
||||
text: 'Its alerts, incidents, schedule and integrations go with it. This cannot be undone.',
|
||||
confirmLabel: 'Delete',
|
||||
danger: true,
|
||||
}))) return;
|
||||
try {
|
||||
await api.deleteTeam(team.id);
|
||||
} catch (err) {
|
||||
error = err.message;
|
||||
}
|
||||
refresh();
|
||||
}
|
||||
|
||||
// --- users -----------------------------------------------------------------
|
||||
|
||||
function usersCard() {
|
||||
const rows = data.users.map((u) => {
|
||||
const self = u.id === myID();
|
||||
return h('tr', { class: u.disabled_at ? 'disabled-row' : '' },
|
||||
h('td', {},
|
||||
h('strong', { text: u.username }),
|
||||
u.disabled_at && h('span', { class: 'row-team', text: 'disabled' }),
|
||||
self && h('span', { class: 'you', text: 'you' })),
|
||||
h('td', { class: 'muted', text: u.email }),
|
||||
h('td', {}, u.is_admin ? h('span', { class: 'row-team', text: 'admin' }) : null),
|
||||
h('td', {},
|
||||
// Neither action is offered for your own account: the server refuses
|
||||
// both, and an enabled-looking button that always fails is worse than
|
||||
// no button.
|
||||
!self && h('button', {
|
||||
class: 'btn-sm',
|
||||
type: 'button',
|
||||
text: u.is_admin ? 'Revoke admin' : 'Make admin',
|
||||
onclick: () => setAdmin(u, !u.is_admin),
|
||||
}),
|
||||
!self && h('button', {
|
||||
class: 'btn-sm danger',
|
||||
type: 'button',
|
||||
text: u.disabled_at ? 'Enable' : 'Disable',
|
||||
onclick: () => setDisabled(u, !u.disabled_at),
|
||||
}),
|
||||
),
|
||||
);
|
||||
});
|
||||
|
||||
return h('div', { class: 'card' },
|
||||
h('h2', { text: 'Users' }),
|
||||
h('p', { class: 'muted small' },
|
||||
'Disabling an account stops it signing in and stops its API keys, and keeps ',
|
||||
'its acknowledgements and timeline entries. Deleting a user erases those.'),
|
||||
h('table', { class: 'admin-table' },
|
||||
h('thead', {}, h('tr', {},
|
||||
h('th', { text: 'User' }),
|
||||
h('th', { text: 'Email' }),
|
||||
h('th', { text: '' }),
|
||||
h('th', { text: '' }))),
|
||||
h('tbody', {}, rows)),
|
||||
);
|
||||
}
|
||||
|
||||
async function setAdmin(user, next) {
|
||||
if (next && !(await confirm({
|
||||
title: `Make ${user.username} an administrator?`,
|
||||
text: 'They will be able to create and delete users, and grant this to others.',
|
||||
confirmLabel: 'Make admin',
|
||||
}))) return;
|
||||
try {
|
||||
await api.setUserAdmin(user.id, next);
|
||||
} catch (err) {
|
||||
error = err.message;
|
||||
}
|
||||
refresh();
|
||||
}
|
||||
|
||||
async function setDisabled(user, next) {
|
||||
if (next && !(await confirm({
|
||||
title: `Disable ${user.username}?`,
|
||||
text: 'They cannot sign in and their API keys stop working. Their history stays.',
|
||||
confirmLabel: 'Disable',
|
||||
danger: true,
|
||||
}))) return;
|
||||
try {
|
||||
await api.setUserDisabled(user.id, next);
|
||||
} catch (err) {
|
||||
error = err.message;
|
||||
}
|
||||
refresh();
|
||||
}
|
||||
|
||||
// --- settings --------------------------------------------------------------
|
||||
|
||||
// Seconds are what the API speaks; people think in minutes and hours. The two
|
||||
// are converted here rather than in the server, which should keep exactly one
|
||||
// unit.
|
||||
const UNITS = [
|
||||
{ label: 'minutes', seconds: 60 },
|
||||
{ label: 'hours', seconds: 3600 },
|
||||
{ label: 'days', seconds: 86400 },
|
||||
];
|
||||
|
||||
function bestUnit(seconds) {
|
||||
for (const u of [...UNITS].reverse()) {
|
||||
if (seconds > 0 && seconds % u.seconds === 0) return u;
|
||||
}
|
||||
return UNITS[0];
|
||||
}
|
||||
|
||||
function settingsCard() {
|
||||
const editable = data.settings.editable || {};
|
||||
const inputs = new Map();
|
||||
|
||||
const rows = Object.entries(editable).map(([key, s]) => {
|
||||
const unit = bestUnit(s.seconds);
|
||||
const value = h('input', {
|
||||
type: 'number',
|
||||
min: '0',
|
||||
value: String(Math.round(s.seconds / unit.seconds)),
|
||||
class: 'setting-value',
|
||||
});
|
||||
const select = h('select', { class: 'setting-unit' },
|
||||
...UNITS.map((u) => h('option', {
|
||||
value: String(u.seconds),
|
||||
text: u.label,
|
||||
selected: u.seconds === unit.seconds,
|
||||
})));
|
||||
inputs.set(key, () => Number(value.value) * Number(select.value));
|
||||
|
||||
return h('tr', {},
|
||||
h('td', {}, h('strong', { text: key.replace(/_seconds$/, '').replace(/_/g, ' ') })),
|
||||
h('td', { class: 'muted small', text: s.description }),
|
||||
h('td', {}, value, select),
|
||||
);
|
||||
});
|
||||
|
||||
const form = h('form', { class: 'admin-settings' },
|
||||
h('table', { class: 'admin-table' }, h('tbody', {}, rows)),
|
||||
h('button', { class: 'btn', type: 'submit', text: 'Save settings' }));
|
||||
|
||||
form.addEventListener('submit', async (e) => {
|
||||
e.preventDefault();
|
||||
if (busy) return;
|
||||
busy = true;
|
||||
const body = {};
|
||||
for (const [key, read] of inputs) body[key] = read();
|
||||
try {
|
||||
await api.setAdminSettings(body);
|
||||
await refresh();
|
||||
} catch (err) {
|
||||
error = err.message;
|
||||
render();
|
||||
} finally {
|
||||
busy = false;
|
||||
}
|
||||
});
|
||||
|
||||
const env = Object.entries(data.settings.from_env || {}).map(([k, v]) =>
|
||||
h('tr', {},
|
||||
h('td', {}, h('code', { text: k })),
|
||||
h('td', { class: 'muted', text: v === '' ? '(unset)' : v })));
|
||||
|
||||
return h('div', { class: 'card' },
|
||||
h('h2', { text: 'Settings' }),
|
||||
h('p', { class: 'muted small', text: 'Saved changes take effect on the next sweep — no restart.' }),
|
||||
form,
|
||||
h('h3', { text: 'From the environment' }),
|
||||
h('p', { class: 'muted small' },
|
||||
'Where the server is plugged in, rather than how it behaves. These are set ',
|
||||
'in the deployment and are read-only here. Credentials are never shown.'),
|
||||
h('table', { class: 'admin-table' }, h('tbody', {}, env)),
|
||||
);
|
||||
}
|
||||
@@ -88,6 +88,19 @@ export const alerts = (query, opts) => call('GET', '/alerts', { query, ...opts }
|
||||
|
||||
// schedule
|
||||
export const teams = () => call('GET', '/teams');
|
||||
export const createTeam = (name) => call('POST', '/teams', { body: { name } });
|
||||
export const renameTeam = (id, name) => call('PUT', `/teams/${id}`, { body: { name } });
|
||||
export const deleteTeam = (id) => call('DELETE', `/teams/${id}`);
|
||||
|
||||
// Administration. Every one of these is refused with 403 for anybody without
|
||||
// the flag, so the UI hides the section rather than guarding it.
|
||||
export const adminTeams = () => call('GET', '/admin/teams');
|
||||
export const adminSettings = () => call('GET', '/admin/settings');
|
||||
export const setAdminSettings = (body) => call('PUT', '/admin/settings', { body });
|
||||
export const setUserAdmin = (id, isAdmin) =>
|
||||
call('PUT', `/users/${id}/admin`, { body: { is_admin: isAdmin } });
|
||||
export const setUserDisabled = (id, disabled) =>
|
||||
call('PUT', `/users/${id}/disabled`, { body: { disabled } });
|
||||
export const schedule = (teamID, from, to) =>
|
||||
call('GET', `/teams/${teamID}/schedule`, { query: { from, to } });
|
||||
|
||||
|
||||
@@ -9,6 +9,7 @@ import * as incident from './incident.js';
|
||||
import * as oncall from './oncall.js';
|
||||
import * as alerts from './alerts.js';
|
||||
import * as account from './account.js';
|
||||
import * as admin from './admin.js';
|
||||
|
||||
const $ = (id) => document.getElementById(id);
|
||||
|
||||
@@ -17,6 +18,7 @@ const SECTIONS = {
|
||||
queue: { title: 'Queue', view: queue },
|
||||
oncall: { title: 'On-call', view: oncall },
|
||||
alerts: { title: 'Alerts', view: alerts },
|
||||
admin: { title: 'Admin', view: admin },
|
||||
more: { title: 'Account', view: account },
|
||||
};
|
||||
|
||||
@@ -24,7 +26,7 @@ function parseRoute(pathname) {
|
||||
const m = pathname.match(/^\/incidents\/(\d+)\/?$/);
|
||||
if (m) return { section: 'queue', incident: Number(m[1]) };
|
||||
const name = pathname.replace(/^\/|\/$/g, '');
|
||||
if (name === 'oncall' || name === 'alerts' || name === 'more') return { section: name };
|
||||
if (name === 'oncall' || name === 'alerts' || name === 'admin' || name === 'more') return { section: name };
|
||||
return { section: 'queue', incident: null };
|
||||
}
|
||||
|
||||
@@ -142,6 +144,9 @@ async function boot() {
|
||||
try {
|
||||
state.me = await api.me();
|
||||
await loadTeams();
|
||||
// The Admin tab exists only for an administrator. Somebody who types /admin
|
||||
// anyway gets the view's own "ask an administrator" card, not a blank page.
|
||||
$('nav-admin').hidden = !state.me?.user?.is_admin;
|
||||
showApp();
|
||||
} catch (err) {
|
||||
if (err.status === 401) showLogin();
|
||||
|
||||
Reference in New Issue
Block a user