Compare commits

...

10 Commits

Author SHA1 Message Date
Niklas Ye 2b396d22d6 Set the chart's placeholder version to 0.26.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
CI / test (push) Successful in 2m56s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 22s
Release / image (push) Successful in 57s
Release / scan-image (push) Successful in 4s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as dc92f51 and e8d45f9, so a
tree heading for v0.26.0 doesn't say 0.25.0.
2026-09-26 08:28:24 +02:00
Niklas Ye 1f1faa437c Show the escalation ladder as a list, with who it would page and where it is
Team -> Escalation was the draft form on the page, which showed the
ladder only as inputs. It is now a table in the style of Switches and
Sources: a row per level with a status badge, who it pages, the wait
before the next level, and the open incidents currently waiting on it.
Below it, the repeat count, the fallback topic and when the ladder last
escalated (linking the incident). The editor moved into an "Edit ladder"
sheet, so a poll of the page underneath can no longer throw away half an
edit, and the page-level draft state went with it.

Targets are resolved to who they mean today, and the badge says what
would actually happen: Ready, Escalating (an unanswered incident has
climbed to level 2 or higher), or Pages nobody. The last is the one worth
seeing before an incident finds it: an empty rota, a person with no ntfy
topic or a disabled account each make a rung a silence with a number on
it, and the target says which. The rules are pageLevel's own, so the
page cannot promise a page the notifier would skip.

"Last escalated" comes from the escalated timeline events that already
exist, so there is no migration. Acknowledging or resolving takes an
incident off the ladder, so Escalating clears then while the history
stays.

API: GET /escalation gains status and waiting per level, username,
reachable and problem per target, and last_escalated_at and
last_escalated_incident_id. Output only and additive; PUT is unchanged
and terdut-tui needs nothing.
2026-09-26 08:28:24 +02:00
Niklas Ye dc92f51cf8 Set the chart's placeholder version to 0.25.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 31s
CI / test (push) Successful in 3m20s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 20s
Release / image (push) Successful in 1m2s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as e8d45f9 and 3ee8583, so a
tree heading for v0.25.0 doesn't say 0.24.0.
2026-09-26 07:52:07 +02:00
Niklas Ye d675f8ec9b List alert sources with their status and last arrival on Team -> Sources
Like Team -> Switches, the page is now a table: a status badge (Active
if the key posted within a day, Quiet if it has but not lately, Never
used), when it last posted a webhook, when an alert last arrived on it,
how many distinct alerts it refreshed in the last 24 hours, and when it
was created. Adding a source moved into a "New source" sheet, and owners
can rename one from its row.

"Last alert" and the count needed alerts to remember which source they
came in on, which they never did, so migration 010 adds
alerts.integration_id and every accepted payload stamps it. Last sender
wins when two sources post the same fingerprint. It is not backfilled: a
NULL says "before this was recorded" rather than guessing, and it heals
by itself as Alertmanager re-sends each alert every repeat_interval.
Revoking a source keeps its alerts, unattributed.

Last webhook and last alert are separate on purpose: a payload with
nothing usable in it stamps the first and not the second. The Quiet
threshold is a fixed day, a colour and not an alarm, since silence that
should page is what dead man's switches are for.

The counts are indexed subqueries (alerts_integration_idx) rather than a
join, which would read every alert a source ever delivered.

API: the integrations list gains status, last_alert_at and alerts_24h,
and PATCH /api/teams/{id}/integrations/{id} renames. Both are additive;
terdut-tui needs nothing.
2026-09-26 07:52:07 +02:00
Niklas Ye e8d45f9d3d Set the chart's placeholder version to 0.24.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 2m55s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 24s
Release / image (push) Successful in 1m0s
Release / scan-image (push) Successful in 4s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 3ee8583 and d2cdcc9, so a
tree heading for v0.24.0 doesn't say 0.23.0.
2026-09-25 23:40:55 +02:00
Niklas Ye f3918b863c List dead man's switches with their status on Team -> Switches
The page was a bare form: it did not say which switches existed or
whether they were alive. It now lists them, each with a Healthy, Dead or
Dormant badge, when its heartbeat was last heard and when it last opened
an incident (linked while that incident is open). A matcher that several
clusters satisfy is broken down per cluster, since a live cluster must
not hide a dead one. The form moved into a "New switch" sheet, and each
row has a Remove with a confirm.

That needed a switch to be a thing, so switches are rows now
(migration 009) with their own name, matcher, timeout and severity,
instead of one string with one team-wide timeout in deadman_configs.
Existing configuration is split into one row per matcher; a team whose
timeout was zero simply has none. The sweeper and the status endpoint
share one death rule (deadmanAlert.dead), so the page cannot disagree
with the pager. Incident group keys are unchanged, so incidents that
are open across the upgrade keep working.

The environment defaults (TERDUT_DEADMAN_*) are seeded into teams once
per install, recorded in settings, so a team that deletes its last
switch does not get it back on the next restart. Installs that already
had per-team rows are marked as seeded by the migration.

Removing a switch stops the watching but leaves an incident it already
opened open until someone resolves it.

API: GET/PUT /api/teams/{id}/deadman are replaced by
GET/POST /deadman/switches and DELETE /deadman/switches/{switchID}.
terdut-tui does not call them, so nothing to mirror there.
2026-09-25 23:40:55 +02:00
Niklas Ye 3ee8583f6f Set the chart's placeholder version to 0.23.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 14s
CI / test (push) Successful in 2m35s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 21s
Release / image (push) Successful in 57s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as d2cdcc9 and 71d7e18, so a
tree heading for v0.23.0 doesn't say 0.22.1.
2026-09-25 19:12:33 +02:00
Niklas Ye 591d5b8df0 Copy an incident to the clipboard as Markdown
A button in the incident header (also `y`, and "Copy incident" in the
more menu) puts everything the page knows on the clipboard, for pasting
into a chat or an agent prompt with no integration involved.

The text carries the facts, every alert with all its labels and
annotations (the page only shows summary or description), the timeline
with notes in full, and the "Seen before" resolution notes. Times are
ISO 8601 and users are named rather than "you", since relative and
first-person wording is ambiguous once pasted elsewhere.

The async clipboard API needs a secure context and this server is often
reached over plain HTTP, so it falls back to execCommand.

Web UI only: no endpoint or JSON shape changed, so nothing to mirror in
terdut-tui.
2026-09-25 19:12:33 +02:00
Niklas Ye d2cdcc9776 Set the chart's placeholder version to 0.22.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 2m45s
Release / test (push) Successful in 4s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 21s
Release / image (push) Successful in 59s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package sets the published version and appVersion
from the tag, so these two fields decide nothing (see the comment
above them). Kept in step anyway, same as 71d7e18 and 734cd9c, so a
tree heading for v0.22.1 doesn't say 0.22.0.

Claude-Session: https://claude.ai/code/session_01MMados3BD1oSjevHxbmVqU
2026-09-25 17:25:17 +02:00
Niklas Ye 8b2789b9b2 Let the filter chips wrap in the desktop incident list
The list pane is 340-420px wide and its chip row scrolled sideways with the
scrollbar hidden. That works by swipe on a phone, but a mouse has nothing to
grab, so Archived (the last chip) could not be reached on a wide screen. In
the desktop layout the row now wraps instead, and the divider between the
status and team chips is hidden there, since it would sit mid-line.

Phones keep the sideways scroll: the rule is inside the min-width: 900px
block.

Claude-Session: https://claude.ai/code/session_01MMados3BD1oSjevHxbmVqU
2026-09-25 17:25:17 +02:00
19 changed files with 1887 additions and 445 deletions
+27 -17
View File
@@ -513,19 +513,27 @@ nothing unless something downstream notices it stop. That is what
`TERDUT_DEADMAN_MATCHERS` defaults to.
**Switches belong to a team**, which decides which of its own alerts are
heartbeats and how long a silence has to last. An owner sets them through
`PUT /api/teams/{teamID}/deadman`; a missed heartbeat opens an incident in the
team whose integration received it.
heartbeats and how long a silence has to last. Each **switch** is a row of its
own — a name, one matcher, a timeout and a severity — so switches in one team
can have different deadlines. An owner adds and removes them on **Team →
Switches**, which lists each with a status (**healthy**, **dead**, or
**dormant** until its first heartbeat), when it was last heard from, and when it
last opened an incident; a matcher that several clusters satisfy is broken down
per cluster. The API is `POST`/`DELETE /api/teams/{teamID}/deadman/switches`. A
missed heartbeat opens an incident in the team whose integration received it.
Removing a switch stops the watching; an incident it already opened stays open
until somebody resolves it.
The environment variables are the starting point, not the setting: at startup
every team **without** a configuration of its own is given one from them, and an
owner's later edit is never overwritten by a redeploy. A team created after
that starts watching nothing until its owner says otherwise — inheriting an
install-wide heartbeat would page a new team about a source it has never heard
of.
The environment variables are the starting point, not the setting: the **first**
time the server starts, every team is given a switch per default matcher from
them, once. After that a team's switches are its own — an owner's edit or
deletion is never put back by a redeploy. A team created later starts watching
nothing until its owner says otherwise — inheriting an install-wide heartbeat
would page a new team about a source it has never heard of.
A matcher is a set of exact label conditions, one of which must be the
`alertname`, in the same format the environment variable uses:
`alertname`, in the format the environment variable uses (one matcher per switch; the
variable takes several, separated by `;`):
```
alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat
@@ -710,16 +718,18 @@ administrator who is not in the team gets the same `404` as anybody else.
| `GET` | `/api/teams/{teamID}/members` | member | Who is in the team |
| `POST` | `/api/teams/{teamID}/members` | **owner** | Add a member, or change their role `{"user_id","role"}` |
| `DELETE` | `/api/teams/{teamID}/members/{userID}` | **owner** | Remove a member. `409` for the last owner |
| `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys |
| `GET` | `/api/teams/{teamID}/integrations` | member | List integrations. Never returns keys. Each carries `status` (`active` if its key posted within 24h, `quiet` if it has but not lately, `never`), `last_used_at` (last webhook, usable or not), `last_alert_at` (when an alert last arrived on it) and `alerts_24h` (distinct alerts it refreshed in the last day). Alerts delivered before the source was recorded (migration 010) have none, so the last two fill in as Alertmanager re-sends them |
| `PATCH` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Rename `{"name"}`. The key does not change |
| `POST` | `/api/teams/{teamID}/integrations` | **owner** | Mint an integration `{"name","kind"}` — key and URL shown once |
| `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration |
| `DELETE` | `/api/teams/{teamID}/integrations/{integrationID}` | **owner** | Revoke an integration. Alerts it delivered stay, unattributed |
| `GET` | `/api/teams/{teamID}/invites` | **owner** | The team's invite links, with their uses and expiry. Never the tokens |
| `POST` | `/api/teams/{teamID}/invites` | **owner** | Mint one `{"role","max_uses"}` — the full URL is returned once |
| `DELETE` | `/api/teams/{teamID}/invites/{inviteID}` | **owner** | Revoke a link before it expires |
| `GET` | `/api/teams/{teamID}/escalation` | member | The team's [escalation ladder](#escalation) `{repeat_count, fallback_topic, levels[]}`. Empty levels means the team has none |
| `GET` | `/api/teams/{teamID}/escalation` | member | The team's [escalation ladder](#escalation) `{repeat_count, fallback_topic, levels[], last_escalated_at?, last_escalated_incident_id?}`. Empty levels means the team has none. Each level also carries `status` (`ready`, `escalating` when an unanswered incident has climbed to it, `unreachable` when nobody on it could be woken), `waiting` (ids of the open incidents on it) and, per target, `username` (who it means today — the person on call, for a rota target), `reachable` and `problem`. The extra fields are output only; `PUT` takes the plain shape |
| `PUT` | `/api/teams/{teamID}/escalation` | **owner** | Replace it wholesale. `400` for a level with no targets or no timeout — a rung that pages nobody is a silence with a number on it |
| `GET` | `/api/teams/{teamID}/deadman` | member | The team's [dead man's switch](#dead-mans-switch) configuration `{matchers, timeout_seconds, severity}` |
| `PUT` | `/api/teams/{teamID}/deadman` | **owner** | Replace it. `400` when no matcher names an `alertname`, because a switch that silently watches nothing is the failure this feature exists to prevent |
| `GET` | `/api/teams/{teamID}/deadman/switches` | member | The team's [dead man's switches](#dead-mans-switch), each `{id, name, matcher, timeout_seconds, severity, status, last_heartbeat_at, last_triggered_at, open_incident_id, sources[]}`. `status` is `healthy`, `dead` or `dormant`; `sources` has one entry per heartbeat fingerprint. Empty when the team watches nothing |
| `POST` | `/api/teams/{teamID}/deadman/switches` | **owner** | Add one: `{name?, matcher, timeout_seconds, severity?}`. `400` when the matcher names no `alertname` or holds several, or the timeout is not positive — a switch that silently watches nothing is the failure this feature exists to prevent |
| `DELETE` | `/api/teams/{teamID}/deadman/switches/{switchID}` | **owner** | Stop watching. An incident it opened stays open. `404` for a switch of another team |
### Notifications
@@ -964,8 +974,8 @@ What changes, and will need attention:
**Dead man's switches moved too.** `TERDUT_DEADMAN_MATCHERS`, `_TIMEOUT` and
`_SEVERITY` are no longer the setting; they are the default each existing team
is seeded with at startup, after which an owner edits them per team through
`PUT /api/teams/{teamID}/deadman` and a redeploy never overwrites that.
is seeded with at startup, after which an owner manages them per team through
`/api/teams/{teamID}/deadman/switches` and a redeploy never overwrites that.
Nothing else about an incident changes, and incidents never move between teams:
an alert belongs to whichever team's key it arrived on.
+2 -2
View File
@@ -15,5 +15,5 @@ type: application
# appVersion and image.tag in values.yaml no longer agree, and that is not an oversight:
# image.tag stays "latest", which is what a local install actually pulls. appVersion is
# metadata and drives nothing.
version: 0.22.0
appVersion: "v0.22.0"
version: 0.26.0
appVersion: "v0.26.0"
+16 -11
View File
@@ -78,7 +78,7 @@ type ingested struct {
// post, and which team the alerts belong to.
func handleIntegrationWebhook(db *sql.DB, notify NotifyConfig) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, err := teamIDForKey(r.Context(), db, chi.URLParam(r, "key"))
src, err := sourceForKey(r.Context(), db, chi.URLParam(r, "key"))
if err != nil {
if errors.Is(err, errUnknownIntegration) {
// 401 and not 404: the path is real, the key is not, and a
@@ -90,11 +90,12 @@ func handleIntegrationWebhook(db *sql.DB, notify NotifyConfig) http.HandlerFunc
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
receiveWebhook(w, r, db, notify, teamID)
receiveWebhook(w, r, db, notify, src)
}
}
func receiveWebhook(w http.ResponseWriter, r *http.Request, db *sql.DB, notify NotifyConfig, teamID int64) {
func receiveWebhook(w http.ResponseWriter, r *http.Request, db *sql.DB, notify NotifyConfig, src alertSource) {
teamID := src.teamID
var payload amPayload
if err := decodeJSON(r, &payload); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid payload"))
@@ -104,7 +105,7 @@ func receiveWebhook(w http.ResponseWriter, r *http.Request, db *sql.DB, notify N
// Alertmanager retries anything that is not 2xx, and a retry of a payload
// we failed to store is more useful than an error it cannot act on — so
// failures are logged, not surfaced.
if err := ingest(r.Context(), db, notify, teamID, payload); err != nil {
if err := ingest(r.Context(), db, notify, src, payload); err != nil {
log.Printf("webhook ingest (team %d, group %q): %v", teamID, payload.GroupKey, err)
}
@@ -114,7 +115,8 @@ func receiveWebhook(w http.ResponseWriter, r *http.Request, db *sql.DB, notify N
// ingest stores a payload's alerts and reconciles the incident for its group.
// The whole payload is one transaction: an incident that opened but whose alerts
// failed to link would be a work item nobody could act on.
func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, teamID int64, payload amPayload) error {
func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, src alertSource, payload amPayload) error {
teamID := src.teamID
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return err
@@ -124,12 +126,12 @@ func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, teamID int64,
// Which arriving alerts are heartbeats is the team's own answer, read
// inside the transaction so an owner editing it mid-payload cannot split
// one webhook across two interpretations.
deadman, err := deadmanConfigForTeam(ctx, tx, teamID)
deadman, err := deadmanSetForTeam(ctx, tx, teamID)
if err != nil {
return err
}
accepted, err := upsertAlerts(ctx, tx, deadman, teamID, payload.Alerts)
accepted, err := upsertAlerts(ctx, tx, deadman, src, payload.Alerts)
if err != nil {
return err
}
@@ -186,7 +188,8 @@ func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, teamID int64,
// upsertAlerts stores each alert of a payload and reports what changed. Payloads
// the ordering guard rejected are left out entirely.
func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, teamID int64, alerts []amAlert) ([]ingested, error) {
func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman deadmanSet, src alertSource, alerts []amAlert) ([]ingested, error) {
teamID := src.teamID
now := time.Now().Unix()
accepted := make([]ingested, 0, len(alerts))
@@ -242,8 +245,8 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, teamID
if _, err := tx.ExecContext(ctx, `
INSERT INTO alerts
(team_id, fingerprint, name, status, labels, annotations, starts_at, ends_at,
generator_url, received_at, resolution_source)
VALUES ($1, $2, $3, $4, $5::jsonb, $6::jsonb, $7, $8, $9, $10, $11)
generator_url, received_at, resolution_source, integration_id)
VALUES ($1, $2, $3, $4, $5::jsonb, $6::jsonb, $7, $8, $9, $10, $11, $12)
ON CONFLICT (team_id, fingerprint) DO UPDATE SET
status = excluded.status,
labels = excluded.labels,
@@ -257,6 +260,8 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, teamID
-- is a breaking API change — see models.Alert.ReceivedAt.
received_at = excluded.received_at,
resolution_source = excluded.resolution_source,
-- Last sender wins; see migration 010.
integration_id = excluded.integration_id,
-- A re-fire makes the alert current again, so it leaves the archive.
archived_at = CASE WHEN excluded.status = 'firing'
THEN NULL ELSE alerts.archived_at END
@@ -267,7 +272,7 @@ func upsertAlerts(ctx context.Context, tx *sql.Tx, deadman DeadmanConfig, teamID
teamID, a.Fingerprint, name, a.Status,
string(labelsJSON), string(annotationsJSON),
a.StartsAt.Unix(), endsAtUnix,
a.GeneratorURL, now, resolutionSource,
a.GeneratorURL, now, resolutionSource, src.integrationID,
); err != nil {
return nil, err
}
+11 -13
View File
@@ -88,27 +88,25 @@ func newDeadmanTS(t *testing.T, deadman api.DeadmanConfig, notify ...api.NotifyC
return s
}
// setTeamDeadman configures the default team's switches over the API, rendering
// the matchers back into the string form the endpoint takes.
// setTeamDeadman gives the default team one switch per configured matcher, over
// the API, the way an owner would add them.
func setTeamDeadman(t *testing.T, s *ts, cfg api.DeadmanConfig) {
t.Helper()
matchers := make([]string, 0, len(cfg.Matchers))
for _, m := range cfg.Matchers {
parts := []string{"alertname=" + m.Name}
for k, v := range m.Labels {
parts = append(parts, k+"="+v)
}
sort.Strings(parts[1:])
matchers = append(matchers, strings.Join(parts, ","))
}
resp := s.req(t, http.MethodPut, "/api/teams/"+defaultTeam+"/deadman", map[string]any{
"matchers": strings.Join(matchers, "; "),
"timeout_seconds": int64(cfg.Timeout.Seconds()),
"severity": cfg.Severity,
})
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("configure the team's dead man's switches: %d", resp.StatusCode)
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/deadman/switches", map[string]any{
"matcher": strings.Join(parts, ","),
"timeout_seconds": int64(cfg.Timeout.Seconds()),
"severity": cfg.Severity,
})
resp.Body.Close()
if resp.StatusCode != http.StatusCreated {
t.Fatalf("add a dead man's switch: %d", resp.StatusCode)
}
}
}
+354 -174
View File
@@ -4,6 +4,8 @@ import (
"context"
"database/sql"
"encoding/json"
"errors"
"fmt"
"log"
"sort"
"strings"
@@ -39,6 +41,17 @@ func (m DeadmanMatcher) String() string {
return m.Name + " (" + strings.Join(parts, ", ") + ")"
}
// config renders the matcher in the form parseDeadmanMatcher reads, which is
// what a switch row stores: `alertname=Watchdog,cluster=prod`.
func (m DeadmanMatcher) config() string {
parts := make([]string, 0, len(m.Labels))
for k, v := range m.Labels {
parts = append(parts, k+"="+v)
}
sort.Strings(parts)
return strings.Join(append([]string{"alertname=" + m.Name}, parts...), ",")
}
// matches reports whether an alert's labels satisfy every condition.
func (m DeadmanMatcher) matches(labels map[string]string) bool {
if labels["alertname"] != m.Name {
@@ -52,12 +65,10 @@ func (m DeadmanMatcher) matches(labels map[string]string) bool {
return true
}
// DeadmanConfig inverts the handling of the alerts it matches: receiving one
// opens nothing, and the absence of one opens an incident.
//
// The unit of monitoring is the fingerprint, not the matcher — two clusters
// sending the same heartbeat alertname are two independent switches, so one
// healthy cluster cannot mask a dead one.
// DeadmanConfig is the server-wide default a team's switches are seeded from:
// the environment's matchers, timeout and severity. Switches themselves are rows
// of a team's own — see DeadmanSwitch — and this is only how a fresh install
// starts out.
type DeadmanConfig struct {
Matchers []DeadmanMatcher
@@ -76,41 +87,81 @@ type DeadmanConfig struct {
// enabled reports whether there is anything to watch.
func (c DeadmanConfig) enabled() bool { return c.Timeout > 0 && len(c.Matchers) > 0 }
// match returns the first matcher an alert satisfies.
func (c DeadmanConfig) match(labels map[string]string) (DeadmanMatcher, bool) {
if !c.enabled() {
return DeadmanMatcher{}, false
}
for _, m := range c.Matchers {
if m.matches(labels) {
return m, true
}
}
return DeadmanMatcher{}, false
// DeadmanSwitch inverts the handling of the alerts it matches: receiving one
// opens nothing, and the absence of one opens an incident.
//
// The unit of monitoring is the fingerprint, not the switch — two clusters
// sending the same heartbeat alertname are two independent heartbeats under one
// switch, so one healthy cluster cannot mask a dead one.
type DeadmanSwitch struct {
ID int64
Name string
Matcher DeadmanMatcher
// Timeout is how long a heartbeat may go unheard before it is declared dead.
Timeout time.Duration
// Severity is what the incident opens at.
Severity string
}
// isDeadman is match without the matcher, for the ingest path.
func (c DeadmanConfig) isDeadman(labels map[string]string) bool {
_, ok := c.match(labels)
// deadmanSet is one team's switches.
type deadmanSet []DeadmanSwitch
// match returns the first switch an alert satisfies.
func (d deadmanSet) match(labels map[string]string) (DeadmanSwitch, bool) {
for _, sw := range d {
if sw.Matcher.matches(labels) {
return sw, true
}
}
return DeadmanSwitch{}, false
}
// isDeadman is match without the switch, for the ingest path.
func (d deadmanSet) isDeadman(labels map[string]string) bool {
_, ok := d.match(labels)
return ok
}
// names lists the distinct alertnames worth loading from the database.
func (c DeadmanConfig) names() []string {
func (d deadmanSet) names() []string {
seen := map[string]bool{}
out := make([]string, 0, len(c.Matchers))
for _, m := range c.Matchers {
if !seen[m.Name] {
seen[m.Name] = true
out = append(out, m.Name)
out := make([]string, 0, len(d))
for _, sw := range d {
if !seen[sw.Matcher.Name] {
seen[sw.Matcher.Name] = true
out = append(out, sw.Matcher.Name)
}
}
return out
}
// parseDeadmanMatcher reads one matcher from its configured form: "," separates
// the conditions and "=" is exact label equality — `alertname=Watchdog,cluster=prod`.
// The error says what is wrong with it, in words a form can show.
func parseDeadmanMatcher(entry string) (DeadmanMatcher, error) {
m := DeadmanMatcher{Labels: map[string]string{}}
for _, cond := range strings.Split(strings.TrimSpace(entry), ",") {
k, v, ok := strings.Cut(cond, "=")
k, v = strings.TrimSpace(k), strings.TrimSpace(v)
if !ok || k == "" || v == "" {
return DeadmanMatcher{}, fmt.Errorf("%q is not label=value", strings.TrimSpace(cond))
}
if k == "alertname" {
m.Name = v
continue
}
m.Labels[k] = v
}
if m.Name == "" {
return DeadmanMatcher{}, errors.New("no alertname condition")
}
return m, nil
}
// ParseDeadmanConfig reads the matcher list from its configured form:
// ";" separates matchers, "," separates the conditions within one, and "=" is
// exact label equality — `alertname=Watchdog,cluster=prod; alertname=Heartbeat`.
// ";" separates matchers, and each is parsed as parseDeadmanMatcher does.
//
// A malformed or alertname-less entry is dropped rather than fatal, following
// config.duration's rule that one bad tuning knob should not take the server
@@ -125,28 +176,9 @@ func ParseDeadmanConfig(matchers string, timeout time.Duration, severity string)
if entry == "" {
continue
}
m := DeadmanMatcher{Labels: map[string]string{}}
malformed := false
for _, cond := range strings.Split(entry, ",") {
k, v, ok := strings.Cut(cond, "=")
k, v = strings.TrimSpace(k), strings.TrimSpace(v)
if !ok || k == "" || v == "" {
log.Printf("deadman: ignoring matcher %q: %q is not label=value", entry, strings.TrimSpace(cond))
malformed = true
break
}
if k == "alertname" {
m.Name = v
continue
}
m.Labels[k] = v
}
if malformed {
continue
}
if m.Name == "" {
log.Printf("deadman: ignoring matcher %q: no alertname condition", entry)
m, err := parseDeadmanMatcher(entry)
if err != nil {
log.Printf("deadman: ignoring matcher %q: %v", entry, err)
continue
}
cfg.Matchers = append(cfg.Matchers, m)
@@ -162,23 +194,35 @@ func ParseDeadmanConfig(matchers string, timeout time.Duration, severity string)
for _, m := range cfg.Matchers {
rendered = append(rendered, m.String())
}
log.Printf("deadman: watching %s, timeout %s, severity %s",
log.Printf("deadman: default for new teams: %s, timeout %s, severity %s",
strings.Join(rendered, "; "), timeout, severity)
}
return cfg
}
// deadmanAlert is one switch: the alert row carrying its last heartbeat.
// deadmanAlert is one heartbeat: the alert row carrying its last sighting, and
// the switch that claimed it.
type deadmanAlert struct {
id int64
teamID int64
fingerprint string
labels map[string]string
matcher DeadmanMatcher
sw DeadmanSwitch
resolved bool
receivedAt int64
}
// dead is the one rule for a silent heartbeat, shared by the sweeper that pages
// on it and the status the Switches page shows, so the page cannot disagree
// with the pager.
//
// An explicit resolved from Alertmanager is a stronger death signal than mere
// absence: the sender is telling us the heartbeat stopped, so there is nothing
// left to wait out.
func (a deadmanAlert) dead(now time.Time) bool {
return a.resolved || a.receivedAt < now.Add(-a.sw.Timeout).Unix()
}
// groupKey is the switch's identity as an incident. Per fingerprint, so each
// source is tracked on its own.
func (a deadmanAlert) groupKey() string { return deadmanGroupPrefix + a.fingerprint }
@@ -189,13 +233,13 @@ func (a deadmanAlert) groupKey() string { return deadmanGroupPrefix + a.fingerpr
// It returns the ids of the alerts it owns, because the generic staleness
// expiry must leave them alone — staleAfter and ends_at would otherwise resolve
// a heartbeat long before its own, much tighter, timeout ever fired.
// Each team is swept against its own configuration: its own matchers, its own
// timeout, its own severity. A team watching nothing is skipped entirely, which
// is most of them.
// Each team is swept against its own switches, each with its own matcher,
// timeout and severity. A team watching nothing is skipped entirely, which is
// most of them.
func sweepDeadman(ctx context.Context, db *sql.DB, notify NotifyConfig) map[int64]bool {
owned := map[int64]bool{}
configs, err := deadmanConfigs(ctx, db)
configs, err := deadmanSets(ctx, db)
if err != nil {
log.Printf("deadman: load configs: %v", err)
return owned
@@ -203,39 +247,35 @@ func sweepDeadman(ctx context.Context, db *sql.DB, notify NotifyConfig) map[int6
now := time.Now()
for teamID, cfg := range configs {
switches, err := deadmanAlerts(ctx, db, teamID, cfg)
heartbeats, err := deadmanAlerts(ctx, db, teamID, cfg)
if err != nil {
log.Printf("deadman: load switches for team %d: %v", teamID, err)
log.Printf("deadman: load heartbeats for team %d: %v", teamID, err)
continue
}
cutoff := now.Add(-cfg.Timeout).Unix()
for _, sw := range switches {
owned[sw.id] = true
for _, hb := range heartbeats {
owned[hb.id] = true
// An explicit resolved from Alertmanager is a stronger death signal
// than mere absence: the sender is telling us the heartbeat
// stopped, so there is nothing left to wait out.
if sw.resolved || sw.receivedAt < cutoff {
if err := deadmanDied(ctx, db, cfg, notify, sw, now); err != nil {
log.Printf("deadman: open incident for %s: %v", sw.matcher.Name, err)
if hb.dead(now) {
if err := deadmanDied(ctx, db, notify, hb, now); err != nil {
log.Printf("deadman: open incident for %s: %v", hb.sw.Matcher.Name, err)
}
continue
}
if err := deadmanRecovered(ctx, db, sw); err != nil {
log.Printf("deadman: resolve incident for %s: %v", sw.matcher.Name, err)
if err := deadmanRecovered(ctx, db, hb); err != nil {
log.Printf("deadman: resolve incident for %s: %v", hb.sw.Matcher.Name, err)
}
}
}
return owned
}
// deadmanAlerts loads every alert row that a matcher claims. The candidate query
// deadmanAlerts loads every alert row that one of a team's switches claims. The candidate query
// is narrowed by alertname so it rides alerts_name_idx; the rest of the matching
// happens in Go, which keeps one implementation of the rules. The rows are read
// in full before the caller writes, so the writes do not run against an open
// cursor over the same table.
func deadmanAlerts(ctx context.Context, db *sql.DB, teamID int64, cfg DeadmanConfig) ([]deadmanAlert, error) {
func deadmanAlerts(ctx context.Context, db *sql.DB, teamID int64, cfg deadmanSet) ([]deadmanAlert, error) {
names := cfg.names()
args := &sqlArgs{}
nameList := make([]any, len(names))
@@ -263,11 +303,11 @@ func deadmanAlerts(ctx context.Context, db *sql.DB, teamID int64, cfg DeadmanCon
}
json.Unmarshal([]byte(labelsJSON), &a.labels) //nolint:errcheck
m, ok := cfg.match(a.labels)
sw, ok := cfg.match(a.labels)
if !ok {
continue
}
a.matcher = m
a.sw = sw
a.resolved = status == "resolved"
out = append(out, a)
}
@@ -284,16 +324,16 @@ func deadmanAlerts(ctx context.Context, db *sql.DB, teamID int64, cfg DeadmanCon
// incidentForGroup), and a source that is gone for good is a one-time page
// rather than a nag. Only a heartbeat that comes back and dies again earns a new
// incident.
func deadmanDied(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify NotifyConfig, sw deadmanAlert, now time.Time) error {
func deadmanDied(ctx context.Context, db *sql.DB, notify NotifyConfig, hb deadmanAlert, now time.Time) error {
var lastTriggered, open int64
if err := db.QueryRowContext(ctx, `
SELECT COALESCE(MAX(triggered_at), 0),
COUNT(*) FILTER (WHERE resolved_at IS NULL)
FROM incidents WHERE team_id = $1 AND group_key = $2`,
sw.teamID, sw.groupKey()).Scan(&lastTriggered, &open); err != nil {
hb.teamID, hb.groupKey()).Scan(&lastTriggered, &open); err != nil {
return err
}
if open > 0 || sw.receivedAt <= lastTriggered {
if open > 0 || hb.receivedAt <= lastTriggered {
return nil
}
@@ -306,18 +346,18 @@ func deadmanDied(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify Noti
// A heartbeat nobody has heard from is not firing, and saying otherwise in
// the alert list would be a lie. An Alertmanager-sourced resolution keeps its
// own source: it told us the truth first.
if !sw.resolved {
if !hb.resolved {
if _, err := tx.ExecContext(ctx, `
UPDATE alerts
SET status = 'resolved',
resolution_source = $1,
ends_at = COALESCE(ends_at, `+nowEpoch+`)
WHERE id = $2 AND status = 'firing'`, resolutionDeadman, sw.id); err != nil {
WHERE id = $2 AND status = 'firing'`, resolutionDeadman, hb.id); err != nil {
return err
}
}
severity := cfg.Severity
severity := hb.sw.Severity
var sev *string
if severity != "" {
sev = &severity
@@ -325,14 +365,14 @@ func deadmanDied(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify Noti
// The incident opens in the team whose integration received the heartbeat:
// the switch belongs to whoever is watching that source, not to the install.
incidentID, err := openIncident(ctx, tx, notify, sw.teamID, sw.groupKey(),
"No heartbeat from "+sw.matcher.String(), sw.labels, sev)
incidentID, err := openIncident(ctx, tx, notify, hb.teamID, hb.groupKey(),
"No heartbeat from "+hb.sw.Matcher.String(), hb.labels, sev)
if err != nil {
return err
}
alertID := sw.id
detail := "last heartbeat " + humanDuration(now.Sub(time.Unix(sw.receivedAt, 0))) + " ago"
alertID := hb.id
detail := "last heartbeat " + humanDuration(now.Sub(time.Unix(hb.receivedAt, 0))) + " ago"
if err := logEvent(ctx, tx, incidentID, evDeadmanSilent, nil, &alertID, &detail); err != nil {
return err
}
@@ -340,7 +380,7 @@ func deadmanDied(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify Noti
if err := tx.Commit(); err != nil {
return err
}
log.Printf("deadman: %s went silent, opened incident %d", sw.matcher.String(), incidentID)
log.Printf("deadman: %s went silent, opened incident %d", hb.sw.Matcher.String(), incidentID)
return nil
}
@@ -350,12 +390,12 @@ func deadmanDied(ctx context.Context, db *sql.DB, cfg DeadmanConfig, notify Noti
// member alerts (linking the heartbeat would have the settled-incident cascade
// close it on the very same sweep that opened it), so the alert-driven cascade
// ignores it entirely and recovery is the only automatic way out.
func deadmanRecovered(ctx context.Context, db *sql.DB, sw deadmanAlert) error {
func deadmanRecovered(ctx context.Context, db *sql.DB, hb deadmanAlert) error {
var incidentID int64
switch err := db.QueryRowContext(ctx, `
SELECT id FROM incidents
WHERE team_id = $1 AND group_key = $2 AND resolved_at IS NULL`,
sw.teamID, sw.groupKey()).Scan(&incidentID); {
hb.teamID, hb.groupKey()).Scan(&incidentID); {
case err == sql.ErrNoRows:
return nil
case err != nil:
@@ -387,116 +427,256 @@ func deadmanRecovered(ctx context.Context, db *sql.DB, sw deadmanAlert) error {
if err := tx.Commit(); err != nil {
return err
}
log.Printf("deadman: %s is back, resolved incident %d", sw.matcher.String(), incidentID)
log.Printf("deadman: %s is back, resolved incident %d", hb.sw.Matcher.String(), incidentID)
return nil
}
// ---------------------------------------------------------------------------
// Per-team configuration
// A team's switches
// ---------------------------------------------------------------------------
// deadmanConfigForTeam reads one team's switches. A team with no row, or with
// nothing configured, gets a disabled config — which is the right answer rather
// than an error: most teams watch no heartbeat at all.
func deadmanConfigForTeam(ctx context.Context, q querier, teamID int64) (DeadmanConfig, error) {
var matchers, severity string
var timeout int64
err := q.QueryRowContext(ctx,
"SELECT matchers, timeout_seconds, severity FROM deadman_configs WHERE team_id = $1",
teamID).Scan(&matchers, &timeout, &severity)
if err == sql.ErrNoRows {
return DeadmanConfig{}, nil
}
if err != nil {
return DeadmanConfig{}, err
}
return parseDeadmanQuietly(matchers, time.Duration(timeout)*time.Second, severity), nil
}
const deadmanSwitchColumns = "id, team_id, name, matcher, timeout_seconds, severity"
// deadmanConfigs reads every team's switches in one query, for the sweeper.
func deadmanConfigs(ctx context.Context, db *sql.DB) (map[int64]DeadmanConfig, error) {
rows, err := db.QueryContext(ctx,
"SELECT team_id, matchers, timeout_seconds, severity FROM deadman_configs")
if err != nil {
return nil, err
}
// scanDeadmanSwitches reads switch rows into per-team sets. A row whose matcher
// no longer parses is skipped rather than fatal: the API refuses to store one,
// so it can only mean a hand edit, and one bad row must not stop the others
// from being watched.
func scanDeadmanSwitches(rows *sql.Rows) (map[int64]deadmanSet, error) {
defer rows.Close()
out := map[int64]DeadmanConfig{}
out := map[int64]deadmanSet{}
for rows.Next() {
var sw DeadmanSwitch
var teamID, timeout int64
var matchers, severity string
if err := rows.Scan(&teamID, &matchers, &timeout, &severity); err != nil {
var matcher string
if err := rows.Scan(&sw.ID, &teamID, &sw.Name, &matcher, &timeout, &sw.Severity); err != nil {
return nil, err
}
cfg := parseDeadmanQuietly(matchers, time.Duration(timeout)*time.Second, severity)
if cfg.enabled() {
out[teamID] = cfg
m, err := parseDeadmanMatcher(matcher)
if err != nil {
log.Printf("deadman: switch %d has an unusable matcher %q: %v", sw.ID, matcher, err)
continue
}
sw.Matcher = m
sw.Timeout = time.Duration(timeout) * time.Second
out[teamID] = append(out[teamID], sw)
}
return out, rows.Err()
}
// SeedDeadmanConfigs gives every team without a row the server's environment
// configuration, so the install that upgrades into per-team switches keeps
// watching exactly what it was watching before.
// deadmanSetForTeam reads one team's switches. A team with none gets an empty
// set — which is the right answer rather than an error: most teams watch no
// heartbeat at all.
func deadmanSetForTeam(ctx context.Context, q querier, teamID int64) (deadmanSet, error) {
rows, err := q.QueryContext(ctx,
"SELECT "+deadmanSwitchColumns+" FROM deadman_switches WHERE team_id = $1 ORDER BY id", teamID)
if err != nil {
return nil, err
}
sets, err := scanDeadmanSwitches(rows)
return sets[teamID], err
}
// deadmanSets reads every team's switches in one query, for the sweeper.
func deadmanSets(ctx context.Context, db *sql.DB) (map[int64]deadmanSet, error) {
rows, err := db.QueryContext(ctx,
"SELECT "+deadmanSwitchColumns+" FROM deadman_switches ORDER BY id")
if err != nil {
return nil, err
}
return scanDeadmanSwitches(rows)
}
// deadmanSeededKey is the settings row that records the environment defaults
// were handed out. Without it, a team that deleted its last switch would get
// the default back on the next restart.
const deadmanSeededKey = "deadman_seeded"
// SeedDeadmanConfigs gives every team the server's environment defaults as
// switches, exactly once per install, so a fresh install watches Watchdog
// without anybody setting it up.
//
// Idempotent, and never overwrites: once a team has a row it owns its own
// configuration, and a redeploy must not quietly put the environment's value
// back over an owner's edit.
// Once seeded it never runs again: a team's switches are its own, and a redeploy
// must not quietly put the environment's value back over an owner's edit or
// deletion. Installs that upgraded from per-team configuration were already
// seeded, which migration 009 records.
//
// A team created after startup gets no row and therefore watches nothing until
// its owner says otherwise. That is deliberate: inheriting an install-wide
// heartbeat would page a new team about a source it has never heard of, and a
// switch nobody chose is the kind that gets muted rather than fixed.
// A team created after that gets none and watches nothing until its owner says
// otherwise. That is deliberate: inheriting an install-wide heartbeat would page
// a new team about a source it has never heard of, and a switch nobody chose is
// the kind that gets muted rather than fixed.
func SeedDeadmanConfigs(ctx context.Context, db *sql.DB, cfg DeadmanConfig) error {
matchers := make([]string, 0, len(cfg.Matchers))
if !cfg.enabled() {
return nil
}
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return err
}
defer tx.Rollback() //nolint:errcheck
res, err := tx.ExecContext(ctx,
"INSERT INTO settings (key, value) VALUES ($1, '1') ON CONFLICT (key) DO NOTHING",
deadmanSeededKey)
if err != nil {
return err
}
if n, _ := res.RowsAffected(); n == 0 {
return nil
}
for _, m := range cfg.Matchers {
parts := []string{"alertname=" + m.Name}
for k, v := range m.Labels {
parts = append(parts, k+"="+v)
if _, err := tx.ExecContext(ctx, `
INSERT INTO deadman_switches (team_id, name, matcher, timeout_seconds, severity)
SELECT id, $1, $1, $2, $3 FROM teams`,
m.config(), int64(cfg.Timeout.Seconds()), cfg.Severity); err != nil {
return err
}
sort.Strings(parts[1:])
matchers = append(matchers, strings.Join(parts, ","))
}
_, err := db.ExecContext(ctx, `
INSERT INTO deadman_configs (team_id, matchers, timeout_seconds, severity)
SELECT id, $1, $2, $3 FROM teams
ON CONFLICT (team_id) DO NOTHING`,
strings.Join(matchers, "; "), int64(cfg.Timeout.Seconds()), cfg.Severity)
return err
return tx.Commit()
}
// parseDeadmanQuietly is ParseDeadmanConfig without the startup logging: a
// team's configuration is read on every sweep and every webhook, and logging it
// each time would bury everything else.
func parseDeadmanQuietly(matchers string, timeout time.Duration, severity string) DeadmanConfig {
cfg := DeadmanConfig{Timeout: timeout, Severity: severity}
for _, entry := range strings.Split(matchers, ";") {
entry = strings.TrimSpace(entry)
if entry == "" {
continue
}
m := DeadmanMatcher{Labels: map[string]string{}}
malformed := false
for _, cond := range strings.Split(entry, ",") {
k, v, ok := strings.Cut(cond, "=")
k, v = strings.TrimSpace(k), strings.TrimSpace(v)
if !ok || k == "" || v == "" {
malformed = true
break
}
if k == "alertname" {
m.Name = v
continue
}
m.Labels[k] = v
}
if malformed || m.Name == "" {
continue
}
cfg.Matchers = append(cfg.Matchers, m)
}
return cfg
// ---------------------------------------------------------------------------
// Status
// ---------------------------------------------------------------------------
const (
switchHealthy = "healthy"
switchDead = "dead"
switchDormant = "dormant"
)
// deadmanSource is one heartbeat under a switch: a fingerprint that matched.
type deadmanSource struct {
Fingerprint string `json:"fingerprint"`
Labels map[string]string `json:"labels"`
Status string `json:"status"`
LastHeartbeatAt time.Time `json:"last_heartbeat_at"`
LastTriggeredAt *time.Time `json:"last_triggered_at"`
IncidentID *int64 `json:"incident_id"`
}
// deadmanSwitchStatus is a switch as the Switches page shows it.
type deadmanSwitchStatus struct {
ID int64 `json:"id"`
Name string `json:"name"`
Matcher string `json:"matcher"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Severity string `json:"severity"`
// Status is dead when any source is, dormant when none has ever been heard
// from, healthy otherwise — a live cluster must not hide a dead one.
Status string `json:"status"`
LastHeartbeatAt *time.Time `json:"last_heartbeat_at"`
LastTriggeredAt *time.Time `json:"last_triggered_at"`
OpenIncidentID *int64 `json:"open_incident_id"`
Sources []deadmanSource `json:"sources"`
}
// deadmanStatuses reports every switch of a team with what its heartbeats are
// doing. The liveness verdict is deadmanAlert.dead, the sweeper's own.
func deadmanStatuses(ctx context.Context, db *sql.DB, teamID int64, set deadmanSet, now time.Time) ([]deadmanSwitchStatus, error) {
out := make([]deadmanSwitchStatus, 0, len(set))
if len(set) == 0 {
return out, nil
}
heartbeats, err := deadmanAlerts(ctx, db, teamID, set)
if err != nil {
return nil, err
}
// One query for every switch's incident history, keyed the way the sweeper
// keys it.
type history struct {
triggeredAt int64
openID int64
}
incidents := map[string]history{}
rows, err := db.QueryContext(ctx, `
SELECT group_key, MAX(triggered_at), COALESCE(MAX(id) FILTER (WHERE resolved_at IS NULL), 0)
FROM incidents
WHERE team_id = $1 AND group_key LIKE $2
GROUP BY group_key`, teamID, deadmanGroupPrefix+"%")
if err != nil {
return nil, err
}
defer rows.Close()
for rows.Next() {
var key string
var h history
if err := rows.Scan(&key, &h.triggeredAt, &h.openID); err != nil {
return nil, err
}
incidents[key] = h
}
if err := rows.Err(); err != nil {
return nil, err
}
bySwitch := map[int64][]deadmanAlert{}
for _, hb := range heartbeats {
bySwitch[hb.sw.ID] = append(bySwitch[hb.sw.ID], hb)
}
later := func(cur *time.Time, unix int64) *time.Time {
t := time.Unix(unix, 0).UTC()
if cur == nil || t.After(*cur) {
return &t
}
return cur
}
for _, sw := range set {
st := deadmanSwitchStatus{
ID: sw.ID, Name: sw.Name, Matcher: sw.Matcher.config(),
TimeoutSeconds: int64(sw.Timeout.Seconds()), Severity: sw.Severity,
Status: switchDormant, Sources: []deadmanSource{},
}
for _, hb := range bySwitch[sw.ID] {
src := deadmanSource{
Fingerprint: hb.fingerprint,
Labels: hb.labels,
Status: switchHealthy,
LastHeartbeatAt: time.Unix(hb.receivedAt, 0).UTC(),
}
if hb.dead(now) {
src.Status = switchDead
}
if h, ok := incidents[hb.groupKey()]; ok {
t := time.Unix(h.triggeredAt, 0).UTC()
src.LastTriggeredAt = &t
st.LastTriggeredAt = later(st.LastTriggeredAt, h.triggeredAt)
if h.openID != 0 {
id := h.openID
src.IncidentID = &id
if st.OpenIncidentID == nil || id > *st.OpenIncidentID {
st.OpenIncidentID = &id
}
}
}
st.LastHeartbeatAt = later(st.LastHeartbeatAt, hb.receivedAt)
st.Sources = append(st.Sources, src)
switch {
case src.Status == switchDead:
st.Status = switchDead
case st.Status == switchDormant:
st.Status = switchHealthy
}
}
// Dead ones first, then by fingerprint: what needs attention leads, and
// the order does not shuffle between refreshes.
sort.Slice(st.Sources, func(i, j int) bool {
a, b := st.Sources[i], st.Sources[j]
if (a.Status == switchDead) != (b.Status == switchDead) {
return a.Status == switchDead
}
return a.Fingerprint < b.Fingerprint
})
out = append(out, st)
}
return out, nil
}
+166 -16
View File
@@ -1,6 +1,7 @@
package api_test
import (
"context"
"net/http"
"strings"
"testing"
@@ -483,13 +484,13 @@ func TestDeadman_ConfigurationIsPerTeam(t *testing.T) {
unwatched := newTeam(t, s, "unwatched")
// Only the first team calls Watchdog a heartbeat.
resp := s.req(t, http.MethodPut, "/api/teams/"+id64(watched.id)+"/deadman", map[string]any{
"matchers": "alertname=Watchdog",
resp := s.req(t, http.MethodPost, "/api/teams/"+id64(watched.id)+"/deadman/switches", map[string]any{
"matcher": "alertname=Watchdog",
"timeout_seconds": 3600,
"severity": "critical",
})
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
if resp.StatusCode != http.StatusCreated {
t.Fatalf("configure the watched team: %d", resp.StatusCode)
}
@@ -550,9 +551,9 @@ func TestDeadman_ConfigurationIsOwnerOnly(t *testing.T) {
decode(t, s.req(t, http.MethodPost, "/api/users/"+id64(user.ID)+"/api-keys",
map[string]string{"name": "test"}), &key)
req, _ := http.NewRequest(http.MethodPut,
s.URL+"/api/teams/"+id64(team.id)+"/deadman",
strings.NewReader(`{"matchers":"alertname=Watchdog","timeout_seconds":60}`))
req, _ := http.NewRequest(http.MethodPost,
s.URL+"/api/teams/"+id64(team.id)+"/deadman/switches",
strings.NewReader(`{"matcher":"alertname=Watchdog","timeout_seconds":60}`))
req.Header.Set("Authorization", "Bearer "+key.Key)
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
@@ -564,7 +565,7 @@ func TestDeadman_ConfigurationIsOwnerOnly(t *testing.T) {
t.Errorf("a member editing the switches: expected 403, got %d", resp.StatusCode)
}
read, _ := http.NewRequest(http.MethodGet, s.URL+"/api/teams/"+id64(team.id)+"/deadman", nil)
read, _ := http.NewRequest(http.MethodGet, s.URL+"/api/teams/"+id64(team.id)+"/deadman/switches", nil)
read.Header.Set("Authorization", "Bearer "+key.Key)
got, err := http.DefaultClient.Do(read)
if err != nil {
@@ -577,16 +578,165 @@ func TestDeadman_ConfigurationIsOwnerOnly(t *testing.T) {
}
// A matcher with no alertname watches nothing, silently, which is the failure
// this feature exists to prevent — so it is refused at the door.
func TestDeadman_UnusableMatchersAreRejected(t *testing.T) {
// this feature exists to prevent — so it is refused at the door, along with the
// other things that would make a switch unable to fire.
func TestDeadman_UnusableSwitchesAreRejected(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
resp := s.req(t, http.MethodPut, "/api/teams/"+defaultTeam+"/deadman", map[string]any{
"matchers": "cluster=prod",
"timeout_seconds": 900,
})
resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("expected 400 for a matcher with no alertname, got %d", resp.StatusCode)
for name, body := range map[string]map[string]any{
"no alertname": {"matcher": "cluster=prod", "timeout_seconds": 900},
"malformed": {"matcher": "alertname=Watchdog,garbage", "timeout_seconds": 900},
"several": {"matcher": "alertname=A; alertname=B", "timeout_seconds": 900},
"zero timeout": {"matcher": "alertname=Watchdog", "timeout_seconds": 0},
"bad severity": {"matcher": "alertname=Watchdog", "timeout_seconds": 900, "severity": "loud"},
"empty matcher": {"matcher": "", "timeout_seconds": 900},
} {
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/deadman/switches", body)
resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("%s: expected 400, got %d", name, resp.StatusCode)
}
}
}
// ---------------------------------------------------------------------------
// The switch list
// ---------------------------------------------------------------------------
// listSwitches reads the default team's switches as the Switches page does.
func listSwitches(t *testing.T, s *ts) []map[string]any {
t.Helper()
return list(t, s.req(t, http.MethodGet, "/api/teams/"+defaultTeam+"/deadman/switches", nil))
}
// A switch is healthy while its heartbeat is fresh, dead once it is silent, and
// dormant until the first one arrives.
func TestDeadman_ListReportsStatus(t *testing.T) {
s, _ := deadmanTS(t, api.ParseDeadmanConfig("alertname=Watchdog; alertname=NeverSent", time.Hour, "critical"))
got := listSwitches(t, s)
if len(got) != 2 {
t.Fatalf("expected 2 switches, got %d", len(got))
}
for _, sw := range got {
if sw["status"] != "dormant" || sw["last_heartbeat_at"] != nil || sw["last_triggered_at"] != nil {
t.Errorf("a switch nobody has heard from should be dormant and blank, got %v", sw)
}
}
heartbeat(t, s, "fp-watchdog", nil)
got = listSwitches(t, s)
if got[0]["status"] != "healthy" || got[0]["last_heartbeat_at"] == nil {
t.Errorf("a fresh heartbeat should be healthy with a timestamp, got %v", got[0])
}
if got[1]["status"] != "dormant" {
t.Errorf("the other switch is still dormant, got %v", got[1]["status"])
}
silence(t, s, "fp-watchdog", 2*time.Hour)
sweep(t, s, noArchive)
got = listSwitches(t, s)
if got[0]["status"] != "dead" {
t.Fatalf("a silent heartbeat should be dead, got %v", got[0]["status"])
}
if got[0]["last_triggered_at"] == nil || got[0]["open_incident_id"] == nil {
t.Errorf("a dead switch should show when it triggered and its open incident, got %v", got[0])
}
}
// One matcher, several clusters: the switch is as bad as its worst heartbeat and
// each heartbeat is listed on its own.
func TestDeadman_ListBreaksDownByFingerprint(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
heartbeat(t, s, "fp-a", map[string]string{"cluster": "a"})
heartbeat(t, s, "fp-b", map[string]string{"cluster": "b"})
silence(t, s, "fp-b", 2*time.Hour)
sw := listSwitches(t, s)[0]
if sw["status"] != "dead" {
t.Errorf("one dead cluster makes the switch dead, got %v", sw["status"])
}
sources := sw["sources"].([]any)
if len(sources) != 2 {
t.Fatalf("expected 2 sources, got %d", len(sources))
}
first, second := sources[0].(map[string]any), sources[1].(map[string]any)
if first["fingerprint"] != "fp-b" || first["status"] != "dead" || second["status"] != "healthy" {
t.Errorf("the dead source should lead, got %v then %v", first, second)
}
}
// Every switch keeps its own deadline.
func TestDeadman_TimeoutsArePerSwitch(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
resp := s.req(t, http.MethodPost, "/api/teams/"+defaultTeam+"/deadman/switches", map[string]any{
"matcher": "alertname=Edge", "timeout_seconds": 300,
})
resp.Body.Close()
heartbeat(t, s, "fp-watchdog", nil)
postWebhook(t, s, []map[string]any{
amAlert("fp-edge", "Edge", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
}, `{}:{alertname="Edge"}`)
// Ten minutes of silence: past the Edge switch's five, inside Watchdog's hour.
silence(t, s, "fp-watchdog", 10*time.Minute)
silence(t, s, "fp-edge", 10*time.Minute)
got := listSwitches(t, s)
if got[0]["status"] != "healthy" || got[1]["status"] != "dead" {
t.Errorf("want Watchdog healthy and Edge dead, got %v and %v", got[0]["status"], got[1]["status"])
}
}
// Deleting is an owner's, is scoped to the team, and leaves what the switch
// already opened alone.
func TestDeadman_DeleteIsScopedToTheTeam(t *testing.T) {
s, _ := deadmanTS(t, deadmanCfg())
other := newTeam(t, s, "other")
id := int64(listSwitches(t, s)[0]["id"].(float64))
// Another team's owner cannot reach it.
resp := other.call(http.MethodDelete, "/api/teams/"+id64(other.id)+"/deadman/switches/"+id64(id), nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("deleting another team's switch: expected 404, got %d", resp.StatusCode)
}
if got := len(listSwitches(t, s)); got != 1 {
t.Fatalf("the switch should have survived, %d left", got)
}
resp = s.req(t, http.MethodDelete, "/api/teams/"+defaultTeam+"/deadman/switches/"+id64(id), nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Fatalf("deleting: expected 204, got %d", resp.StatusCode)
}
if got := len(listSwitches(t, s)); got != 0 {
t.Errorf("expected no switches, got %d", got)
}
}
// The environment's defaults are handed out once and then belong to the teams.
func TestDeadman_SeedRunsOnce(t *testing.T) {
s := newTS(t)
cfg := api.ParseDeadmanConfig("alertname=Watchdog", time.Hour, "critical")
if err := api.SeedDeadmanConfigs(context.Background(), s.db, cfg); err != nil {
t.Fatalf("seed: %v", err)
}
if got := len(listSwitches(t, s)); got != 1 {
t.Fatalf("the first seed should add the default, got %d switches", got)
}
// The owner deletes it; a restart must not put it back.
id := int64(listSwitches(t, s)[0]["id"].(float64))
s.req(t, http.MethodDelete, "/api/teams/"+defaultTeam+"/deadman/switches/"+id64(id), nil).Body.Close()
if err := api.SeedDeadmanConfigs(context.Background(), s.db, cfg); err != nil {
t.Fatalf("seed again: %v", err)
}
if got := len(listSwitches(t, s)); got != 0 {
t.Errorf("a second seed resurrected %d switch(es)", got)
}
}
+185 -1
View File
@@ -346,10 +346,194 @@ func handleGetEscalation(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, escalationResponse(policy, teamID))
view, err := escalationStatus(r.Context(), db, teamID, policy)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, view)
}
}
// Level statuses, as the Escalation page colours them.
const (
levelReady = "ready"
levelEscalating = "escalating"
levelUnreachable = "unreachable"
)
// escalationTargetView is a target with who it means today and whether that
// person can actually be woken. The extra fields are output only: the PUT body
// is the plain escalationTargetJSON, and anything else in it is ignored.
type escalationTargetView struct {
escalationTargetJSON
// Username is who the target resolves to right now: the named person, or
// whoever the rota says is on call today. Empty when nobody is.
Username string `json:"username,omitempty"`
// Reachable is whether a page to this target would go anywhere, and Problem
// says why not when it would not — the same conditions pageLevel skips on.
Reachable bool `json:"reachable"`
Problem string `json:"problem,omitempty"`
}
type escalationLevelView struct {
Position int64 `json:"position"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Targets []escalationTargetView `json:"targets"`
// Status is unreachable when no target of the level could be woken — a rung
// that looks configured and pages nobody, which is worth seeing before an
// incident finds it — escalating when an unanswered incident has climbed to
// it, and ready otherwise.
Status string `json:"status"`
// Waiting lists the open, unacknowledged incidents currently on this level.
Waiting []int64 `json:"waiting"`
}
type escalationView struct {
TeamID int64 `json:"team_id"`
RepeatCount int64 `json:"repeat_count"`
FallbackTopic string `json:"fallback_topic"`
Levels []escalationLevelView `json:"levels"`
// LastEscalatedAt is when an incident of this team last moved up the ladder,
// or ran off the end of it, and LastEscalatedIncidentID which one. Absent
// when nothing ever has: a ladder nobody has needed yet.
LastEscalatedAt *time.Time `json:"last_escalated_at,omitempty"`
LastEscalatedIncidentID *int64 `json:"last_escalated_incident_id,omitempty"`
}
// escalationStatus is a team's ladder together with what it would do right now
// and what it has been doing. The resolution follows pageLevel's rules, so the
// page cannot promise a page that the notifier would skip.
func escalationStatus(ctx context.Context, db *sql.DB, teamID int64, policy *escalationPolicy) (escalationView, error) {
base := escalationResponse(policy, teamID)
out := escalationView{
TeamID: teamID, RepeatCount: base.RepeatCount, FallbackTopic: base.FallbackTopic,
Levels: []escalationLevelView{},
}
if !policy.configured() {
return out, nil
}
onCall, err := currentOnCall(ctx, db, teamID)
if err != nil {
return out, err
}
type account struct {
username string
topic bool
disabled bool
}
accounts := map[int64]account{}
lookup := func(id int64) (account, error) {
if a, ok := accounts[id]; ok {
return a, nil
}
var a account
var topic *string
var disabledAt *int64
if err := db.QueryRowContext(ctx,
"SELECT username, ntfy_topic, disabled_at FROM users WHERE id = $1", id).
Scan(&a.username, &topic, &disabledAt); err != nil {
return a, err
}
a.topic = topic != nil && *topic != ""
a.disabled = disabledAt != nil
accounts[id] = a
return a, nil
}
waiting := map[int64][]int64{}
rows, err := db.QueryContext(ctx, `
SELECT id, escalation_level FROM incidents
WHERE team_id = $1 AND resolved_at IS NULL AND archived_at IS NULL
AND status = 'triggered' AND escalation_level > 0
ORDER BY id`, teamID)
if err != nil {
return out, err
}
for rows.Next() {
var id, level int64
if err := rows.Scan(&id, &level); err != nil {
rows.Close()
return out, err
}
waiting[level] = append(waiting[level], id)
}
rows.Close()
if err := rows.Err(); err != nil {
return out, err
}
for _, l := range base.Levels {
level := escalationLevelView{
Position: l.Position, TimeoutSeconds: l.TimeoutSeconds,
Targets: []escalationTargetView{}, Waiting: []int64{},
}
if w := waiting[l.Position]; w != nil {
level.Waiting = w
}
anyReachable := false
for _, t := range l.Targets {
view := escalationTargetView{escalationTargetJSON: t}
userID := t.UserID
if t.Kind == "oncall" {
userID = onCall
}
switch {
case userID == nil:
view.Problem = "nobody is on call today"
default:
a, err := lookup(*userID)
switch {
case err != nil:
view.Problem = "account not found"
case a.disabled:
view.Username, view.Problem = a.username, "account is disabled"
case !a.topic:
view.Username, view.Problem = a.username, "has no ntfy topic"
default:
view.Username, view.Reachable = a.username, true
}
}
anyReachable = anyReachable || view.Reachable
level.Targets = append(level.Targets, view)
}
switch {
case !anyReachable:
level.Status = levelUnreachable
case l.Position >= 2 && len(level.Waiting) > 0:
level.Status = levelEscalating
default:
level.Status = levelReady
}
out.Levels = append(out.Levels, level)
}
var incidentID, at int64
switch err := db.QueryRowContext(ctx, `
SELECT e.incident_id, e.created_at
FROM incident_events e JOIN incidents i ON i.id = e.incident_id
WHERE i.team_id = $1 AND e.type = $2
ORDER BY e.created_at DESC, e.id DESC LIMIT 1`, teamID, evEscalated).
Scan(&incidentID, &at); {
case err == sql.ErrNoRows:
case err != nil:
return out, err
default:
t := time.Unix(at, 0).UTC()
out.LastEscalatedAt, out.LastEscalatedIncidentID = &t, &incidentID
}
return out, nil
}
type escalationLevelJSON struct {
Position int64 `json:"position"`
TimeoutSeconds int64 `json:"timeout_seconds"`
+119
View File
@@ -383,3 +383,122 @@ func TestEscalation_SkipsUnreachableTargets(t *testing.T) {
t.Errorf("a target with no topic should page nothing, paged %v", got)
}
}
// ---------------------------------------------------------------------------
// The ladder as the Escalation page reads it
// ---------------------------------------------------------------------------
type ladderLevel struct {
Status string `json:"status"`
Waiting []int64 `json:"waiting"`
Targets []struct {
Kind string `json:"kind"`
Username string `json:"username"`
Reachable bool `json:"reachable"`
Problem string `json:"problem"`
} `json:"targets"`
}
type ladderView struct {
Levels []ladderLevel `json:"levels"`
LastEscalatedAt *string `json:"last_escalated_at"`
LastEscalatedIncidentID *int64 `json:"last_escalated_incident_id"`
}
func readLadder(t *testing.T, s *ts) ladderView {
t.Helper()
var v ladderView
decode(t, s.req(t, http.MethodGet, "/api/teams/"+defaultTeam+"/escalation", nil), &v)
return v
}
// Targets say who they mean today, so "whoever is on call" is a name and not a
// promise.
func TestEscalation_StatusResolvesTargets(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
second := teamUser(t, s, "second", "terdut-second")
ladder(t, s, second, 0, "terdut-fallback")
v := readLadder(t, s)
if len(v.Levels) != 2 {
t.Fatalf("expected 2 levels, got %d", len(v.Levels))
}
if got := v.Levels[0].Targets[0]; got.Kind != "oncall" || got.Username != "admin" || !got.Reachable {
t.Errorf("the rota target should resolve to the person on call, got %+v", got)
}
if got := v.Levels[1].Targets[0]; got.Username != "second" || !got.Reachable {
t.Errorf("the named target should be reachable, got %+v", got)
}
if v.Levels[0].Status != "ready" || v.Levels[1].Status != "ready" || v.LastEscalatedAt != nil {
t.Errorf("an idle, healthy ladder is ready and has never escalated, got %+v", v)
}
}
// A rung that would page nobody is called out before an incident finds it.
func TestEscalation_StatusFlagsUnreachableLevels(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
silent := teamUser(t, s, "silent", "terdut-silent")
ladder(t, s, silent, 0, "terdut-fallback")
// Nobody on call today, and the named person loses their topic.
s.exec(t, "DELETE FROM schedule_entries")
s.exec(t, "UPDATE users SET ntfy_topic = NULL WHERE id = $1", silent)
v := readLadder(t, s)
if v.Levels[0].Status != "unreachable" || v.Levels[0].Targets[0].Problem != "nobody is on call today" {
t.Errorf("an empty rota should make level 1 unreachable, got %+v", v.Levels[0])
}
if v.Levels[1].Status != "unreachable" || v.Levels[1].Targets[0].Problem != "has no ntfy topic" {
t.Errorf("a person with no topic should make level 2 unreachable, got %+v", v.Levels[1])
}
s.exec(t, "UPDATE users SET disabled_at = 1 WHERE id = $1", silent)
if p := readLadder(t, s).Levels[1].Targets[0].Problem; p != "account is disabled" {
t.Errorf("a disabled account should say so, got %q", p)
}
}
// Where unanswered incidents are right now, and when the ladder last did its
// job.
func TestEscalation_StatusShowsWhoIsWaitingAndLastEscalation(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
second := teamUser(t, s, "second", "terdut-second")
ladder(t, s, second, 0, "terdut-fallback")
postWebhook(t, s, []map[string]any{
amAlert("fp-wait", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.sweepNotify(t)
// On level 1 it is waiting, which is normal and not yet an escalation.
v := readLadder(t, s)
if len(v.Levels[0].Waiting) != 1 || v.Levels[0].Status != "ready" || v.LastEscalatedAt != nil {
t.Fatalf("a fresh incident waits on level 1 quietly, got %+v", v)
}
overdue(t, s, 1)
s.sweepNotify(t)
v = readLadder(t, s)
if v.Levels[1].Status != "escalating" || len(v.Levels[1].Waiting) != 1 || v.Levels[1].Waiting[0] != 1 {
t.Errorf("level 2 should be escalating with the incident on it, got %+v", v.Levels[1])
}
if v.LastEscalatedAt == nil || v.LastEscalatedIncidentID == nil || *v.LastEscalatedIncidentID != 1 {
t.Errorf("the escalation should be recorded, got %+v", v)
}
// Somebody answers: nothing is waiting, but the history stays.
s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil).Body.Close()
v = readLadder(t, s)
if v.Levels[1].Status != "ready" || len(v.Levels[1].Waiting) != 0 || v.LastEscalatedAt == nil {
t.Errorf("an acknowledged incident stops waiting but stays in the history, got %+v", v)
}
}
// No ladder is a real answer, not an error.
func TestEscalation_StatusWithoutALadder(t *testing.T) {
s, _ := notifyTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com", RepeatEvery: 15 * time.Minute})
v := readLadder(t, s)
if len(v.Levels) != 0 || v.LastEscalatedAt != nil {
t.Errorf("a team with no ladder should read as empty, got %+v", v)
}
}
+4 -2
View File
@@ -144,12 +144,14 @@ func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config) http.Handler
// A team's own dead man's switches: which of its alerts are heartbeats,
// and how long a silence has to last before somebody is paged.
r.Get("/api/teams/{teamID}/deadman", handleGetTeamDeadman(db))
r.Put("/api/teams/{teamID}/deadman", handleSetTeamDeadman(db))
r.Get("/api/teams/{teamID}/deadman/switches", handleListTeamDeadman(db))
r.Post("/api/teams/{teamID}/deadman/switches", handleCreateTeamDeadman(db))
r.Delete("/api/teams/{teamID}/deadman/switches/{switchID}", handleDeleteTeamDeadman(db))
// Integrations: where a team's alerts come in, and the key that says so.
r.Get("/api/teams/{teamID}/integrations", handleListIntegrations(db))
r.Post("/api/teams/{teamID}/integrations", handleCreateIntegration(db, notify.PublicURL))
r.Patch("/api/teams/{teamID}/integrations/{integrationID}", handleRenameIntegration(db))
r.Delete("/api/teams/{teamID}/integrations/{integrationID}", handleDeleteIntegration(db))
// The rota is per team. /api/schedule/current is the exception: it
+171
View File
@@ -0,0 +1,171 @@
package api_test
import (
"bytes"
"net/http"
"strings"
"testing"
"time"
)
// listSources reads a team's alert sources as the Sources page does.
func listSources(t *testing.T, tm teamFixture) []map[string]any {
t.Helper()
return list(t, tm.call(http.MethodGet, "/api/teams/"+id64(tm.id)+"/integrations", nil))
}
// addSource mints a second source in a team and returns its key.
func addSource(t *testing.T, tm teamFixture, name string) string {
t.Helper()
var out struct {
Key string `json:"key"`
}
decode(t, tm.call(http.MethodPost, "/api/teams/"+id64(tm.id)+"/integrations",
map[string]string{"name": name}), &out)
return out.Key
}
// A source that has never posted is "never", with nothing to say about alerts.
func TestSources_NeverUsedIsBlank(t *testing.T) {
s := newTS(t)
tm := newTeam(t, s, "red")
got := listSources(t, tm)
if len(got) != 1 {
t.Fatalf("expected 1 source, got %d", len(got))
}
src := got[0]
if src["status"] != "never" || src["last_used_at"] != nil || src["last_alert_at"] != nil {
t.Errorf("a source nobody has posted on should be blank, got %v", src)
}
if src["alerts_24h"].(float64) != 0 {
t.Errorf("alerts_24h = %v, want 0", src["alerts_24h"])
}
}
// Each source is credited with what arrived on its own key, and only that.
func TestSources_AlertsAreAttributedToTheirSource(t *testing.T) {
s := newTS(t)
tm := newTeam(t, s, "red")
second := addSource(t, tm, "staging")
postToIntegration(t, s, tm.key, "fp-1", "DiskFull")
postToIntegration(t, s, tm.key, "fp-2", "CPUHot")
got := listSources(t, tm)
first, other := got[0], got[1]
if first["status"] != "active" || first["last_used_at"] == nil || first["last_alert_at"] == nil {
t.Errorf("the source that posted should be active with timestamps, got %v", first)
}
if first["alerts_24h"].(float64) != 2 {
t.Errorf("alerts_24h = %v, want 2", first["alerts_24h"])
}
if other["status"] != "never" || other["alerts_24h"].(float64) != 0 {
t.Errorf("the other source should be untouched, got %v", other)
}
// Re-sending the same alert on the other key moves it: last sender wins.
postToIntegration(t, s, second, "fp-1", "DiskFull")
got = listSources(t, tm)
if got[0]["alerts_24h"].(float64) != 1 || got[1]["alerts_24h"].(float64) != 1 {
t.Errorf("fp-1 should have moved to the second source, got %v and %v",
got[0]["alerts_24h"], got[1]["alerts_24h"])
}
}
// A payload with no alerts in it is a webhook, not an alert: the source was
// heard from, and nothing arrived.
func TestSources_EmptyPayloadStampsUseButNotAlert(t *testing.T) {
s := newTS(t)
tm := newTeam(t, s, "red")
resp, err := http.Post(s.URL+"/api/integrations/"+tm.key+"/alertmanager",
"application/json", bytes.NewReader([]byte(`{"version":"4","status":"firing","alerts":[]}`)))
if err != nil {
t.Fatalf("post: %v", err)
}
resp.Body.Close()
src := listSources(t, tm)[0]
if src["status"] != "active" || src["last_alert_at"] != nil {
t.Errorf("want active with no alert yet, got %v", src)
}
}
// Quiet is "has posted, not lately"; the alert counter forgets after a day but
// the last alert's timestamp is kept.
func TestSources_QuietAfterADay(t *testing.T) {
s := newTS(t)
tm := newTeam(t, s, "red")
postToIntegration(t, s, tm.key, "fp-1", "DiskFull")
old := time.Now().Add(-48 * time.Hour).Unix()
s.exec(t, "UPDATE integrations SET last_used_at = $1", old)
s.exec(t, "UPDATE alerts SET received_at = $1 WHERE fingerprint = 'fp-1'", old)
src := listSources(t, tm)[0]
if src["status"] != "quiet" {
t.Errorf("status = %v, want quiet", src["status"])
}
if src["alerts_24h"].(float64) != 0 {
t.Errorf("alerts_24h = %v, want 0", src["alerts_24h"])
}
if src["last_alert_at"] == nil {
t.Error("last_alert_at should survive the day")
}
}
// Revoking a source does not take its alerts with it.
func TestSources_RevokeKeepsTheAlerts(t *testing.T) {
s := newTS(t)
tm := newTeam(t, s, "red")
postToIntegration(t, s, tm.key, "fp-1", "DiskFull")
id := int64(listSources(t, tm)[0]["id"].(float64))
resp := tm.call(http.MethodDelete, "/api/teams/"+id64(tm.id)+"/integrations/"+id64(id), nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Fatalf("revoke: %d", resp.StatusCode)
}
if got := len(list(t, tm.call(http.MethodGet, "/api/alerts", nil))); got != 1 {
t.Errorf("the alert should outlive its source, got %d alerts", got)
}
}
// Renaming is an owner's, scoped to the team, and does not touch the key.
func TestSources_Rename(t *testing.T) {
s := newTS(t)
tm := newTeam(t, s, "red")
other := newTeam(t, s, "blue")
id := int64(listSources(t, tm)[0]["id"].(float64))
path := "/api/teams/" + id64(tm.id) + "/integrations/" + id64(id)
resp := tm.call(http.MethodPatch, path, map[string]string{"name": " prod "})
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Fatalf("rename: %d", resp.StatusCode)
}
if name := listSources(t, tm)[0]["name"]; name != "prod" {
t.Errorf("name = %q, want it trimmed to prod", name)
}
postToIntegration(t, s, tm.key, "fp-1", "DiskFull") // the old key still works
for name, body := range map[string]map[string]string{
"empty": {"name": " "},
"too long": {"name": strings.Repeat("x", 101)},
} {
resp := tm.call(http.MethodPatch, path, body)
resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("%s name: expected 400, got %d", name, resp.StatusCode)
}
}
// Another team's owner cannot reach it.
resp = other.call(http.MethodPatch, "/api/teams/"+id64(other.id)+"/integrations/"+id64(id),
map[string]string{"name": "mine now"})
resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("renaming another team's source: expected 404, got %d", resp.StatusCode)
}
}
+218 -68
View File
@@ -355,8 +355,42 @@ func isLastTeamOwner(ctx context.Context, db *sql.DB, teamID, userID int64) (boo
// Integrations
// ---------------------------------------------------------------------------
// handleListIntegrations lists a team's integrations. Never the keys: those
// exist in plaintext only in the response that created them.
// sourceQuietAfter is how long a source may go without posting before the
// Sources page calls it quiet rather than active. A day is longer than any
// repeat_interval worth having, so an Alertmanager that is up and has anything
// firing never crosses it; a source with nothing firing may, and that is a
// reason to look, not proof of a fault — which is why this is a colour and not
// an alarm. Dead man's switches are where silence pages.
const sourceQuietAfter = 24 * time.Hour
const (
sourceActive = "active"
sourceQuiet = "quiet"
sourceNever = "never"
)
// integrationStatus is an integration as the Sources page shows it.
type integrationStatus struct {
models.Integration
// Status is active when the key posted within sourceQuietAfter, quiet when
// it has posted but not lately, never when it has not posted at all.
Status string `json:"status"`
// LastAlertAt is when an alert last arrived on this source, which is not the
// same as when it last posted: a payload with nothing usable in it stamps
// last_used_at and not this. Absent until an alert has arrived since
// migration 010 started recording it.
LastAlertAt *time.Time `json:"last_alert_at,omitempty"`
// Alerts24h counts the distinct alerts this source refreshed in the last
// day. An alert re-sent every few hours counts once, not once per re-send.
Alerts24h int64 `json:"alerts_24h"`
}
// handleListIntegrations lists a team's integrations with what each has been
// delivering. Never the keys: those exist in plaintext only in the response that
// created them.
func handleListIntegrations(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
@@ -367,28 +401,45 @@ func handleListIntegrations(db *sql.DB) http.HandlerFunc {
return
}
now := time.Now()
rows, err := db.QueryContext(r.Context(), `
SELECT id, team_id, kind, name, created_at, last_used_at
FROM integrations
WHERE team_id = $1
ORDER BY id`, teamID)
SELECT i.id, i.team_id, i.kind, i.name, i.created_at, i.last_used_at,
-- Scalar subqueries, not a join and GROUP BY: each is a
-- single range over alerts_integration_idx, where the join
-- would read every alert a source ever delivered.
(SELECT MAX(received_at) FROM alerts WHERE integration_id = i.id),
(SELECT COUNT(*) FROM alerts
WHERE integration_id = i.id AND received_at >= $2)
FROM integrations i
WHERE i.team_id = $1
ORDER BY i.id`, teamID, now.Add(-sourceQuietAfter).Unix())
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer rows.Close()
integrations := []models.Integration{}
integrations := []integrationStatus{}
for rows.Next() {
var i models.Integration
var i integrationStatus
var created int64
var lastUsed *int64
if err := rows.Scan(&i.ID, &i.TeamID, &i.Kind, &i.Name, &created, &lastUsed); err != nil {
var lastUsed, lastAlert *int64
if err := rows.Scan(&i.ID, &i.TeamID, &i.Kind, &i.Name, &created, &lastUsed,
&lastAlert, &i.Alerts24h); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
i.CreatedAt = time.Unix(created, 0).UTC()
i.LastUsedAt = unixPtr(lastUsed)
i.LastAlertAt = unixPtr(lastAlert)
switch {
case i.LastUsedAt == nil:
i.Status = sourceNever
case now.Sub(*i.LastUsedAt) > sourceQuietAfter:
i.Status = sourceQuiet
default:
i.Status = sourceActive
}
integrations = append(integrations, i)
}
if err := rows.Err(); err != nil {
@@ -457,6 +508,54 @@ func handleCreateIntegration(db *sql.DB, publicURL string) http.HandlerFunc {
}
}
// handleRenameIntegration renames a source. The key is untouched, so nothing
// posting with it notices.
func handleRenameIntegration(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
if !ok {
return
}
if !requireTeamOwner(w, r, teamID) {
return
}
id, err := strconv.ParseInt(chi.URLParam(r, "integrationID"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid integration id"))
return
}
var req struct {
Name string `json:"name"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
req.Name = strings.TrimSpace(req.Name)
if req.Name == "" {
respond(w, http.StatusBadRequest, errResp("name is required"))
return
}
if len(req.Name) > 100 {
respond(w, http.StatusBadRequest, errResp("name is too long"))
return
}
res, err := db.ExecContext(r.Context(),
"UPDATE integrations SET name = $1 WHERE id = $2 AND team_id = $3", req.Name, id, teamID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("not found"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
func handleDeleteIntegration(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
@@ -493,24 +592,32 @@ func integrationPath(key, kind string) string {
return "/api/integrations/" + key + "/" + kind
}
// teamIDForKey resolves an integration key to its team, and stamps the key's
// alertSource is who an arriving webhook is from: the integration whose key it
// used, and the team that integration puts its alerts in.
type alertSource struct {
integrationID int64
teamID int64
}
// sourceForKey resolves an integration key to its source, and stamps the key's
// last use. An unknown key is not an error worth distinguishing: the caller is
// told nothing beyond "no".
func teamIDForKey(ctx context.Context, db *sql.DB, key string) (int64, error) {
var teamID int64
func sourceForKey(ctx context.Context, db *sql.DB, key string) (alertSource, error) {
var src alertSource
err := db.QueryRowContext(ctx,
"SELECT team_id FROM integrations WHERE key_hash = $1", hashToken(key)).Scan(&teamID)
"SELECT id, team_id FROM integrations WHERE key_hash = $1", hashToken(key)).
Scan(&src.integrationID, &src.teamID)
if errors.Is(err, sql.ErrNoRows) {
return 0, errUnknownIntegration
return alertSource{}, errUnknownIntegration
}
if err != nil {
return 0, err
return alertSource{}, err
}
// Best effort, like an API key's: a failed stamp must not reject an alert.
db.ExecContext(ctx, //nolint:errcheck
"UPDATE integrations SET last_used_at = $1 WHERE key_hash = $2",
time.Now().Unix(), hashToken(key))
return teamID, nil
"UPDATE integrations SET last_used_at = $1 WHERE id = $2",
time.Now().Unix(), src.integrationID)
return src, nil
}
var errUnknownIntegration = errors.New("unknown integration key")
@@ -539,17 +646,24 @@ func defaultTeamID(ctx context.Context, db *sql.DB) (int64, error) {
// A team's dead man's switches
// ---------------------------------------------------------------------------
// deadmanResponse is the wire shape of a team's switch configuration. The
// timeout is seconds rather than a duration string, because that is what the
// column holds and what arithmetic is done on; a client renders it.
type deadmanResponse struct {
TeamID int64 `json:"team_id"`
Matchers string `json:"matchers"`
// deadmanSwitchRequest is what creating a switch takes. The timeout is seconds,
// because that is what the column holds and what arithmetic is done on; a client
// renders it.
type deadmanSwitchRequest struct {
Name string `json:"name"`
Matcher string `json:"matcher"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Severity string `json:"severity"`
}
func handleGetTeamDeadman(db *sql.DB) http.HandlerFunc {
// deadmanSeverities are the severities an incident can open at.
var deadmanSeverities = map[string]bool{"critical": true, "error": true, "warning": true, "info": true}
// handleListTeamDeadman lists a team's switches with what each one's heartbeats
// are doing. A team with none gets an empty list, which is a configuration and
// not an absence: answering 404 would make "off" indistinguishable from "this
// server does not do this".
func handleListTeamDeadman(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
if !ok {
@@ -559,27 +673,26 @@ func handleGetTeamDeadman(db *sql.DB) http.HandlerFunc {
return
}
out := deadmanResponse{TeamID: teamID, Severity: "critical"}
err := db.QueryRowContext(r.Context(),
"SELECT matchers, timeout_seconds, severity FROM deadman_configs WHERE team_id = $1",
teamID).Scan(&out.Matchers, &out.TimeoutSeconds, &out.Severity)
if err != nil && !errors.Is(err, sql.ErrNoRows) {
set, err := deadmanSetForTeam(r.Context(), db, teamID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
out, err := deadmanStatuses(r.Context(), db, teamID, set, time.Now())
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
// A team with no row watches nothing, which is a configuration and not
// an absence: answering 404 would make "off" indistinguishable from
// "this server does not do this".
respond(w, http.StatusOK, out)
}
}
// handleSetTeamDeadman replaces a team's switch configuration.
// handleCreateTeamDeadman adds one switch.
//
// Validated by parsing: a matcher string that survives ParseDeadmanConfig with
// nothing usable in it is rejected rather than stored, because a switch that
// silently watches nothing is the failure this feature exists to prevent.
func handleSetTeamDeadman(db *sql.DB) http.HandlerFunc {
// Validated by parsing: a matcher with no alertname is rejected rather than
// stored, because a switch that silently watches nothing is the failure this
// feature exists to prevent.
func handleCreateTeamDeadman(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
if !ok {
@@ -589,50 +702,87 @@ func handleSetTeamDeadman(db *sql.DB) http.HandlerFunc {
return
}
var req struct {
Matchers string `json:"matchers"`
TimeoutSeconds int64 `json:"timeout_seconds"`
Severity string `json:"severity"`
}
var req deadmanSwitchRequest
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
req.Matchers = strings.TrimSpace(req.Matchers)
req.Matcher = strings.TrimSpace(req.Matcher)
req.Name = strings.TrimSpace(req.Name)
if req.Severity == "" {
req.Severity = "critical"
}
if req.TimeoutSeconds < 0 {
respond(w, http.StatusBadRequest, errResp("timeout_seconds must not be negative"))
if !deadmanSeverities[req.Severity] {
respond(w, http.StatusBadRequest, errResp("severity must be critical, error, warning or info"))
return
}
if req.Matchers != "" {
parsed := parseDeadmanQuietly(req.Matchers, time.Duration(req.TimeoutSeconds)*time.Second, req.Severity)
if len(parsed.Matchers) == 0 {
respond(w, http.StatusBadRequest, errResp(
"no usable matchers: each must name an alertname, as in alertname=Watchdog,cluster=prod"))
return
}
if req.TimeoutSeconds <= 0 {
respond(w, http.StatusBadRequest, errResp("timeout_seconds must be positive"))
return
}
if strings.Contains(req.Matcher, ";") {
respond(w, http.StatusBadRequest, errResp("one matcher per switch: add another switch instead of separating with ;"))
return
}
m, err := parseDeadmanMatcher(req.Matcher)
if err != nil {
respond(w, http.StatusBadRequest, errResp(
"unusable matcher ("+err.Error()+"): each must name an alertname, as in alertname=Watchdog,cluster=prod"))
return
}
if req.Name == "" {
req.Name = m.config()
}
if len(req.Name) > 100 {
respond(w, http.StatusBadRequest, errResp("name is too long"))
return
}
if _, err := db.ExecContext(r.Context(), `
INSERT INTO deadman_configs (team_id, matchers, timeout_seconds, severity, updated_at)
VALUES ($1, $2, $3, $4, `+nowEpoch+`)
ON CONFLICT (team_id) DO UPDATE SET
matchers = excluded.matchers,
timeout_seconds = excluded.timeout_seconds,
severity = excluded.severity,
updated_at = excluded.updated_at`,
teamID, req.Matchers, req.TimeoutSeconds, req.Severity); err != nil {
var id int64
if err := db.QueryRowContext(r.Context(), `
INSERT INTO deadman_switches (team_id, name, matcher, timeout_seconds, severity)
VALUES ($1, $2, $3, $4, $5) RETURNING id`,
teamID, req.Name, m.config(), req.TimeoutSeconds, req.Severity).Scan(&id); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, deadmanResponse{
TeamID: teamID,
Matchers: req.Matchers,
TimeoutSeconds: req.TimeoutSeconds,
Severity: req.Severity,
respond(w, http.StatusCreated, deadmanSwitchStatus{
ID: id, Name: req.Name, Matcher: m.config(),
TimeoutSeconds: req.TimeoutSeconds, Severity: req.Severity,
Status: switchDormant, Sources: []deadmanSource{},
})
}
}
// handleDeleteTeamDeadman removes a switch. An incident it already opened stays
// open until somebody resolves it: deleting the switch says "stop watching", not
// "the problem is gone".
func handleDeleteTeamDeadman(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
teamID, ok := teamParam(w, r)
if !ok {
return
}
if !requireTeamOwner(w, r, teamID) {
return
}
switchID, err := strconv.ParseInt(chi.URLParam(r, "switchID"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid switch id"))
return
}
res, err := db.ExecContext(r.Context(),
"DELETE FROM deadman_switches WHERE id = $1 AND team_id = $2", switchID, teamID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("switch not found"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
@@ -0,0 +1,54 @@
-- Dead man's switches become rows of their own.
--
-- 004 kept a team's switches in one string with one timeout and one severity,
-- which was enough to configure them and not enough to show them: there was no
-- thing to list, nothing to hang a status on, and every switch in a team had to
-- share a deadline. A row per switch gives each its own name, matcher, timeout
-- and severity, and gives the Team → Switches page something to be a list of.
--
-- The matcher keeps the syntax the string used, one matcher per row:
-- `alertname=Watchdog,cluster=prod`. The unit of monitoring is still the
-- fingerprint, so a matcher that many clusters satisfy is still one switch row
-- watching several independent heartbeats.
CREATE TABLE deadman_switches (
id BIGSERIAL PRIMARY KEY,
team_id BIGINT NOT NULL REFERENCES teams(id) ON DELETE CASCADE,
-- What the owner calls it. Defaults to the matcher when they do not say.
name TEXT NOT NULL,
-- "," separates the label conditions, "=" is exact equality, and alertname is
-- mandatory: it is what keeps the sweeper's candidate query on an index.
matcher TEXT NOT NULL,
-- Seconds of silence before the switch is declared dead. Never zero: a switch
-- that cannot fire is deleted, not disabled.
timeout_seconds BIGINT NOT NULL CHECK (timeout_seconds > 0),
-- The severity its incidents open at. See 004 for why they carry their own.
severity TEXT NOT NULL DEFAULT 'critical',
created_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
);
CREATE INDEX deadman_switches_team_idx ON deadman_switches (team_id);
-- Carry every team's configuration over, one row per matcher. A team whose
-- timeout was zero had switches turned off, which is now "no rows".
INSERT INTO deadman_switches (team_id, name, matcher, timeout_seconds, severity)
SELECT c.team_id, btrim(m), btrim(m), c.timeout_seconds, c.severity
FROM deadman_configs c,
LATERAL regexp_split_to_table(c.matchers, ';') AS m
WHERE c.timeout_seconds > 0
AND btrim(m) <> ''
ORDER BY c.team_id;
-- The server seeds environment defaults into teams once, and remembers that it
-- did. An install that had a row per team was already seeded; without this
-- marker the first start after upgrading would seed teams that had switched
-- theirs off.
INSERT INTO settings (key, value)
SELECT 'deadman_seeded', '1'
WHERE EXISTS (SELECT 1 FROM deadman_configs);
DROP TABLE deadman_configs;
@@ -0,0 +1,21 @@
-- Which alert source an alert last arrived on.
--
-- Team -> Sources shows when each source last posted, which integrations
-- already knew (last_used_at, stamped on every webhook). What it could not say
-- was what a source delivered: an alert never recorded the key it came in on, so
-- "prod alertmanager" and "staging alertmanager" were indistinguishable once
-- inside. This column is that link, and lets the page show each source's last
-- alert and how many alerts it has kept fresh over the past day.
--
-- Last sender wins: every accepted payload restamps it, the way it advances
-- received_at. Two sources posting the same fingerprint into one team is
-- already one alert, and it is attributed to whichever spoke last.
--
-- Nullable, and not backfilled. Alerts that arrived before this migration have
-- no source, and NULL says so honestly rather than guessing. It heals by itself:
-- Alertmanager re-sends every alert each repeat_interval, and each re-send is an
-- accepted payload. Deleting a source keeps its alerts, unattributed.
ALTER TABLE alerts ADD COLUMN integration_id BIGINT REFERENCES integrations(id) ON DELETE SET NULL;
CREATE INDEX alerts_integration_idx ON alerts (integration_id, received_at)
WHERE integration_id IS NOT NULL;
+26
View File
@@ -0,0 +1,26 @@
package web
import (
"io/fs"
"strings"
"testing"
)
func TestCopyIncidentIsEmbedded(t *testing.T) {
sub, err := fs.Sub(files, "static")
if err != nil {
t.Fatal(err)
}
for file, want := range map[string]string{
"js/incident.js": "copyIncident",
"js/ui.js": "copy:",
} {
b, err := fs.ReadFile(sub, file)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(b), want) {
t.Errorf("%s lacks %s", file, want)
}
}
}
+32 -1
View File
@@ -359,7 +359,15 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.badge.st-triggered, .badge.st-firing { background: var(--crit-soft); color: var(--crit); }
.badge.st-acknowledged { background: var(--warn-soft); color: var(--warn); }
.badge.st-snoozed { background: var(--snooze-soft); color: var(--snooze); }
.badge.st-resolved { background: var(--ok-soft); color: var(--ok); }
.badge.st-resolved, .badge.st-healthy { background: var(--ok-soft); color: var(--ok); }
.badge.st-dead { background: var(--crit-soft); color: var(--crit); }
/* Dormant is the plain badge on purpose: nothing has gone wrong and nothing has
gone right, which is what the muted default already says. */
.badge.st-dormant, .badge.st-never { background: var(--surface-2); color: var(--muted); }
.badge.st-active { background: var(--ok-soft); color: var(--ok); }
.badge.st-quiet, .badge.st-escalating { background: var(--warn-soft); color: var(--warn); }
.badge.st-ready { background: var(--ok-soft); color: var(--ok); }
.badge.st-unreachable { background: var(--crit-soft); color: var(--crit); }
.badge.sev-critical { background: var(--crit-soft); color: var(--crit); }
.badge.sev-warning { background: var(--warn-soft); color: var(--warn); }
.badge.sev-info { background: var(--info-soft); color: var(--info); }
@@ -377,6 +385,8 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
-webkit-backdrop-filter: saturate(1.4) blur(12px);
border-bottom: 1px solid var(--border);
}
.detail-head .copy { margin-left: auto; }
.clip-buffer { position: fixed; top: 0; left: 0; opacity: 0; pointer-events: none; }
.detail-head .crumb { font-weight: 600; color: var(--muted); font-size: 14px; }
.detail-title { font-size: 21px; font-weight: 750; letter-spacing: -0.01em; margin: 16px 0 8px; overflow-wrap: anywhere; }
.detail-badges { display: flex; flex-wrap: wrap; gap: 6px; margin-bottom: 14px; }
@@ -638,6 +648,10 @@ kbd {
.view-queue .pane { overflow: auto; height: 100dvh; }
.pane-list { border-right: 1px solid var(--border); }
.pane-list .chips { position: sticky; top: 0; z-index: 2; background: var(--bg); padding-top: 16px; }
/* The pane is 340-420px wide and a mouse cannot scroll a row whose scrollbar
is hidden, so the chips wrap here instead: Archived stays reachable. */
.pane-list .chips { flex-wrap: wrap; overflow-x: visible; }
.pane-list .chip-sep { display: none; }
.view-queue:not(.has-detail) .pane-detail { display: block; }
/* On desktop the list stays visible next to the detail. */
@@ -685,6 +699,18 @@ kbd {
.disabled-row td { opacity: 0.55; }
.btn-sm.danger { color: var(--crit); border-color: var(--crit-soft); }
/* --- status lists: alert sources and dead man's switches -----------------
Six columns do not fit a phone, so the table scrolls inside its card rather
than the page. A heartbeat under a switch with several is indented, the way
the escalation ladder indents its levels. */
.card-head { display: flex; align-items: center; justify-content: space-between; gap: 12px; flex-wrap: wrap; }
.table-scroll { overflow-x: auto; margin-top: 12px; }
.status-table th, .status-table td { white-space: nowrap; }
.status-table td.wrap { white-space: normal; min-width: 12em; }
.status-table .source-row td { border-bottom-style: dashed; }
.status-table .source-row td:first-child { padding-left: 16px; }
.source-labels { display: flex; flex-wrap: wrap; gap: 4px; align-items: center; }
.inline-form { display: flex; gap: 8px; margin-top: 12px; }
.inline-form input { flex: 1; min-width: 0; }
@@ -803,6 +829,11 @@ button.rota-day:hover { background: var(--surface-2); }
.ladder-head { display: flex; align-items: center; gap: 10px; margin-bottom: 6px; }
.ladder-targets { display: flex; flex-direction: column; gap: 6px; margin-top: 8px; }
.target-row { display: flex; gap: 6px; align-items: center; flex-wrap: wrap; }
.ladder-editor { display: flex; flex-direction: column; gap: 10px; align-items: flex-start; margin-top: 12px; }
/* A target that would not wake anybody says why, in place: it is the reason a
level is red, and the thing to go and fix. */
.target-line { display: flex; gap: 8px; align-items: baseline; flex-wrap: wrap; }
.target-problem { color: var(--crit); font-size: 12px; font-weight: 600; }
/* An integration key is shown exactly once, so it should look like something
to act on rather than another row of text. */
+7 -2
View File
@@ -133,11 +133,16 @@ export const removeTeamMember = (id, userID) => call('DELETE', `/teams/${id}/mem
export const integrations = (id) => call('GET', `/teams/${id}/integrations`);
export const createIntegration = (id, name) =>
call('POST', `/teams/${id}/integrations`, { body: { name } });
export const renameIntegration = (id, integrationID, name) =>
call('PATCH', `/teams/${id}/integrations/${integrationID}`, { body: { name } });
export const deleteIntegration = (id, integrationID) =>
call('DELETE', `/teams/${id}/integrations/${integrationID}`);
export const deadman = (id) => call('GET', `/teams/${id}/deadman`);
export const setDeadman = (id, body) => call('PUT', `/teams/${id}/deadman`, { body });
export const deadmanSwitches = (id) => call('GET', `/teams/${id}/deadman/switches`);
export const createDeadmanSwitch = (id, body) =>
call('POST', `/teams/${id}/deadman/switches`, { body });
export const deleteDeadmanSwitch = (id, switchID) =>
call('DELETE', `/teams/${id}/deadman/switches/${switchID}`);
export const escalation = (id) => call('GET', `/teams/${id}/escalation`);
export const setEscalation = (id, body) => call('PUT', `/teams/${id}/escalation`, { body });
+102 -2
View File
@@ -66,6 +66,8 @@ function render() {
h('button', { class: 'btn btn-ghost btn-icon back', type: 'button', 'aria-label': 'Back to queue', onclick: back },
icon('back')),
h('span', { class: 'crumb', text: currentID != null ? `Incident #${currentID}` : '' }),
inc && h('button', { class: 'btn btn-ghost btn-icon copy', type: 'button', 'aria-label': 'Copy incident', title: 'Copy incident (y)', onclick: copyIncident },
icon('copy')),
);
if (!inc) {
@@ -186,8 +188,9 @@ function alertItem(a) {
// ---------- timeline ----------
function eventText(ev) {
const person = ev.user_id != null ? who(ev.user_id, ev.username) : null;
// named spells users out instead of "you", for text that leaves this page.
function eventText(ev, named = false) {
const person = ev.user_id != null ? (named ? ev.username || 'someone' : who(ev.user_id, ev.username)) : null;
const strong = (t) => h('span', { class: 'who', text: t || 'someone' });
const alertName = () => {
const a = (inc.alerts || []).find((x) => x.id === ev.alert_id);
@@ -263,6 +266,100 @@ function timelineItem(ev) {
);
}
// ---------- copy ----------
const fence = (rows) => ['```', ...rows, '```'];
const pairs = (obj) => Object.entries(obj || {}).sort(([a], [b]) => a.localeCompare(b)).map(([k, v]) => `${k}=${v}`);
// incidentMarkdown is everything on this page as text that reads well in a chat
// or an agent prompt. Times are ISO 8601, since "3 min ago" means nothing once
// it has been pasted somewhere else.
function incidentMarkdown() {
const out = [`# Incident #${inc.id}: ${inc.title}`, ''];
const add = (k, v) => { if (v != null && v !== '') out.push(`- ${k}: ${v}`); };
add('Status', inc.status);
add('Severity', inc.severity);
add('Team', inc.team_name);
add('Assigned to', inc.assigned_to_id != null ? inc.assigned_to || 'someone' : 'unassigned');
add('Triggered', inc.triggered_at);
if (inc.acknowledged_at) add('Acknowledged', `${inc.acknowledged_at} by ${inc.acknowledged_by || 'someone'}`);
if (isOpen() && isFuture(inc.snoozed_until)) add('Snoozed until', inc.snoozed_until);
if (inc.escalation_level > 0) add('Escalation level', inc.escalation_level);
if (inc.resolved_at) add('Resolved', `${inc.resolved_at} (${inc.resolution_source === 'manual' ? 'manually' : 'all alerts stopped firing'})`);
if (inc.archived_at) add('Archived', inc.archived_at);
const group = pairs(inc.group_labels);
if (group.length) out.push('- Grouped by:', ...group.map((g) => ` - ${g}`));
const alerts = inc.alerts || [];
out.push('', `## Alerts (${alerts.length})`);
for (const a of alerts) {
out.push('', `### ${a.name} (${a.status})`);
out.push(`- Started: ${a.starts_at}`);
if (a.status === 'resolved' && a.ends_at) out.push(`- Ended: ${a.ends_at}`);
if (a.generator_url) out.push(`- Source: ${a.generator_url}`);
const labels = pairs(a.labels);
if (labels.length) out.push('', 'Labels:', ...fence(labels));
const annotations = Object.entries(a.annotations || {}).sort(([x], [y]) => x.localeCompare(y));
if (annotations.length) out.push('', 'Annotations:', ...fence(annotations.map(([k, v]) => `${k}: ${v}`)));
}
const sorted = [...events].sort((a, b) => Date.parse(a.created_at) - Date.parse(b.created_at) || a.id - b.id);
if (sorted.length) {
out.push('', '## Timeline', '');
for (const ev of sorted) {
const text = eventText(ev, true).map((f) => (f instanceof Node ? f.textContent : f)).join('');
out.push(`- ${ev.created_at} ${text}`);
if (isNote(ev) && ev.detail) {
const label = ev.type === 'resolution_note' ? ' (what fixed it)' : '';
out.push(...(label ? [label] : []), ...ev.detail.split('\n').map((l) => ` > ${l}`));
}
}
}
if (similarList.length) {
out.push('', '## Seen before', '', 'Earlier incidents with the same signature:');
for (const s of similarList) {
out.push(`- #${s.id} ${s.title} (resolved ${s.resolved_at})`);
for (const n of s.resolution_notes || []) {
out.push(' - What fixed it:', ...(n.detail || '').split('\n').map((l) => ` > ${l}`));
}
}
}
out.push('', `_Copied from Terminal Duty at ${new Date().toISOString()}_`, '');
return out.join('\n');
}
// writeClipboard falls back to execCommand: the async API needs a secure
// context, and this server is often reached over plain HTTP.
async function writeClipboard(text) {
try {
await navigator.clipboard.writeText(text);
return;
} catch {
// fall through
}
const ta = h('textarea', { readonly: true, 'aria-hidden': 'true', class: 'clip-buffer' });
ta.value = text;
document.body.append(ta);
ta.select();
try {
if (!document.execCommand('copy')) throw new Error('copy refused');
} finally {
ta.remove();
}
}
async function copyIncident() {
if (!inc) return;
try {
await writeClipboard(incidentMarkdown());
toast('Copied incident');
} catch {
toast('Could not copy', 'error');
}
}
// ---------- actions ----------
const isOpen = () => inc.status !== 'resolved';
@@ -470,10 +567,12 @@ async function moreMenu() {
items.push(item('user', 'Assign…', assign));
items.push(isSnoozed() ? item('bell', 'End snooze', unsnooze) : item('clock', 'Snooze…', snooze));
items.push(item('note', 'Add note…', addNote));
items.push(item('copy', 'Copy incident', copyIncident));
items.push(h('li', { class: 'menu-sep', role: 'separator' }));
items.push(item('checkCircle', 'Resolve…', resolve, 'danger'));
} else {
items.push(item('note', 'Add note…', addNote));
items.push(item('copy', 'Copy incident', copyIncident));
items.push(inc.archived_at ? item('undo', 'Unarchive', unarchive) : item('archive', 'Archive', archive));
}
@@ -502,6 +601,7 @@ export function key(e) {
case 'z': if (isOpen() && !isSnoozed()) snooze(); return true;
case 'Z': if (isSnoozed()) unsnooze(); return true;
case 'c': addNote(); return true;
case 'y': copyIncident(); return true;
case 'x': if (!isOpen()) (inc.archived_at ? unarchive() : archive()); return true;
default: return false;
}
+371 -136
View File
@@ -18,9 +18,9 @@
// than no form, but it is not the thing enforcing anything.
import * as api from './api.js';
import { h, clear, spinner, confirm, icon, openSheet, closeSheet, menuCard } from './ui.js';
import { h, clear, spinner, confirm, icon, openSheet, closeSheet, menuCard, badge, labelChip } from './ui.js';
import { state, currentTeam, users as allUsers, myID } from './state.js';
import { isoDate, addDays, mondayOf, initial } from './format.js';
import { isoDate, addDays, mondayOf, initial, ago, when, duration } from './format.js';
const view = () => document.getElementById('view-team');
@@ -52,12 +52,10 @@ let freshKey = null; // an integration key, shown once, until the view is left
export function show(route) {
const next = route?.tab ?? null;
// A different sub-section wants different data, so the old answer goes
// rather than being shown under the new heading until the fetch lands. The
// ladder draft goes with it: it is an edit of the page being left.
// rather than being shown under the new heading until the fetch lands.
if (next !== tab) {
tab = next;
data = null;
draft = null;
}
if (!data) clear(view(), subnav(), spinner());
refresh();
@@ -111,14 +109,14 @@ async function load(id) {
return { members, escalation };
}
if (tab === 'sources') return { integrations: await api.integrations(id) };
if (tab === 'deadman') return { deadman: await api.deadman(id) };
if (tab === 'deadman') return { deadman: await api.deadmanSwitches(id) };
const grid = gridDays();
const [members, integrations, escalation, deadman, schedule] = await Promise.all([
api.teamMembers(id),
api.integrations(id),
api.escalation(id),
api.deadman(id),
api.deadmanSwitches(id),
api.schedule(id, isoDate(grid.start), isoDate(addDays(grid.start, grid.count - 1))),
]);
return { members, integrations, escalation, deadman, schedule };
@@ -185,7 +183,6 @@ function teamPicker() {
teamID = Number(select.value);
data = null;
freshKey = null;
draft = null;
refresh();
});
return h('div', { class: 'card' }, h('h2', { text: 'Team' }), select);
@@ -203,8 +200,8 @@ function overview() {
const levels = (data.escalation?.levels || []).length;
const keys = (data.integrations || []).length;
const unused = (data.integrations || []).filter((i) => !i.last_used_at).length;
const switches = (data.deadman?.matchers || '')
.split(';').map((x) => x.trim()).filter(Boolean).length;
const switches = (data.deadman || []).length;
const dead = (data.deadman || []).filter((s) => s.status === 'dead').length;
return h('div', { class: 'overview-menu' },
menuCard('/team/rota', 'Rota', null,
@@ -220,7 +217,9 @@ function overview() {
? (unused ? `${unused} of them never used.` : 'All in use.')
: 'No key yet, so nothing can reach this team.'),
menuCard('/team/deadman', 'Dead man’s switches', switches || null,
switches ? 'Alerts whose absence opens an incident.' : 'Nothing watched.'),
switches
? (dead ? `${dead} of them silent.` : 'All quiet, as they should be.')
: 'Nothing watched.'),
);
}
@@ -455,57 +454,124 @@ function memberSelect(selected) {
// --- escalation ------------------------------------------------------------
// The ladder is edited as a whole and sent as a whole, because the API replaces
// it wholesale: the levels are an order, and patching one rung would leave the
// numbering of the others undecided.
let draft = null;
const LEVEL_STATUS = {
ready: { label: 'Ready', hint: 'Somebody here can be woken.' },
escalating: { label: 'Escalating', hint: 'An unanswered incident has climbed to this level.' },
unreachable: { label: 'Pages nobody', hint: 'Nobody on this level can be woken right now.' },
};
const levelBadge = (status) => statusBadge(LEVEL_STATUS, status, 'ready');
// One target as the list shows it: who it means today, and why it would not
// wake them if it would not.
function targetLine(t) {
const label = t.kind === 'oncall'
? `On call${t.username ? ` · ${t.username}` : ''}`
: (t.username || 'Unknown person');
return h('div', { class: 'target-line' },
h('span', { text: label }),
t.problem && h('span', { class: 'target-problem', text: t.problem }));
}
function escalationCard() {
const esc = data.escalation;
if (!draft) {
draft = {
repeat_count: esc.repeat_count || 0,
fallback_topic: esc.fallback_topic || '',
levels: (esc.levels || []).map((l) => ({
timeout_seconds: l.timeout_seconds,
targets: (l.targets || []).map((t) => ({ kind: t.kind, user_id: t.user_id })),
const esc = data.escalation || {};
const levels = esc.levels || [];
const rows = levels.map((l) => h('tr', {},
h('td', {}, h('strong', { text: `Level ${l.position}` })),
h('td', {}, levelBadge(l.status)),
h('td', { class: 'wrap' }, ...l.targets.map(targetLine)),
h('td', { class: 'muted small', text: duration(l.timeout_seconds * 1000) }),
h('td', { class: 'small' }, l.waiting?.length
? l.waiting.flatMap((id, i) => [i > 0 && ', ', h('a', { href: `/incidents/${id}`, text: `#${id}` })])
: h('span', { class: 'muted', text: '—' })),
));
const facts = [];
if (levels.length) {
const n = esc.repeat_count || 0;
if (n) facts.push(`Then the whole ladder repeats ${n} more ${n === 1 ? 'time' : 'times'}.`);
facts.push(esc.fallback_topic
? ['Finally the ntfy topic ', h('code', { text: esc.fallback_topic }), ' is paged once.']
: 'No fallback topic: after the last level the chain just ends.');
facts.push(esc.last_escalated_at
? ['Last escalated ',
h('span', { title: when(esc.last_escalated_at), text: ago(esc.last_escalated_at) }),
' on ', h('a', { href: `/incidents/${esc.last_escalated_incident_id}`, text: `#${esc.last_escalated_incident_id}` }), '.']
: 'Nothing has needed to escalate yet.');
}
return h('div', { class: 'card' },
h('div', { class: 'card-head' },
h('h2', { text: 'Escalation' }),
isOwner() && h('button', {
class: 'btn', type: 'button', onclick: openLadderEditor,
text: levels.length ? 'Edit ladder' : 'Set up ladder',
})),
};
}
h('p', { class: 'muted small' },
'When a level’s wait passes and nobody has acknowledged, the next level is ',
'paged. Acknowledging or resolving stops it; snoozing pauses it.'),
levels.length
? h('div', { class: 'table-scroll' },
h('table', { class: 'admin-table status-table' },
h('thead', {}, h('tr', {},
h('th', { text: 'Level' }), h('th', { text: 'Status' }), h('th', { text: 'Pages' }),
h('th', { text: 'Then after' }), h('th', { text: 'Waiting now' }))),
h('tbody', {}, rows)))
: h('p', { class: 'muted' },
'No ladder. An unacknowledged incident re-pages the same person every ',
'reminder interval and nobody else is woken.'),
...facts.map((f) => h('p', { class: 'muted small' }, f)),
);
}
const body = [];
if (!draft.levels.length) {
body.push(h('p', { class: 'muted' },
'No ladder. An unacknowledged incident re-pages the same person every ',
'reminder interval and nobody else is woken.'));
}
// The ladder is edited as a whole and sent as a whole, because the API replaces
// it wholesale: the levels are an order, and patching one rung would leave the
// numbering of the others undecided. The draft lives in the sheet, so a poll of
// the page underneath cannot throw away half an edit.
function openLadderEditor() {
const esc = data.escalation || {};
const draft = {
repeat_count: esc.repeat_count || 0,
fallback_topic: esc.fallback_topic || '',
levels: (esc.levels || []).map((l) => ({
timeout_seconds: l.timeout_seconds,
targets: (l.targets || []).map((t) => ({ kind: t.kind, user_id: t.user_id })),
})),
};
draft.levels.forEach((level, i) => {
body.push(h('div', { class: 'ladder-level' },
h('div', { class: 'ladder-head' },
h('strong', { text: `Level ${i + 1}` }),
isOwner() && h('button', {
class: 'btn-sm danger', type: 'button', text: 'Remove',
onclick: () => { draft.levels.splice(i, 1); render(); },
})),
h('label', {}, 'Wait ', minutesInput(level.timeout_seconds, (secs) => {
level.timeout_seconds = secs;
}), ' before the next level'),
h('div', { class: 'ladder-targets' },
...level.targets.map((t, ti) => targetRow(level, t, ti)),
isOwner() && h('button', {
class: 'btn-sm', type: 'button', text: '+ target',
onclick: () => { level.targets.push({ kind: 'oncall' }); render(); },
})),
));
});
const body = h('div', { class: 'ladder-editor' });
const problem = h('p', { class: 'load-error', hidden: true });
if (isOwner()) {
body.push(h('button', {
const paint = () => {
const parts = [];
if (!draft.levels.length) {
parts.push(h('p', { class: 'muted small' }, 'No levels yet. Add the first one.'));
}
draft.levels.forEach((level, i) => {
parts.push(h('div', { class: 'ladder-level' },
h('div', { class: 'ladder-head' },
h('strong', { text: `Level ${i + 1}` }),
h('button', {
class: 'btn-sm danger', type: 'button', text: 'Remove',
onclick: () => { draft.levels.splice(i, 1); paint(); },
})),
h('label', {}, 'Wait ', minutesInput(level.timeout_seconds, (secs) => {
level.timeout_seconds = secs;
}), ' before the next level'),
h('div', { class: 'ladder-targets' },
...level.targets.map((t, ti) => targetRow(level, t, ti, paint)),
h('button', {
class: 'btn-sm', type: 'button', text: '+ target',
onclick: () => { level.targets.push({ kind: 'oncall' }); paint(); },
})),
));
});
parts.push(h('button', {
class: 'btn-sm', type: 'button', text: '+ level',
onclick: () => {
draft.levels.push({ timeout_seconds: 300, targets: [{ kind: 'oncall' }] });
render();
paint();
},
}));
@@ -518,31 +584,43 @@ function escalationCard() {
type: 'text', value: draft.fallback_topic, placeholder: 'terdut-oncall-all',
oninput: (e) => { draft.fallback_topic = e.target.value; },
});
body.push(h('label', {}, 'Repeat the whole ladder ', repeat, ' more times'));
body.push(h('label', {}, 'Then page this ntfy topic once ', fallback));
body.push(h('button', {
class: 'btn', type: 'button', text: 'Save ladder',
onclick: () => act(() => api.setEscalation(teamID, draft), { resetDraft: true }),
}));
}
parts.push(h('label', {}, 'Repeat the whole ladder ', repeat, ' more times'));
parts.push(h('label', {}, 'Then page this ntfy topic once ', fallback));
clear(body, ...parts);
};
paint();
return h('div', { class: 'card' },
h('h2', { text: 'Escalation' }),
h('p', { class: 'muted small' },
'When a level’s wait passes and nobody has acknowledged, the next level is ',
'paged. Acknowledging or resolving stops it; snoozing pauses it.'),
...body,
);
const save = h('button', { class: 'btn btn-primary', type: 'button', text: 'Save ladder' });
save.addEventListener('click', async () => {
try {
await api.setEscalation(teamID, draft);
} catch (err) {
problem.textContent = err.message;
problem.hidden = false;
return;
}
closeSheet(true);
refresh();
});
openSheet(() => [
h('h2', { class: 'sheet-title', text: 'Edit ladder' }),
body,
problem,
h('div', { class: 'sheet-actions' },
h('button', { class: 'btn', type: 'button', text: 'Cancel', onclick: () => closeSheet(false) }),
save),
]);
}
function targetRow(level, target, index) {
function targetRow(level, target, index, repaint) {
const kind = h('select', {},
h('option', { value: 'oncall', text: 'Whoever is on call', selected: target.kind === 'oncall' }),
h('option', { value: 'user', text: 'A specific person', selected: target.kind === 'user' }));
kind.addEventListener('change', () => {
target.kind = kind.value;
target.user_id = kind.value === 'user' ? (data.members[0] || {}).user_id : undefined;
render();
repaint();
});
const who = target.kind === 'user'
@@ -553,10 +631,10 @@ function targetRow(level, target, index) {
}
return h('div', { class: 'target-row' }, kind, who,
isOwner() && h('button', {
h('button', {
class: 'btn-sm danger', type: 'button', text: '×',
title: 'Remove this target',
onclick: () => { level.targets.splice(index, 1); render(); },
onclick: () => { level.targets.splice(index, 1); repaint(); },
}));
}
@@ -574,33 +652,54 @@ function minutesInput(seconds, onChange) {
function integrationsCard() {
const rows = (data.integrations || []).map((i) =>
h('tr', {},
h('td', {}, h('strong', { text: i.name })),
h('td', { class: 'muted small', text: i.kind }),
h('td', { class: 'muted small', text: i.last_used_at ? 'in use' : 'never used' }),
h('td', {}, isOwner() && h('button', {
class: 'btn-sm danger', type: 'button', text: 'Revoke',
onclick: async () => {
if (!(await confirm({
title: `Revoke ${i.name}?`,
text: 'Anything posting with this key stops delivering immediately.',
confirmLabel: 'Revoke',
danger: true,
}))) return;
act(() => api.deleteIntegration(teamID, i.id));
},
})),
h('td', {}, sourceBadge(i.status)),
h('td', { class: 'wrap' },
h('strong', { text: i.name }),
h('div', { class: 'muted small', text: i.kind })),
// When the key last posted, and when an alert last arrived on it. They
// differ: a payload with nothing usable in it stamps only the first.
h('td', { class: 'muted small' }, timeCell(i.last_used_at)),
h('td', { class: 'muted small' }, timeCell(i.last_alert_at)),
h('td', { class: 'muted small num', title: 'Distinct alerts refreshed in the last 24 hours',
text: String(i.alerts_24h ?? 0) }),
h('td', { class: 'muted small' }, h('span', { title: when(i.created_at), text: ago(i.created_at) })),
h('td', {}, isOwner() && h('div', { class: 'row-actions' },
h('button', {
class: 'btn-sm', type: 'button', text: 'Rename', onclick: () => openRenameSource(i),
}),
h('button', {
class: 'btn-sm danger', type: 'button', text: 'Revoke',
onclick: async () => {
if (!(await confirm({
title: `Revoke ${i.name}?`,
text: 'Anything posting with this key stops delivering immediately. Alerts it already delivered stay.',
confirmLabel: 'Revoke',
danger: true,
}))) return;
act(() => api.deleteIntegration(teamID, i.id));
},
}))),
));
return h('div', { class: 'card' },
h('h2', { text: 'Alert sources' }),
h('div', { class: 'card-head' },
h('h2', { text: 'Alert sources' }),
isOwner() && h('button', {
class: 'btn', type: 'button', text: 'New source', onclick: openNewSource,
})),
h('p', { class: 'muted small' },
'Alerts arrive on an integration key, which says both that the sender may ',
'post and which team the alerts belong to.'),
rows.length
? h('table', { class: 'admin-table' }, h('tbody', {}, rows))
: h('p', { class: 'muted', text: 'No alert source yet, so nothing can reach this team.' }),
freshKey && newKeyPanel(),
isOwner() && !freshKey && newIntegrationForm(),
rows.length
? h('div', { class: 'table-scroll' },
h('table', { class: 'admin-table status-table' },
h('thead', {}, h('tr', {},
h('th', { text: 'Status' }), h('th', { text: 'Source' }),
h('th', { text: 'Last webhook' }), h('th', { text: 'Last alert' }),
h('th', { class: 'num', text: 'Alerts 24h' }), h('th', { text: 'Created' }), h('th'))),
h('tbody', {}, rows)))
: h('p', { class: 'muted', text: 'No alert source yet, so nothing can reach this team.' }),
);
}
@@ -631,66 +730,204 @@ function newKeyPanel() {
);
}
function newIntegrationForm() {
const name = h('input', { type: 'text', placeholder: 'prod alertmanager', required: true });
const form = h('form', { class: 'inline-form' }, name,
h('button', { class: 'btn', type: 'submit', text: 'Add' }));
// A sheet with one name field, for adding a source and for renaming one: the two
// differ only in what they call and what they put in the box.
function openNameSheet({ title, submit, value, run }) {
const name = h('input', {
type: 'text', placeholder: 'prod alertmanager', required: true, value, autofocus: true,
});
const problem = h('p', { class: 'load-error', hidden: true });
const form = h('form', { class: 'stacked-form' },
h('label', {}, 'Name ', name),
problem,
h('div', { class: 'sheet-actions' },
h('button', { class: 'btn', type: 'button', text: 'Cancel', onclick: () => closeSheet(false) }),
h('button', { class: 'btn btn-primary', type: 'submit', text: submit })));
form.addEventListener('submit', async (e) => {
e.preventDefault();
try {
freshKey = await api.createIntegration(teamID, name.value.trim());
await refresh();
await run(name.value.trim());
} catch (err) {
error = err.message;
render();
problem.textContent = err.message;
problem.hidden = false;
return;
}
closeSheet(true);
refresh();
});
openSheet(() => [h('h2', { class: 'sheet-title', text: title }), form]);
}
function openNewSource() {
openNameSheet({
title: 'New source', submit: 'Add source', value: '',
// The key comes back once, and the card shows it until dismissed.
run: async (name) => { freshKey = await api.createIntegration(teamID, name); },
});
}
function openRenameSource(i) {
openNameSheet({
title: `Rename ${i.name}`, submit: 'Rename', value: i.name,
run: (name) => api.renameIntegration(teamID, i.id, name),
});
return form;
}
// --- dead man's switches ---------------------------------------------------
// Status badges, shared by the Sources and Switches lists: a table of label and
// hint per status, and one function to draw it. Module-level, so the two cards
// can be defined in either order.
const SWITCH_STATUS = {
healthy: { label: 'Healthy', hint: 'Heard from within its timeout.' },
dead: { label: 'Dead', hint: 'Silent for longer than its timeout.' },
dormant: { label: 'Dormant', hint: 'Nothing has matched yet, so there is nothing to lose.' },
};
const SOURCE_STATUS = {
active: { label: 'Active', hint: 'Posted within the last day.' },
quiet: { label: 'Quiet', hint: 'Has posted, but not in the last day. Nothing firing is a fine reason.' },
never: { label: 'Never used', hint: 'Nothing has been posted with this key yet.' },
};
function statusBadge(table, status, fallback) {
const s = table[status] || table[fallback];
const el = badge(s.label, `st-${status}`);
el.title = s.hint;
return el;
}
const switchBadge = (status) => statusBadge(SWITCH_STATUS, status, 'dormant');
const sourceBadge = (status) => statusBadge(SOURCE_STATUS, status, 'never');
const timeCell = (iso) => iso
? h('span', { title: when(iso), text: ago(iso) })
: h('span', { class: 'muted', text: 'never' });
// When it last opened an incident. An incident that is still open is a link,
// because that is the thing somebody looking at a red row wants next.
const triggeredCell = (iso, incidentID) => {
if (!iso) return h('span', { class: 'muted', text: 'never' });
return incidentID
? h('a', { href: `/incidents/${incidentID}`, title: when(iso) }, `#${incidentID} · ${ago(iso)}`)
: h('span', { title: when(iso), text: ago(iso) });
};
function switchRows(sw) {
const main = h('tr', {},
h('td', {}, switchBadge(sw.status)),
h('td', { class: 'wrap' },
h('strong', { text: sw.name }),
sw.name !== sw.matcher && h('div', { class: 'muted small' }, h('code', { text: sw.matcher }))),
h('td', { class: 'muted small' }, timeCell(sw.last_heartbeat_at)),
h('td', { class: 'muted small' }, triggeredCell(sw.last_triggered_at, sw.open_incident_id)),
h('td', { class: 'muted small', text: duration(sw.timeout_seconds * 1000) }),
h('td', {}, isOwner() && h('button', {
class: 'btn-sm danger', type: 'button', text: 'Remove',
onclick: async () => {
if (!(await confirm({
title: `Remove ${sw.name}?`,
text: 'It stops being watched. An incident it already opened stays open until it is resolved.',
confirmLabel: 'Remove',
danger: true,
}))) return;
act(() => api.deleteDeadmanSwitch(teamID, sw.id));
},
})),
);
// One heartbeat is the switch's own times; several are worth telling apart,
// since a live cluster must not hide a dead one.
const sources = sw.sources.length > 1
? sw.sources.map((src) => h('tr', { class: 'source-row' },
h('td', {}, switchBadge(src.status)),
h('td', { class: 'source-labels' },
...Object.entries(src.labels || {})
.filter(([k]) => k !== 'alertname')
.map(([k, v]) => labelChip(k, v)),
!Object.keys(src.labels || {}).some((k) => k !== 'alertname')
&& h('code', { class: 'small', text: src.fingerprint })),
h('td', { class: 'muted small' }, timeCell(src.last_heartbeat_at)),
h('td', { class: 'muted small' }, triggeredCell(src.last_triggered_at, src.incident_id)),
h('td'), h('td')))
: [];
return [main, ...sources];
}
function deadmanCard() {
const d = data.deadman || {};
const matchers = h('input', {
type: 'text', value: d.matchers || '', placeholder: 'alertname=Watchdog',
class: 'wide',
});
const timeout = h('input', {
type: 'number', min: '0', class: 'setting-value',
value: String(Math.round((d.timeout_seconds || 0) / 60)),
});
const severity = h('select', {},
...['critical', 'error', 'warning', 'info'].map((s) =>
h('option', { value: s, text: s, selected: (d.severity || 'critical') === s })));
const form = h('form', { class: 'stacked-form' },
h('label', {}, 'Heartbeat alerts ', matchers),
h('label', {}, 'Declare dead after ', timeout, ' minutes of silence'),
h('label', {}, 'Open the incident at severity ', severity),
h('button', { class: 'btn', type: 'submit', text: 'Save switches' }));
form.addEventListener('submit', (e) => {
e.preventDefault();
act(() => api.setDeadman(teamID, {
matchers: matchers.value.trim(),
timeout_seconds: Number(timeout.value) * 60,
severity: severity.value,
}));
});
const switches = data.deadman || [];
return h('div', { class: 'card' },
h('h2', { text: 'Dead man’s switches' }),
h('div', { class: 'card-head' },
h('h2', { text: 'Dead man’s switches' }),
isOwner() && h('button', {
class: 'btn', type: 'button', text: 'New switch', onclick: openNewSwitch,
})),
h('p', { class: 'muted small' },
'Alerts whose ABSENCE is the signal. Receiving one opens nothing; going ',
'quiet for longer than the timeout opens an incident. ',
h('code', { text: 'alertname=Watchdog,cluster=prod; alertname=EdgeHeartbeat' }),
' — semicolons separate switches, commas separate conditions, and every ',
'switch must name an alertname. Leave empty to watch nothing.'),
isOwner() ? form : h('p', { class: 'muted', text: d.matchers || 'Nothing watched.' }),
'quiet for longer than the switch’s timeout opens an incident.'),
switches.length
? h('div', { class: 'table-scroll' },
h('table', { class: 'admin-table status-table' },
h('thead', {}, h('tr', {},
h('th', { text: 'Status' }), h('th', { text: 'Switch' }),
h('th', { text: 'Last heartbeat' }), h('th', { text: 'Last triggered' }),
h('th', { text: 'Silent after' }), h('th'))),
h('tbody', {}, switches.flatMap(switchRows))))
: h('p', { class: 'muted', text: 'Nothing watched.' }),
);
}
// The form lives in the sheet, not on the page: most visits are to look at the
// list, and a form that is always open is the page this replaced.
function openNewSwitch() {
const name = h('input', { type: 'text', placeholder: 'Prod Watchdog', autofocus: true });
const matcher = h('input', {
type: 'text', placeholder: 'alertname=Watchdog,cluster=prod', class: 'wide', required: true,
});
const timeout = h('input', {
type: 'number', min: '1', value: '15', class: 'setting-value', required: true,
});
const severity = h('select', {},
...['critical', 'error', 'warning', 'info'].map((s) => h('option', { value: s, text: s })));
const problem = h('p', { class: 'load-error', hidden: true });
const form = h('form', { class: 'stacked-form' },
h('label', {}, 'Name (optional) ', name),
h('label', {}, 'Heartbeat alert ', matcher),
h('p', { class: 'muted small' },
'Conditions are ', h('code', { text: 'label=value' }), ' separated by commas, and one ',
'must be ', h('code', { text: 'alertname' }), '. Every distinct label set that ',
'matches is watched on its own.'),
h('label', {}, 'Declare dead after ', timeout, ' minutes of silence'),
h('label', {}, 'Open the incident at severity ', severity),
problem,
h('div', { class: 'sheet-actions' },
h('button', { class: 'btn', type: 'button', text: 'Cancel', onclick: () => closeSheet(false) }),
h('button', { class: 'btn btn-primary', type: 'submit', text: 'Add switch' })));
form.addEventListener('submit', async (e) => {
e.preventDefault();
try {
await api.createDeadmanSwitch(teamID, {
name: name.value.trim(),
matcher: matcher.value.trim(),
timeout_seconds: Math.round(Number(timeout.value) * 60),
severity: severity.value,
});
} catch (err) {
problem.textContent = err.message;
problem.hidden = false;
return;
}
closeSheet(true);
refresh();
});
openSheet(() => [h('h2', { class: 'sheet-title', text: 'New switch' }), form]);
}
// --- members ---------------------------------------------------------------
function membersCard() {
@@ -735,14 +972,12 @@ function membersCard() {
// act runs a write and reloads. Errors are shown rather than thrown away: a
// 409 from the last-owner guard or the schedule's conflict rule is the server
// explaining itself, and the reader needs to see it.
async function act(fn, { resetDraft = false } = {}) {
async function act(fn) {
try {
await fn();
error = null;
if (resetDraft) draft = null;
} catch (err) {
error = err.message;
}
if (!resetDraft) draft = null;
await refresh();
}
+1
View File
@@ -48,6 +48,7 @@ const ICONS = {
note: ['M5 4h14v12l-4 4H5z', 'M15 20v-4h4', 'M9 9h6M9 13h4'],
archive: ['M3.5 5h17v4h-17z', 'M5 9v10h14V9', 'M10 13h4'],
flag: ['M5 21V4', 'M5 4h11l-2 4 2 4H5'],
copy: ['rect:9,9,11,11,2', 'M5 15V6a2 2 0 0 1 2-2h9'],
trash: ['M4 7h16', 'M9 7V4h6v3', 'M6 7l1 13h10l1-13'],
chevronLeft: ['M15 18l-6-6 6-6'],
chevronRight: ['M9 6l6 6-6 6'],