Turn incoming alerts into incidents
Release / build (amd64, linux) (push) Failing after 11s
Release / build (amd64, darwin) (push) Failing after 12s
Release / build (arm64, darwin) (push) Failing after 11s
Release / build (arm64, linux) (push) Failing after 11s
Release / release (push) Has been skipped
Release / chart (push) Failing after 13s
Release / docker (push) Failing after 19s

The alerts row was both Alertmanager's record and the human work queue, and
the two have different owners. The webhook upsert rewrites that row on every
notification; acknowledgement, comments and archiving were columns on it that
the upsert happened not to touch. So an alert that resolved and re-fired days
later still read as acknowledged by whoever acked the first occurrence — the
ack outlived the thing it referred to. Nothing recorded transitions either:
rows are mutated in place, so there was no timeline and no way to compute how
long anything took.

Alerts are now read-only signal records with two states, and incidents are
the work item: triggered, acknowledged or resolved, with an assignee, a
snooze, notes and an append-only timeline. Many alerts map to one incident,
and a new occurrence opens a new incident, which is what makes a stale ack
impossible rather than merely unlikely.

Correlation uses Alertmanager's own groupKey. It already grouped the alerts
according to the group_by routing tree the operator configured and sends the
result on every webhook, where it was being discarded; adopting it means
changing group_by in alertmanager.yml changes correlation here, with no
second grouping scheme to configure and keep in sync.

An incident opens only when an alert transitions into firing — an unseen
fingerprint, a newer startsAt, or a resolved alert starting again. The
unchanged notifications Alertmanager re-sends every repeat_interval are none
of those. That rule is what lets manual resolution be terminal: without it,
closing an incident by hand would be undone by the next re-send of an alert
that never stopped firing, and the button would be a lie. Snooze covers the
"not now" case instead. Incidents otherwise resolve by cascade, once every
alert under them has stopped firing, whether by webhook or by expiry.

New incidents are assigned to whoever holds today's schedule entry. The
schedule table has existed since the first release with nothing reading it.

Also here, following from the split:

  - Incident severity is a high-water mark over its alerts, never lowered.
    An incident that hit critical was a critical incident, and downgrading a
    live one would demote it in the queue while the work is still open.
  - /api/stats/incidents reports MTTA and MTTR, null rather than zero until
    there is something to average. Neither was computable before.
  - Alert archiving becomes sweeper-only housekeeping; the archive people
    interact with is the incident's.

Breaking: the alert acknowledge, archive and comment endpoints are gone, and
the alert object drops the acknowledgement fields and gains incident_id. The
README maps each removed endpoint to its replacement. Migration 008 backfills
an incident per existing alert, archived ones included so no comment is
orphaned, carrying acknowledgements across and turning comments into timeline
notes.

Both documented alert contracts are untouched: received_at still advances on
every accepted payload, re-sends included, and resolution_source still says
how much to trust ends_at. The upsert is byte-for-byte what it was, now
running inside the ingest transaction.
This commit is contained in:
Niklas Ye
2026-07-30 17:02:13 +02:00
parent a602ff3efc
commit 279ef6cf8b
15 changed files with 2537 additions and 477 deletions
+310 -62
View File
@@ -1,6 +1,7 @@
package api
import (
"context"
"database/sql"
"encoding/json"
"log"
@@ -17,9 +18,17 @@ const (
// amPayload mirrors the Alertmanager webhook v4 payload.
type amPayload struct {
Version string `json:"version"`
Status string `json:"status"`
Alerts []amAlert `json:"alerts"`
Version string `json:"version"`
Status string `json:"status"`
// GroupKey and GroupLabels are how alerts get correlated into incidents.
// Alertmanager has already done the grouping work according to the group_by
// routing tree the operator configured, so we adopt its answer instead of
// inventing a second grouping scheme here.
GroupKey string `json:"groupKey"`
GroupLabels map[string]string `json:"groupLabels"`
Alerts []amAlert `json:"alerts"`
}
type amAlert struct {
@@ -32,6 +41,23 @@ type amAlert struct {
Fingerprint string `json:"fingerprint"`
}
// ingested records what actually happened to one alert of a payload, which is
// what decides whether an incident opens.
type ingested struct {
id int64
name string
firing bool
// newOccurrence marks an alert that transitioned *into* firing: a
// fingerprint we had never seen, a newer startsAt, or a resolved alert that
// started again. A repeat_interval re-send of an already-firing alert is
// none of these, which is what keeps a manually resolved incident closed.
newOccurrence bool
// justResolved marks the firing → resolved edge, worth a timeline entry.
justResolved bool
}
func handleAlertmanagerWebhook(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
var payload amPayload
@@ -40,67 +66,289 @@ func handleAlertmanagerWebhook(db *sql.DB) http.HandlerFunc {
return
}
now := time.Now().Unix()
for _, a := range payload.Alerts {
name := a.Labels["alertname"]
labelsJSON, _ := json.Marshal(a.Labels)
annotationsJSON, _ := json.Marshal(a.Annotations)
// Zero time ("0001-01-01T00:00:00Z") means "no end known" — that is the
// convention of Alertmanager's ingest API. Outgoing notifications
// normally carry a real future endsAt instead, which is the watermark
// the sweeper uses to expire alerts that stop being refreshed.
var endsAtUnix *int64
if a.EndsAt.Year() > 1 {
t := a.EndsAt.Unix()
endsAtUnix = &t
}
var resolutionSource *string
if a.Status == "resolved" {
s := resolutionAlertmanager
resolutionSource = &s
}
// The WHERE clause discards payloads that describe an alert instance
// older than the stored one. Alertmanager retries failed notifications,
// so a stale firing retry can arrive after the resolved one; it carries
// the same startsAt, whereas a genuine re-fire carries a newer one.
// Within a single instance, resolution is terminal.
_, err := db.ExecContext(r.Context(), `
INSERT INTO alerts
(fingerprint, name, status, labels, annotations, starts_at, ends_at,
generator_url, received_at, resolution_source)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(fingerprint) DO UPDATE SET
status = excluded.status,
labels = excluded.labels,
annotations = excluded.annotations,
starts_at = excluded.starts_at,
ends_at = excluded.ends_at,
generator_url = excluded.generator_url,
-- Load-bearing: advancing received_at on every accepted
-- payload, re-sends included, is the documented liveness
-- heartbeat clients and the sweeper both read. Removing it
-- is a breaking API change — see models.Alert.ReceivedAt.
received_at = excluded.received_at,
resolution_source = excluded.resolution_source,
-- A re-fire makes the alert current again, so it leaves the archive.
archived_at = CASE WHEN excluded.status = 'firing'
THEN NULL ELSE alerts.archived_at END
WHERE excluded.starts_at > alerts.starts_at
OR (excluded.starts_at = alerts.starts_at
AND NOT (alerts.status = 'resolved' AND excluded.status = 'firing'))`,
a.Fingerprint, name, a.Status,
string(labelsJSON), string(annotationsJSON),
a.StartsAt.Unix(), endsAtUnix,
a.GeneratorURL, now, resolutionSource,
)
if err != nil {
log.Printf("upsert alert %s: %v", a.Fingerprint, err)
}
// Alertmanager retries anything that is not 2xx, and a retry of a payload
// we failed to store is more useful than an error it cannot act on — so
// failures are logged, not surfaced.
if err := ingest(r.Context(), db, payload); err != nil {
log.Printf("webhook ingest (group %q): %v", payload.GroupKey, err)
}
w.WriteHeader(http.StatusOK)
}
}
// ingest stores a payload's alerts and reconciles the incident for its group.
// The whole payload is one transaction: an incident that opened but whose alerts
// failed to link would be a work item nobody could act on.
func ingest(ctx context.Context, db *sql.DB, payload amPayload) error {
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return err
}
defer tx.Rollback() //nolint:errcheck
accepted, err := upsertAlerts(ctx, tx, payload.Alerts)
if err != nil {
return err
}
// touched collects every incident this payload affected, so severity and the
// resolution cascade are recomputed once per incident at the end.
touched := map[int64]bool{}
incidentID, err := incidentForGroup(ctx, tx, payload, accepted)
if err != nil {
return err
}
if incidentID != 0 {
touched[incidentID] = true
for _, a := range accepted {
if !a.firing {
continue
}
if err := linkAlert(ctx, tx, incidentID, a.id); err != nil {
return err
}
}
}
for _, a := range accepted {
if !a.justResolved {
continue
}
id, err := openIncidentForAlert(ctx, tx, a.id)
if err != nil {
return err
}
if id == 0 {
continue
}
touched[id] = true
alertID := a.id
if err := logEvent(ctx, tx, id, evAlertResolved, nil, &alertID, nil); err != nil {
return err
}
}
for id := range touched {
if err := refreshSeverity(ctx, tx, id); err != nil {
return err
}
if _, err := resolveIfSettled(ctx, tx, id); err != nil {
return err
}
}
return tx.Commit()
}
// upsertAlerts stores each alert of a payload and reports what changed. Payloads
// the ordering guard rejected are left out entirely.
func upsertAlerts(ctx context.Context, tx *sql.Tx, alerts []amAlert) ([]ingested, error) {
now := time.Now().Unix()
accepted := make([]ingested, 0, len(alerts))
for _, a := range alerts {
name := a.Labels["alertname"]
labelsJSON, _ := json.Marshal(a.Labels)
annotationsJSON, _ := json.Marshal(a.Annotations)
// The stored state has to be read before the upsert overwrites it: it is
// the only way to tell a genuine new occurrence from a re-send.
var prevStatus string
var prevStartsAt int64
existed := true
switch err := tx.QueryRowContext(ctx,
"SELECT status, starts_at FROM alerts WHERE fingerprint = ?", a.Fingerprint,
).Scan(&prevStatus, &prevStartsAt); {
case err == sql.ErrNoRows:
existed = false
case err != nil:
return nil, err
}
// Zero time ("0001-01-01T00:00:00Z") means "no end known" — that is the
// convention of Alertmanager's ingest API. Outgoing notifications
// normally carry a real future endsAt instead, which is the watermark
// the sweeper uses to expire alerts that stop being refreshed.
var endsAtUnix *int64
if a.EndsAt.Year() > 1 {
t := a.EndsAt.Unix()
endsAtUnix = &t
}
var resolutionSource *string
if a.Status == "resolved" {
s := resolutionAlertmanager
resolutionSource = &s
}
// The WHERE clause discards payloads that describe an alert instance
// older than the stored one. Alertmanager retries failed notifications,
// so a stale firing retry can arrive after the resolved one; it carries
// the same startsAt, whereas a genuine re-fire carries a newer one.
// Within a single instance, resolution is terminal.
if _, err := tx.ExecContext(ctx, `
INSERT INTO alerts
(fingerprint, name, status, labels, annotations, starts_at, ends_at,
generator_url, received_at, resolution_source)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(fingerprint) DO UPDATE SET
status = excluded.status,
labels = excluded.labels,
annotations = excluded.annotations,
starts_at = excluded.starts_at,
ends_at = excluded.ends_at,
generator_url = excluded.generator_url,
-- Load-bearing: advancing received_at on every accepted
-- payload, re-sends included, is the documented liveness
-- heartbeat clients and the sweeper both read. Removing it
-- is a breaking API change — see models.Alert.ReceivedAt.
received_at = excluded.received_at,
resolution_source = excluded.resolution_source,
-- A re-fire makes the alert current again, so it leaves the archive.
archived_at = CASE WHEN excluded.status = 'firing'
THEN NULL ELSE alerts.archived_at END
WHERE excluded.starts_at > alerts.starts_at
OR (excluded.starts_at = alerts.starts_at
AND NOT (alerts.status = 'resolved' AND excluded.status = 'firing'))`,
a.Fingerprint, name, a.Status,
string(labelsJSON), string(annotationsJSON),
a.StartsAt.Unix(), endsAtUnix,
a.GeneratorURL, now, resolutionSource,
); err != nil {
return nil, err
}
var id int64
var curStatus string
var curStartsAt int64
if err := tx.QueryRowContext(ctx,
"SELECT id, status, starts_at FROM alerts WHERE fingerprint = ?", a.Fingerprint,
).Scan(&id, &curStatus, &curStartsAt); err != nil {
return nil, err
}
// The upsert copies status and starts_at straight from the payload, so a
// row that does not match it is one the ordering guard rejected. A
// discarded payload describes a past instance and must not touch the
// incident state either.
if existed && (curStatus != a.Status || curStartsAt != a.StartsAt.Unix()) {
continue
}
firing := a.Status == "firing"
accepted = append(accepted, ingested{
id: id,
name: name,
firing: firing,
newOccurrence: firing && (!existed || a.StartsAt.Unix() > prevStartsAt || prevStatus == "resolved"),
justResolved: !firing && existed && prevStatus == "firing",
})
}
return accepted, nil
}
// incidentForGroup returns the open incident that this payload's firing alerts
// belong to, opening one if the group has none. It returns 0 when the payload
// warrants no incident at all.
//
// The rule that matters: a group with no open incident gets a new one only if
// something actually started firing. Without that, a manually resolved incident
// would reappear on the next repeat_interval re-send of an alert that never
// stopped, and manual resolution would be meaningless.
func incidentForGroup(ctx context.Context, tx *sql.Tx, payload amPayload, accepted []ingested) (int64, error) {
var firstName string
anyFiring, anyNew := false, false
for _, a := range accepted {
if a.firing {
if !anyFiring {
firstName = a.name
}
anyFiring = true
}
if a.newOccurrence {
anyNew = true
}
}
if !anyFiring {
// A payload of nothing but resolutions never opens an incident.
return 0, nil
}
groupKey := payload.GroupKey
if groupKey == "" {
// Alertmanager always sends groupKey; a sender that does not still gets
// one incident per alert name rather than one giant shared incident.
groupKey = "groupless:" + firstName
}
var id int64
switch err := tx.QueryRowContext(ctx,
"SELECT id FROM incidents WHERE group_key = ? AND resolved_at IS NULL", groupKey,
).Scan(&id); {
case err == nil:
return id, nil
case err != sql.ErrNoRows:
return 0, err
}
if !anyNew {
return 0, nil
}
return openIncident(ctx, tx, groupKey, payload.GroupLabels, firstName)
}
// openIncident creates an incident for a group and assigns it to whoever is on
// call today, which is the point at which the schedule stops being decorative.
func openIncident(ctx context.Context, tx *sql.Tx, groupKey string, groupLabels map[string]string, fallbackName string) (int64, error) {
onCall, err := currentOnCall(ctx, tx)
if err != nil {
return 0, err
}
labelsJSON, _ := json.Marshal(groupLabels)
if groupLabels == nil {
labelsJSON = []byte("{}")
}
res, err := tx.ExecContext(ctx, `
INSERT INTO incidents (group_key, title, group_labels, status, triggered_at, assigned_to)
VALUES (?, ?, ?, 'triggered', ?, ?)`,
groupKey, incidentTitle(groupLabels, fallbackName), string(labelsJSON),
time.Now().Unix(), onCall)
if err != nil {
return 0, err
}
id, err := res.LastInsertId()
if err != nil {
return 0, err
}
if err := logEvent(ctx, tx, id, evTriggered, nil, nil, nil); err != nil {
return 0, err
}
if onCall != nil {
// On an "assigned" event user_id is the assignee, not the actor.
if err := logEvent(ctx, tx, id, evAssigned, onCall, nil, nil); err != nil {
return 0, err
}
}
return id, nil
}
// linkAlert adds an alert to an incident, emitting a timeline entry only the
// first time. Re-sends of an already-linked alert are silent.
func linkAlert(ctx context.Context, tx *sql.Tx, incidentID, alertID int64) error {
res, err := tx.ExecContext(ctx, `
INSERT OR IGNORE INTO incident_alerts (incident_id, alert_id, added_at)
VALUES (?, ?, ?)`, incidentID, alertID, time.Now().Unix())
if err != nil {
return err
}
if n, _ := res.RowsAffected(); n == 0 {
return nil
}
return logEvent(ctx, tx, incidentID, evAlertAdded, nil, &alertID, nil)
}
+22 -115
View File
@@ -15,15 +15,21 @@ import (
)
// alertSelectFrom is the shared SELECT … FROM … clause used by all alert queries.
// It LEFT JOINs users so acknowledged_by username is always available.
// The subquery resolves the alert's most recent incident: membership is kept in
// incident_alerts rather than as a column here, because one alert row is reused
// across occurrences and belongs to a different incident each time.
const alertSelectFrom = `
SELECT a.id, a.fingerprint, a.name, a.status,
a.labels, a.annotations,
a.starts_at, a.ends_at, a.generator_url, a.received_at,
a.acknowledged_by, a.acknowledged_at, u.username,
(SELECT ia.incident_id
FROM incident_alerts ia
JOIN incidents i ON i.id = ia.incident_id
WHERE ia.alert_id = a.id
ORDER BY i.triggered_at DESC, i.id DESC
LIMIT 1),
a.resolution_source, a.archived_at
FROM alerts a
LEFT JOIN users u ON u.id = a.acknowledged_by`
FROM alerts a`
func handleListAlerts(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
@@ -45,6 +51,12 @@ func handleListAlerts(db *sql.DB) http.HandlerFunc {
} else {
where = append(where, "a.archived_at IS NULL")
}
if incidentID := q.Get("incident_id"); incidentID != "" {
if n, err := strconv.ParseInt(incidentID, 10, 64); err == nil {
where = append(where, "a.id IN (SELECT alert_id FROM incident_alerts WHERE incident_id = ?)")
args = append(args, n)
}
}
if from := q.Get("from"); from != "" {
if t, err := time.Parse("2006-01-02", from); err == nil {
@@ -114,55 +126,7 @@ func handleGetAlert(db *sql.DB) http.HandlerFunc {
}
}
func handleAcknowledge(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
return
}
user, _ := userFromContext(r.Context())
res, err := db.ExecContext(r.Context(),
"UPDATE alerts SET acknowledged_by = ?, acknowledged_at = ? WHERE id = ?",
user.ID, time.Now().Unix(), id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("alert not found"))
return
}
a, _ := fetchAlert(r.Context(), db, id)
respond(w, http.StatusOK, a)
}
}
func handleUnacknowledge(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
return
}
res, err := db.ExecContext(r.Context(),
"UPDATE alerts SET acknowledged_by = NULL, acknowledged_at = NULL WHERE id = ?", id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("alert not found"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
// fetchAlert loads a single alert by ID using the shared JOIN query.
// fetchAlert loads a single alert by ID using the shared query.
func fetchAlert(ctx context.Context, db *sql.DB, id int64) (models.Alert, error) {
return scanAlert(db.QueryRowContext(ctx, alertSelectFrom+" WHERE a.id = ?", id))
}
@@ -176,81 +140,24 @@ func scanAlert(s scanner) (models.Alert, error) {
var a models.Alert
var labelsJSON, annotationsJSON string
var startsAtUnix, receivedAtUnix int64
var endsAtUnix, ackAtUnix, archivedAtUnix *int64
var ackByID *int64
var ackByUser *string
var endsAtUnix, archivedAtUnix *int64
if err := s.Scan(
&a.ID, &a.Fingerprint, &a.Name, &a.Status,
&labelsJSON, &annotationsJSON,
&startsAtUnix, &endsAtUnix,
&a.GeneratorURL, &receivedAtUnix,
&ackByID, &ackAtUnix, &ackByUser,
&a.IncidentID,
&a.ResolutionSource, &archivedAtUnix,
); err != nil {
return a, err
}
json.Unmarshal([]byte(labelsJSON), &a.Labels) //nolint:errcheck
json.Unmarshal([]byte(labelsJSON), &a.Labels) //nolint:errcheck
json.Unmarshal([]byte(annotationsJSON), &a.Annotations) //nolint:errcheck
a.StartsAt = time.Unix(startsAtUnix, 0).UTC()
a.ReceivedAt = time.Unix(receivedAtUnix, 0).UTC()
if endsAtUnix != nil {
t := time.Unix(*endsAtUnix, 0).UTC()
a.EndsAt = &t
}
if ackByID != nil {
t := time.Unix(*ackAtUnix, 0).UTC()
a.AcknowledgedByID = ackByID
a.AcknowledgedByUser = ackByUser
a.AcknowledgedAt = &t
}
if archivedAtUnix != nil {
t := time.Unix(*archivedAtUnix, 0).UTC()
a.ArchivedAt = &t
}
a.EndsAt = unixPtr(endsAtUnix)
a.ArchivedAt = unixPtr(archivedAtUnix)
return a, nil
}
func handleArchive(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
return
}
res, err := db.ExecContext(r.Context(),
"UPDATE alerts SET archived_at = unixepoch() WHERE id = ?", id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("alert not found"))
return
}
a, _ := fetchAlert(r.Context(), db, id)
respond(w, http.StatusOK, a)
}
}
func handleUnarchive(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
return
}
res, err := db.ExecContext(r.Context(),
"UPDATE alerts SET archived_at = NULL WHERE id = ?", id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("alert not found"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
+24 -108
View File
@@ -171,9 +171,22 @@ func TestBootstrap_SecondCallForbidden(t *testing.T) {
// Alert upsert by fingerprint
// ---------------------------------------------------------------------------
func postWebhook(t *testing.T, s *ts, alerts []map[string]any) {
// postWebhook sends an Alertmanager v4 payload. groupKey is optional: omitting
// it exercises the fallback for senders that do not group, which is what most of
// these tests want.
func postWebhook(t *testing.T, s *ts, alerts []map[string]any, groupKey ...string) {
t.Helper()
payload := map[string]any{"version": "4", "status": "firing", "alerts": alerts}
if len(groupKey) > 0 {
payload["groupKey"] = groupKey[0]
// Alertmanager groups by alertname by default, so the group labels echo
// the first alert's name.
if len(alerts) > 0 {
if labels, ok := alerts[0]["labels"].(map[string]string); ok {
payload["groupLabels"] = map[string]string{"alertname": labels["alertname"]}
}
}
}
data, _ := json.Marshal(payload)
resp, err := http.Post(s.URL+"/api/alertmanager/webhook", "application/json", bytes.NewReader(data))
if err != nil {
@@ -243,84 +256,6 @@ func TestAlertUpsert_DifferentFingerprintsStored(t *testing.T) {
}
}
// ---------------------------------------------------------------------------
// Alert acknowledge
// ---------------------------------------------------------------------------
func TestAcknowledge(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{{
"status": "firing", "labels": map[string]string{"alertname": "X"},
"annotations": map[string]string{}, "startsAt": "2026-05-20T10:00:00Z",
"endsAt": "0001-01-01T00:00:00Z", "generatorURL": "", "fingerprint": "fp-ack",
}})
resp := s.req(t, http.MethodPost, "/api/alerts/1/acknowledge", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("acknowledge returned %d", resp.StatusCode)
}
var alert map[string]any
decode(t, resp, &alert)
if alert["acknowledged_by"] == nil {
t.Error("expected acknowledged_by to be set")
}
// Clear it.
resp = s.req(t, http.MethodDelete, "/api/alerts/1/acknowledge", nil)
if resp.StatusCode != http.StatusNoContent {
t.Errorf("unacknowledge returned %d", resp.StatusCode)
}
resp = s.req(t, http.MethodGet, "/api/alerts/1", nil)
var alert2 map[string]any
decode(t, resp, &alert2)
if alert2["acknowledged_by"] != nil {
t.Error("expected acknowledged_by to be cleared")
}
}
// ---------------------------------------------------------------------------
// Comments — own-only deletion
// ---------------------------------------------------------------------------
func TestComment_DeleteOwnOnly(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{{
"status": "firing", "labels": map[string]string{"alertname": "Y"},
"annotations": map[string]string{}, "startsAt": "2026-05-20T10:00:00Z",
"endsAt": "0001-01-01T00:00:00Z", "generatorURL": "", "fingerprint": "fp-comment",
}})
// Create a second user and their own key.
s.req(t, http.MethodPost, "/api/users",
map[string]string{"username": "alice", "email": "alice@test.com"})
keyResp := s.req(t, http.MethodPost, "/api/users/2/api-keys",
map[string]string{"name": "alice-key"})
var keyData map[string]any
decode(t, keyResp, &keyData)
aliceKey := keyData["key"].(string)
// Admin posts a comment.
s.req(t, http.MethodPost, "/api/alerts/1/comments",
map[string]string{"content": "admin note"})
// Alice tries to delete admin's comment (should 404).
req, _ := http.NewRequest(http.MethodDelete, s.URL+"/api/alerts/1/comments/1", nil)
req.Header.Set("Authorization", "Bearer "+aliceKey)
resp, _ := http.DefaultClient.Do(req)
resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("expected 404 when deleting another user's comment, got %d", resp.StatusCode)
}
// Admin deletes own comment (should 204).
resp = s.req(t, http.MethodDelete, "/api/alerts/1/comments/1", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Errorf("expected 204 when deleting own comment, got %d", resp.StatusCode)
}
}
// ---------------------------------------------------------------------------
// Schedule conflict
// ---------------------------------------------------------------------------
@@ -409,35 +344,29 @@ func TestStats_ByHourReturnsTwentyFourSlots(t *testing.T) {
// Archive
// ---------------------------------------------------------------------------
func TestArchive_RoundTrip(t *testing.T) {
// Alert archiving is sweeper-only housekeeping now — nobody archives an alert by
// hand — but the list filter it drives is still part of the API.
func TestArchive_AlertListFilter(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{{
"status": "resolved", "fingerprint": "arch1",
"labels": map[string]string{"alertname": "Archivable"},
"annotations": map[string]string{},
"startsAt": "2026-05-20T10:00:00Z", "endsAt": "2026-05-20T11:00:00Z",
"labels": map[string]string{"alertname": "Archivable"},
"annotations": map[string]string{},
"startsAt": "2026-05-20T10:00:00Z",
"endsAt": "2026-05-20T11:00:00Z",
"generatorURL": "",
}})
// 1. Alert appears in default list (not archived).
// 1. Alert appears in the default list.
var alerts []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
if len(alerts) != 1 {
t.Fatalf("expected 1 alert in default list, got %d", len(alerts))
}
id := int(alerts[0]["id"].(float64))
// 2. Archive it.
resp := s.req(t, http.MethodPost, fmt.Sprintf("/api/alerts/%d/archive", id), nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("archive: expected 200, got %d", resp.StatusCode)
}
var archived map[string]any
decode(t, resp, &archived)
if archived["archived_at"] == nil {
t.Error("expected archived_at to be set in response")
}
// 2. Let the sweeper archive it: ends_at is already well past archiveAfter.
api.Sweep(context.Background(), s.db, time.Hour, 6*time.Hour)
// 3. Default list excludes it.
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
@@ -450,19 +379,6 @@ func TestArchive_RoundTrip(t *testing.T) {
if len(alerts) != 1 {
t.Fatalf("expected 1 archived alert, got %d", len(alerts))
}
// 5. Un-archive.
resp = s.req(t, http.MethodDelete, fmt.Sprintf("/api/alerts/%d/archive", id), nil)
if resp.StatusCode != http.StatusNoContent {
t.Fatalf("unarchive: expected 204, got %d", resp.StatusCode)
}
resp.Body.Close()
// 6. Back in default list.
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
if len(alerts) != 1 {
t.Errorf("expected unarchived alert to reappear, got %d results", len(alerts))
}
}
// ---------------------------------------------------------------------------
+143 -10
View File
@@ -4,6 +4,7 @@ import (
"context"
"database/sql"
"log"
"strings"
"time"
)
@@ -33,12 +34,16 @@ func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter tim
}
}
// Sweep runs a single pass: expire stale firing alerts, then archive resolved
// ones. Expiry runs first so an alert can expire and be archived in one pass.
// Sweep runs a single pass, in dependency order: expire stale firing alerts,
// close the incidents that leaves with nothing firing, then archive whatever has
// been settled long enough. Running them in one pass means an alert can go stale
// and its incident can close and archive without waiting three ticks.
// Exported so tests can drive a pass without waiting on the ticker.
func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration) {
expireStale(ctx, db, staleAfter)
resolveSettledIncidents(ctx, db)
archiveResolved(ctx, db, archiveAfter)
archiveResolvedIncidents(ctx, db, archiveAfter)
}
// expireStale resolves firing alerts that Alertmanager has stopped refreshing.
@@ -53,26 +58,132 @@ func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Durati
// - received_at is older than staleAfter. Alertmanager re-sends firing
// notifications every repeat_interval, making received_at a liveness
// heartbeat — provided staleAfter exceeds that interval.
//
// The matching rows are collected before the update rather than updated in bulk,
// because each one owes its incident a timeline entry.
func expireStale(ctx context.Context, db *sql.DB, staleAfter time.Duration) {
now := time.Now()
res, err := db.ExecContext(ctx, `
ids, err := staleAlertIDs(ctx, db, now, staleAfter)
if err != nil {
log.Printf("sweeper: find stale: %v", err)
return
}
if len(ids) == 0 {
return
}
args := make([]any, 0, len(ids)+1)
args = append(args, resolutionExpiry)
for _, id := range ids {
args = append(args, id)
}
if _, err := db.ExecContext(ctx, `
UPDATE alerts
SET status = 'resolved',
resolution_source = ?,
ends_at = COALESCE(ends_at, unixepoch())
WHERE status = 'firing'
AND archived_at IS NULL
AND ((ends_at IS NOT NULL AND ends_at < ?) OR received_at < ?)`,
resolutionExpiry, now.Add(-expiryGrace).Unix(), now.Add(-staleAfter).Unix())
if err != nil {
WHERE id IN (`+placeholders(len(ids))+`)`, args...); err != nil {
log.Printf("sweeper: expire stale: %v", err)
return
}
if n, _ := res.RowsAffected(); n > 0 {
log.Printf("sweeper: expired %d stale firing alert(s)", n)
log.Printf("sweeper: expired %d stale firing alert(s)", len(ids))
for _, id := range ids {
incidentID, err := openIncidentForAlert(ctx, db, id)
if err != nil {
log.Printf("sweeper: incident for alert %d: %v", id, err)
continue
}
if incidentID == 0 {
continue
}
alertID := id
if err := logEvent(ctx, db, incidentID, evAlertResolved, nil, &alertID, nil); err != nil {
log.Printf("sweeper: log expiry event: %v", err)
}
}
}
// staleAlertIDs reads the ids in one go and closes the cursor before the caller
// writes: the pool is limited to a single connection, so an open read would
// block the update behind it.
func staleAlertIDs(ctx context.Context, db *sql.DB, now time.Time, staleAfter time.Duration) ([]int64, error) {
rows, err := db.QueryContext(ctx, `
SELECT id FROM alerts
WHERE status = 'firing'
AND archived_at IS NULL
AND ((ends_at IS NOT NULL AND ends_at < ?) OR received_at < ?)`,
now.Add(-expiryGrace).Unix(), now.Add(-staleAfter).Unix())
if err != nil {
return nil, err
}
defer rows.Close()
var ids []int64
for rows.Next() {
var id int64
if err := rows.Scan(&id); err != nil {
return nil, err
}
ids = append(ids, id)
}
return ids, rows.Err()
}
// resolveSettledIncidents closes incidents whose alerts have all stopped firing.
// This is the cascade from alerts up to the work item, and it is what turns an
// expiry into a closed incident rather than one that sits open forever.
func resolveSettledIncidents(ctx context.Context, db *sql.DB) {
ids, err := settledIncidentIDs(ctx, db)
if err != nil {
log.Printf("sweeper: find settled incidents: %v", err)
return
}
resolved := 0
for _, id := range ids {
ok, err := resolveIfSettled(ctx, db, id)
if err != nil {
log.Printf("sweeper: resolve incident %d: %v", id, err)
continue
}
if ok {
resolved++
}
}
if resolved > 0 {
log.Printf("sweeper: resolved %d settled incident(s)", resolved)
}
}
func settledIncidentIDs(ctx context.Context, db *sql.DB) ([]int64, error) {
rows, err := db.QueryContext(ctx, `
SELECT i.id
FROM incidents i
WHERE i.resolved_at IS NULL
AND EXISTS (SELECT 1 FROM incident_alerts ia WHERE ia.incident_id = i.id)
AND NOT EXISTS (SELECT 1
FROM incident_alerts ia
JOIN alerts a ON a.id = ia.alert_id
WHERE ia.incident_id = i.id
AND a.status = 'firing')`)
if err != nil {
return nil, err
}
defer rows.Close()
var ids []int64
for rows.Next() {
var id int64
if err := rows.Scan(&id); err != nil {
return nil, err
}
ids = append(ids, id)
}
return ids, rows.Err()
}
// archiveResolved hides resolved alerts that have been settled for archiveAfter.
func archiveResolved(ctx context.Context, db *sql.DB, archiveAfter time.Duration) {
cutoff := time.Now().Add(-archiveAfter).Unix()
@@ -89,3 +200,25 @@ func archiveResolved(ctx context.Context, db *sql.DB, archiveAfter time.Duration
log.Printf("archiver: archived %d resolved alert(s)", n)
}
}
// archiveResolvedIncidents does the same for the work items, on the same clock.
func archiveResolvedIncidents(ctx context.Context, db *sql.DB, archiveAfter time.Duration) {
cutoff := time.Now().Add(-archiveAfter).Unix()
res, err := db.ExecContext(ctx,
`UPDATE incidents SET archived_at = unixepoch()
WHERE resolved_at IS NOT NULL
AND archived_at IS NULL
AND resolved_at < ?`, cutoff)
if err != nil {
log.Printf("archiver: incidents: %v", err)
return
}
if n, _ := res.RowsAffected(); n > 0 {
log.Printf("archiver: archived %d resolved incident(s)", n)
}
}
// placeholders builds "?, ?, …" for an IN clause of n values.
func placeholders(n int) string {
return strings.TrimSuffix(strings.Repeat("?, ", n), ", ")
}
-131
View File
@@ -1,131 +0,0 @@
package api
import (
"database/sql"
"net/http"
"strconv"
"time"
"github.com/go-chi/chi/v5"
"github.com/yeniklas/terdut-server/internal/models"
)
func handleListComments(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
alertID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
return
}
// Verify the alert exists.
var exists int
if err := db.QueryRowContext(r.Context(), "SELECT 1 FROM alerts WHERE id = ?", alertID).Scan(&exists); err != nil {
respond(w, http.StatusNotFound, errResp("alert not found"))
return
}
rows, err := db.QueryContext(r.Context(), `
SELECT c.id, c.alert_id, c.user_id, u.username, c.content, c.created_at
FROM alert_comments c
JOIN users u ON u.id = c.user_id
WHERE c.alert_id = ?
ORDER BY c.created_at ASC`, alertID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer rows.Close()
comments := []models.Comment{}
for rows.Next() {
var c models.Comment
var ts int64
if err := rows.Scan(&c.ID, &c.AlertID, &c.UserID, &c.Username, &c.Content, &ts); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
c.CreatedAt = time.Unix(ts, 0).UTC()
comments = append(comments, c)
}
respond(w, http.StatusOK, comments)
}
}
func handleCreateComment(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
alertID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
return
}
var req struct {
Content string `json:"content"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
if req.Content == "" {
respond(w, http.StatusBadRequest, errResp("content is required"))
return
}
// Verify the alert exists.
var exists int
if err := db.QueryRowContext(r.Context(), "SELECT 1 FROM alerts WHERE id = ?", alertID).Scan(&exists); err != nil {
respond(w, http.StatusNotFound, errResp("alert not found"))
return
}
user, _ := userFromContext(r.Context())
res, err := db.ExecContext(r.Context(),
"INSERT INTO alert_comments (alert_id, user_id, content) VALUES (?, ?, ?)",
alertID, user.ID, req.Content)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
commentID, _ := res.LastInsertId()
comment := models.Comment{
ID: commentID,
AlertID: alertID,
UserID: user.ID,
Username: user.Username,
Content: req.Content,
CreatedAt: time.Now().UTC(),
}
respond(w, http.StatusCreated, comment)
}
}
func handleDeleteComment(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
alertID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
return
}
commentID, err := strconv.ParseInt(chi.URLParam(r, "commentID"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid comment id"))
return
}
user, _ := userFromContext(r.Context())
res, err := db.ExecContext(r.Context(),
"DELETE FROM alert_comments WHERE id = ? AND alert_id = ? AND user_id = ?",
commentID, alertID, user.ID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("comment not found"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
+269
View File
@@ -0,0 +1,269 @@
package api
import (
"context"
"database/sql"
"encoding/json"
"sort"
"strings"
"time"
"github.com/yeniklas/terdut-server/internal/models"
)
// Values for incidents.resolution_source, recording who closed the incident:
// every member alert stopped firing, or a person decided it was done.
const (
incidentResolutionAlerts = "alerts"
incidentResolutionManual = "manual"
)
// Incident timeline event types. Stored as free text so adding one later is not
// a migration, but these are the ones the server writes.
const (
evTriggered = "triggered"
evAlertAdded = "alert_added"
evAlertResolved = "alert_resolved"
evAcknowledged = "acknowledged"
evUnacknowledged = "unacknowledged"
evAssigned = "assigned"
evSnoozed = "snoozed"
evUnsnoozed = "unsnoozed"
evResolved = "resolved"
evNote = "note"
)
// severityLabel is the Alertmanager label an incident's severity is derived from.
const severityLabel = "severity"
// querier is satisfied by both *sql.DB and *sql.Tx, so the helpers below work
// inside the webhook's transaction and standalone from handlers and the sweeper.
type querier interface {
ExecContext(ctx context.Context, query string, args ...any) (sql.Result, error)
QueryContext(ctx context.Context, query string, args ...any) (*sql.Rows, error)
QueryRowContext(ctx context.Context, query string, args ...any) *sql.Row
}
const incidentSelectFrom = `
SELECT i.id, i.group_key, i.title, i.group_labels, i.status, i.severity,
i.triggered_at,
i.acknowledged_by, i.acknowledged_at, ack.username,
i.assigned_to, asg.username, i.snoozed_until,
i.resolved_at, i.resolution_source, i.archived_at
FROM incidents i
LEFT JOIN users ack ON ack.id = i.acknowledged_by
LEFT JOIN users asg ON asg.id = i.assigned_to`
func scanIncident(s scanner) (models.Incident, error) {
var i models.Incident
var groupLabelsJSON string
var triggeredAt int64
var ackAt, snoozedUntil, resolvedAt, archivedAt *int64
if err := s.Scan(
&i.ID, &i.GroupKey, &i.Title, &groupLabelsJSON, &i.Status, &i.Severity,
&triggeredAt,
&i.AcknowledgedByID, &ackAt, &i.AcknowledgedByUser,
&i.AssignedToID, &i.AssignedToUser, &snoozedUntil,
&resolvedAt, &i.ResolutionSource, &archivedAt,
); err != nil {
return i, err
}
json.Unmarshal([]byte(groupLabelsJSON), &i.GroupLabels) //nolint:errcheck
i.TriggeredAt = time.Unix(triggeredAt, 0).UTC()
i.AcknowledgedAt = unixPtr(ackAt)
i.SnoozedUntil = unixPtr(snoozedUntil)
i.ResolvedAt = unixPtr(resolvedAt)
i.ArchivedAt = unixPtr(archivedAt)
return i, nil
}
// unixPtr converts a nullable Unix-second column to a nullable UTC time.
func unixPtr(sec *int64) *time.Time {
if sec == nil {
return nil
}
t := time.Unix(*sec, 0).UTC()
return &t
}
func fetchIncident(ctx context.Context, q querier, id int64) (models.Incident, error) {
return scanIncident(q.QueryRowContext(ctx, incidentSelectFrom+" WHERE i.id = ?", id))
}
// logEvent appends one entry to an incident's timeline. A nil userID means the
// server acted rather than a person.
func logEvent(ctx context.Context, q querier, incidentID int64, evType string, userID, alertID *int64, detail *string) error {
_, err := q.ExecContext(ctx, `
INSERT INTO incident_events (incident_id, type, user_id, alert_id, detail, created_at)
VALUES (?, ?, ?, ?, ?, ?)`,
incidentID, evType, userID, alertID, detail, time.Now().Unix())
return err
}
// todayUTC is the schedule's day key. The schedule's smallest unit is one UTC day.
func todayUTC() string {
return time.Now().UTC().Format("2006-01-02")
}
// currentOnCall returns today's on-call user, or nil when nobody is scheduled.
// A missing schedule entry is not an error — incidents just open unassigned.
func currentOnCall(ctx context.Context, q querier) (*int64, error) {
var userID int64
err := q.QueryRowContext(ctx,
"SELECT user_id FROM schedule_entries WHERE date = ?", todayUTC()).Scan(&userID)
if err == sql.ErrNoRows {
return nil, nil
}
if err != nil {
return nil, err
}
return &userID, nil
}
// severityRank orders the conventional Alertmanager severity label values.
// Anything unrecognised sorts below all of them rather than being dropped.
func severityRank(s string) int {
switch strings.ToLower(s) {
case "critical":
return 4
case "error":
return 3
case "warning":
return 2
case "info":
return 1
default:
return 0
}
}
// refreshSeverity raises an incident's severity to the highest `severity` label
// seen across its alerts.
//
// It is a high-water mark, never lowered: an incident that hit critical was a
// critical incident, even after the critical alert clears and a warning is all
// that is left firing. Downgrading a live incident would also quietly demote it
// in the queue while the work is still open.
func refreshSeverity(ctx context.Context, q querier, incidentID int64) error {
rows, err := q.QueryContext(ctx, `
SELECT json_extract(a.labels, '$.'||?)
FROM incident_alerts ia
JOIN alerts a ON a.id = ia.alert_id
WHERE ia.incident_id = ?`, severityLabel, incidentID)
if err != nil {
return err
}
best := ""
for rows.Next() {
var sev *string
if err := rows.Scan(&sev); err != nil {
rows.Close()
return err
}
if sev != nil && severityRank(*sev) > severityRank(best) {
best = *sev
}
}
if err := rows.Err(); err != nil {
rows.Close()
return err
}
rows.Close()
if best == "" {
return nil
}
// The comparison lives in SQL so an unrelated concurrent update cannot be
// clobbered by a stale read.
_, err = q.ExecContext(ctx, `
UPDATE incidents SET severity = ?
WHERE id = ?
AND (severity IS NULL OR `+severityRankSQL("severity")+` < ?)`,
best, incidentID, severityRank(best))
return err
}
// severityRankSQL mirrors severityRank for use inside a statement. SQL cannot
// order these strings meaningfully on its own.
func severityRankSQL(col string) string {
return `CASE lower(COALESCE(` + col + `, ''))
WHEN 'critical' THEN 4
WHEN 'error' THEN 3
WHEN 'warning' THEN 2
WHEN 'info' THEN 1
ELSE 0 END`
}
// resolveIfSettled closes an incident once every alert under it has stopped
// firing — PagerDuty's cascade, and the only automatic route out of the open
// state. Reports whether it actually resolved anything.
func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, error) {
res, err := q.ExecContext(ctx, `
UPDATE incidents
SET status = 'resolved',
resolved_at = ?,
resolution_source = ?
WHERE id = ?
AND resolved_at IS NULL
-- An incident with no members yet is mid-creation, not settled.
AND EXISTS (SELECT 1 FROM incident_alerts ia WHERE ia.incident_id = incidents.id)
AND NOT EXISTS (SELECT 1
FROM incident_alerts ia
JOIN alerts a ON a.id = ia.alert_id
WHERE ia.incident_id = incidents.id
AND a.status = 'firing')`,
time.Now().Unix(), incidentResolutionAlerts, incidentID)
if err != nil {
return false, err
}
n, _ := res.RowsAffected()
if n == 0 {
return false, nil
}
return true, logEvent(ctx, q, incidentID, evResolved, nil, nil, nil)
}
// openIncidentForAlert returns the open incident an alert currently belongs to,
// or 0 when it has none. Used when an alert resolves or expires so the event
// lands on the right timeline.
func openIncidentForAlert(ctx context.Context, q querier, alertID int64) (int64, error) {
var id int64
err := q.QueryRowContext(ctx, `
SELECT i.id
FROM incident_alerts ia
JOIN incidents i ON i.id = ia.incident_id
WHERE ia.alert_id = ? AND i.resolved_at IS NULL`, alertID).Scan(&id)
if err == sql.ErrNoRows {
return 0, nil
}
return id, err
}
// incidentTitle renders a human-readable title from Alertmanager's groupLabels,
// leading with the alert name and appending whatever else the operator grouped
// by. Falls back to the alert's own name when the payload carried no groupLabels.
func incidentTitle(groupLabels map[string]string, fallback string) string {
name := groupLabels["alertname"]
if name == "" {
name = fallback
}
if name == "" {
name = "Incident"
}
rest := make([]string, 0, len(groupLabels))
for k, v := range groupLabels {
if k == "alertname" {
continue
}
rest = append(rest, k+"="+v)
}
if len(rest) == 0 {
return name
}
sort.Strings(rest)
return name + " (" + strings.Join(rest, ", ") + ")"
}
+554
View File
@@ -0,0 +1,554 @@
package api
import (
"database/sql"
"fmt"
"net/http"
"strconv"
"strings"
"time"
"github.com/go-chi/chi/v5"
"github.com/yeniklas/terdut-server/internal/models"
)
func handleListIncidents(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
q := r.URL.Query()
where := []string{}
args := []any{}
// Without an explicit status the queue shows open work, which is what an
// on-call person opens the tool to see.
if status := q.Get("status"); status != "" {
where = append(where, "i.status = ?")
args = append(args, status)
} else {
where = append(where, "i.resolved_at IS NULL")
}
if q.Get("archived") == "true" {
where = append(where, "i.archived_at IS NOT NULL")
} else {
where = append(where, "i.archived_at IS NULL")
}
// A snooze expires by simply falling into the past; nothing sweeps it.
if q.Get("snoozed") == "true" {
where = append(where, "i.snoozed_until > ?")
args = append(args, time.Now().Unix())
} else {
where = append(where, "(i.snoozed_until IS NULL OR i.snoozed_until <= ?)")
args = append(args, time.Now().Unix())
}
if severity := q.Get("severity"); severity != "" {
where = append(where, "i.severity = ?")
args = append(args, severity)
}
if assignee := q.Get("assigned_to"); assignee != "" {
if n, err := strconv.ParseInt(assignee, 10, 64); err == nil {
where = append(where, "i.assigned_to = ?")
args = append(args, n)
}
}
if from := q.Get("from"); from != "" {
if t, err := time.Parse("2006-01-02", from); err == nil {
where = append(where, "i.triggered_at >= ?")
args = append(args, t.UTC().Unix())
}
}
if to := q.Get("to"); to != "" {
if t, err := time.Parse("2006-01-02", to); err == nil {
where = append(where, "i.triggered_at < ?")
args = append(args, t.UTC().AddDate(0, 0, 1).Unix())
}
}
limit := 50
if l := q.Get("limit"); l != "" {
if n, err := strconv.Atoi(l); err == nil && n > 0 && n <= 500 {
limit = n
}
}
order := "i.triggered_at DESC"
if q.Get("sort") == "severity" {
order = severityRankSQL("i.severity") + " DESC, i.triggered_at DESC"
}
args = append(args, limit)
rows, err := db.QueryContext(r.Context(),
fmt.Sprintf("%s WHERE %s ORDER BY %s LIMIT ?",
incidentSelectFrom, strings.Join(where, " AND "), order),
args...)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer rows.Close()
incidents := []models.Incident{}
for rows.Next() {
i, err := scanIncident(rows)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
incidents = append(incidents, i)
}
respond(w, http.StatusOK, incidents)
}
}
func handleGetIncident(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
inc, err := fetchIncident(r.Context(), db, id)
if err == sql.ErrNoRows {
respond(w, http.StatusNotFound, errResp("incident not found"))
return
}
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if inc.Alerts, err = incidentAlerts(r, db, id); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, inc)
}
}
func handleIncidentAlerts(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
if !incidentExists(w, r, db, id) {
return
}
alerts, err := incidentAlerts(r, db, id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, alerts)
}
}
func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
if !incidentExists(w, r, db, id) {
return
}
rows, err := db.QueryContext(r.Context(), `
SELECT e.id, e.incident_id, e.type, e.user_id, u.username,
e.alert_id, e.detail, e.created_at
FROM incident_events e
LEFT JOIN users u ON u.id = e.user_id
WHERE e.incident_id = ?
ORDER BY e.created_at ASC, e.id ASC`, id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer rows.Close()
events := []models.IncidentEvent{}
for rows.Next() {
var e models.IncidentEvent
var ts int64
if err := rows.Scan(&e.ID, &e.IncidentID, &e.Type, &e.UserID, &e.Username,
&e.AlertID, &e.Detail, &ts); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
e.CreatedAt = time.Unix(ts, 0).UTC()
events = append(events, e)
}
respond(w, http.StatusOK, events)
}
}
func handleIncidentAcknowledge(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
user, _ := userFromContext(r.Context())
if !updateOpenIncident(w, r, db, id,
`UPDATE incidents SET status = 'acknowledged', acknowledged_by = ?, acknowledged_at = ?
WHERE id = ? AND resolved_at IS NULL`, user.ID, time.Now().Unix(), id) {
return
}
if err := logEvent(r.Context(), db, id, evAcknowledged, &user.ID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respondIncident(w, r, db, id)
}
}
func handleIncidentUnacknowledge(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
user, _ := userFromContext(r.Context())
if !updateOpenIncident(w, r, db, id,
`UPDATE incidents SET status = 'triggered', acknowledged_by = NULL, acknowledged_at = NULL
WHERE id = ? AND resolved_at IS NULL`, id) {
return
}
if err := logEvent(r.Context(), db, id, evUnacknowledged, &user.ID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
// handleIncidentResolve closes an incident by hand. This is terminal: a later
// occurrence opens a new incident rather than reopening this one, which is what
// stops a resolved incident from reappearing on the next repeat_interval
// re-send of an alert that never stopped firing. Use snooze for "not now".
func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
user, _ := userFromContext(r.Context())
if !updateOpenIncident(w, r, db, id,
`UPDATE incidents SET status = 'resolved', resolved_at = ?, resolution_source = ?
WHERE id = ? AND resolved_at IS NULL`,
time.Now().Unix(), incidentResolutionManual, id) {
return
}
if err := logEvent(r.Context(), db, id, evResolved, &user.ID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respondIncident(w, r, db, id)
}
}
func handleIncidentAssign(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
var req struct {
UserID int64 `json:"user_id"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
if req.UserID == 0 {
respond(w, http.StatusBadRequest, errResp("user_id is required"))
return
}
var exists int
if err := db.QueryRowContext(r.Context(),
"SELECT 1 FROM users WHERE id = ?", req.UserID).Scan(&exists); err != nil {
respond(w, http.StatusNotFound, errResp("user not found"))
return
}
if !updateOpenIncident(w, r, db, id,
"UPDATE incidents SET assigned_to = ? WHERE id = ? AND resolved_at IS NULL",
req.UserID, id) {
return
}
// On an "assigned" event user_id is the assignee, not the actor.
if err := logEvent(r.Context(), db, id, evAssigned, &req.UserID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respondIncident(w, r, db, id)
}
}
// handleIncidentSnooze hides an incident from the default queue without closing
// it. Accepts either an absolute {"until": RFC3339} or a relative
// {"duration": "2h"}.
func handleIncidentSnooze(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
var req struct {
Until string `json:"until"`
Duration string `json:"duration"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
var until time.Time
switch {
case req.Until != "":
t, err := time.Parse(time.RFC3339, req.Until)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid until (expected RFC3339)"))
return
}
until = t
case req.Duration != "":
d, err := time.ParseDuration(req.Duration)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid duration"))
return
}
until = time.Now().Add(d)
default:
respond(w, http.StatusBadRequest, errResp("until or duration is required"))
return
}
if !until.After(time.Now()) {
respond(w, http.StatusBadRequest, errResp("snooze must end in the future"))
return
}
user, _ := userFromContext(r.Context())
if !updateOpenIncident(w, r, db, id,
"UPDATE incidents SET snoozed_until = ? WHERE id = ? AND resolved_at IS NULL",
until.Unix(), id) {
return
}
detail := until.UTC().Format(time.RFC3339)
if err := logEvent(r.Context(), db, id, evSnoozed, &user.ID, nil, &detail); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respondIncident(w, r, db, id)
}
}
func handleIncidentUnsnooze(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
user, _ := userFromContext(r.Context())
if !updateOpenIncident(w, r, db, id,
"UPDATE incidents SET snoozed_until = NULL WHERE id = ? AND resolved_at IS NULL", id) {
return
}
if err := logEvent(r.Context(), db, id, evUnsnoozed, &user.ID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
func handleIncidentArchive(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
res, err := db.ExecContext(r.Context(),
"UPDATE incidents SET archived_at = unixepoch() WHERE id = ?", id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("incident not found"))
return
}
respondIncident(w, r, db, id)
}
}
func handleIncidentUnarchive(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
res, err := db.ExecContext(r.Context(),
"UPDATE incidents SET archived_at = NULL WHERE id = ?", id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("incident not found"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
// handleCreateNote adds a note to the timeline. Notes are ordinary events, so a
// single query renders the whole story of an incident in order.
func handleCreateNote(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
var req struct {
Content string `json:"content"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
return
}
if req.Content == "" {
respond(w, http.StatusBadRequest, errResp("content is required"))
return
}
if !incidentExists(w, r, db, id) {
return
}
user, _ := userFromContext(r.Context())
now := time.Now()
res, err := db.ExecContext(r.Context(), `
INSERT INTO incident_events (incident_id, type, user_id, detail, created_at)
VALUES (?, ?, ?, ?, ?)`, id, evNote, user.ID, req.Content, now.Unix())
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
eventID, _ := res.LastInsertId()
respond(w, http.StatusCreated, models.IncidentEvent{
ID: eventID,
IncidentID: id,
Type: evNote,
UserID: &user.ID,
Username: &user.Username,
Detail: &req.Content,
CreatedAt: now.UTC().Truncate(time.Second),
})
}
}
// handleDeleteNote removes one of your own notes. Only notes are deletable — the
// rest of the timeline is what actually happened, and is not editable.
func handleDeleteNote(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
id, ok := incidentIDParam(w, r)
if !ok {
return
}
eventID, err := strconv.ParseInt(chi.URLParam(r, "eventID"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid note id"))
return
}
user, _ := userFromContext(r.Context())
res, err := db.ExecContext(r.Context(), `
DELETE FROM incident_events
WHERE id = ? AND incident_id = ? AND type = ? AND user_id = ?`,
eventID, id, evNote, user.ID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if n, _ := res.RowsAffected(); n == 0 {
respond(w, http.StatusNotFound, errResp("note not found"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
// ---------------------------------------------------------------------------
// Shared handler plumbing
// ---------------------------------------------------------------------------
func incidentIDParam(w http.ResponseWriter, r *http.Request) (int64, bool) {
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid incident id"))
return 0, false
}
return id, true
}
func incidentExists(w http.ResponseWriter, r *http.Request, db *sql.DB, id int64) bool {
var exists int
if err := db.QueryRowContext(r.Context(),
"SELECT 1 FROM incidents WHERE id = ?", id).Scan(&exists); err != nil {
respond(w, http.StatusNotFound, errResp("incident not found"))
return false
}
return true
}
// updateOpenIncident runs a mutation that is only valid while an incident is
// open. The query must be constrained to `resolved_at IS NULL`, so no rows means
// either the incident does not exist or it is already closed — two different
// answers the caller should not have to distinguish itself.
func updateOpenIncident(w http.ResponseWriter, r *http.Request, db *sql.DB, id int64, query string, args ...any) bool {
res, err := db.ExecContext(r.Context(), query, args...)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return false
}
if n, _ := res.RowsAffected(); n > 0 {
return true
}
if !incidentExists(w, r, db, id) {
return false
}
respond(w, http.StatusConflict, errResp("incident is resolved"))
return false
}
func respondIncident(w http.ResponseWriter, r *http.Request, db *sql.DB, id int64) {
inc, err := fetchIncident(r.Context(), db, id)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, inc)
}
// incidentAlerts loads the alerts under an incident, newest signal first.
func incidentAlerts(r *http.Request, db *sql.DB, id int64) ([]models.Alert, error) {
rows, err := db.QueryContext(r.Context(), alertSelectFrom+`
JOIN incident_alerts m ON m.alert_id = a.id
WHERE m.incident_id = ?
ORDER BY a.received_at DESC`, id)
if err != nil {
return nil, err
}
defer rows.Close()
alerts := []models.Alert{}
for rows.Next() {
a, err := scanAlert(rows)
if err != nil {
return nil, err
}
alerts = append(alerts, a)
}
return alerts, rows.Err()
}
+772
View File
@@ -0,0 +1,772 @@
package api_test
import (
"context"
"fmt"
"net/http"
"os"
"path/filepath"
"sort"
"testing"
"time"
"github.com/yeniklas/terdut-server/internal/api"
"github.com/yeniklas/terdut-server/internal/db"
)
// amAlert builds one alert of a webhook payload.
func amAlert(fingerprint, name, status, startsAt, endsAt string, labels map[string]string) map[string]any {
l := map[string]string{"alertname": name}
for k, v := range labels {
l[k] = v
}
return map[string]any{
"status": status,
"labels": l,
"annotations": map[string]string{},
"startsAt": startsAt,
"endsAt": endsAt,
"generatorURL": "",
"fingerprint": fingerprint,
}
}
func listIncidents(t *testing.T, s *ts, query string) []map[string]any {
t.Helper()
var out []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/incidents"+query, nil), &out)
return out
}
func getIncident(t *testing.T, s *ts, id int) map[string]any {
t.Helper()
var out map[string]any
decode(t, s.req(t, http.MethodGet, fmt.Sprintf("/api/incidents/%d", id), nil), &out)
return out
}
func timeline(t *testing.T, s *ts, id int) []map[string]any {
t.Helper()
var out []map[string]any
decode(t, s.req(t, http.MethodGet, fmt.Sprintf("/api/incidents/%d/timeline", id), nil), &out)
return out
}
// eventTypes flattens a timeline to its event types, which is what the ordering
// assertions actually care about.
func eventTypes(events []map[string]any) []string {
types := make([]string, len(events))
for i, e := range events {
types[i] = e["type"].(string)
}
return types
}
// countIncidents counts rows directly, including resolved and archived ones that
// no list view returns.
func (s *ts) countIncidents(t *testing.T) int {
t.Helper()
var n int
if err := s.db.QueryRow("SELECT COUNT(*) FROM incidents").Scan(&n); err != nil {
t.Fatalf("count incidents: %v", err)
}
return n
}
// ---------------------------------------------------------------------------
// Ingest: alerts becoming incidents
// ---------------------------------------------------------------------------
func TestWebhook_FiringOpensIncident(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-1", "HighCPU", "firing", "2026-05-20T10:00:00Z", zeroTime,
map[string]string{"severity": "critical"}),
}, "{}:{alertname=\"HighCPU\"}")
incidents := listIncidents(t, s, "")
if len(incidents) != 1 {
t.Fatalf("expected 1 incident, got %d", len(incidents))
}
inc := incidents[0]
if inc["status"] != "triggered" {
t.Errorf("expected status triggered, got %v", inc["status"])
}
if inc["severity"] != "critical" {
t.Errorf("expected severity critical, got %v", inc["severity"])
}
if inc["title"] != "HighCPU" {
t.Errorf("expected title from groupLabels, got %v", inc["title"])
}
// The alert points back at the incident it opened.
var alerts []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
if len(alerts) != 1 || alerts[0]["incident_id"] == nil {
t.Fatalf("expected the alert to carry an incident_id, got %v", alerts)
}
}
// Alertmanager already grouped these; we adopt its answer rather than
// correlating again.
func TestWebhook_SameGroupKeyJoinsOneIncident(t *testing.T) {
s := newTS(t)
const groupKey = "{}:{alertname=\"DiskFull\"}"
postWebhook(t, s, []map[string]any{
amAlert("fp-a", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
amAlert("fp-b", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
}, groupKey)
incidents := listIncidents(t, s, "")
if len(incidents) != 1 {
t.Fatalf("expected 1 incident for one groupKey, got %d", len(incidents))
}
id := int(incidents[0]["id"].(float64))
inc := getIncident(t, s, id)
members, _ := inc["alerts"].([]any)
if len(members) != 2 {
t.Fatalf("expected 2 alerts under the incident, got %d", len(members))
}
added := 0
for _, ty := range eventTypes(timeline(t, s, id)) {
if ty == "alert_added" {
added++
}
}
if added != 2 {
t.Errorf("expected 2 alert_added events, got %d", added)
}
}
// The load-bearing rule. Alertmanager re-sends firing notifications every
// repeat_interval; if those re-sends reopened incidents, resolving one by hand
// would mean nothing.
func TestWebhook_HeartbeatDoesNotReopenResolvedIncident(t *testing.T) {
s := newTS(t)
const groupKey = "{}:{alertname=\"Flapper\"}"
alert := amAlert("fp-hb", "Flapper", "firing", "2026-05-20T10:00:00Z", zeroTime, nil)
postWebhook(t, s, []map[string]any{alert}, groupKey)
resp := s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("resolve returned %d", resp.StatusCode)
}
resp.Body.Close()
// Same startsAt, same fingerprint: a re-send, not a new occurrence.
postWebhook(t, s, []map[string]any{alert}, groupKey)
if n := s.countIncidents(t); n != 1 {
t.Fatalf("expected the heartbeat to open no incident, got %d total", n)
}
if inc := getIncident(t, s, 1); inc["resolved_at"] == nil {
t.Error("expected incident 1 to stay resolved")
}
// The alert itself is still firing and still being tracked — only the work
// item is closed.
status, _, _ := s.alertRow(t, "fp-hb")
if status != "firing" {
t.Errorf("expected the alert to still be firing, got %q", status)
}
}
func TestWebhook_NewOccurrenceOpensNewIncident(t *testing.T) {
s := newTS(t)
const groupKey = "{}:{alertname=\"Recurring\"}"
postWebhook(t, s, []map[string]any{
amAlert("fp-new", "Recurring", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
}, groupKey)
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
// A newer startsAt is a genuinely new occurrence, not a re-send.
postWebhook(t, s, []map[string]any{
amAlert("fp-new", "Recurring", "firing", "2026-05-21T09:00:00Z", zeroTime, nil),
}, groupKey)
if n := s.countIncidents(t); n != 2 {
t.Fatalf("expected a second incident for the new occurrence, got %d total", n)
}
open := listIncidents(t, s, "")
if len(open) != 1 || int(open[0]["id"].(float64)) != 2 {
t.Fatalf("expected incident 2 to be the open one, got %v", open)
}
// The new incident starts unacknowledged: that is the point of the split.
if open[0]["acknowledged_by"] != nil {
t.Error("expected a fresh occurrence to start unacknowledged")
}
}
func TestWebhook_ResolvedOnlyPayloadOpensNothing(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-res", "AlreadyOver", "resolved", "2026-05-20T10:00:00Z", "2026-05-20T11:00:00Z", nil),
}, "{}:{alertname=\"AlreadyOver\"}")
if n := s.countIncidents(t); n != 0 {
t.Errorf("expected no incident from a resolved-only payload, got %d", n)
}
if status, _, _ := s.alertRow(t, "fp-res"); status != "resolved" {
t.Errorf("expected the alert itself to be stored, got %q", status)
}
}
// An incident that hit critical was a critical incident, even once the critical
// alert clears and only a warning is left.
func TestIncident_SeverityIsHighWaterMark(t *testing.T) {
s := newTS(t)
const groupKey = "{}:{alertname=\"Mixed\"}"
postWebhook(t, s, []map[string]any{
amAlert("fp-warn", "Mixed", "firing", "2026-05-20T10:00:00Z", zeroTime,
map[string]string{"severity": "warning"}),
amAlert("fp-crit", "Mixed", "firing", "2026-05-20T10:00:00Z", zeroTime,
map[string]string{"severity": "critical"}),
}, groupKey)
if inc := getIncident(t, s, 1); inc["severity"] != "critical" {
t.Fatalf("expected severity critical, got %v", inc["severity"])
}
// The critical alert clears; the warning keeps the incident open.
postWebhook(t, s, []map[string]any{
amAlert("fp-crit", "Mixed", "resolved", "2026-05-20T10:00:00Z", "2026-05-20T11:00:00Z",
map[string]string{"severity": "critical"}),
}, groupKey)
inc := getIncident(t, s, 1)
if inc["resolved_at"] != nil {
t.Fatal("expected the incident to stay open")
}
if inc["severity"] != "critical" {
t.Errorf("expected severity to stay critical, got %v", inc["severity"])
}
}
// ---------------------------------------------------------------------------
// Resolution cascade
// ---------------------------------------------------------------------------
func TestIncident_AllAlertsResolvedAutoResolves(t *testing.T) {
s := newTS(t)
const groupKey = "{}:{alertname=\"Pair\"}"
postWebhook(t, s, []map[string]any{
amAlert("fp-p1", "Pair", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
amAlert("fp-p2", "Pair", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
}, groupKey)
// One down, one still firing: the work is not done.
postWebhook(t, s, []map[string]any{
amAlert("fp-p1", "Pair", "resolved", "2026-05-20T10:00:00Z", "2026-05-20T11:00:00Z", nil),
}, groupKey)
if inc := getIncident(t, s, 1); inc["resolved_at"] != nil {
t.Fatal("expected the incident to stay open while an alert is firing")
}
postWebhook(t, s, []map[string]any{
amAlert("fp-p2", "Pair", "resolved", "2026-05-20T10:00:00Z", "2026-05-20T11:30:00Z", nil),
}, groupKey)
inc := getIncident(t, s, 1)
if inc["status"] != "resolved" {
t.Errorf("expected status resolved, got %v", inc["status"])
}
if inc["resolution_source"] != "alerts" {
t.Errorf("expected resolution_source alerts, got %v", inc["resolution_source"])
}
}
// Expiry is inference, not observation, but it still has to close the work item
// — otherwise a lost resolved notification leaves an incident open forever.
func TestExpiry_CascadesToIncidentResolution(t *testing.T) {
s := newTS(t)
postAlert(t, s, "fp-exp", "firing", time.Now().Add(-24*time.Hour).Format(time.RFC3339), zeroTime)
s.exec(t, "UPDATE alerts SET received_at = ? WHERE fingerprint = 'fp-exp'",
time.Now().Add(-10*time.Hour).Unix())
sweep(t, s, 6*time.Hour)
inc := getIncident(t, s, 1)
if inc["status"] != "resolved" {
t.Errorf("expected the incident to resolve after expiry, got %v", inc["status"])
}
if inc["resolution_source"] != "alerts" {
t.Errorf("expected resolution_source alerts, got %v", inc["resolution_source"])
}
// The expiry is recorded against the alert, not the incident.
if _, source, _ := s.alertRow(t, "fp-exp"); source == nil || *source != "expiry" {
t.Errorf("expected the alert's resolution_source to stay expiry, got %v", source)
}
if types := eventTypes(timeline(t, s, 1)); !contains(types, "alert_resolved") {
t.Errorf("expected an alert_resolved event on the timeline, got %v", types)
}
}
// ---------------------------------------------------------------------------
// Workflow actions
// ---------------------------------------------------------------------------
func TestIncident_Acknowledge(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-ack", "X", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
resp := s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("acknowledge returned %d", resp.StatusCode)
}
var inc map[string]any
decode(t, resp, &inc)
if inc["acknowledged_by"] == nil {
t.Error("expected acknowledged_by to be set")
}
if inc["status"] != "acknowledged" {
t.Errorf("expected status acknowledged, got %v", inc["status"])
}
resp = s.req(t, http.MethodDelete, "/api/incidents/1/acknowledge", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Fatalf("unacknowledge returned %d", resp.StatusCode)
}
if inc := getIncident(t, s, 1); inc["status"] != "triggered" {
t.Errorf("expected status back to triggered, got %v", inc["status"])
}
}
func TestIncident_ManualResolveIsTerminal(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-term", "Terminal", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
resp := s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("first resolve returned %d", resp.StatusCode)
}
var inc map[string]any
decode(t, resp, &inc)
if inc["resolution_source"] != "manual" {
t.Errorf("expected resolution_source manual, got %v", inc["resolution_source"])
}
resp = s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusConflict {
t.Errorf("expected 409 on re-resolve, got %d", resp.StatusCode)
}
// Acknowledging a closed incident is equally meaningless.
resp = s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusConflict {
t.Errorf("expected 409 acknowledging a resolved incident, got %d", resp.StatusCode)
}
}
func TestIncident_SnoozeHiddenFromDefaultList(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-snz", "Noisy", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
resp := s.req(t, http.MethodPost, "/api/incidents/1/snooze", map[string]string{"duration": "2h"})
if resp.StatusCode != http.StatusOK {
t.Fatalf("snooze returned %d", resp.StatusCode)
}
var inc map[string]any
decode(t, resp, &inc)
if inc["snoozed_until"] == nil {
t.Error("expected snoozed_until to be set")
}
if got := listIncidents(t, s, ""); len(got) != 0 {
t.Errorf("expected the snoozed incident to be hidden, got %d", len(got))
}
if got := listIncidents(t, s, "?snoozed=true"); len(got) != 1 {
t.Errorf("expected snoozed=true to show it, got %d", len(got))
}
// A snooze is not a resolution: the incident is still open work.
if inc := getIncident(t, s, 1); inc["resolved_at"] != nil {
t.Error("expected a snoozed incident to stay open")
}
resp = s.req(t, http.MethodDelete, "/api/incidents/1/snooze", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Fatalf("unsnooze returned %d", resp.StatusCode)
}
if got := listIncidents(t, s, ""); len(got) != 1 {
t.Errorf("expected the incident back in the default list, got %d", len(got))
}
}
func TestIncident_SnoozeRejectsPastDeadline(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-past", "Past", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
resp := s.req(t, http.MethodPost, "/api/incidents/1/snooze",
map[string]string{"until": time.Now().Add(-time.Hour).UTC().Format(time.RFC3339)})
resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("expected 400 for a snooze in the past, got %d", resp.StatusCode)
}
}
// The schedule stops being decorative here: it is read at trigger time.
func TestIncident_AutoAssignedToCurrentOnCall(t *testing.T) {
s := newTS(t)
today := time.Now().UTC().Format("2006-01-02")
resp := s.req(t, http.MethodPost, "/api/schedule",
map[string]any{"user_id": 1, "dates": []string{today}})
if resp.StatusCode != http.StatusCreated {
t.Fatalf("schedule assignment returned %d", resp.StatusCode)
}
resp.Body.Close()
postWebhook(t, s, []map[string]any{
amAlert("fp-oncall", "PageMe", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
inc := getIncident(t, s, 1)
if inc["assigned_to"] != "admin" {
t.Errorf("expected the incident assigned to today's on-call, got %v", inc["assigned_to"])
}
if types := eventTypes(timeline(t, s, 1)); !contains(types, "assigned") {
t.Errorf("expected an assigned event, got %v", types)
}
}
func TestIncident_AssignToUser(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-asg", "Assignable", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.req(t, http.MethodPost, "/api/users",
map[string]string{"username": "alice", "email": "alice@test.com"}).Body.Close()
resp := s.req(t, http.MethodPost, "/api/incidents/1/assign", map[string]any{"user_id": 2})
if resp.StatusCode != http.StatusOK {
t.Fatalf("assign returned %d", resp.StatusCode)
}
var inc map[string]any
decode(t, resp, &inc)
if inc["assigned_to"] != "alice" {
t.Errorf("expected assigned_to alice, got %v", inc["assigned_to"])
}
resp = s.req(t, http.MethodPost, "/api/incidents/1/assign", map[string]any{"user_id": 99})
resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("expected 404 assigning an unknown user, got %d", resp.StatusCode)
}
}
// ---------------------------------------------------------------------------
// Timeline and notes
// ---------------------------------------------------------------------------
func TestIncident_TimelineOrdering(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-tl", "Storyline", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil).Body.Close()
s.req(t, http.MethodPost, "/api/incidents/1/notes",
map[string]string{"content": "looking into it"}).Body.Close()
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
events := timeline(t, s, 1)
want := []string{"triggered", "alert_added", "acknowledged", "note", "resolved"}
got := eventTypes(events)
if len(got) != len(want) {
t.Fatalf("expected timeline %v, got %v", want, got)
}
for i := range want {
if got[i] != want[i] {
t.Fatalf("expected timeline %v, got %v", want, got)
}
}
// The note carries its author; system events do not.
for _, e := range events {
if e["type"] == "note" {
if e["username"] != "admin" || e["detail"] != "looking into it" {
t.Errorf("unexpected note event: %v", e)
}
}
if e["type"] == "triggered" && e["username"] != nil {
t.Errorf("expected the triggered event to have no author, got %v", e["username"])
}
}
}
func TestIncident_NoteDeleteOwnOnly(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-note", "Y", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.req(t, http.MethodPost, "/api/users",
map[string]string{"username": "alice", "email": "alice@test.com"}).Body.Close()
var keyData map[string]any
decode(t, s.req(t, http.MethodPost, "/api/users/2/api-keys",
map[string]string{"name": "alice-key"}), &keyData)
aliceKey := keyData["key"].(string)
var note map[string]any
decode(t, s.req(t, http.MethodPost, "/api/incidents/1/notes",
map[string]string{"content": "admin note"}), &note)
noteID := int(note["id"].(float64))
// Alice cannot delete admin's note.
req, _ := http.NewRequest(http.MethodDelete,
fmt.Sprintf("%s/api/incidents/1/notes/%d", s.URL, noteID), nil)
req.Header.Set("Authorization", "Bearer "+aliceKey)
resp, _ := http.DefaultClient.Do(req)
resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("expected 404 deleting another user's note, got %d", resp.StatusCode)
}
resp = s.req(t, http.MethodDelete, fmt.Sprintf("/api/incidents/1/notes/%d", noteID), nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Errorf("expected 204 deleting own note, got %d", resp.StatusCode)
}
}
// Only notes are deletable — the rest of the timeline is what happened.
func TestIncident_CannotDeleteSystemEvent(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-sys", "System", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
events := timeline(t, s, 1)
id := int(events[0]["id"].(float64))
resp := s.req(t, http.MethodDelete, fmt.Sprintf("/api/incidents/1/notes/%d", id), nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("expected 404 deleting a system event, got %d", resp.StatusCode)
}
}
// ---------------------------------------------------------------------------
// Archive
// ---------------------------------------------------------------------------
func TestIncident_ArchiveRoundTrip(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-arc", "Archivable", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
resp := s.req(t, http.MethodPost, "/api/incidents/1/archive", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("archive returned %d", resp.StatusCode)
}
var inc map[string]any
decode(t, resp, &inc)
if inc["archived_at"] == nil {
t.Error("expected archived_at to be set")
}
if got := listIncidents(t, s, "?status=resolved"); len(got) != 0 {
t.Errorf("expected the archived incident to be hidden, got %d", len(got))
}
if got := listIncidents(t, s, "?status=resolved&archived=true"); len(got) != 1 {
t.Errorf("expected archived=true to show it, got %d", len(got))
}
resp = s.req(t, http.MethodDelete, "/api/incidents/1/archive", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Fatalf("unarchive returned %d", resp.StatusCode)
}
if got := listIncidents(t, s, "?status=resolved"); len(got) != 1 {
t.Errorf("expected the incident back, got %d", len(got))
}
}
func TestSweeper_ArchivesResolvedIncidents(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-swp", "Old", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
s.exec(t, "UPDATE incidents SET resolved_at = ? WHERE id = 1",
time.Now().Add(-30*24*time.Hour).Unix())
api.Sweep(context.Background(), s.db, 7*24*time.Hour, 6*time.Hour)
if inc := getIncident(t, s, 1); inc["archived_at"] == nil {
t.Error("expected the sweeper to archive a long-resolved incident")
}
}
// ---------------------------------------------------------------------------
// Stats
// ---------------------------------------------------------------------------
func TestStats_Incidents(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-s1", "One", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
}, "g1")
postWebhook(t, s, []map[string]any{
amAlert("fp-s2", "Two", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
}, "g2")
s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil).Body.Close()
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
var stats map[string]any
decode(t, s.req(t, http.MethodGet, "/api/stats/incidents", nil), &stats)
if stats["total"].(float64) != 2 {
t.Errorf("expected total 2, got %v", stats["total"])
}
if stats["resolved"].(float64) != 1 {
t.Errorf("expected resolved 1, got %v", stats["resolved"])
}
if stats["triggered"].(float64) != 1 {
t.Errorf("expected triggered 1, got %v", stats["triggered"])
}
// One incident has been acknowledged and resolved, so both averages exist.
if stats["mtta_seconds"] == nil || stats["mttr_seconds"] == nil {
t.Errorf("expected mtta and mttr to be computable, got %v", stats)
}
}
// Nothing acknowledged yet means "no data", which is not the same claim as zero.
func TestStats_IncidentsNullMTTAWhenNothingAcknowledged(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-s3", "Untouched", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
var stats map[string]any
decode(t, s.req(t, http.MethodGet, "/api/stats/incidents", nil), &stats)
if stats["mtta_seconds"] != nil {
t.Errorf("expected mtta_seconds null, got %v", stats["mtta_seconds"])
}
}
// ---------------------------------------------------------------------------
// Migration backfill
// ---------------------------------------------------------------------------
// An upgrade must not drop the acknowledgements and comments people already
// have, so 008 is replayed here over a database left at 007.
func TestMigration_BackfillCarriesAckAndComments(t *testing.T) {
database, err := db.Open(":memory:")
if err != nil {
t.Fatalf("open db: %v", err)
}
t.Cleanup(func() { database.Close() })
files, err := filepath.Glob("../db/migrations/*.sql")
if err != nil || len(files) == 0 {
t.Fatalf("find migrations: %v", err)
}
sort.Strings(files)
var incidentsMigration string
for _, f := range files {
if filepath.Base(f) >= "008" {
incidentsMigration = f
break
}
data, err := os.ReadFile(f)
if err != nil {
t.Fatalf("read %s: %v", f, err)
}
if _, err := database.Exec(string(data)); err != nil {
t.Fatalf("apply %s: %v", f, err)
}
}
if incidentsMigration == "" {
t.Fatal("008 migration not found")
}
// A database as it would look on the old schema: an acknowledged firing
// alert with a comment on it.
if _, err := database.Exec(`
INSERT INTO users (id, username, email) VALUES (1, 'admin', 'admin@test.com');
INSERT INTO alerts (id, fingerprint, name, status, labels, annotations,
starts_at, received_at, acknowledged_by, acknowledged_at)
VALUES (1, 'legacy-fp', 'LegacyAlert', 'firing',
'{"severity":"warning"}', '{}', 1000, 1000, 1, 1500);
INSERT INTO alert_comments (alert_id, user_id, content, created_at)
VALUES (1, 1, 'legacy comment', 1600);`); err != nil {
t.Fatalf("seed pre-008 data: %v", err)
}
data, err := os.ReadFile(incidentsMigration)
if err != nil {
t.Fatalf("read 008: %v", err)
}
if _, err := database.Exec(string(data)); err != nil {
t.Fatalf("apply 008: %v", err)
}
var status, groupKey string
var ackBy int64
var severity string
if err := database.QueryRow(
"SELECT status, group_key, acknowledged_by, severity FROM incidents WHERE id = 1",
).Scan(&status, &groupKey, &ackBy, &severity); err != nil {
t.Fatalf("read backfilled incident: %v", err)
}
if status != "acknowledged" {
t.Errorf("expected the ack to carry over as status, got %q", status)
}
if groupKey != "backfill:legacy-fp" {
t.Errorf("unexpected group_key %q", groupKey)
}
if ackBy != 1 {
t.Errorf("expected acknowledged_by 1, got %d", ackBy)
}
if severity != "warning" {
t.Errorf("expected severity carried from labels, got %q", severity)
}
var notes int
if err := database.QueryRow(
"SELECT COUNT(*) FROM incident_events WHERE type = 'note' AND detail = 'legacy comment'",
).Scan(&notes); err != nil {
t.Fatalf("count notes: %v", err)
}
if notes != 1 {
t.Errorf("expected the comment to become a note, got %d", notes)
}
// And the columns that caused the ack-survives-a-re-fire bug are gone.
if _, err := database.Exec("SELECT acknowledged_by FROM alerts"); err == nil {
t.Error("expected alerts.acknowledged_by to be dropped")
}
}
func contains(haystack []string, needle string) bool {
for _, s := range haystack {
if s == needle {
return true
}
}
return false
}
+18 -7
View File
@@ -31,21 +31,32 @@ func NewRouter(db *sql.DB) http.Handler {
r.Post("/api/users/{id}/api-keys", handleCreateAPIKey(db))
r.Delete("/api/users/{id}/api-keys/{keyID}", handleDeleteAPIKey(db))
// Alerts are read-only: they are Alertmanager's record, not a work
// queue. Everything a person does happens on the incident instead.
r.Get("/api/alerts", handleListAlerts(db))
r.Get("/api/alerts/{id}", handleGetAlert(db))
r.Post("/api/alerts/{id}/acknowledge", handleAcknowledge(db))
r.Delete("/api/alerts/{id}/acknowledge", handleUnacknowledge(db))
r.Post("/api/alerts/{id}/archive", handleArchive(db))
r.Delete("/api/alerts/{id}/archive", handleUnarchive(db))
r.Get("/api/alerts/{id}/comments", handleListComments(db))
r.Post("/api/alerts/{id}/comments", handleCreateComment(db))
r.Delete("/api/alerts/{id}/comments/{commentID}", handleDeleteComment(db))
r.Get("/api/incidents", handleListIncidents(db))
r.Get("/api/incidents/{id}", handleGetIncident(db))
r.Get("/api/incidents/{id}/alerts", handleIncidentAlerts(db))
r.Get("/api/incidents/{id}/timeline", handleIncidentTimeline(db))
r.Post("/api/incidents/{id}/acknowledge", handleIncidentAcknowledge(db))
r.Delete("/api/incidents/{id}/acknowledge", handleIncidentUnacknowledge(db))
r.Post("/api/incidents/{id}/resolve", handleIncidentResolve(db))
r.Post("/api/incidents/{id}/assign", handleIncidentAssign(db))
r.Post("/api/incidents/{id}/snooze", handleIncidentSnooze(db))
r.Delete("/api/incidents/{id}/snooze", handleIncidentUnsnooze(db))
r.Post("/api/incidents/{id}/archive", handleIncidentArchive(db))
r.Delete("/api/incidents/{id}/archive", handleIncidentUnarchive(db))
r.Post("/api/incidents/{id}/notes", handleCreateNote(db))
r.Delete("/api/incidents/{id}/notes/{eventID}", handleDeleteNote(db))
r.Post("/api/schedule", handleCreateSchedule(db))
r.Get("/api/schedule/current", handleCurrentSchedule(db)) // must be before /{id}
r.Get("/api/schedule", handleListSchedule(db))
r.Delete("/api/schedule/{id}", handleDeleteSchedule(db))
r.Get("/api/stats/incidents", handleStatsIncidents(db))
r.Get("/api/stats/alerts", handleStatsAlerts(db))
r.Get("/api/stats/alerts/top", handleStatsTop(db))
r.Get("/api/stats/alerts/by-hour", handleStatsByHour(db))
+49 -9
View File
@@ -11,7 +11,7 @@ import (
func handleStatsAlerts(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
where, args := statsFilter(r.URL.Query())
where, args := statsFilter(r.URL.Query(), "received_at")
var total, firing, resolved int64
err := db.QueryRowContext(r.Context(), fmt.Sprintf(`
@@ -34,7 +34,7 @@ func handleStatsAlerts(db *sql.DB) http.HandlerFunc {
func handleStatsTop(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
where, args := statsFilter(r.URL.Query())
where, args := statsFilter(r.URL.Query(), "received_at")
limit := 10
if l := r.URL.Query().Get("limit"); l != "" {
@@ -78,7 +78,7 @@ func handleStatsTop(db *sql.DB) http.HandlerFunc {
func handleStatsByHour(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
where, args := statsFilter(r.URL.Query())
where, args := statsFilter(r.URL.Query(), "received_at")
rows, err := db.QueryContext(r.Context(), fmt.Sprintf(`
SELECT CAST(strftime('%%H', datetime(received_at, 'unixepoch')) AS INTEGER) AS hr,
@@ -118,7 +118,7 @@ func handleStatsByHour(db *sql.DB) http.HandlerFunc {
func handleStatsByDay(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
where, args := statsFilter(r.URL.Query())
where, args := statsFilter(r.URL.Query(), "received_at")
// SQLite strftime('%w') → 0=Sunday … 6=Saturday
rows, err := db.QueryContext(r.Context(), fmt.Sprintf(`
@@ -159,19 +159,59 @@ func handleStatsByDay(db *sql.DB) http.HandlerFunc {
}
}
// statsFilter builds a WHERE clause and args from optional ?from and ?to query params.
// Archived alerts are always excluded, matching the default GET /api/alerts view.
func statsFilter(q url.Values) (where string, args []any) {
// handleStatsIncidents reports the queue and the two numbers a rota actually
// cares about: how long it takes someone to pick work up, and how long it takes
// to finish. Neither was computable before incidents existed — alert rows are
// mutated in place and carry no acknowledgement or closure time.
func handleStatsIncidents(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
where, args := statsFilter(r.URL.Query(), "triggered_at")
var total, triggered, acknowledged, resolved int64
var mtta, mttr *float64
err := db.QueryRowContext(r.Context(), fmt.Sprintf(`
SELECT COUNT(*),
SUM(CASE WHEN status = 'triggered' THEN 1 ELSE 0 END),
SUM(CASE WHEN status = 'acknowledged' THEN 1 ELSE 0 END),
SUM(CASE WHEN status = 'resolved' THEN 1 ELSE 0 END),
AVG(CASE WHEN acknowledged_at IS NOT NULL
THEN acknowledged_at - triggered_at END),
AVG(CASE WHEN resolved_at IS NOT NULL
THEN resolved_at - triggered_at END)
FROM incidents WHERE %s`, where), args...,
).Scan(&total, &triggered, &acknowledged, &resolved, &mtta, &mttr)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, map[string]any{
"total": total,
"triggered": triggered,
"acknowledged": acknowledged,
"resolved": resolved,
// Null until something has actually been acknowledged or resolved —
// zero would read as "instant", which is a different claim.
"mtta_seconds": mtta,
"mttr_seconds": mttr,
})
}
}
// statsFilter builds a WHERE clause and args from optional ?from and ?to query
// params, filtering on timeCol. Archived rows are always excluded, matching the
// default list views.
func statsFilter(q url.Values, timeCol string) (where string, args []any) {
clauses := []string{"archived_at IS NULL"}
if from := q.Get("from"); from != "" {
if t, err := time.Parse("2006-01-02", from); err == nil {
clauses = append(clauses, "received_at >= ?")
clauses = append(clauses, timeCol+" >= ?")
args = append(args, t.UTC().Unix())
}
}
if to := q.Get("to"); to != "" {
if t, err := time.Parse("2006-01-02", to); err == nil {
clauses = append(clauses, "received_at < ?")
clauses = append(clauses, timeCol+" < ?")
args = append(args, t.UTC().AddDate(0, 0, 1).Unix())
}
}