Turn incoming alerts into incidents
Release / build (amd64, linux) (push) Failing after 11s
Release / build (amd64, darwin) (push) Failing after 12s
Release / build (arm64, darwin) (push) Failing after 11s
Release / build (arm64, linux) (push) Failing after 11s
Release / release (push) Has been skipped
Release / chart (push) Failing after 13s
Release / docker (push) Failing after 19s
Release / build (amd64, linux) (push) Failing after 11s
Release / build (amd64, darwin) (push) Failing after 12s
Release / build (arm64, darwin) (push) Failing after 11s
Release / build (arm64, linux) (push) Failing after 11s
Release / release (push) Has been skipped
Release / chart (push) Failing after 13s
Release / docker (push) Failing after 19s
The alerts row was both Alertmanager's record and the human work queue, and
the two have different owners. The webhook upsert rewrites that row on every
notification; acknowledgement, comments and archiving were columns on it that
the upsert happened not to touch. So an alert that resolved and re-fired days
later still read as acknowledged by whoever acked the first occurrence — the
ack outlived the thing it referred to. Nothing recorded transitions either:
rows are mutated in place, so there was no timeline and no way to compute how
long anything took.
Alerts are now read-only signal records with two states, and incidents are
the work item: triggered, acknowledged or resolved, with an assignee, a
snooze, notes and an append-only timeline. Many alerts map to one incident,
and a new occurrence opens a new incident, which is what makes a stale ack
impossible rather than merely unlikely.
Correlation uses Alertmanager's own groupKey. It already grouped the alerts
according to the group_by routing tree the operator configured and sends the
result on every webhook, where it was being discarded; adopting it means
changing group_by in alertmanager.yml changes correlation here, with no
second grouping scheme to configure and keep in sync.
An incident opens only when an alert transitions into firing — an unseen
fingerprint, a newer startsAt, or a resolved alert starting again. The
unchanged notifications Alertmanager re-sends every repeat_interval are none
of those. That rule is what lets manual resolution be terminal: without it,
closing an incident by hand would be undone by the next re-send of an alert
that never stopped firing, and the button would be a lie. Snooze covers the
"not now" case instead. Incidents otherwise resolve by cascade, once every
alert under them has stopped firing, whether by webhook or by expiry.
New incidents are assigned to whoever holds today's schedule entry. The
schedule table has existed since the first release with nothing reading it.
Also here, following from the split:
- Incident severity is a high-water mark over its alerts, never lowered.
An incident that hit critical was a critical incident, and downgrading a
live one would demote it in the queue while the work is still open.
- /api/stats/incidents reports MTTA and MTTR, null rather than zero until
there is something to average. Neither was computable before.
- Alert archiving becomes sweeper-only housekeeping; the archive people
interact with is the incident's.
Breaking: the alert acknowledge, archive and comment endpoints are gone, and
the alert object drops the acknowledgement fields and gains incident_id. The
README maps each removed endpoint to its replacement. Migration 008 backfills
an incident per existing alert, archived ones included so no comment is
orphaned, carrying acknowledgements across and turning comments into timeline
notes.
Both documented alert contracts are untouched: received_at still advances on
every accepted payload, re-sends included, and resolution_source still says
how much to trust ends_at. The upsert is byte-for-byte what it was, now
running inside the ingest transaction.
This commit is contained in:
+310
-62
@@ -1,6 +1,7 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"encoding/json"
|
||||
"log"
|
||||
@@ -17,9 +18,17 @@ const (
|
||||
|
||||
// amPayload mirrors the Alertmanager webhook v4 payload.
|
||||
type amPayload struct {
|
||||
Version string `json:"version"`
|
||||
Status string `json:"status"`
|
||||
Alerts []amAlert `json:"alerts"`
|
||||
Version string `json:"version"`
|
||||
Status string `json:"status"`
|
||||
|
||||
// GroupKey and GroupLabels are how alerts get correlated into incidents.
|
||||
// Alertmanager has already done the grouping work according to the group_by
|
||||
// routing tree the operator configured, so we adopt its answer instead of
|
||||
// inventing a second grouping scheme here.
|
||||
GroupKey string `json:"groupKey"`
|
||||
GroupLabels map[string]string `json:"groupLabels"`
|
||||
|
||||
Alerts []amAlert `json:"alerts"`
|
||||
}
|
||||
|
||||
type amAlert struct {
|
||||
@@ -32,6 +41,23 @@ type amAlert struct {
|
||||
Fingerprint string `json:"fingerprint"`
|
||||
}
|
||||
|
||||
// ingested records what actually happened to one alert of a payload, which is
|
||||
// what decides whether an incident opens.
|
||||
type ingested struct {
|
||||
id int64
|
||||
name string
|
||||
firing bool
|
||||
|
||||
// newOccurrence marks an alert that transitioned *into* firing: a
|
||||
// fingerprint we had never seen, a newer startsAt, or a resolved alert that
|
||||
// started again. A repeat_interval re-send of an already-firing alert is
|
||||
// none of these, which is what keeps a manually resolved incident closed.
|
||||
newOccurrence bool
|
||||
|
||||
// justResolved marks the firing → resolved edge, worth a timeline entry.
|
||||
justResolved bool
|
||||
}
|
||||
|
||||
func handleAlertmanagerWebhook(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
var payload amPayload
|
||||
@@ -40,67 +66,289 @@ func handleAlertmanagerWebhook(db *sql.DB) http.HandlerFunc {
|
||||
return
|
||||
}
|
||||
|
||||
now := time.Now().Unix()
|
||||
for _, a := range payload.Alerts {
|
||||
name := a.Labels["alertname"]
|
||||
labelsJSON, _ := json.Marshal(a.Labels)
|
||||
annotationsJSON, _ := json.Marshal(a.Annotations)
|
||||
|
||||
// Zero time ("0001-01-01T00:00:00Z") means "no end known" — that is the
|
||||
// convention of Alertmanager's ingest API. Outgoing notifications
|
||||
// normally carry a real future endsAt instead, which is the watermark
|
||||
// the sweeper uses to expire alerts that stop being refreshed.
|
||||
var endsAtUnix *int64
|
||||
if a.EndsAt.Year() > 1 {
|
||||
t := a.EndsAt.Unix()
|
||||
endsAtUnix = &t
|
||||
}
|
||||
|
||||
var resolutionSource *string
|
||||
if a.Status == "resolved" {
|
||||
s := resolutionAlertmanager
|
||||
resolutionSource = &s
|
||||
}
|
||||
|
||||
// The WHERE clause discards payloads that describe an alert instance
|
||||
// older than the stored one. Alertmanager retries failed notifications,
|
||||
// so a stale firing retry can arrive after the resolved one; it carries
|
||||
// the same startsAt, whereas a genuine re-fire carries a newer one.
|
||||
// Within a single instance, resolution is terminal.
|
||||
_, err := db.ExecContext(r.Context(), `
|
||||
INSERT INTO alerts
|
||||
(fingerprint, name, status, labels, annotations, starts_at, ends_at,
|
||||
generator_url, received_at, resolution_source)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
ON CONFLICT(fingerprint) DO UPDATE SET
|
||||
status = excluded.status,
|
||||
labels = excluded.labels,
|
||||
annotations = excluded.annotations,
|
||||
starts_at = excluded.starts_at,
|
||||
ends_at = excluded.ends_at,
|
||||
generator_url = excluded.generator_url,
|
||||
-- Load-bearing: advancing received_at on every accepted
|
||||
-- payload, re-sends included, is the documented liveness
|
||||
-- heartbeat clients and the sweeper both read. Removing it
|
||||
-- is a breaking API change — see models.Alert.ReceivedAt.
|
||||
received_at = excluded.received_at,
|
||||
resolution_source = excluded.resolution_source,
|
||||
-- A re-fire makes the alert current again, so it leaves the archive.
|
||||
archived_at = CASE WHEN excluded.status = 'firing'
|
||||
THEN NULL ELSE alerts.archived_at END
|
||||
WHERE excluded.starts_at > alerts.starts_at
|
||||
OR (excluded.starts_at = alerts.starts_at
|
||||
AND NOT (alerts.status = 'resolved' AND excluded.status = 'firing'))`,
|
||||
a.Fingerprint, name, a.Status,
|
||||
string(labelsJSON), string(annotationsJSON),
|
||||
a.StartsAt.Unix(), endsAtUnix,
|
||||
a.GeneratorURL, now, resolutionSource,
|
||||
)
|
||||
if err != nil {
|
||||
log.Printf("upsert alert %s: %v", a.Fingerprint, err)
|
||||
}
|
||||
// Alertmanager retries anything that is not 2xx, and a retry of a payload
|
||||
// we failed to store is more useful than an error it cannot act on — so
|
||||
// failures are logged, not surfaced.
|
||||
if err := ingest(r.Context(), db, payload); err != nil {
|
||||
log.Printf("webhook ingest (group %q): %v", payload.GroupKey, err)
|
||||
}
|
||||
|
||||
w.WriteHeader(http.StatusOK)
|
||||
}
|
||||
}
|
||||
|
||||
// ingest stores a payload's alerts and reconciles the incident for its group.
|
||||
// The whole payload is one transaction: an incident that opened but whose alerts
|
||||
// failed to link would be a work item nobody could act on.
|
||||
func ingest(ctx context.Context, db *sql.DB, payload amPayload) error {
|
||||
tx, err := db.BeginTx(ctx, nil)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer tx.Rollback() //nolint:errcheck
|
||||
|
||||
accepted, err := upsertAlerts(ctx, tx, payload.Alerts)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
// touched collects every incident this payload affected, so severity and the
|
||||
// resolution cascade are recomputed once per incident at the end.
|
||||
touched := map[int64]bool{}
|
||||
|
||||
incidentID, err := incidentForGroup(ctx, tx, payload, accepted)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if incidentID != 0 {
|
||||
touched[incidentID] = true
|
||||
for _, a := range accepted {
|
||||
if !a.firing {
|
||||
continue
|
||||
}
|
||||
if err := linkAlert(ctx, tx, incidentID, a.id); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
for _, a := range accepted {
|
||||
if !a.justResolved {
|
||||
continue
|
||||
}
|
||||
id, err := openIncidentForAlert(ctx, tx, a.id)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if id == 0 {
|
||||
continue
|
||||
}
|
||||
touched[id] = true
|
||||
alertID := a.id
|
||||
if err := logEvent(ctx, tx, id, evAlertResolved, nil, &alertID, nil); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
|
||||
for id := range touched {
|
||||
if err := refreshSeverity(ctx, tx, id); err != nil {
|
||||
return err
|
||||
}
|
||||
if _, err := resolveIfSettled(ctx, tx, id); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
|
||||
return tx.Commit()
|
||||
}
|
||||
|
||||
// upsertAlerts stores each alert of a payload and reports what changed. Payloads
|
||||
// the ordering guard rejected are left out entirely.
|
||||
func upsertAlerts(ctx context.Context, tx *sql.Tx, alerts []amAlert) ([]ingested, error) {
|
||||
now := time.Now().Unix()
|
||||
accepted := make([]ingested, 0, len(alerts))
|
||||
|
||||
for _, a := range alerts {
|
||||
name := a.Labels["alertname"]
|
||||
labelsJSON, _ := json.Marshal(a.Labels)
|
||||
annotationsJSON, _ := json.Marshal(a.Annotations)
|
||||
|
||||
// The stored state has to be read before the upsert overwrites it: it is
|
||||
// the only way to tell a genuine new occurrence from a re-send.
|
||||
var prevStatus string
|
||||
var prevStartsAt int64
|
||||
existed := true
|
||||
switch err := tx.QueryRowContext(ctx,
|
||||
"SELECT status, starts_at FROM alerts WHERE fingerprint = ?", a.Fingerprint,
|
||||
).Scan(&prevStatus, &prevStartsAt); {
|
||||
case err == sql.ErrNoRows:
|
||||
existed = false
|
||||
case err != nil:
|
||||
return nil, err
|
||||
}
|
||||
|
||||
// Zero time ("0001-01-01T00:00:00Z") means "no end known" — that is the
|
||||
// convention of Alertmanager's ingest API. Outgoing notifications
|
||||
// normally carry a real future endsAt instead, which is the watermark
|
||||
// the sweeper uses to expire alerts that stop being refreshed.
|
||||
var endsAtUnix *int64
|
||||
if a.EndsAt.Year() > 1 {
|
||||
t := a.EndsAt.Unix()
|
||||
endsAtUnix = &t
|
||||
}
|
||||
|
||||
var resolutionSource *string
|
||||
if a.Status == "resolved" {
|
||||
s := resolutionAlertmanager
|
||||
resolutionSource = &s
|
||||
}
|
||||
|
||||
// The WHERE clause discards payloads that describe an alert instance
|
||||
// older than the stored one. Alertmanager retries failed notifications,
|
||||
// so a stale firing retry can arrive after the resolved one; it carries
|
||||
// the same startsAt, whereas a genuine re-fire carries a newer one.
|
||||
// Within a single instance, resolution is terminal.
|
||||
if _, err := tx.ExecContext(ctx, `
|
||||
INSERT INTO alerts
|
||||
(fingerprint, name, status, labels, annotations, starts_at, ends_at,
|
||||
generator_url, received_at, resolution_source)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
ON CONFLICT(fingerprint) DO UPDATE SET
|
||||
status = excluded.status,
|
||||
labels = excluded.labels,
|
||||
annotations = excluded.annotations,
|
||||
starts_at = excluded.starts_at,
|
||||
ends_at = excluded.ends_at,
|
||||
generator_url = excluded.generator_url,
|
||||
-- Load-bearing: advancing received_at on every accepted
|
||||
-- payload, re-sends included, is the documented liveness
|
||||
-- heartbeat clients and the sweeper both read. Removing it
|
||||
-- is a breaking API change — see models.Alert.ReceivedAt.
|
||||
received_at = excluded.received_at,
|
||||
resolution_source = excluded.resolution_source,
|
||||
-- A re-fire makes the alert current again, so it leaves the archive.
|
||||
archived_at = CASE WHEN excluded.status = 'firing'
|
||||
THEN NULL ELSE alerts.archived_at END
|
||||
WHERE excluded.starts_at > alerts.starts_at
|
||||
OR (excluded.starts_at = alerts.starts_at
|
||||
AND NOT (alerts.status = 'resolved' AND excluded.status = 'firing'))`,
|
||||
a.Fingerprint, name, a.Status,
|
||||
string(labelsJSON), string(annotationsJSON),
|
||||
a.StartsAt.Unix(), endsAtUnix,
|
||||
a.GeneratorURL, now, resolutionSource,
|
||||
); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
var id int64
|
||||
var curStatus string
|
||||
var curStartsAt int64
|
||||
if err := tx.QueryRowContext(ctx,
|
||||
"SELECT id, status, starts_at FROM alerts WHERE fingerprint = ?", a.Fingerprint,
|
||||
).Scan(&id, &curStatus, &curStartsAt); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
// The upsert copies status and starts_at straight from the payload, so a
|
||||
// row that does not match it is one the ordering guard rejected. A
|
||||
// discarded payload describes a past instance and must not touch the
|
||||
// incident state either.
|
||||
if existed && (curStatus != a.Status || curStartsAt != a.StartsAt.Unix()) {
|
||||
continue
|
||||
}
|
||||
|
||||
firing := a.Status == "firing"
|
||||
accepted = append(accepted, ingested{
|
||||
id: id,
|
||||
name: name,
|
||||
firing: firing,
|
||||
newOccurrence: firing && (!existed || a.StartsAt.Unix() > prevStartsAt || prevStatus == "resolved"),
|
||||
justResolved: !firing && existed && prevStatus == "firing",
|
||||
})
|
||||
}
|
||||
|
||||
return accepted, nil
|
||||
}
|
||||
|
||||
// incidentForGroup returns the open incident that this payload's firing alerts
|
||||
// belong to, opening one if the group has none. It returns 0 when the payload
|
||||
// warrants no incident at all.
|
||||
//
|
||||
// The rule that matters: a group with no open incident gets a new one only if
|
||||
// something actually started firing. Without that, a manually resolved incident
|
||||
// would reappear on the next repeat_interval re-send of an alert that never
|
||||
// stopped, and manual resolution would be meaningless.
|
||||
func incidentForGroup(ctx context.Context, tx *sql.Tx, payload amPayload, accepted []ingested) (int64, error) {
|
||||
var firstName string
|
||||
anyFiring, anyNew := false, false
|
||||
for _, a := range accepted {
|
||||
if a.firing {
|
||||
if !anyFiring {
|
||||
firstName = a.name
|
||||
}
|
||||
anyFiring = true
|
||||
}
|
||||
if a.newOccurrence {
|
||||
anyNew = true
|
||||
}
|
||||
}
|
||||
if !anyFiring {
|
||||
// A payload of nothing but resolutions never opens an incident.
|
||||
return 0, nil
|
||||
}
|
||||
|
||||
groupKey := payload.GroupKey
|
||||
if groupKey == "" {
|
||||
// Alertmanager always sends groupKey; a sender that does not still gets
|
||||
// one incident per alert name rather than one giant shared incident.
|
||||
groupKey = "groupless:" + firstName
|
||||
}
|
||||
|
||||
var id int64
|
||||
switch err := tx.QueryRowContext(ctx,
|
||||
"SELECT id FROM incidents WHERE group_key = ? AND resolved_at IS NULL", groupKey,
|
||||
).Scan(&id); {
|
||||
case err == nil:
|
||||
return id, nil
|
||||
case err != sql.ErrNoRows:
|
||||
return 0, err
|
||||
}
|
||||
|
||||
if !anyNew {
|
||||
return 0, nil
|
||||
}
|
||||
return openIncident(ctx, tx, groupKey, payload.GroupLabels, firstName)
|
||||
}
|
||||
|
||||
// openIncident creates an incident for a group and assigns it to whoever is on
|
||||
// call today, which is the point at which the schedule stops being decorative.
|
||||
func openIncident(ctx context.Context, tx *sql.Tx, groupKey string, groupLabels map[string]string, fallbackName string) (int64, error) {
|
||||
onCall, err := currentOnCall(ctx, tx)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
|
||||
labelsJSON, _ := json.Marshal(groupLabels)
|
||||
if groupLabels == nil {
|
||||
labelsJSON = []byte("{}")
|
||||
}
|
||||
|
||||
res, err := tx.ExecContext(ctx, `
|
||||
INSERT INTO incidents (group_key, title, group_labels, status, triggered_at, assigned_to)
|
||||
VALUES (?, ?, ?, 'triggered', ?, ?)`,
|
||||
groupKey, incidentTitle(groupLabels, fallbackName), string(labelsJSON),
|
||||
time.Now().Unix(), onCall)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
id, err := res.LastInsertId()
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
|
||||
if err := logEvent(ctx, tx, id, evTriggered, nil, nil, nil); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
if onCall != nil {
|
||||
// On an "assigned" event user_id is the assignee, not the actor.
|
||||
if err := logEvent(ctx, tx, id, evAssigned, onCall, nil, nil); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
return id, nil
|
||||
}
|
||||
|
||||
// linkAlert adds an alert to an incident, emitting a timeline entry only the
|
||||
// first time. Re-sends of an already-linked alert are silent.
|
||||
func linkAlert(ctx context.Context, tx *sql.Tx, incidentID, alertID int64) error {
|
||||
res, err := tx.ExecContext(ctx, `
|
||||
INSERT OR IGNORE INTO incident_alerts (incident_id, alert_id, added_at)
|
||||
VALUES (?, ?, ?)`, incidentID, alertID, time.Now().Unix())
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
return nil
|
||||
}
|
||||
return logEvent(ctx, tx, incidentID, evAlertAdded, nil, &alertID, nil)
|
||||
}
|
||||
|
||||
+22
-115
@@ -15,15 +15,21 @@ import (
|
||||
)
|
||||
|
||||
// alertSelectFrom is the shared SELECT … FROM … clause used by all alert queries.
|
||||
// It LEFT JOINs users so acknowledged_by username is always available.
|
||||
// The subquery resolves the alert's most recent incident: membership is kept in
|
||||
// incident_alerts rather than as a column here, because one alert row is reused
|
||||
// across occurrences and belongs to a different incident each time.
|
||||
const alertSelectFrom = `
|
||||
SELECT a.id, a.fingerprint, a.name, a.status,
|
||||
a.labels, a.annotations,
|
||||
a.starts_at, a.ends_at, a.generator_url, a.received_at,
|
||||
a.acknowledged_by, a.acknowledged_at, u.username,
|
||||
(SELECT ia.incident_id
|
||||
FROM incident_alerts ia
|
||||
JOIN incidents i ON i.id = ia.incident_id
|
||||
WHERE ia.alert_id = a.id
|
||||
ORDER BY i.triggered_at DESC, i.id DESC
|
||||
LIMIT 1),
|
||||
a.resolution_source, a.archived_at
|
||||
FROM alerts a
|
||||
LEFT JOIN users u ON u.id = a.acknowledged_by`
|
||||
FROM alerts a`
|
||||
|
||||
func handleListAlerts(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
@@ -45,6 +51,12 @@ func handleListAlerts(db *sql.DB) http.HandlerFunc {
|
||||
} else {
|
||||
where = append(where, "a.archived_at IS NULL")
|
||||
}
|
||||
if incidentID := q.Get("incident_id"); incidentID != "" {
|
||||
if n, err := strconv.ParseInt(incidentID, 10, 64); err == nil {
|
||||
where = append(where, "a.id IN (SELECT alert_id FROM incident_alerts WHERE incident_id = ?)")
|
||||
args = append(args, n)
|
||||
}
|
||||
}
|
||||
|
||||
if from := q.Get("from"); from != "" {
|
||||
if t, err := time.Parse("2006-01-02", from); err == nil {
|
||||
@@ -114,55 +126,7 @@ func handleGetAlert(db *sql.DB) http.HandlerFunc {
|
||||
}
|
||||
}
|
||||
|
||||
func handleAcknowledge(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
|
||||
return
|
||||
}
|
||||
user, _ := userFromContext(r.Context())
|
||||
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"UPDATE alerts SET acknowledged_by = ?, acknowledged_at = ? WHERE id = ?",
|
||||
user.ID, time.Now().Unix(), id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("alert not found"))
|
||||
return
|
||||
}
|
||||
|
||||
a, _ := fetchAlert(r.Context(), db, id)
|
||||
respond(w, http.StatusOK, a)
|
||||
}
|
||||
}
|
||||
|
||||
func handleUnacknowledge(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
|
||||
return
|
||||
}
|
||||
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"UPDATE alerts SET acknowledged_by = NULL, acknowledged_at = NULL WHERE id = ?", id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("alert not found"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
// fetchAlert loads a single alert by ID using the shared JOIN query.
|
||||
// fetchAlert loads a single alert by ID using the shared query.
|
||||
func fetchAlert(ctx context.Context, db *sql.DB, id int64) (models.Alert, error) {
|
||||
return scanAlert(db.QueryRowContext(ctx, alertSelectFrom+" WHERE a.id = ?", id))
|
||||
}
|
||||
@@ -176,81 +140,24 @@ func scanAlert(s scanner) (models.Alert, error) {
|
||||
var a models.Alert
|
||||
var labelsJSON, annotationsJSON string
|
||||
var startsAtUnix, receivedAtUnix int64
|
||||
var endsAtUnix, ackAtUnix, archivedAtUnix *int64
|
||||
var ackByID *int64
|
||||
var ackByUser *string
|
||||
var endsAtUnix, archivedAtUnix *int64
|
||||
|
||||
if err := s.Scan(
|
||||
&a.ID, &a.Fingerprint, &a.Name, &a.Status,
|
||||
&labelsJSON, &annotationsJSON,
|
||||
&startsAtUnix, &endsAtUnix,
|
||||
&a.GeneratorURL, &receivedAtUnix,
|
||||
&ackByID, &ackAtUnix, &ackByUser,
|
||||
&a.IncidentID,
|
||||
&a.ResolutionSource, &archivedAtUnix,
|
||||
); err != nil {
|
||||
return a, err
|
||||
}
|
||||
|
||||
json.Unmarshal([]byte(labelsJSON), &a.Labels) //nolint:errcheck
|
||||
json.Unmarshal([]byte(labelsJSON), &a.Labels) //nolint:errcheck
|
||||
json.Unmarshal([]byte(annotationsJSON), &a.Annotations) //nolint:errcheck
|
||||
a.StartsAt = time.Unix(startsAtUnix, 0).UTC()
|
||||
a.ReceivedAt = time.Unix(receivedAtUnix, 0).UTC()
|
||||
if endsAtUnix != nil {
|
||||
t := time.Unix(*endsAtUnix, 0).UTC()
|
||||
a.EndsAt = &t
|
||||
}
|
||||
if ackByID != nil {
|
||||
t := time.Unix(*ackAtUnix, 0).UTC()
|
||||
a.AcknowledgedByID = ackByID
|
||||
a.AcknowledgedByUser = ackByUser
|
||||
a.AcknowledgedAt = &t
|
||||
}
|
||||
if archivedAtUnix != nil {
|
||||
t := time.Unix(*archivedAtUnix, 0).UTC()
|
||||
a.ArchivedAt = &t
|
||||
}
|
||||
a.EndsAt = unixPtr(endsAtUnix)
|
||||
a.ArchivedAt = unixPtr(archivedAtUnix)
|
||||
return a, nil
|
||||
}
|
||||
|
||||
func handleArchive(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
|
||||
return
|
||||
}
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"UPDATE alerts SET archived_at = unixepoch() WHERE id = ?", id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("alert not found"))
|
||||
return
|
||||
}
|
||||
a, _ := fetchAlert(r.Context(), db, id)
|
||||
respond(w, http.StatusOK, a)
|
||||
}
|
||||
}
|
||||
|
||||
func handleUnarchive(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
|
||||
return
|
||||
}
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"UPDATE alerts SET archived_at = NULL WHERE id = ?", id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("alert not found"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
+24
-108
@@ -171,9 +171,22 @@ func TestBootstrap_SecondCallForbidden(t *testing.T) {
|
||||
// Alert upsert by fingerprint
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func postWebhook(t *testing.T, s *ts, alerts []map[string]any) {
|
||||
// postWebhook sends an Alertmanager v4 payload. groupKey is optional: omitting
|
||||
// it exercises the fallback for senders that do not group, which is what most of
|
||||
// these tests want.
|
||||
func postWebhook(t *testing.T, s *ts, alerts []map[string]any, groupKey ...string) {
|
||||
t.Helper()
|
||||
payload := map[string]any{"version": "4", "status": "firing", "alerts": alerts}
|
||||
if len(groupKey) > 0 {
|
||||
payload["groupKey"] = groupKey[0]
|
||||
// Alertmanager groups by alertname by default, so the group labels echo
|
||||
// the first alert's name.
|
||||
if len(alerts) > 0 {
|
||||
if labels, ok := alerts[0]["labels"].(map[string]string); ok {
|
||||
payload["groupLabels"] = map[string]string{"alertname": labels["alertname"]}
|
||||
}
|
||||
}
|
||||
}
|
||||
data, _ := json.Marshal(payload)
|
||||
resp, err := http.Post(s.URL+"/api/alertmanager/webhook", "application/json", bytes.NewReader(data))
|
||||
if err != nil {
|
||||
@@ -243,84 +256,6 @@ func TestAlertUpsert_DifferentFingerprintsStored(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Alert acknowledge
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestAcknowledge(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{{
|
||||
"status": "firing", "labels": map[string]string{"alertname": "X"},
|
||||
"annotations": map[string]string{}, "startsAt": "2026-05-20T10:00:00Z",
|
||||
"endsAt": "0001-01-01T00:00:00Z", "generatorURL": "", "fingerprint": "fp-ack",
|
||||
}})
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/alerts/1/acknowledge", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("acknowledge returned %d", resp.StatusCode)
|
||||
}
|
||||
var alert map[string]any
|
||||
decode(t, resp, &alert)
|
||||
if alert["acknowledged_by"] == nil {
|
||||
t.Error("expected acknowledged_by to be set")
|
||||
}
|
||||
|
||||
// Clear it.
|
||||
resp = s.req(t, http.MethodDelete, "/api/alerts/1/acknowledge", nil)
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Errorf("unacknowledge returned %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodGet, "/api/alerts/1", nil)
|
||||
var alert2 map[string]any
|
||||
decode(t, resp, &alert2)
|
||||
if alert2["acknowledged_by"] != nil {
|
||||
t.Error("expected acknowledged_by to be cleared")
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Comments — own-only deletion
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestComment_DeleteOwnOnly(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{{
|
||||
"status": "firing", "labels": map[string]string{"alertname": "Y"},
|
||||
"annotations": map[string]string{}, "startsAt": "2026-05-20T10:00:00Z",
|
||||
"endsAt": "0001-01-01T00:00:00Z", "generatorURL": "", "fingerprint": "fp-comment",
|
||||
}})
|
||||
|
||||
// Create a second user and their own key.
|
||||
s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]string{"username": "alice", "email": "alice@test.com"})
|
||||
keyResp := s.req(t, http.MethodPost, "/api/users/2/api-keys",
|
||||
map[string]string{"name": "alice-key"})
|
||||
var keyData map[string]any
|
||||
decode(t, keyResp, &keyData)
|
||||
aliceKey := keyData["key"].(string)
|
||||
|
||||
// Admin posts a comment.
|
||||
s.req(t, http.MethodPost, "/api/alerts/1/comments",
|
||||
map[string]string{"content": "admin note"})
|
||||
|
||||
// Alice tries to delete admin's comment (should 404).
|
||||
req, _ := http.NewRequest(http.MethodDelete, s.URL+"/api/alerts/1/comments/1", nil)
|
||||
req.Header.Set("Authorization", "Bearer "+aliceKey)
|
||||
resp, _ := http.DefaultClient.Do(req)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNotFound {
|
||||
t.Errorf("expected 404 when deleting another user's comment, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
// Admin deletes own comment (should 204).
|
||||
resp = s.req(t, http.MethodDelete, "/api/alerts/1/comments/1", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Errorf("expected 204 when deleting own comment, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Schedule conflict
|
||||
// ---------------------------------------------------------------------------
|
||||
@@ -409,35 +344,29 @@ func TestStats_ByHourReturnsTwentyFourSlots(t *testing.T) {
|
||||
// Archive
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestArchive_RoundTrip(t *testing.T) {
|
||||
// Alert archiving is sweeper-only housekeeping now — nobody archives an alert by
|
||||
// hand — but the list filter it drives is still part of the API.
|
||||
func TestArchive_AlertListFilter(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
postWebhook(t, s, []map[string]any{{
|
||||
"status": "resolved", "fingerprint": "arch1",
|
||||
"labels": map[string]string{"alertname": "Archivable"},
|
||||
"annotations": map[string]string{},
|
||||
"startsAt": "2026-05-20T10:00:00Z", "endsAt": "2026-05-20T11:00:00Z",
|
||||
"labels": map[string]string{"alertname": "Archivable"},
|
||||
"annotations": map[string]string{},
|
||||
"startsAt": "2026-05-20T10:00:00Z",
|
||||
"endsAt": "2026-05-20T11:00:00Z",
|
||||
"generatorURL": "",
|
||||
}})
|
||||
|
||||
// 1. Alert appears in default list (not archived).
|
||||
// 1. Alert appears in the default list.
|
||||
var alerts []map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
|
||||
if len(alerts) != 1 {
|
||||
t.Fatalf("expected 1 alert in default list, got %d", len(alerts))
|
||||
}
|
||||
id := int(alerts[0]["id"].(float64))
|
||||
|
||||
// 2. Archive it.
|
||||
resp := s.req(t, http.MethodPost, fmt.Sprintf("/api/alerts/%d/archive", id), nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("archive: expected 200, got %d", resp.StatusCode)
|
||||
}
|
||||
var archived map[string]any
|
||||
decode(t, resp, &archived)
|
||||
if archived["archived_at"] == nil {
|
||||
t.Error("expected archived_at to be set in response")
|
||||
}
|
||||
// 2. Let the sweeper archive it: ends_at is already well past archiveAfter.
|
||||
api.Sweep(context.Background(), s.db, time.Hour, 6*time.Hour)
|
||||
|
||||
// 3. Default list excludes it.
|
||||
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
|
||||
@@ -450,19 +379,6 @@ func TestArchive_RoundTrip(t *testing.T) {
|
||||
if len(alerts) != 1 {
|
||||
t.Fatalf("expected 1 archived alert, got %d", len(alerts))
|
||||
}
|
||||
|
||||
// 5. Un-archive.
|
||||
resp = s.req(t, http.MethodDelete, fmt.Sprintf("/api/alerts/%d/archive", id), nil)
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Fatalf("unarchive: expected 204, got %d", resp.StatusCode)
|
||||
}
|
||||
resp.Body.Close()
|
||||
|
||||
// 6. Back in default list.
|
||||
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
|
||||
if len(alerts) != 1 {
|
||||
t.Errorf("expected unarchived alert to reappear, got %d results", len(alerts))
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
+143
-10
@@ -4,6 +4,7 @@ import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"log"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
@@ -33,12 +34,16 @@ func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter tim
|
||||
}
|
||||
}
|
||||
|
||||
// Sweep runs a single pass: expire stale firing alerts, then archive resolved
|
||||
// ones. Expiry runs first so an alert can expire and be archived in one pass.
|
||||
// Sweep runs a single pass, in dependency order: expire stale firing alerts,
|
||||
// close the incidents that leaves with nothing firing, then archive whatever has
|
||||
// been settled long enough. Running them in one pass means an alert can go stale
|
||||
// and its incident can close and archive without waiting three ticks.
|
||||
// Exported so tests can drive a pass without waiting on the ticker.
|
||||
func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration) {
|
||||
expireStale(ctx, db, staleAfter)
|
||||
resolveSettledIncidents(ctx, db)
|
||||
archiveResolved(ctx, db, archiveAfter)
|
||||
archiveResolvedIncidents(ctx, db, archiveAfter)
|
||||
}
|
||||
|
||||
// expireStale resolves firing alerts that Alertmanager has stopped refreshing.
|
||||
@@ -53,26 +58,132 @@ func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Durati
|
||||
// - received_at is older than staleAfter. Alertmanager re-sends firing
|
||||
// notifications every repeat_interval, making received_at a liveness
|
||||
// heartbeat — provided staleAfter exceeds that interval.
|
||||
//
|
||||
// The matching rows are collected before the update rather than updated in bulk,
|
||||
// because each one owes its incident a timeline entry.
|
||||
func expireStale(ctx context.Context, db *sql.DB, staleAfter time.Duration) {
|
||||
now := time.Now()
|
||||
res, err := db.ExecContext(ctx, `
|
||||
|
||||
ids, err := staleAlertIDs(ctx, db, now, staleAfter)
|
||||
if err != nil {
|
||||
log.Printf("sweeper: find stale: %v", err)
|
||||
return
|
||||
}
|
||||
if len(ids) == 0 {
|
||||
return
|
||||
}
|
||||
|
||||
args := make([]any, 0, len(ids)+1)
|
||||
args = append(args, resolutionExpiry)
|
||||
for _, id := range ids {
|
||||
args = append(args, id)
|
||||
}
|
||||
if _, err := db.ExecContext(ctx, `
|
||||
UPDATE alerts
|
||||
SET status = 'resolved',
|
||||
resolution_source = ?,
|
||||
ends_at = COALESCE(ends_at, unixepoch())
|
||||
WHERE status = 'firing'
|
||||
AND archived_at IS NULL
|
||||
AND ((ends_at IS NOT NULL AND ends_at < ?) OR received_at < ?)`,
|
||||
resolutionExpiry, now.Add(-expiryGrace).Unix(), now.Add(-staleAfter).Unix())
|
||||
if err != nil {
|
||||
WHERE id IN (`+placeholders(len(ids))+`)`, args...); err != nil {
|
||||
log.Printf("sweeper: expire stale: %v", err)
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n > 0 {
|
||||
log.Printf("sweeper: expired %d stale firing alert(s)", n)
|
||||
log.Printf("sweeper: expired %d stale firing alert(s)", len(ids))
|
||||
|
||||
for _, id := range ids {
|
||||
incidentID, err := openIncidentForAlert(ctx, db, id)
|
||||
if err != nil {
|
||||
log.Printf("sweeper: incident for alert %d: %v", id, err)
|
||||
continue
|
||||
}
|
||||
if incidentID == 0 {
|
||||
continue
|
||||
}
|
||||
alertID := id
|
||||
if err := logEvent(ctx, db, incidentID, evAlertResolved, nil, &alertID, nil); err != nil {
|
||||
log.Printf("sweeper: log expiry event: %v", err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// staleAlertIDs reads the ids in one go and closes the cursor before the caller
|
||||
// writes: the pool is limited to a single connection, so an open read would
|
||||
// block the update behind it.
|
||||
func staleAlertIDs(ctx context.Context, db *sql.DB, now time.Time, staleAfter time.Duration) ([]int64, error) {
|
||||
rows, err := db.QueryContext(ctx, `
|
||||
SELECT id FROM alerts
|
||||
WHERE status = 'firing'
|
||||
AND archived_at IS NULL
|
||||
AND ((ends_at IS NOT NULL AND ends_at < ?) OR received_at < ?)`,
|
||||
now.Add(-expiryGrace).Unix(), now.Add(-staleAfter).Unix())
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
var ids []int64
|
||||
for rows.Next() {
|
||||
var id int64
|
||||
if err := rows.Scan(&id); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
ids = append(ids, id)
|
||||
}
|
||||
return ids, rows.Err()
|
||||
}
|
||||
|
||||
// resolveSettledIncidents closes incidents whose alerts have all stopped firing.
|
||||
// This is the cascade from alerts up to the work item, and it is what turns an
|
||||
// expiry into a closed incident rather than one that sits open forever.
|
||||
func resolveSettledIncidents(ctx context.Context, db *sql.DB) {
|
||||
ids, err := settledIncidentIDs(ctx, db)
|
||||
if err != nil {
|
||||
log.Printf("sweeper: find settled incidents: %v", err)
|
||||
return
|
||||
}
|
||||
|
||||
resolved := 0
|
||||
for _, id := range ids {
|
||||
ok, err := resolveIfSettled(ctx, db, id)
|
||||
if err != nil {
|
||||
log.Printf("sweeper: resolve incident %d: %v", id, err)
|
||||
continue
|
||||
}
|
||||
if ok {
|
||||
resolved++
|
||||
}
|
||||
}
|
||||
if resolved > 0 {
|
||||
log.Printf("sweeper: resolved %d settled incident(s)", resolved)
|
||||
}
|
||||
}
|
||||
|
||||
func settledIncidentIDs(ctx context.Context, db *sql.DB) ([]int64, error) {
|
||||
rows, err := db.QueryContext(ctx, `
|
||||
SELECT i.id
|
||||
FROM incidents i
|
||||
WHERE i.resolved_at IS NULL
|
||||
AND EXISTS (SELECT 1 FROM incident_alerts ia WHERE ia.incident_id = i.id)
|
||||
AND NOT EXISTS (SELECT 1
|
||||
FROM incident_alerts ia
|
||||
JOIN alerts a ON a.id = ia.alert_id
|
||||
WHERE ia.incident_id = i.id
|
||||
AND a.status = 'firing')`)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
var ids []int64
|
||||
for rows.Next() {
|
||||
var id int64
|
||||
if err := rows.Scan(&id); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
ids = append(ids, id)
|
||||
}
|
||||
return ids, rows.Err()
|
||||
}
|
||||
|
||||
// archiveResolved hides resolved alerts that have been settled for archiveAfter.
|
||||
func archiveResolved(ctx context.Context, db *sql.DB, archiveAfter time.Duration) {
|
||||
cutoff := time.Now().Add(-archiveAfter).Unix()
|
||||
@@ -89,3 +200,25 @@ func archiveResolved(ctx context.Context, db *sql.DB, archiveAfter time.Duration
|
||||
log.Printf("archiver: archived %d resolved alert(s)", n)
|
||||
}
|
||||
}
|
||||
|
||||
// archiveResolvedIncidents does the same for the work items, on the same clock.
|
||||
func archiveResolvedIncidents(ctx context.Context, db *sql.DB, archiveAfter time.Duration) {
|
||||
cutoff := time.Now().Add(-archiveAfter).Unix()
|
||||
res, err := db.ExecContext(ctx,
|
||||
`UPDATE incidents SET archived_at = unixepoch()
|
||||
WHERE resolved_at IS NOT NULL
|
||||
AND archived_at IS NULL
|
||||
AND resolved_at < ?`, cutoff)
|
||||
if err != nil {
|
||||
log.Printf("archiver: incidents: %v", err)
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n > 0 {
|
||||
log.Printf("archiver: archived %d resolved incident(s)", n)
|
||||
}
|
||||
}
|
||||
|
||||
// placeholders builds "?, ?, …" for an IN clause of n values.
|
||||
func placeholders(n int) string {
|
||||
return strings.TrimSuffix(strings.Repeat("?, ", n), ", ")
|
||||
}
|
||||
|
||||
@@ -1,131 +0,0 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"net/http"
|
||||
"strconv"
|
||||
"time"
|
||||
|
||||
"github.com/go-chi/chi/v5"
|
||||
"github.com/yeniklas/terdut-server/internal/models"
|
||||
)
|
||||
|
||||
func handleListComments(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
alertID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
|
||||
return
|
||||
}
|
||||
|
||||
// Verify the alert exists.
|
||||
var exists int
|
||||
if err := db.QueryRowContext(r.Context(), "SELECT 1 FROM alerts WHERE id = ?", alertID).Scan(&exists); err != nil {
|
||||
respond(w, http.StatusNotFound, errResp("alert not found"))
|
||||
return
|
||||
}
|
||||
|
||||
rows, err := db.QueryContext(r.Context(), `
|
||||
SELECT c.id, c.alert_id, c.user_id, u.username, c.content, c.created_at
|
||||
FROM alert_comments c
|
||||
JOIN users u ON u.id = c.user_id
|
||||
WHERE c.alert_id = ?
|
||||
ORDER BY c.created_at ASC`, alertID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
comments := []models.Comment{}
|
||||
for rows.Next() {
|
||||
var c models.Comment
|
||||
var ts int64
|
||||
if err := rows.Scan(&c.ID, &c.AlertID, &c.UserID, &c.Username, &c.Content, &ts); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
c.CreatedAt = time.Unix(ts, 0).UTC()
|
||||
comments = append(comments, c)
|
||||
}
|
||||
respond(w, http.StatusOK, comments)
|
||||
}
|
||||
}
|
||||
|
||||
func handleCreateComment(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
alertID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
|
||||
return
|
||||
}
|
||||
|
||||
var req struct {
|
||||
Content string `json:"content"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid request body"))
|
||||
return
|
||||
}
|
||||
if req.Content == "" {
|
||||
respond(w, http.StatusBadRequest, errResp("content is required"))
|
||||
return
|
||||
}
|
||||
|
||||
// Verify the alert exists.
|
||||
var exists int
|
||||
if err := db.QueryRowContext(r.Context(), "SELECT 1 FROM alerts WHERE id = ?", alertID).Scan(&exists); err != nil {
|
||||
respond(w, http.StatusNotFound, errResp("alert not found"))
|
||||
return
|
||||
}
|
||||
|
||||
user, _ := userFromContext(r.Context())
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"INSERT INTO alert_comments (alert_id, user_id, content) VALUES (?, ?, ?)",
|
||||
alertID, user.ID, req.Content)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
commentID, _ := res.LastInsertId()
|
||||
|
||||
comment := models.Comment{
|
||||
ID: commentID,
|
||||
AlertID: alertID,
|
||||
UserID: user.ID,
|
||||
Username: user.Username,
|
||||
Content: req.Content,
|
||||
CreatedAt: time.Now().UTC(),
|
||||
}
|
||||
respond(w, http.StatusCreated, comment)
|
||||
}
|
||||
}
|
||||
|
||||
func handleDeleteComment(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
alertID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid alert id"))
|
||||
return
|
||||
}
|
||||
commentID, err := strconv.ParseInt(chi.URLParam(r, "commentID"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid comment id"))
|
||||
return
|
||||
}
|
||||
|
||||
user, _ := userFromContext(r.Context())
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"DELETE FROM alert_comments WHERE id = ? AND alert_id = ? AND user_id = ?",
|
||||
commentID, alertID, user.ID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("comment not found"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,269 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"encoding/json"
|
||||
"sort"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/yeniklas/terdut-server/internal/models"
|
||||
)
|
||||
|
||||
// Values for incidents.resolution_source, recording who closed the incident:
|
||||
// every member alert stopped firing, or a person decided it was done.
|
||||
const (
|
||||
incidentResolutionAlerts = "alerts"
|
||||
incidentResolutionManual = "manual"
|
||||
)
|
||||
|
||||
// Incident timeline event types. Stored as free text so adding one later is not
|
||||
// a migration, but these are the ones the server writes.
|
||||
const (
|
||||
evTriggered = "triggered"
|
||||
evAlertAdded = "alert_added"
|
||||
evAlertResolved = "alert_resolved"
|
||||
evAcknowledged = "acknowledged"
|
||||
evUnacknowledged = "unacknowledged"
|
||||
evAssigned = "assigned"
|
||||
evSnoozed = "snoozed"
|
||||
evUnsnoozed = "unsnoozed"
|
||||
evResolved = "resolved"
|
||||
evNote = "note"
|
||||
)
|
||||
|
||||
// severityLabel is the Alertmanager label an incident's severity is derived from.
|
||||
const severityLabel = "severity"
|
||||
|
||||
// querier is satisfied by both *sql.DB and *sql.Tx, so the helpers below work
|
||||
// inside the webhook's transaction and standalone from handlers and the sweeper.
|
||||
type querier interface {
|
||||
ExecContext(ctx context.Context, query string, args ...any) (sql.Result, error)
|
||||
QueryContext(ctx context.Context, query string, args ...any) (*sql.Rows, error)
|
||||
QueryRowContext(ctx context.Context, query string, args ...any) *sql.Row
|
||||
}
|
||||
|
||||
const incidentSelectFrom = `
|
||||
SELECT i.id, i.group_key, i.title, i.group_labels, i.status, i.severity,
|
||||
i.triggered_at,
|
||||
i.acknowledged_by, i.acknowledged_at, ack.username,
|
||||
i.assigned_to, asg.username, i.snoozed_until,
|
||||
i.resolved_at, i.resolution_source, i.archived_at
|
||||
FROM incidents i
|
||||
LEFT JOIN users ack ON ack.id = i.acknowledged_by
|
||||
LEFT JOIN users asg ON asg.id = i.assigned_to`
|
||||
|
||||
func scanIncident(s scanner) (models.Incident, error) {
|
||||
var i models.Incident
|
||||
var groupLabelsJSON string
|
||||
var triggeredAt int64
|
||||
var ackAt, snoozedUntil, resolvedAt, archivedAt *int64
|
||||
|
||||
if err := s.Scan(
|
||||
&i.ID, &i.GroupKey, &i.Title, &groupLabelsJSON, &i.Status, &i.Severity,
|
||||
&triggeredAt,
|
||||
&i.AcknowledgedByID, &ackAt, &i.AcknowledgedByUser,
|
||||
&i.AssignedToID, &i.AssignedToUser, &snoozedUntil,
|
||||
&resolvedAt, &i.ResolutionSource, &archivedAt,
|
||||
); err != nil {
|
||||
return i, err
|
||||
}
|
||||
|
||||
json.Unmarshal([]byte(groupLabelsJSON), &i.GroupLabels) //nolint:errcheck
|
||||
i.TriggeredAt = time.Unix(triggeredAt, 0).UTC()
|
||||
i.AcknowledgedAt = unixPtr(ackAt)
|
||||
i.SnoozedUntil = unixPtr(snoozedUntil)
|
||||
i.ResolvedAt = unixPtr(resolvedAt)
|
||||
i.ArchivedAt = unixPtr(archivedAt)
|
||||
return i, nil
|
||||
}
|
||||
|
||||
// unixPtr converts a nullable Unix-second column to a nullable UTC time.
|
||||
func unixPtr(sec *int64) *time.Time {
|
||||
if sec == nil {
|
||||
return nil
|
||||
}
|
||||
t := time.Unix(*sec, 0).UTC()
|
||||
return &t
|
||||
}
|
||||
|
||||
func fetchIncident(ctx context.Context, q querier, id int64) (models.Incident, error) {
|
||||
return scanIncident(q.QueryRowContext(ctx, incidentSelectFrom+" WHERE i.id = ?", id))
|
||||
}
|
||||
|
||||
// logEvent appends one entry to an incident's timeline. A nil userID means the
|
||||
// server acted rather than a person.
|
||||
func logEvent(ctx context.Context, q querier, incidentID int64, evType string, userID, alertID *int64, detail *string) error {
|
||||
_, err := q.ExecContext(ctx, `
|
||||
INSERT INTO incident_events (incident_id, type, user_id, alert_id, detail, created_at)
|
||||
VALUES (?, ?, ?, ?, ?, ?)`,
|
||||
incidentID, evType, userID, alertID, detail, time.Now().Unix())
|
||||
return err
|
||||
}
|
||||
|
||||
// todayUTC is the schedule's day key. The schedule's smallest unit is one UTC day.
|
||||
func todayUTC() string {
|
||||
return time.Now().UTC().Format("2006-01-02")
|
||||
}
|
||||
|
||||
// currentOnCall returns today's on-call user, or nil when nobody is scheduled.
|
||||
// A missing schedule entry is not an error — incidents just open unassigned.
|
||||
func currentOnCall(ctx context.Context, q querier) (*int64, error) {
|
||||
var userID int64
|
||||
err := q.QueryRowContext(ctx,
|
||||
"SELECT user_id FROM schedule_entries WHERE date = ?", todayUTC()).Scan(&userID)
|
||||
if err == sql.ErrNoRows {
|
||||
return nil, nil
|
||||
}
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return &userID, nil
|
||||
}
|
||||
|
||||
// severityRank orders the conventional Alertmanager severity label values.
|
||||
// Anything unrecognised sorts below all of them rather than being dropped.
|
||||
func severityRank(s string) int {
|
||||
switch strings.ToLower(s) {
|
||||
case "critical":
|
||||
return 4
|
||||
case "error":
|
||||
return 3
|
||||
case "warning":
|
||||
return 2
|
||||
case "info":
|
||||
return 1
|
||||
default:
|
||||
return 0
|
||||
}
|
||||
}
|
||||
|
||||
// refreshSeverity raises an incident's severity to the highest `severity` label
|
||||
// seen across its alerts.
|
||||
//
|
||||
// It is a high-water mark, never lowered: an incident that hit critical was a
|
||||
// critical incident, even after the critical alert clears and a warning is all
|
||||
// that is left firing. Downgrading a live incident would also quietly demote it
|
||||
// in the queue while the work is still open.
|
||||
func refreshSeverity(ctx context.Context, q querier, incidentID int64) error {
|
||||
rows, err := q.QueryContext(ctx, `
|
||||
SELECT json_extract(a.labels, '$.'||?)
|
||||
FROM incident_alerts ia
|
||||
JOIN alerts a ON a.id = ia.alert_id
|
||||
WHERE ia.incident_id = ?`, severityLabel, incidentID)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
best := ""
|
||||
for rows.Next() {
|
||||
var sev *string
|
||||
if err := rows.Scan(&sev); err != nil {
|
||||
rows.Close()
|
||||
return err
|
||||
}
|
||||
if sev != nil && severityRank(*sev) > severityRank(best) {
|
||||
best = *sev
|
||||
}
|
||||
}
|
||||
if err := rows.Err(); err != nil {
|
||||
rows.Close()
|
||||
return err
|
||||
}
|
||||
rows.Close()
|
||||
|
||||
if best == "" {
|
||||
return nil
|
||||
}
|
||||
// The comparison lives in SQL so an unrelated concurrent update cannot be
|
||||
// clobbered by a stale read.
|
||||
_, err = q.ExecContext(ctx, `
|
||||
UPDATE incidents SET severity = ?
|
||||
WHERE id = ?
|
||||
AND (severity IS NULL OR `+severityRankSQL("severity")+` < ?)`,
|
||||
best, incidentID, severityRank(best))
|
||||
return err
|
||||
}
|
||||
|
||||
// severityRankSQL mirrors severityRank for use inside a statement. SQL cannot
|
||||
// order these strings meaningfully on its own.
|
||||
func severityRankSQL(col string) string {
|
||||
return `CASE lower(COALESCE(` + col + `, ''))
|
||||
WHEN 'critical' THEN 4
|
||||
WHEN 'error' THEN 3
|
||||
WHEN 'warning' THEN 2
|
||||
WHEN 'info' THEN 1
|
||||
ELSE 0 END`
|
||||
}
|
||||
|
||||
// resolveIfSettled closes an incident once every alert under it has stopped
|
||||
// firing — PagerDuty's cascade, and the only automatic route out of the open
|
||||
// state. Reports whether it actually resolved anything.
|
||||
func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, error) {
|
||||
res, err := q.ExecContext(ctx, `
|
||||
UPDATE incidents
|
||||
SET status = 'resolved',
|
||||
resolved_at = ?,
|
||||
resolution_source = ?
|
||||
WHERE id = ?
|
||||
AND resolved_at IS NULL
|
||||
-- An incident with no members yet is mid-creation, not settled.
|
||||
AND EXISTS (SELECT 1 FROM incident_alerts ia WHERE ia.incident_id = incidents.id)
|
||||
AND NOT EXISTS (SELECT 1
|
||||
FROM incident_alerts ia
|
||||
JOIN alerts a ON a.id = ia.alert_id
|
||||
WHERE ia.incident_id = incidents.id
|
||||
AND a.status = 'firing')`,
|
||||
time.Now().Unix(), incidentResolutionAlerts, incidentID)
|
||||
if err != nil {
|
||||
return false, err
|
||||
}
|
||||
n, _ := res.RowsAffected()
|
||||
if n == 0 {
|
||||
return false, nil
|
||||
}
|
||||
return true, logEvent(ctx, q, incidentID, evResolved, nil, nil, nil)
|
||||
}
|
||||
|
||||
// openIncidentForAlert returns the open incident an alert currently belongs to,
|
||||
// or 0 when it has none. Used when an alert resolves or expires so the event
|
||||
// lands on the right timeline.
|
||||
func openIncidentForAlert(ctx context.Context, q querier, alertID int64) (int64, error) {
|
||||
var id int64
|
||||
err := q.QueryRowContext(ctx, `
|
||||
SELECT i.id
|
||||
FROM incident_alerts ia
|
||||
JOIN incidents i ON i.id = ia.incident_id
|
||||
WHERE ia.alert_id = ? AND i.resolved_at IS NULL`, alertID).Scan(&id)
|
||||
if err == sql.ErrNoRows {
|
||||
return 0, nil
|
||||
}
|
||||
return id, err
|
||||
}
|
||||
|
||||
// incidentTitle renders a human-readable title from Alertmanager's groupLabels,
|
||||
// leading with the alert name and appending whatever else the operator grouped
|
||||
// by. Falls back to the alert's own name when the payload carried no groupLabels.
|
||||
func incidentTitle(groupLabels map[string]string, fallback string) string {
|
||||
name := groupLabels["alertname"]
|
||||
if name == "" {
|
||||
name = fallback
|
||||
}
|
||||
if name == "" {
|
||||
name = "Incident"
|
||||
}
|
||||
|
||||
rest := make([]string, 0, len(groupLabels))
|
||||
for k, v := range groupLabels {
|
||||
if k == "alertname" {
|
||||
continue
|
||||
}
|
||||
rest = append(rest, k+"="+v)
|
||||
}
|
||||
if len(rest) == 0 {
|
||||
return name
|
||||
}
|
||||
sort.Strings(rest)
|
||||
return name + " (" + strings.Join(rest, ", ") + ")"
|
||||
}
|
||||
@@ -0,0 +1,554 @@
|
||||
package api
|
||||
|
||||
import (
|
||||
"database/sql"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/go-chi/chi/v5"
|
||||
"github.com/yeniklas/terdut-server/internal/models"
|
||||
)
|
||||
|
||||
func handleListIncidents(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
q := r.URL.Query()
|
||||
|
||||
where := []string{}
|
||||
args := []any{}
|
||||
|
||||
// Without an explicit status the queue shows open work, which is what an
|
||||
// on-call person opens the tool to see.
|
||||
if status := q.Get("status"); status != "" {
|
||||
where = append(where, "i.status = ?")
|
||||
args = append(args, status)
|
||||
} else {
|
||||
where = append(where, "i.resolved_at IS NULL")
|
||||
}
|
||||
|
||||
if q.Get("archived") == "true" {
|
||||
where = append(where, "i.archived_at IS NOT NULL")
|
||||
} else {
|
||||
where = append(where, "i.archived_at IS NULL")
|
||||
}
|
||||
|
||||
// A snooze expires by simply falling into the past; nothing sweeps it.
|
||||
if q.Get("snoozed") == "true" {
|
||||
where = append(where, "i.snoozed_until > ?")
|
||||
args = append(args, time.Now().Unix())
|
||||
} else {
|
||||
where = append(where, "(i.snoozed_until IS NULL OR i.snoozed_until <= ?)")
|
||||
args = append(args, time.Now().Unix())
|
||||
}
|
||||
|
||||
if severity := q.Get("severity"); severity != "" {
|
||||
where = append(where, "i.severity = ?")
|
||||
args = append(args, severity)
|
||||
}
|
||||
if assignee := q.Get("assigned_to"); assignee != "" {
|
||||
if n, err := strconv.ParseInt(assignee, 10, 64); err == nil {
|
||||
where = append(where, "i.assigned_to = ?")
|
||||
args = append(args, n)
|
||||
}
|
||||
}
|
||||
if from := q.Get("from"); from != "" {
|
||||
if t, err := time.Parse("2006-01-02", from); err == nil {
|
||||
where = append(where, "i.triggered_at >= ?")
|
||||
args = append(args, t.UTC().Unix())
|
||||
}
|
||||
}
|
||||
if to := q.Get("to"); to != "" {
|
||||
if t, err := time.Parse("2006-01-02", to); err == nil {
|
||||
where = append(where, "i.triggered_at < ?")
|
||||
args = append(args, t.UTC().AddDate(0, 0, 1).Unix())
|
||||
}
|
||||
}
|
||||
|
||||
limit := 50
|
||||
if l := q.Get("limit"); l != "" {
|
||||
if n, err := strconv.Atoi(l); err == nil && n > 0 && n <= 500 {
|
||||
limit = n
|
||||
}
|
||||
}
|
||||
|
||||
order := "i.triggered_at DESC"
|
||||
if q.Get("sort") == "severity" {
|
||||
order = severityRankSQL("i.severity") + " DESC, i.triggered_at DESC"
|
||||
}
|
||||
args = append(args, limit)
|
||||
|
||||
rows, err := db.QueryContext(r.Context(),
|
||||
fmt.Sprintf("%s WHERE %s ORDER BY %s LIMIT ?",
|
||||
incidentSelectFrom, strings.Join(where, " AND "), order),
|
||||
args...)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
incidents := []models.Incident{}
|
||||
for rows.Next() {
|
||||
i, err := scanIncident(rows)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
incidents = append(incidents, i)
|
||||
}
|
||||
respond(w, http.StatusOK, incidents)
|
||||
}
|
||||
}
|
||||
|
||||
func handleGetIncident(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
inc, err := fetchIncident(r.Context(), db, id)
|
||||
if err == sql.ErrNoRows {
|
||||
respond(w, http.StatusNotFound, errResp("incident not found"))
|
||||
return
|
||||
}
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if inc.Alerts, err = incidentAlerts(r, db, id); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, inc)
|
||||
}
|
||||
}
|
||||
|
||||
func handleIncidentAlerts(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !incidentExists(w, r, db, id) {
|
||||
return
|
||||
}
|
||||
alerts, err := incidentAlerts(r, db, id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, alerts)
|
||||
}
|
||||
}
|
||||
|
||||
func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
if !incidentExists(w, r, db, id) {
|
||||
return
|
||||
}
|
||||
|
||||
rows, err := db.QueryContext(r.Context(), `
|
||||
SELECT e.id, e.incident_id, e.type, e.user_id, u.username,
|
||||
e.alert_id, e.detail, e.created_at
|
||||
FROM incident_events e
|
||||
LEFT JOIN users u ON u.id = e.user_id
|
||||
WHERE e.incident_id = ?
|
||||
ORDER BY e.created_at ASC, e.id ASC`, id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
events := []models.IncidentEvent{}
|
||||
for rows.Next() {
|
||||
var e models.IncidentEvent
|
||||
var ts int64
|
||||
if err := rows.Scan(&e.ID, &e.IncidentID, &e.Type, &e.UserID, &e.Username,
|
||||
&e.AlertID, &e.Detail, &ts); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
e.CreatedAt = time.Unix(ts, 0).UTC()
|
||||
events = append(events, e)
|
||||
}
|
||||
respond(w, http.StatusOK, events)
|
||||
}
|
||||
}
|
||||
|
||||
func handleIncidentAcknowledge(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
user, _ := userFromContext(r.Context())
|
||||
if !updateOpenIncident(w, r, db, id,
|
||||
`UPDATE incidents SET status = 'acknowledged', acknowledged_by = ?, acknowledged_at = ?
|
||||
WHERE id = ? AND resolved_at IS NULL`, user.ID, time.Now().Unix(), id) {
|
||||
return
|
||||
}
|
||||
if err := logEvent(r.Context(), db, id, evAcknowledged, &user.ID, nil, nil); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respondIncident(w, r, db, id)
|
||||
}
|
||||
}
|
||||
|
||||
func handleIncidentUnacknowledge(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
user, _ := userFromContext(r.Context())
|
||||
if !updateOpenIncident(w, r, db, id,
|
||||
`UPDATE incidents SET status = 'triggered', acknowledged_by = NULL, acknowledged_at = NULL
|
||||
WHERE id = ? AND resolved_at IS NULL`, id) {
|
||||
return
|
||||
}
|
||||
if err := logEvent(r.Context(), db, id, evUnacknowledged, &user.ID, nil, nil); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
// handleIncidentResolve closes an incident by hand. This is terminal: a later
|
||||
// occurrence opens a new incident rather than reopening this one, which is what
|
||||
// stops a resolved incident from reappearing on the next repeat_interval
|
||||
// re-send of an alert that never stopped firing. Use snooze for "not now".
|
||||
func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
user, _ := userFromContext(r.Context())
|
||||
if !updateOpenIncident(w, r, db, id,
|
||||
`UPDATE incidents SET status = 'resolved', resolved_at = ?, resolution_source = ?
|
||||
WHERE id = ? AND resolved_at IS NULL`,
|
||||
time.Now().Unix(), incidentResolutionManual, id) {
|
||||
return
|
||||
}
|
||||
if err := logEvent(r.Context(), db, id, evResolved, &user.ID, nil, nil); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respondIncident(w, r, db, id)
|
||||
}
|
||||
}
|
||||
|
||||
func handleIncidentAssign(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
var req struct {
|
||||
UserID int64 `json:"user_id"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid request body"))
|
||||
return
|
||||
}
|
||||
if req.UserID == 0 {
|
||||
respond(w, http.StatusBadRequest, errResp("user_id is required"))
|
||||
return
|
||||
}
|
||||
var exists int
|
||||
if err := db.QueryRowContext(r.Context(),
|
||||
"SELECT 1 FROM users WHERE id = ?", req.UserID).Scan(&exists); err != nil {
|
||||
respond(w, http.StatusNotFound, errResp("user not found"))
|
||||
return
|
||||
}
|
||||
|
||||
if !updateOpenIncident(w, r, db, id,
|
||||
"UPDATE incidents SET assigned_to = ? WHERE id = ? AND resolved_at IS NULL",
|
||||
req.UserID, id) {
|
||||
return
|
||||
}
|
||||
// On an "assigned" event user_id is the assignee, not the actor.
|
||||
if err := logEvent(r.Context(), db, id, evAssigned, &req.UserID, nil, nil); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respondIncident(w, r, db, id)
|
||||
}
|
||||
}
|
||||
|
||||
// handleIncidentSnooze hides an incident from the default queue without closing
|
||||
// it. Accepts either an absolute {"until": RFC3339} or a relative
|
||||
// {"duration": "2h"}.
|
||||
func handleIncidentSnooze(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
var req struct {
|
||||
Until string `json:"until"`
|
||||
Duration string `json:"duration"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid request body"))
|
||||
return
|
||||
}
|
||||
|
||||
var until time.Time
|
||||
switch {
|
||||
case req.Until != "":
|
||||
t, err := time.Parse(time.RFC3339, req.Until)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid until (expected RFC3339)"))
|
||||
return
|
||||
}
|
||||
until = t
|
||||
case req.Duration != "":
|
||||
d, err := time.ParseDuration(req.Duration)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid duration"))
|
||||
return
|
||||
}
|
||||
until = time.Now().Add(d)
|
||||
default:
|
||||
respond(w, http.StatusBadRequest, errResp("until or duration is required"))
|
||||
return
|
||||
}
|
||||
if !until.After(time.Now()) {
|
||||
respond(w, http.StatusBadRequest, errResp("snooze must end in the future"))
|
||||
return
|
||||
}
|
||||
|
||||
user, _ := userFromContext(r.Context())
|
||||
if !updateOpenIncident(w, r, db, id,
|
||||
"UPDATE incidents SET snoozed_until = ? WHERE id = ? AND resolved_at IS NULL",
|
||||
until.Unix(), id) {
|
||||
return
|
||||
}
|
||||
detail := until.UTC().Format(time.RFC3339)
|
||||
if err := logEvent(r.Context(), db, id, evSnoozed, &user.ID, nil, &detail); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respondIncident(w, r, db, id)
|
||||
}
|
||||
}
|
||||
|
||||
func handleIncidentUnsnooze(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
user, _ := userFromContext(r.Context())
|
||||
if !updateOpenIncident(w, r, db, id,
|
||||
"UPDATE incidents SET snoozed_until = NULL WHERE id = ? AND resolved_at IS NULL", id) {
|
||||
return
|
||||
}
|
||||
if err := logEvent(r.Context(), db, id, evUnsnoozed, &user.ID, nil, nil); err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
func handleIncidentArchive(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"UPDATE incidents SET archived_at = unixepoch() WHERE id = ?", id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("incident not found"))
|
||||
return
|
||||
}
|
||||
respondIncident(w, r, db, id)
|
||||
}
|
||||
}
|
||||
|
||||
func handleIncidentUnarchive(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
res, err := db.ExecContext(r.Context(),
|
||||
"UPDATE incidents SET archived_at = NULL WHERE id = ?", id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("incident not found"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
// handleCreateNote adds a note to the timeline. Notes are ordinary events, so a
|
||||
// single query renders the whole story of an incident in order.
|
||||
func handleCreateNote(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
var req struct {
|
||||
Content string `json:"content"`
|
||||
}
|
||||
if err := decodeJSON(r, &req); err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid request body"))
|
||||
return
|
||||
}
|
||||
if req.Content == "" {
|
||||
respond(w, http.StatusBadRequest, errResp("content is required"))
|
||||
return
|
||||
}
|
||||
if !incidentExists(w, r, db, id) {
|
||||
return
|
||||
}
|
||||
|
||||
user, _ := userFromContext(r.Context())
|
||||
now := time.Now()
|
||||
res, err := db.ExecContext(r.Context(), `
|
||||
INSERT INTO incident_events (incident_id, type, user_id, detail, created_at)
|
||||
VALUES (?, ?, ?, ?, ?)`, id, evNote, user.ID, req.Content, now.Unix())
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
eventID, _ := res.LastInsertId()
|
||||
|
||||
respond(w, http.StatusCreated, models.IncidentEvent{
|
||||
ID: eventID,
|
||||
IncidentID: id,
|
||||
Type: evNote,
|
||||
UserID: &user.ID,
|
||||
Username: &user.Username,
|
||||
Detail: &req.Content,
|
||||
CreatedAt: now.UTC().Truncate(time.Second),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// handleDeleteNote removes one of your own notes. Only notes are deletable — the
|
||||
// rest of the timeline is what actually happened, and is not editable.
|
||||
func handleDeleteNote(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
id, ok := incidentIDParam(w, r)
|
||||
if !ok {
|
||||
return
|
||||
}
|
||||
eventID, err := strconv.ParseInt(chi.URLParam(r, "eventID"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid note id"))
|
||||
return
|
||||
}
|
||||
|
||||
user, _ := userFromContext(r.Context())
|
||||
res, err := db.ExecContext(r.Context(), `
|
||||
DELETE FROM incident_events
|
||||
WHERE id = ? AND incident_id = ? AND type = ? AND user_id = ?`,
|
||||
eventID, id, evNote, user.ID)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n == 0 {
|
||||
respond(w, http.StatusNotFound, errResp("note not found"))
|
||||
return
|
||||
}
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Shared handler plumbing
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func incidentIDParam(w http.ResponseWriter, r *http.Request) (int64, bool) {
|
||||
id, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
|
||||
if err != nil {
|
||||
respond(w, http.StatusBadRequest, errResp("invalid incident id"))
|
||||
return 0, false
|
||||
}
|
||||
return id, true
|
||||
}
|
||||
|
||||
func incidentExists(w http.ResponseWriter, r *http.Request, db *sql.DB, id int64) bool {
|
||||
var exists int
|
||||
if err := db.QueryRowContext(r.Context(),
|
||||
"SELECT 1 FROM incidents WHERE id = ?", id).Scan(&exists); err != nil {
|
||||
respond(w, http.StatusNotFound, errResp("incident not found"))
|
||||
return false
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// updateOpenIncident runs a mutation that is only valid while an incident is
|
||||
// open. The query must be constrained to `resolved_at IS NULL`, so no rows means
|
||||
// either the incident does not exist or it is already closed — two different
|
||||
// answers the caller should not have to distinguish itself.
|
||||
func updateOpenIncident(w http.ResponseWriter, r *http.Request, db *sql.DB, id int64, query string, args ...any) bool {
|
||||
res, err := db.ExecContext(r.Context(), query, args...)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return false
|
||||
}
|
||||
if n, _ := res.RowsAffected(); n > 0 {
|
||||
return true
|
||||
}
|
||||
if !incidentExists(w, r, db, id) {
|
||||
return false
|
||||
}
|
||||
respond(w, http.StatusConflict, errResp("incident is resolved"))
|
||||
return false
|
||||
}
|
||||
|
||||
func respondIncident(w http.ResponseWriter, r *http.Request, db *sql.DB, id int64) {
|
||||
inc, err := fetchIncident(r.Context(), db, id)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
respond(w, http.StatusOK, inc)
|
||||
}
|
||||
|
||||
// incidentAlerts loads the alerts under an incident, newest signal first.
|
||||
func incidentAlerts(r *http.Request, db *sql.DB, id int64) ([]models.Alert, error) {
|
||||
rows, err := db.QueryContext(r.Context(), alertSelectFrom+`
|
||||
JOIN incident_alerts m ON m.alert_id = a.id
|
||||
WHERE m.incident_id = ?
|
||||
ORDER BY a.received_at DESC`, id)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer rows.Close()
|
||||
|
||||
alerts := []models.Alert{}
|
||||
for rows.Next() {
|
||||
a, err := scanAlert(rows)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
alerts = append(alerts, a)
|
||||
}
|
||||
return alerts, rows.Err()
|
||||
}
|
||||
@@ -0,0 +1,772 @@
|
||||
package api_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/yeniklas/terdut-server/internal/api"
|
||||
"github.com/yeniklas/terdut-server/internal/db"
|
||||
)
|
||||
|
||||
// amAlert builds one alert of a webhook payload.
|
||||
func amAlert(fingerprint, name, status, startsAt, endsAt string, labels map[string]string) map[string]any {
|
||||
l := map[string]string{"alertname": name}
|
||||
for k, v := range labels {
|
||||
l[k] = v
|
||||
}
|
||||
return map[string]any{
|
||||
"status": status,
|
||||
"labels": l,
|
||||
"annotations": map[string]string{},
|
||||
"startsAt": startsAt,
|
||||
"endsAt": endsAt,
|
||||
"generatorURL": "",
|
||||
"fingerprint": fingerprint,
|
||||
}
|
||||
}
|
||||
|
||||
func listIncidents(t *testing.T, s *ts, query string) []map[string]any {
|
||||
t.Helper()
|
||||
var out []map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, "/api/incidents"+query, nil), &out)
|
||||
return out
|
||||
}
|
||||
|
||||
func getIncident(t *testing.T, s *ts, id int) map[string]any {
|
||||
t.Helper()
|
||||
var out map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, fmt.Sprintf("/api/incidents/%d", id), nil), &out)
|
||||
return out
|
||||
}
|
||||
|
||||
func timeline(t *testing.T, s *ts, id int) []map[string]any {
|
||||
t.Helper()
|
||||
var out []map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, fmt.Sprintf("/api/incidents/%d/timeline", id), nil), &out)
|
||||
return out
|
||||
}
|
||||
|
||||
// eventTypes flattens a timeline to its event types, which is what the ordering
|
||||
// assertions actually care about.
|
||||
func eventTypes(events []map[string]any) []string {
|
||||
types := make([]string, len(events))
|
||||
for i, e := range events {
|
||||
types[i] = e["type"].(string)
|
||||
}
|
||||
return types
|
||||
}
|
||||
|
||||
// countIncidents counts rows directly, including resolved and archived ones that
|
||||
// no list view returns.
|
||||
func (s *ts) countIncidents(t *testing.T) int {
|
||||
t.Helper()
|
||||
var n int
|
||||
if err := s.db.QueryRow("SELECT COUNT(*) FROM incidents").Scan(&n); err != nil {
|
||||
t.Fatalf("count incidents: %v", err)
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Ingest: alerts becoming incidents
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestWebhook_FiringOpensIncident(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-1", "HighCPU", "firing", "2026-05-20T10:00:00Z", zeroTime,
|
||||
map[string]string{"severity": "critical"}),
|
||||
}, "{}:{alertname=\"HighCPU\"}")
|
||||
|
||||
incidents := listIncidents(t, s, "")
|
||||
if len(incidents) != 1 {
|
||||
t.Fatalf("expected 1 incident, got %d", len(incidents))
|
||||
}
|
||||
inc := incidents[0]
|
||||
if inc["status"] != "triggered" {
|
||||
t.Errorf("expected status triggered, got %v", inc["status"])
|
||||
}
|
||||
if inc["severity"] != "critical" {
|
||||
t.Errorf("expected severity critical, got %v", inc["severity"])
|
||||
}
|
||||
if inc["title"] != "HighCPU" {
|
||||
t.Errorf("expected title from groupLabels, got %v", inc["title"])
|
||||
}
|
||||
|
||||
// The alert points back at the incident it opened.
|
||||
var alerts []map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
|
||||
if len(alerts) != 1 || alerts[0]["incident_id"] == nil {
|
||||
t.Fatalf("expected the alert to carry an incident_id, got %v", alerts)
|
||||
}
|
||||
}
|
||||
|
||||
// Alertmanager already grouped these; we adopt its answer rather than
|
||||
// correlating again.
|
||||
func TestWebhook_SameGroupKeyJoinsOneIncident(t *testing.T) {
|
||||
s := newTS(t)
|
||||
const groupKey = "{}:{alertname=\"DiskFull\"}"
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-a", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
amAlert("fp-b", "DiskFull", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
}, groupKey)
|
||||
|
||||
incidents := listIncidents(t, s, "")
|
||||
if len(incidents) != 1 {
|
||||
t.Fatalf("expected 1 incident for one groupKey, got %d", len(incidents))
|
||||
}
|
||||
id := int(incidents[0]["id"].(float64))
|
||||
|
||||
inc := getIncident(t, s, id)
|
||||
members, _ := inc["alerts"].([]any)
|
||||
if len(members) != 2 {
|
||||
t.Fatalf("expected 2 alerts under the incident, got %d", len(members))
|
||||
}
|
||||
|
||||
added := 0
|
||||
for _, ty := range eventTypes(timeline(t, s, id)) {
|
||||
if ty == "alert_added" {
|
||||
added++
|
||||
}
|
||||
}
|
||||
if added != 2 {
|
||||
t.Errorf("expected 2 alert_added events, got %d", added)
|
||||
}
|
||||
}
|
||||
|
||||
// The load-bearing rule. Alertmanager re-sends firing notifications every
|
||||
// repeat_interval; if those re-sends reopened incidents, resolving one by hand
|
||||
// would mean nothing.
|
||||
func TestWebhook_HeartbeatDoesNotReopenResolvedIncident(t *testing.T) {
|
||||
s := newTS(t)
|
||||
const groupKey = "{}:{alertname=\"Flapper\"}"
|
||||
alert := amAlert("fp-hb", "Flapper", "firing", "2026-05-20T10:00:00Z", zeroTime, nil)
|
||||
|
||||
postWebhook(t, s, []map[string]any{alert}, groupKey)
|
||||
resp := s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("resolve returned %d", resp.StatusCode)
|
||||
}
|
||||
resp.Body.Close()
|
||||
|
||||
// Same startsAt, same fingerprint: a re-send, not a new occurrence.
|
||||
postWebhook(t, s, []map[string]any{alert}, groupKey)
|
||||
|
||||
if n := s.countIncidents(t); n != 1 {
|
||||
t.Fatalf("expected the heartbeat to open no incident, got %d total", n)
|
||||
}
|
||||
if inc := getIncident(t, s, 1); inc["resolved_at"] == nil {
|
||||
t.Error("expected incident 1 to stay resolved")
|
||||
}
|
||||
|
||||
// The alert itself is still firing and still being tracked — only the work
|
||||
// item is closed.
|
||||
status, _, _ := s.alertRow(t, "fp-hb")
|
||||
if status != "firing" {
|
||||
t.Errorf("expected the alert to still be firing, got %q", status)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWebhook_NewOccurrenceOpensNewIncident(t *testing.T) {
|
||||
s := newTS(t)
|
||||
const groupKey = "{}:{alertname=\"Recurring\"}"
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-new", "Recurring", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
}, groupKey)
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
|
||||
|
||||
// A newer startsAt is a genuinely new occurrence, not a re-send.
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-new", "Recurring", "firing", "2026-05-21T09:00:00Z", zeroTime, nil),
|
||||
}, groupKey)
|
||||
|
||||
if n := s.countIncidents(t); n != 2 {
|
||||
t.Fatalf("expected a second incident for the new occurrence, got %d total", n)
|
||||
}
|
||||
open := listIncidents(t, s, "")
|
||||
if len(open) != 1 || int(open[0]["id"].(float64)) != 2 {
|
||||
t.Fatalf("expected incident 2 to be the open one, got %v", open)
|
||||
}
|
||||
|
||||
// The new incident starts unacknowledged: that is the point of the split.
|
||||
if open[0]["acknowledged_by"] != nil {
|
||||
t.Error("expected a fresh occurrence to start unacknowledged")
|
||||
}
|
||||
}
|
||||
|
||||
func TestWebhook_ResolvedOnlyPayloadOpensNothing(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-res", "AlreadyOver", "resolved", "2026-05-20T10:00:00Z", "2026-05-20T11:00:00Z", nil),
|
||||
}, "{}:{alertname=\"AlreadyOver\"}")
|
||||
|
||||
if n := s.countIncidents(t); n != 0 {
|
||||
t.Errorf("expected no incident from a resolved-only payload, got %d", n)
|
||||
}
|
||||
if status, _, _ := s.alertRow(t, "fp-res"); status != "resolved" {
|
||||
t.Errorf("expected the alert itself to be stored, got %q", status)
|
||||
}
|
||||
}
|
||||
|
||||
// An incident that hit critical was a critical incident, even once the critical
|
||||
// alert clears and only a warning is left.
|
||||
func TestIncident_SeverityIsHighWaterMark(t *testing.T) {
|
||||
s := newTS(t)
|
||||
const groupKey = "{}:{alertname=\"Mixed\"}"
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-warn", "Mixed", "firing", "2026-05-20T10:00:00Z", zeroTime,
|
||||
map[string]string{"severity": "warning"}),
|
||||
amAlert("fp-crit", "Mixed", "firing", "2026-05-20T10:00:00Z", zeroTime,
|
||||
map[string]string{"severity": "critical"}),
|
||||
}, groupKey)
|
||||
|
||||
if inc := getIncident(t, s, 1); inc["severity"] != "critical" {
|
||||
t.Fatalf("expected severity critical, got %v", inc["severity"])
|
||||
}
|
||||
|
||||
// The critical alert clears; the warning keeps the incident open.
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-crit", "Mixed", "resolved", "2026-05-20T10:00:00Z", "2026-05-20T11:00:00Z",
|
||||
map[string]string{"severity": "critical"}),
|
||||
}, groupKey)
|
||||
|
||||
inc := getIncident(t, s, 1)
|
||||
if inc["resolved_at"] != nil {
|
||||
t.Fatal("expected the incident to stay open")
|
||||
}
|
||||
if inc["severity"] != "critical" {
|
||||
t.Errorf("expected severity to stay critical, got %v", inc["severity"])
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Resolution cascade
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestIncident_AllAlertsResolvedAutoResolves(t *testing.T) {
|
||||
s := newTS(t)
|
||||
const groupKey = "{}:{alertname=\"Pair\"}"
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-p1", "Pair", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
amAlert("fp-p2", "Pair", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
}, groupKey)
|
||||
|
||||
// One down, one still firing: the work is not done.
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-p1", "Pair", "resolved", "2026-05-20T10:00:00Z", "2026-05-20T11:00:00Z", nil),
|
||||
}, groupKey)
|
||||
if inc := getIncident(t, s, 1); inc["resolved_at"] != nil {
|
||||
t.Fatal("expected the incident to stay open while an alert is firing")
|
||||
}
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-p2", "Pair", "resolved", "2026-05-20T10:00:00Z", "2026-05-20T11:30:00Z", nil),
|
||||
}, groupKey)
|
||||
|
||||
inc := getIncident(t, s, 1)
|
||||
if inc["status"] != "resolved" {
|
||||
t.Errorf("expected status resolved, got %v", inc["status"])
|
||||
}
|
||||
if inc["resolution_source"] != "alerts" {
|
||||
t.Errorf("expected resolution_source alerts, got %v", inc["resolution_source"])
|
||||
}
|
||||
}
|
||||
|
||||
// Expiry is inference, not observation, but it still has to close the work item
|
||||
// — otherwise a lost resolved notification leaves an incident open forever.
|
||||
func TestExpiry_CascadesToIncidentResolution(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postAlert(t, s, "fp-exp", "firing", time.Now().Add(-24*time.Hour).Format(time.RFC3339), zeroTime)
|
||||
|
||||
s.exec(t, "UPDATE alerts SET received_at = ? WHERE fingerprint = 'fp-exp'",
|
||||
time.Now().Add(-10*time.Hour).Unix())
|
||||
sweep(t, s, 6*time.Hour)
|
||||
|
||||
inc := getIncident(t, s, 1)
|
||||
if inc["status"] != "resolved" {
|
||||
t.Errorf("expected the incident to resolve after expiry, got %v", inc["status"])
|
||||
}
|
||||
if inc["resolution_source"] != "alerts" {
|
||||
t.Errorf("expected resolution_source alerts, got %v", inc["resolution_source"])
|
||||
}
|
||||
|
||||
// The expiry is recorded against the alert, not the incident.
|
||||
if _, source, _ := s.alertRow(t, "fp-exp"); source == nil || *source != "expiry" {
|
||||
t.Errorf("expected the alert's resolution_source to stay expiry, got %v", source)
|
||||
}
|
||||
if types := eventTypes(timeline(t, s, 1)); !contains(types, "alert_resolved") {
|
||||
t.Errorf("expected an alert_resolved event on the timeline, got %v", types)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Workflow actions
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestIncident_Acknowledge(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-ack", "X", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("acknowledge returned %d", resp.StatusCode)
|
||||
}
|
||||
var inc map[string]any
|
||||
decode(t, resp, &inc)
|
||||
if inc["acknowledged_by"] == nil {
|
||||
t.Error("expected acknowledged_by to be set")
|
||||
}
|
||||
if inc["status"] != "acknowledged" {
|
||||
t.Errorf("expected status acknowledged, got %v", inc["status"])
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodDelete, "/api/incidents/1/acknowledge", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Fatalf("unacknowledge returned %d", resp.StatusCode)
|
||||
}
|
||||
if inc := getIncident(t, s, 1); inc["status"] != "triggered" {
|
||||
t.Errorf("expected status back to triggered, got %v", inc["status"])
|
||||
}
|
||||
}
|
||||
|
||||
func TestIncident_ManualResolveIsTerminal(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-term", "Terminal", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("first resolve returned %d", resp.StatusCode)
|
||||
}
|
||||
var inc map[string]any
|
||||
decode(t, resp, &inc)
|
||||
if inc["resolution_source"] != "manual" {
|
||||
t.Errorf("expected resolution_source manual, got %v", inc["resolution_source"])
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusConflict {
|
||||
t.Errorf("expected 409 on re-resolve, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
// Acknowledging a closed incident is equally meaningless.
|
||||
resp = s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusConflict {
|
||||
t.Errorf("expected 409 acknowledging a resolved incident, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIncident_SnoozeHiddenFromDefaultList(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-snz", "Noisy", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/incidents/1/snooze", map[string]string{"duration": "2h"})
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("snooze returned %d", resp.StatusCode)
|
||||
}
|
||||
var inc map[string]any
|
||||
decode(t, resp, &inc)
|
||||
if inc["snoozed_until"] == nil {
|
||||
t.Error("expected snoozed_until to be set")
|
||||
}
|
||||
|
||||
if got := listIncidents(t, s, ""); len(got) != 0 {
|
||||
t.Errorf("expected the snoozed incident to be hidden, got %d", len(got))
|
||||
}
|
||||
if got := listIncidents(t, s, "?snoozed=true"); len(got) != 1 {
|
||||
t.Errorf("expected snoozed=true to show it, got %d", len(got))
|
||||
}
|
||||
|
||||
// A snooze is not a resolution: the incident is still open work.
|
||||
if inc := getIncident(t, s, 1); inc["resolved_at"] != nil {
|
||||
t.Error("expected a snoozed incident to stay open")
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodDelete, "/api/incidents/1/snooze", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Fatalf("unsnooze returned %d", resp.StatusCode)
|
||||
}
|
||||
if got := listIncidents(t, s, ""); len(got) != 1 {
|
||||
t.Errorf("expected the incident back in the default list, got %d", len(got))
|
||||
}
|
||||
}
|
||||
|
||||
func TestIncident_SnoozeRejectsPastDeadline(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-past", "Past", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/incidents/1/snooze",
|
||||
map[string]string{"until": time.Now().Add(-time.Hour).UTC().Format(time.RFC3339)})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusBadRequest {
|
||||
t.Errorf("expected 400 for a snooze in the past, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// The schedule stops being decorative here: it is read at trigger time.
|
||||
func TestIncident_AutoAssignedToCurrentOnCall(t *testing.T) {
|
||||
s := newTS(t)
|
||||
|
||||
today := time.Now().UTC().Format("2006-01-02")
|
||||
resp := s.req(t, http.MethodPost, "/api/schedule",
|
||||
map[string]any{"user_id": 1, "dates": []string{today}})
|
||||
if resp.StatusCode != http.StatusCreated {
|
||||
t.Fatalf("schedule assignment returned %d", resp.StatusCode)
|
||||
}
|
||||
resp.Body.Close()
|
||||
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-oncall", "PageMe", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
inc := getIncident(t, s, 1)
|
||||
if inc["assigned_to"] != "admin" {
|
||||
t.Errorf("expected the incident assigned to today's on-call, got %v", inc["assigned_to"])
|
||||
}
|
||||
if types := eventTypes(timeline(t, s, 1)); !contains(types, "assigned") {
|
||||
t.Errorf("expected an assigned event, got %v", types)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIncident_AssignToUser(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-asg", "Assignable", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]string{"username": "alice", "email": "alice@test.com"}).Body.Close()
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/incidents/1/assign", map[string]any{"user_id": 2})
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("assign returned %d", resp.StatusCode)
|
||||
}
|
||||
var inc map[string]any
|
||||
decode(t, resp, &inc)
|
||||
if inc["assigned_to"] != "alice" {
|
||||
t.Errorf("expected assigned_to alice, got %v", inc["assigned_to"])
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodPost, "/api/incidents/1/assign", map[string]any{"user_id": 99})
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNotFound {
|
||||
t.Errorf("expected 404 assigning an unknown user, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Timeline and notes
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestIncident_TimelineOrdering(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-tl", "Storyline", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil).Body.Close()
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/notes",
|
||||
map[string]string{"content": "looking into it"}).Body.Close()
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
|
||||
|
||||
events := timeline(t, s, 1)
|
||||
want := []string{"triggered", "alert_added", "acknowledged", "note", "resolved"}
|
||||
got := eventTypes(events)
|
||||
if len(got) != len(want) {
|
||||
t.Fatalf("expected timeline %v, got %v", want, got)
|
||||
}
|
||||
for i := range want {
|
||||
if got[i] != want[i] {
|
||||
t.Fatalf("expected timeline %v, got %v", want, got)
|
||||
}
|
||||
}
|
||||
|
||||
// The note carries its author; system events do not.
|
||||
for _, e := range events {
|
||||
if e["type"] == "note" {
|
||||
if e["username"] != "admin" || e["detail"] != "looking into it" {
|
||||
t.Errorf("unexpected note event: %v", e)
|
||||
}
|
||||
}
|
||||
if e["type"] == "triggered" && e["username"] != nil {
|
||||
t.Errorf("expected the triggered event to have no author, got %v", e["username"])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestIncident_NoteDeleteOwnOnly(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-note", "Y", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
s.req(t, http.MethodPost, "/api/users",
|
||||
map[string]string{"username": "alice", "email": "alice@test.com"}).Body.Close()
|
||||
var keyData map[string]any
|
||||
decode(t, s.req(t, http.MethodPost, "/api/users/2/api-keys",
|
||||
map[string]string{"name": "alice-key"}), &keyData)
|
||||
aliceKey := keyData["key"].(string)
|
||||
|
||||
var note map[string]any
|
||||
decode(t, s.req(t, http.MethodPost, "/api/incidents/1/notes",
|
||||
map[string]string{"content": "admin note"}), ¬e)
|
||||
noteID := int(note["id"].(float64))
|
||||
|
||||
// Alice cannot delete admin's note.
|
||||
req, _ := http.NewRequest(http.MethodDelete,
|
||||
fmt.Sprintf("%s/api/incidents/1/notes/%d", s.URL, noteID), nil)
|
||||
req.Header.Set("Authorization", "Bearer "+aliceKey)
|
||||
resp, _ := http.DefaultClient.Do(req)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNotFound {
|
||||
t.Errorf("expected 404 deleting another user's note, got %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodDelete, fmt.Sprintf("/api/incidents/1/notes/%d", noteID), nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Errorf("expected 204 deleting own note, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// Only notes are deletable — the rest of the timeline is what happened.
|
||||
func TestIncident_CannotDeleteSystemEvent(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-sys", "System", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
events := timeline(t, s, 1)
|
||||
id := int(events[0]["id"].(float64))
|
||||
resp := s.req(t, http.MethodDelete, fmt.Sprintf("/api/incidents/1/notes/%d", id), nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNotFound {
|
||||
t.Errorf("expected 404 deleting a system event, got %d", resp.StatusCode)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Archive
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestIncident_ArchiveRoundTrip(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-arc", "Archivable", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
|
||||
|
||||
resp := s.req(t, http.MethodPost, "/api/incidents/1/archive", nil)
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
t.Fatalf("archive returned %d", resp.StatusCode)
|
||||
}
|
||||
var inc map[string]any
|
||||
decode(t, resp, &inc)
|
||||
if inc["archived_at"] == nil {
|
||||
t.Error("expected archived_at to be set")
|
||||
}
|
||||
|
||||
if got := listIncidents(t, s, "?status=resolved"); len(got) != 0 {
|
||||
t.Errorf("expected the archived incident to be hidden, got %d", len(got))
|
||||
}
|
||||
if got := listIncidents(t, s, "?status=resolved&archived=true"); len(got) != 1 {
|
||||
t.Errorf("expected archived=true to show it, got %d", len(got))
|
||||
}
|
||||
|
||||
resp = s.req(t, http.MethodDelete, "/api/incidents/1/archive", nil)
|
||||
resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusNoContent {
|
||||
t.Fatalf("unarchive returned %d", resp.StatusCode)
|
||||
}
|
||||
if got := listIncidents(t, s, "?status=resolved"); len(got) != 1 {
|
||||
t.Errorf("expected the incident back, got %d", len(got))
|
||||
}
|
||||
}
|
||||
|
||||
func TestSweeper_ArchivesResolvedIncidents(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-swp", "Old", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
|
||||
|
||||
s.exec(t, "UPDATE incidents SET resolved_at = ? WHERE id = 1",
|
||||
time.Now().Add(-30*24*time.Hour).Unix())
|
||||
api.Sweep(context.Background(), s.db, 7*24*time.Hour, 6*time.Hour)
|
||||
|
||||
if inc := getIncident(t, s, 1); inc["archived_at"] == nil {
|
||||
t.Error("expected the sweeper to archive a long-resolved incident")
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Stats
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestStats_Incidents(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-s1", "One", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
}, "g1")
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-s2", "Two", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
}, "g2")
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil).Body.Close()
|
||||
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
|
||||
|
||||
var stats map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, "/api/stats/incidents", nil), &stats)
|
||||
|
||||
if stats["total"].(float64) != 2 {
|
||||
t.Errorf("expected total 2, got %v", stats["total"])
|
||||
}
|
||||
if stats["resolved"].(float64) != 1 {
|
||||
t.Errorf("expected resolved 1, got %v", stats["resolved"])
|
||||
}
|
||||
if stats["triggered"].(float64) != 1 {
|
||||
t.Errorf("expected triggered 1, got %v", stats["triggered"])
|
||||
}
|
||||
// One incident has been acknowledged and resolved, so both averages exist.
|
||||
if stats["mtta_seconds"] == nil || stats["mttr_seconds"] == nil {
|
||||
t.Errorf("expected mtta and mttr to be computable, got %v", stats)
|
||||
}
|
||||
}
|
||||
|
||||
// Nothing acknowledged yet means "no data", which is not the same claim as zero.
|
||||
func TestStats_IncidentsNullMTTAWhenNothingAcknowledged(t *testing.T) {
|
||||
s := newTS(t)
|
||||
postWebhook(t, s, []map[string]any{
|
||||
amAlert("fp-s3", "Untouched", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
|
||||
})
|
||||
|
||||
var stats map[string]any
|
||||
decode(t, s.req(t, http.MethodGet, "/api/stats/incidents", nil), &stats)
|
||||
if stats["mtta_seconds"] != nil {
|
||||
t.Errorf("expected mtta_seconds null, got %v", stats["mtta_seconds"])
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Migration backfill
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
// An upgrade must not drop the acknowledgements and comments people already
|
||||
// have, so 008 is replayed here over a database left at 007.
|
||||
func TestMigration_BackfillCarriesAckAndComments(t *testing.T) {
|
||||
database, err := db.Open(":memory:")
|
||||
if err != nil {
|
||||
t.Fatalf("open db: %v", err)
|
||||
}
|
||||
t.Cleanup(func() { database.Close() })
|
||||
|
||||
files, err := filepath.Glob("../db/migrations/*.sql")
|
||||
if err != nil || len(files) == 0 {
|
||||
t.Fatalf("find migrations: %v", err)
|
||||
}
|
||||
sort.Strings(files)
|
||||
|
||||
var incidentsMigration string
|
||||
for _, f := range files {
|
||||
if filepath.Base(f) >= "008" {
|
||||
incidentsMigration = f
|
||||
break
|
||||
}
|
||||
data, err := os.ReadFile(f)
|
||||
if err != nil {
|
||||
t.Fatalf("read %s: %v", f, err)
|
||||
}
|
||||
if _, err := database.Exec(string(data)); err != nil {
|
||||
t.Fatalf("apply %s: %v", f, err)
|
||||
}
|
||||
}
|
||||
if incidentsMigration == "" {
|
||||
t.Fatal("008 migration not found")
|
||||
}
|
||||
|
||||
// A database as it would look on the old schema: an acknowledged firing
|
||||
// alert with a comment on it.
|
||||
if _, err := database.Exec(`
|
||||
INSERT INTO users (id, username, email) VALUES (1, 'admin', 'admin@test.com');
|
||||
INSERT INTO alerts (id, fingerprint, name, status, labels, annotations,
|
||||
starts_at, received_at, acknowledged_by, acknowledged_at)
|
||||
VALUES (1, 'legacy-fp', 'LegacyAlert', 'firing',
|
||||
'{"severity":"warning"}', '{}', 1000, 1000, 1, 1500);
|
||||
INSERT INTO alert_comments (alert_id, user_id, content, created_at)
|
||||
VALUES (1, 1, 'legacy comment', 1600);`); err != nil {
|
||||
t.Fatalf("seed pre-008 data: %v", err)
|
||||
}
|
||||
|
||||
data, err := os.ReadFile(incidentsMigration)
|
||||
if err != nil {
|
||||
t.Fatalf("read 008: %v", err)
|
||||
}
|
||||
if _, err := database.Exec(string(data)); err != nil {
|
||||
t.Fatalf("apply 008: %v", err)
|
||||
}
|
||||
|
||||
var status, groupKey string
|
||||
var ackBy int64
|
||||
var severity string
|
||||
if err := database.QueryRow(
|
||||
"SELECT status, group_key, acknowledged_by, severity FROM incidents WHERE id = 1",
|
||||
).Scan(&status, &groupKey, &ackBy, &severity); err != nil {
|
||||
t.Fatalf("read backfilled incident: %v", err)
|
||||
}
|
||||
if status != "acknowledged" {
|
||||
t.Errorf("expected the ack to carry over as status, got %q", status)
|
||||
}
|
||||
if groupKey != "backfill:legacy-fp" {
|
||||
t.Errorf("unexpected group_key %q", groupKey)
|
||||
}
|
||||
if ackBy != 1 {
|
||||
t.Errorf("expected acknowledged_by 1, got %d", ackBy)
|
||||
}
|
||||
if severity != "warning" {
|
||||
t.Errorf("expected severity carried from labels, got %q", severity)
|
||||
}
|
||||
|
||||
var notes int
|
||||
if err := database.QueryRow(
|
||||
"SELECT COUNT(*) FROM incident_events WHERE type = 'note' AND detail = 'legacy comment'",
|
||||
).Scan(¬es); err != nil {
|
||||
t.Fatalf("count notes: %v", err)
|
||||
}
|
||||
if notes != 1 {
|
||||
t.Errorf("expected the comment to become a note, got %d", notes)
|
||||
}
|
||||
|
||||
// And the columns that caused the ack-survives-a-re-fire bug are gone.
|
||||
if _, err := database.Exec("SELECT acknowledged_by FROM alerts"); err == nil {
|
||||
t.Error("expected alerts.acknowledged_by to be dropped")
|
||||
}
|
||||
}
|
||||
|
||||
func contains(haystack []string, needle string) bool {
|
||||
for _, s := range haystack {
|
||||
if s == needle {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
+18
-7
@@ -31,21 +31,32 @@ func NewRouter(db *sql.DB) http.Handler {
|
||||
r.Post("/api/users/{id}/api-keys", handleCreateAPIKey(db))
|
||||
r.Delete("/api/users/{id}/api-keys/{keyID}", handleDeleteAPIKey(db))
|
||||
|
||||
// Alerts are read-only: they are Alertmanager's record, not a work
|
||||
// queue. Everything a person does happens on the incident instead.
|
||||
r.Get("/api/alerts", handleListAlerts(db))
|
||||
r.Get("/api/alerts/{id}", handleGetAlert(db))
|
||||
r.Post("/api/alerts/{id}/acknowledge", handleAcknowledge(db))
|
||||
r.Delete("/api/alerts/{id}/acknowledge", handleUnacknowledge(db))
|
||||
r.Post("/api/alerts/{id}/archive", handleArchive(db))
|
||||
r.Delete("/api/alerts/{id}/archive", handleUnarchive(db))
|
||||
r.Get("/api/alerts/{id}/comments", handleListComments(db))
|
||||
r.Post("/api/alerts/{id}/comments", handleCreateComment(db))
|
||||
r.Delete("/api/alerts/{id}/comments/{commentID}", handleDeleteComment(db))
|
||||
|
||||
r.Get("/api/incidents", handleListIncidents(db))
|
||||
r.Get("/api/incidents/{id}", handleGetIncident(db))
|
||||
r.Get("/api/incidents/{id}/alerts", handleIncidentAlerts(db))
|
||||
r.Get("/api/incidents/{id}/timeline", handleIncidentTimeline(db))
|
||||
r.Post("/api/incidents/{id}/acknowledge", handleIncidentAcknowledge(db))
|
||||
r.Delete("/api/incidents/{id}/acknowledge", handleIncidentUnacknowledge(db))
|
||||
r.Post("/api/incidents/{id}/resolve", handleIncidentResolve(db))
|
||||
r.Post("/api/incidents/{id}/assign", handleIncidentAssign(db))
|
||||
r.Post("/api/incidents/{id}/snooze", handleIncidentSnooze(db))
|
||||
r.Delete("/api/incidents/{id}/snooze", handleIncidentUnsnooze(db))
|
||||
r.Post("/api/incidents/{id}/archive", handleIncidentArchive(db))
|
||||
r.Delete("/api/incidents/{id}/archive", handleIncidentUnarchive(db))
|
||||
r.Post("/api/incidents/{id}/notes", handleCreateNote(db))
|
||||
r.Delete("/api/incidents/{id}/notes/{eventID}", handleDeleteNote(db))
|
||||
|
||||
r.Post("/api/schedule", handleCreateSchedule(db))
|
||||
r.Get("/api/schedule/current", handleCurrentSchedule(db)) // must be before /{id}
|
||||
r.Get("/api/schedule", handleListSchedule(db))
|
||||
r.Delete("/api/schedule/{id}", handleDeleteSchedule(db))
|
||||
|
||||
r.Get("/api/stats/incidents", handleStatsIncidents(db))
|
||||
r.Get("/api/stats/alerts", handleStatsAlerts(db))
|
||||
r.Get("/api/stats/alerts/top", handleStatsTop(db))
|
||||
r.Get("/api/stats/alerts/by-hour", handleStatsByHour(db))
|
||||
|
||||
+49
-9
@@ -11,7 +11,7 @@ import (
|
||||
|
||||
func handleStatsAlerts(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query())
|
||||
where, args := statsFilter(r.URL.Query(), "received_at")
|
||||
|
||||
var total, firing, resolved int64
|
||||
err := db.QueryRowContext(r.Context(), fmt.Sprintf(`
|
||||
@@ -34,7 +34,7 @@ func handleStatsAlerts(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleStatsTop(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query())
|
||||
where, args := statsFilter(r.URL.Query(), "received_at")
|
||||
|
||||
limit := 10
|
||||
if l := r.URL.Query().Get("limit"); l != "" {
|
||||
@@ -78,7 +78,7 @@ func handleStatsTop(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleStatsByHour(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query())
|
||||
where, args := statsFilter(r.URL.Query(), "received_at")
|
||||
|
||||
rows, err := db.QueryContext(r.Context(), fmt.Sprintf(`
|
||||
SELECT CAST(strftime('%%H', datetime(received_at, 'unixepoch')) AS INTEGER) AS hr,
|
||||
@@ -118,7 +118,7 @@ func handleStatsByHour(db *sql.DB) http.HandlerFunc {
|
||||
|
||||
func handleStatsByDay(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query())
|
||||
where, args := statsFilter(r.URL.Query(), "received_at")
|
||||
|
||||
// SQLite strftime('%w') → 0=Sunday … 6=Saturday
|
||||
rows, err := db.QueryContext(r.Context(), fmt.Sprintf(`
|
||||
@@ -159,19 +159,59 @@ func handleStatsByDay(db *sql.DB) http.HandlerFunc {
|
||||
}
|
||||
}
|
||||
|
||||
// statsFilter builds a WHERE clause and args from optional ?from and ?to query params.
|
||||
// Archived alerts are always excluded, matching the default GET /api/alerts view.
|
||||
func statsFilter(q url.Values) (where string, args []any) {
|
||||
// handleStatsIncidents reports the queue and the two numbers a rota actually
|
||||
// cares about: how long it takes someone to pick work up, and how long it takes
|
||||
// to finish. Neither was computable before incidents existed — alert rows are
|
||||
// mutated in place and carry no acknowledgement or closure time.
|
||||
func handleStatsIncidents(db *sql.DB) http.HandlerFunc {
|
||||
return func(w http.ResponseWriter, r *http.Request) {
|
||||
where, args := statsFilter(r.URL.Query(), "triggered_at")
|
||||
|
||||
var total, triggered, acknowledged, resolved int64
|
||||
var mtta, mttr *float64
|
||||
err := db.QueryRowContext(r.Context(), fmt.Sprintf(`
|
||||
SELECT COUNT(*),
|
||||
SUM(CASE WHEN status = 'triggered' THEN 1 ELSE 0 END),
|
||||
SUM(CASE WHEN status = 'acknowledged' THEN 1 ELSE 0 END),
|
||||
SUM(CASE WHEN status = 'resolved' THEN 1 ELSE 0 END),
|
||||
AVG(CASE WHEN acknowledged_at IS NOT NULL
|
||||
THEN acknowledged_at - triggered_at END),
|
||||
AVG(CASE WHEN resolved_at IS NOT NULL
|
||||
THEN resolved_at - triggered_at END)
|
||||
FROM incidents WHERE %s`, where), args...,
|
||||
).Scan(&total, &triggered, &acknowledged, &resolved, &mtta, &mttr)
|
||||
if err != nil {
|
||||
respond(w, http.StatusInternalServerError, errResp("internal error"))
|
||||
return
|
||||
}
|
||||
|
||||
respond(w, http.StatusOK, map[string]any{
|
||||
"total": total,
|
||||
"triggered": triggered,
|
||||
"acknowledged": acknowledged,
|
||||
"resolved": resolved,
|
||||
// Null until something has actually been acknowledged or resolved —
|
||||
// zero would read as "instant", which is a different claim.
|
||||
"mtta_seconds": mtta,
|
||||
"mttr_seconds": mttr,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// statsFilter builds a WHERE clause and args from optional ?from and ?to query
|
||||
// params, filtering on timeCol. Archived rows are always excluded, matching the
|
||||
// default list views.
|
||||
func statsFilter(q url.Values, timeCol string) (where string, args []any) {
|
||||
clauses := []string{"archived_at IS NULL"}
|
||||
if from := q.Get("from"); from != "" {
|
||||
if t, err := time.Parse("2006-01-02", from); err == nil {
|
||||
clauses = append(clauses, "received_at >= ?")
|
||||
clauses = append(clauses, timeCol+" >= ?")
|
||||
args = append(args, t.UTC().Unix())
|
||||
}
|
||||
}
|
||||
if to := q.Get("to"); to != "" {
|
||||
if t, err := time.Parse("2006-01-02", to); err == nil {
|
||||
clauses = append(clauses, "received_at < ?")
|
||||
clauses = append(clauses, timeCol+" < ?")
|
||||
args = append(args, t.UTC().AddDate(0, 0, 1).Unix())
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,123 @@
|
||||
-- Splits the single alerts row into two objects, the way an incident management
|
||||
-- tool needs them: alerts stay the machine-owned signal record that Alertmanager
|
||||
-- writes, and incidents become the human work item people acknowledge, assign,
|
||||
-- snooze, discuss and resolve.
|
||||
--
|
||||
-- Correlation uses Alertmanager's own groupKey, so incidents follow the group_by
|
||||
-- routing tree the operator already tuned rather than a second grouping scheme
|
||||
-- invented here.
|
||||
|
||||
CREATE TABLE incidents (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
group_key TEXT NOT NULL, -- Alertmanager groupKey, opaque
|
||||
title TEXT NOT NULL, -- rendered from group_labels
|
||||
group_labels TEXT NOT NULL DEFAULT '{}', -- JSON
|
||||
status TEXT NOT NULL CHECK(status IN ('triggered', 'acknowledged', 'resolved')),
|
||||
severity TEXT, -- highest `severity` label across firing members
|
||||
triggered_at INTEGER NOT NULL,
|
||||
acknowledged_by INTEGER REFERENCES users(id) ON DELETE SET NULL,
|
||||
acknowledged_at INTEGER,
|
||||
assigned_to INTEGER REFERENCES users(id) ON DELETE SET NULL,
|
||||
snoozed_until INTEGER,
|
||||
resolved_at INTEGER,
|
||||
resolution_source TEXT, -- 'alerts' | 'manual'
|
||||
archived_at INTEGER
|
||||
);
|
||||
|
||||
-- Load-bearing: at most one OPEN incident per group_key. This is what makes
|
||||
-- "resolved incident + a new alert occurrence = a new incident" work, and it is
|
||||
-- the constraint the webhook's find-or-open lookup relies on.
|
||||
CREATE UNIQUE INDEX incidents_open_group_key_idx ON incidents(group_key) WHERE resolved_at IS NULL;
|
||||
CREATE INDEX incidents_status_idx ON incidents(status);
|
||||
CREATE INDEX incidents_triggered_at_idx ON incidents(triggered_at DESC);
|
||||
CREATE INDEX incidents_archived_at_idx ON incidents(archived_at);
|
||||
|
||||
-- Membership is historical, not a pointer on alerts: one alert row (one
|
||||
-- fingerprint) resolves and re-fires over time and belongs to a different
|
||||
-- incident each occurrence.
|
||||
CREATE TABLE incident_alerts (
|
||||
incident_id INTEGER NOT NULL REFERENCES incidents(id) ON DELETE CASCADE,
|
||||
alert_id INTEGER NOT NULL REFERENCES alerts(id) ON DELETE CASCADE,
|
||||
added_at INTEGER NOT NULL DEFAULT (strftime('%s', 'now')),
|
||||
PRIMARY KEY (incident_id, alert_id)
|
||||
);
|
||||
|
||||
CREATE INDEX incident_alerts_alert_id_idx ON incident_alerts(alert_id);
|
||||
|
||||
-- The timeline. Append-only, and the only history this server keeps: alert rows
|
||||
-- are mutated in place, so without this there is no record that anything
|
||||
-- happened. Notes are events too, so one query renders the whole story.
|
||||
CREATE TABLE incident_events (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
incident_id INTEGER NOT NULL REFERENCES incidents(id) ON DELETE CASCADE,
|
||||
-- triggered | alert_added | alert_resolved | acknowledged | unacknowledged
|
||||
-- | assigned | snoozed | unsnoozed | resolved | note
|
||||
type TEXT NOT NULL,
|
||||
user_id INTEGER REFERENCES users(id) ON DELETE SET NULL, -- NULL = the server acted
|
||||
alert_id INTEGER REFERENCES alerts(id) ON DELETE SET NULL,
|
||||
detail TEXT,
|
||||
created_at INTEGER NOT NULL DEFAULT (strftime('%s', 'now'))
|
||||
);
|
||||
|
||||
CREATE INDEX incident_events_incident_idx ON incident_events(incident_id, created_at);
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- Backfill
|
||||
--
|
||||
-- Every pre-existing alert gets its own incident, archived ones included, so no
|
||||
-- acknowledgement and no comment is orphaned. There is no historical groupKey to
|
||||
-- correlate on, hence one incident per fingerprint under a 'backfill:' prefix
|
||||
-- that can never collide with a real Alertmanager groupKey.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
INSERT INTO incidents (group_key, title, group_labels, status, severity, triggered_at,
|
||||
acknowledged_by, acknowledged_at, assigned_to,
|
||||
resolved_at, resolution_source, archived_at)
|
||||
SELECT 'backfill:' || a.fingerprint,
|
||||
a.name,
|
||||
json_object('alertname', a.name),
|
||||
CASE WHEN a.status = 'resolved' THEN 'resolved'
|
||||
WHEN a.acknowledged_by IS NOT NULL THEN 'acknowledged'
|
||||
ELSE 'triggered' END,
|
||||
json_extract(a.labels, '$.severity'),
|
||||
a.starts_at,
|
||||
a.acknowledged_by,
|
||||
a.acknowledged_at,
|
||||
a.acknowledged_by,
|
||||
CASE WHEN a.status = 'resolved' THEN COALESCE(a.ends_at, a.received_at) END,
|
||||
CASE WHEN a.status = 'resolved' THEN 'alerts' END,
|
||||
a.archived_at
|
||||
FROM alerts a;
|
||||
|
||||
INSERT INTO incident_alerts (incident_id, alert_id, added_at)
|
||||
SELECT i.id, a.id, a.starts_at
|
||||
FROM alerts a
|
||||
JOIN incidents i ON i.group_key = 'backfill:' || a.fingerprint;
|
||||
|
||||
INSERT INTO incident_events (incident_id, type, alert_id, created_at)
|
||||
SELECT i.id, 'triggered', ia.alert_id, i.triggered_at
|
||||
FROM incidents i JOIN incident_alerts ia ON ia.incident_id = i.id;
|
||||
|
||||
INSERT INTO incident_events (incident_id, type, user_id, created_at)
|
||||
SELECT i.id, 'acknowledged', i.acknowledged_by, i.acknowledged_at
|
||||
FROM incidents i WHERE i.acknowledged_at IS NOT NULL;
|
||||
|
||||
INSERT INTO incident_events (incident_id, type, created_at)
|
||||
SELECT i.id, 'resolved', i.resolved_at
|
||||
FROM incidents i WHERE i.resolved_at IS NOT NULL;
|
||||
|
||||
INSERT INTO incident_events (incident_id, type, user_id, alert_id, detail, created_at)
|
||||
SELECT ia.incident_id, 'note', c.user_id, c.alert_id, c.content, c.created_at
|
||||
FROM alert_comments c
|
||||
JOIN incident_alerts ia ON ia.alert_id = c.alert_id;
|
||||
|
||||
-- ---------------------------------------------------------------------------
|
||||
-- Workflow state now lives on incidents only. Leaving these behind would keep
|
||||
-- the bug they caused: the webhook upsert owns the alerts row and never cleared
|
||||
-- the acknowledgement, so a re-fire days later still read as acknowledged.
|
||||
-- ---------------------------------------------------------------------------
|
||||
|
||||
DROP TABLE alert_comments;
|
||||
|
||||
ALTER TABLE alerts DROP COLUMN acknowledged_by;
|
||||
ALTER TABLE alerts DROP COLUMN acknowledged_at;
|
||||
@@ -2,6 +2,10 @@ package models
|
||||
|
||||
import "time"
|
||||
|
||||
// Alert is the machine-owned signal record: what Alertmanager told us, and
|
||||
// nothing else. It has two states, firing and resolved, and no human ever writes
|
||||
// to it — acknowledgement, assignment, notes and closure all live on the
|
||||
// Incident an alert belongs to.
|
||||
type Alert struct {
|
||||
ID int64 `json:"id"`
|
||||
Fingerprint string `json:"fingerprint"`
|
||||
@@ -27,10 +31,11 @@ type Alert struct {
|
||||
// TestWebhook_ResendBumpsReceivedAt.
|
||||
ReceivedAt time.Time `json:"received_at"`
|
||||
|
||||
// Populated when the alert has been acknowledged.
|
||||
AcknowledgedByID *int64 `json:"acknowledged_by_id,omitempty"`
|
||||
AcknowledgedByUser *string `json:"acknowledged_by,omitempty"`
|
||||
AcknowledgedAt *time.Time `json:"acknowledged_at,omitempty"`
|
||||
// IncidentID is the most recent incident this alert belongs to. An alert row
|
||||
// is reused across occurrences of the same fingerprint, so over its life it
|
||||
// belongs to a series of incidents; incident_alerts keeps the full history
|
||||
// and this is only the newest link.
|
||||
IncidentID *int64 `json:"incident_id,omitempty"`
|
||||
|
||||
// ResolutionSource records why a resolved alert left the firing state:
|
||||
// "alertmanager" for a real resolved webhook, "expiry" when the sweeper
|
||||
|
||||
@@ -1,12 +0,0 @@
|
||||
package models
|
||||
|
||||
import "time"
|
||||
|
||||
type Comment struct {
|
||||
ID int64 `json:"id"`
|
||||
AlertID int64 `json:"alert_id"`
|
||||
UserID int64 `json:"user_id"`
|
||||
Username string `json:"username"`
|
||||
Content string `json:"content"`
|
||||
CreatedAt time.Time `json:"created_at"`
|
||||
}
|
||||
@@ -0,0 +1,68 @@
|
||||
package models
|
||||
|
||||
import "time"
|
||||
|
||||
// Incident is the human work item: the thing that gets acknowledged, assigned,
|
||||
// snoozed, discussed and resolved. Alerts are the machine-owned signal records
|
||||
// underneath it — many alerts map to one incident, correlated by the groupKey
|
||||
// Alertmanager already computed from the operator's group_by configuration.
|
||||
//
|
||||
// Nothing here is ever written by the Alertmanager webhook except Status, which
|
||||
// the webhook and the sweeper may flip to "resolved" once every member alert has
|
||||
// stopped firing.
|
||||
type Incident struct {
|
||||
ID int64 `json:"id"`
|
||||
GroupKey string `json:"group_key"`
|
||||
Title string `json:"title"`
|
||||
GroupLabels map[string]string `json:"group_labels"`
|
||||
|
||||
// Status is "triggered", "acknowledged" or "resolved".
|
||||
Status string `json:"status"`
|
||||
|
||||
// Severity is the highest `severity` label across the alerts that were
|
||||
// firing when it was last recomputed. It is deliberately not cleared when an
|
||||
// incident resolves — a resolved incident should still say how bad it was.
|
||||
Severity *string `json:"severity,omitempty"`
|
||||
|
||||
TriggeredAt time.Time `json:"triggered_at"`
|
||||
|
||||
AcknowledgedByID *int64 `json:"acknowledged_by_id,omitempty"`
|
||||
AcknowledgedByUser *string `json:"acknowledged_by,omitempty"`
|
||||
AcknowledgedAt *time.Time `json:"acknowledged_at,omitempty"`
|
||||
|
||||
AssignedToID *int64 `json:"assigned_to_id,omitempty"`
|
||||
AssignedToUser *string `json:"assigned_to,omitempty"`
|
||||
|
||||
// SnoozedUntil hides the incident from the default queue without closing it.
|
||||
// A timestamp in the past reads as "not snoozed"; nothing sweeps it.
|
||||
SnoozedUntil *time.Time `json:"snoozed_until,omitempty"`
|
||||
|
||||
ResolvedAt *time.Time `json:"resolved_at,omitempty"`
|
||||
|
||||
// ResolutionSource is "alerts" when every member alert stopped firing, or
|
||||
// "manual" when a human closed it. Manual resolution is terminal: a later
|
||||
// occurrence opens a new incident rather than reopening this one.
|
||||
ResolutionSource *string `json:"resolution_source,omitempty"`
|
||||
|
||||
ArchivedAt *time.Time `json:"archived_at,omitempty"`
|
||||
|
||||
// Alerts is populated by GET /api/incidents/{id} only.
|
||||
Alerts []Alert `json:"alerts,omitempty"`
|
||||
}
|
||||
|
||||
// IncidentEvent is one entry in an incident's timeline. The table is append-only
|
||||
// and is the only history this server keeps — alert rows are mutated in place.
|
||||
//
|
||||
// Type is one of: triggered, alert_added, alert_resolved, acknowledged,
|
||||
// unacknowledged, assigned, snoozed, unsnoozed, resolved, note. A nil UserID
|
||||
// means the server acted rather than a person.
|
||||
type IncidentEvent struct {
|
||||
ID int64 `json:"id"`
|
||||
IncidentID int64 `json:"incident_id"`
|
||||
Type string `json:"type"`
|
||||
UserID *int64 `json:"user_id,omitempty"`
|
||||
Username *string `json:"username,omitempty"`
|
||||
AlertID *int64 `json:"alert_id,omitempty"`
|
||||
Detail *string `json:"detail,omitempty"`
|
||||
CreatedAt time.Time `json:"created_at"`
|
||||
}
|
||||
Reference in New Issue
Block a user