Files
terdut-server/internal/api/incident_store.go
T
Niklas Ye dc39e3a5d3
CI / chart (pull_request) Successful in 1s
CI / security (pull_request) Successful in 17s
CI / test (pull_request) Successful in 2m5s
Move the database to Postgres, before teams need the schema
First step of #1, and it goes first for one reason: #4 adds a team_id to
nearly every table, and doing that twice -- once for SQLite, once for
Postgres -- is work nobody gets paid for. The teams migrations now only
have to be written against one database.

The ten SQLite migrations are replaced by a single Postgres baseline
rather than ported one by one. They were incremental in a way that has
no value on a fresh install: 004 adds columns 008 drops again, and 008's
backfill rewrites data a Postgres database never had. The history stays
in git; the schema they add up to is now 001_baseline.sql.

Timestamps stay BIGINT unix seconds and are NOT converted to timestamptz.
Everything in Go already speaks epochs, so converting would have been a
second, larger change riding along inside this one. It is worth doing on
its own. The JSON columns did move to jsonb, because #4 will want to
filter and index on labels.

Most of the port is mechanical -- 170 placeholders from ? to $1 -- but
four things needed more than a search and replace:

  * Dynamically built WHERE clauses cannot keep their numbering straight
    by hand, so they hand out placeholders through sqlArgs instead. A
    filter can now be added or reordered without renumbering anything.

  * SUM(resolved_at IS NULL) was SQLite counting a boolean as 0 or 1.
    Postgres has no sum(boolean), and this was breaking every dead man's
    switch -- silently, since the sweeper only logs. Now COUNT(*) FILTER.

  * unixepoch() became FLOOR(EXTRACT(EPOCH FROM now()))::bigint. The
    FLOOR is load-bearing: a bare cast rounds half up, so a row written
    at .6 of a second claimed a timestamp a second in the future and
    disagreed with the time.Now().Unix() the Go side stamps.

  * The unique-violation check matched SQLite's error text. It matches
    SQLSTATE 23505 now, so a renamed constraint cannot turn a 409 back
    into a 500.

Tests need a real Postgres, because there is no in-memory Postgres the
way there was an in-memory SQLite. Each test gets its own schema on a
shared server -- cheaper than a database each, and still isolated.
TERDUT_TEST_DSN says where it is; `make test-db` starts one locally and
ci.yaml runs one as a service container. An unset DSN fails the suite
rather than skipping it: a run that quietly tests nothing is worse than
one that does not run.

TestMigration_BackfillCarriesAckAndComments is deleted along with the
migrations it replayed. What it protected -- an upgrade not losing
acknowledgements and comments -- now belongs to scripts/sqlite-to-postgres.go,
which is build-tagged so the SQLite driver stays out of the server
binary. Both are meant to be deleted once this install has migrated.

The chart loses the PVC, the data volume and the python backup sidecar,
and requires database.dsnSecret.name: it provisions no database and
cannot guess where the credentials live, so a render without it is meant
to fail. Backups move to where Postgres actually runs. The other half of
that -- the postgresql CR, the k8up pg_dump annotation and the network
policy -- is a change to the wrapper chart in Ryuvia/charts and is not in
here.

Verified rather than assumed: the gate is green with -race against
Postgres 17, govulncheck and gitleaks are clean, and the migration script
was run end to end against a SQLite database built at the old schema and
seeded in every table. Ids survive, so incidents keep their numbers and
every foreign key still points where it did; the identity sequences are
moved past the copied ids, and a webhook after the migration opened
incident 12 rather than colliding at 1.
2026-09-20 10:44:12 +02:00

301 lines
9.7 KiB
Go

package api
import (
"context"
"database/sql"
"encoding/json"
"sort"
"strings"
"time"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// Values for incidents.resolution_source, recording who closed the incident:
// every member alert stopped firing, or a person decided it was done.
const (
incidentResolutionAlerts = "alerts"
incidentResolutionManual = "manual"
// incidentResolutionRecovered closes a dead man's switch incident whose
// heartbeat started arriving again. It cannot be "alerts": these incidents
// have no member alerts for the cascade to work from.
incidentResolutionRecovered = "recovered"
)
// Incident timeline event types. Stored as free text so adding one later is not
// a migration, but these are the ones the server writes.
const (
evTriggered = "triggered"
evAlertAdded = "alert_added"
evAlertResolved = "alert_resolved"
evAcknowledged = "acknowledged"
evUnacknowledged = "unacknowledged"
evAssigned = "assigned"
evSnoozed = "snoozed"
evUnsnoozed = "unsnoozed"
evResolved = "resolved"
evNote = "note"
evDeadmanSilent = "deadman_silent"
)
// severityLabel is the Alertmanager label an incident's severity is derived from.
const severityLabel = "severity"
// querier is satisfied by both *sql.DB and *sql.Tx, so the helpers below work
// inside the webhook's transaction and standalone from handlers and the sweeper.
type querier interface {
ExecContext(ctx context.Context, query string, args ...any) (sql.Result, error)
QueryContext(ctx context.Context, query string, args ...any) (*sql.Rows, error)
QueryRowContext(ctx context.Context, query string, args ...any) *sql.Row
}
const incidentSelectFrom = `
SELECT i.id, i.group_key, i.title, i.group_labels, i.status, i.severity,
i.triggered_at,
i.acknowledged_by, i.acknowledged_at, ack.username,
i.assigned_to, asg.username, i.snoozed_until,
i.resolved_at, i.resolution_source, i.archived_at
FROM incidents i
LEFT JOIN users ack ON ack.id = i.acknowledged_by
LEFT JOIN users asg ON asg.id = i.assigned_to`
func scanIncident(s scanner) (models.Incident, error) {
var i models.Incident
var groupLabelsJSON string
var triggeredAt int64
var ackAt, snoozedUntil, resolvedAt, archivedAt *int64
if err := s.Scan(
&i.ID, &i.GroupKey, &i.Title, &groupLabelsJSON, &i.Status, &i.Severity,
&triggeredAt,
&i.AcknowledgedByID, &ackAt, &i.AcknowledgedByUser,
&i.AssignedToID, &i.AssignedToUser, &snoozedUntil,
&resolvedAt, &i.ResolutionSource, &archivedAt,
); err != nil {
return i, err
}
json.Unmarshal([]byte(groupLabelsJSON), &i.GroupLabels) //nolint:errcheck
i.TriggeredAt = time.Unix(triggeredAt, 0).UTC()
i.AcknowledgedAt = unixPtr(ackAt)
i.SnoozedUntil = unixPtr(snoozedUntil)
i.ResolvedAt = unixPtr(resolvedAt)
i.ArchivedAt = unixPtr(archivedAt)
return i, nil
}
// unixPtr converts a nullable Unix-second column to a nullable UTC time.
func unixPtr(sec *int64) *time.Time {
if sec == nil {
return nil
}
t := time.Unix(*sec, 0).UTC()
return &t
}
func fetchIncident(ctx context.Context, q querier, id int64) (models.Incident, error) {
return scanIncident(q.QueryRowContext(ctx, incidentSelectFrom+" WHERE i.id = $1", id))
}
// logEvent appends one entry to an incident's timeline. A nil userID means the
// server acted rather than a person.
func logEvent(ctx context.Context, q querier, incidentID int64, evType string, userID, alertID *int64, detail *string) error {
_, err := q.ExecContext(ctx, `
INSERT INTO incident_events (incident_id, type, user_id, alert_id, detail, created_at)
VALUES ($1, $2, $3, $4, $5, $6)`,
incidentID, evType, userID, alertID, detail, time.Now().Unix())
return err
}
// todayUTC is the schedule's day key. The schedule's smallest unit is one UTC day.
func todayUTC() string {
return time.Now().UTC().Format("2006-01-02")
}
// currentOnCall returns today's on-call user, or nil when nobody is scheduled.
// A missing schedule entry is not an error — incidents just open unassigned.
func currentOnCall(ctx context.Context, q querier) (*int64, error) {
var userID int64
err := q.QueryRowContext(ctx,
"SELECT user_id FROM schedule_entries WHERE date = $1", todayUTC()).Scan(&userID)
if err == sql.ErrNoRows {
return nil, nil
}
if err != nil {
return nil, err
}
return &userID, nil
}
// severityRank orders the conventional Alertmanager severity label values.
// Anything unrecognised sorts below all of them rather than being dropped.
func severityRank(s string) int {
switch strings.ToLower(s) {
case "critical":
return 4
case "error":
return 3
case "warning":
return 2
case "info":
return 1
default:
return 0
}
}
// refreshSeverity raises an incident's severity to the highest `severity` label
// seen across its alerts.
//
// It is a high-water mark, never lowered: an incident that hit critical was a
// critical incident, even after the critical alert clears and a warning is all
// that is left firing. Downgrading a live incident would also quietly demote it
// in the queue while the work is still open.
func refreshSeverity(ctx context.Context, q querier, incidentID int64) error {
rows, err := q.QueryContext(ctx, `
SELECT a.labels ->> $1
FROM incident_alerts ia
JOIN alerts a ON a.id = ia.alert_id
WHERE ia.incident_id = $2`, severityLabel, incidentID)
if err != nil {
return err
}
best := ""
for rows.Next() {
var sev *string
if err := rows.Scan(&sev); err != nil {
rows.Close()
return err
}
if sev != nil && severityRank(*sev) > severityRank(best) {
best = *sev
}
}
if err := rows.Err(); err != nil {
rows.Close()
return err
}
rows.Close()
if best == "" {
return nil
}
// The comparison lives in SQL so an unrelated concurrent update cannot be
// clobbered by a stale read.
_, err = q.ExecContext(ctx, `
UPDATE incidents SET severity = $1
WHERE id = $2
AND (severity IS NULL OR `+severityRankSQL("severity")+` < $3)`,
best, incidentID, severityRank(best))
return err
}
// severityRankSQL mirrors severityRank for use inside a statement. SQL cannot
// order these strings meaningfully on its own.
func severityRankSQL(col string) string {
return `CASE lower(COALESCE(` + col + `, ''))
WHEN 'critical' THEN 4
WHEN 'error' THEN 3
WHEN 'warning' THEN 2
WHEN 'info' THEN 1
ELSE 0 END`
}
// resolveIfSettled closes an incident once every alert under it has stopped
// firing — PagerDuty's cascade, and the only automatic route out of the open
// state. Reports whether it actually resolved anything.
func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, error) {
res, err := q.ExecContext(ctx, `
UPDATE incidents
SET status = 'resolved',
resolved_at = $1,
resolution_source = $2
WHERE id = $3
AND resolved_at IS NULL
-- An incident with no members yet is mid-creation, not settled.
AND EXISTS (SELECT 1 FROM incident_alerts ia WHERE ia.incident_id = incidents.id)
AND NOT EXISTS (SELECT 1
FROM incident_alerts ia
JOIN alerts a ON a.id = ia.alert_id
WHERE ia.incident_id = incidents.id
AND a.status = 'firing')`,
time.Now().Unix(), incidentResolutionAlerts, incidentID)
if err != nil {
return false, err
}
n, _ := res.RowsAffected()
if n == 0 {
return false, nil
}
if err := logEvent(ctx, q, incidentID, evResolved, nil, nil, nil); err != nil {
return false, err
}
// The all-clear goes only to whoever was paged in the first place, which
// enqueueResolved works out from the incident's own notification history.
// Manual resolution sends nothing: the person who closed it already knows.
return true, enqueueResolved(ctx, q, incidentID)
}
// acknowledgeIncident records that userID has picked an incident up, and reports
// whether it changed anything — an already-resolved incident is left alone.
// Shared by the authenticated handler and the Acknowledge button in a push
// notification, so both write the same state and the same timeline entry.
func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int64) (bool, error) {
res, err := q.ExecContext(ctx, `
UPDATE incidents
SET status = 'acknowledged', acknowledged_by = $1, acknowledged_at = $2
WHERE id = $3 AND resolved_at IS NULL`,
userID, time.Now().Unix(), incidentID)
if err != nil {
return false, err
}
if n, _ := res.RowsAffected(); n == 0 {
return false, nil
}
return true, logEvent(ctx, q, incidentID, evAcknowledged, &userID, nil, nil)
}
// openIncidentForAlert returns the open incident an alert currently belongs to,
// or 0 when it has none. Used when an alert resolves or expires so the event
// lands on the right timeline.
func openIncidentForAlert(ctx context.Context, q querier, alertID int64) (int64, error) {
var id int64
err := q.QueryRowContext(ctx, `
SELECT i.id
FROM incident_alerts ia
JOIN incidents i ON i.id = ia.incident_id
WHERE ia.alert_id = $1 AND i.resolved_at IS NULL`, alertID).Scan(&id)
if err == sql.ErrNoRows {
return 0, nil
}
return id, err
}
// incidentTitle renders a human-readable title from Alertmanager's groupLabels,
// leading with the alert name and appending whatever else the operator grouped
// by. Falls back to the alert's own name when the payload carried no groupLabels.
func incidentTitle(groupLabels map[string]string, fallback string) string {
name := groupLabels["alertname"]
if name == "" {
name = fallback
}
if name == "" {
name = "Incident"
}
rest := make([]string, 0, len(groupLabels))
for k, v := range groupLabels {
if k == "alertname" {
continue
}
rest = append(rest, k+"="+v)
}
if len(rest) == 0 {
return name
}
sort.Strings(rest)
return name + " (" + strings.Join(rest, ", ") + ")"
}