Files
terdut-server/internal/models/alert.go
T
Niklas Ye 9029d48584
CI / chart (pull_request) Successful in 2s
CI / security (pull_request) Failing after 19s
CI / test (pull_request) Successful in 5m34s
Let an operator authenticate with a seeded key, and reset the schema
- TERDUT_OPERATOR_KEY creates or re-keys the instance-scoped service account
  "terdut-operator" at every start, so terdut-operator needs no bootstrap
  handshake. An instance-scoped account now acts as owner of every team's
  configuration, but is not a member of any team.
- POST /api/teams takes an external_id (instance service accounts only) and
  is idempotent on it, so automation finds its own team again after a crash
  instead of adopting by display name. GET /api/teams?name= is removed.
- Integration and dead man's switch names are unique per team (409). The
  escalation PUT accepts usernames and resolves them itself.
- The 18 migrations are squashed into 001_schema.sql, with no Default team.
  TERDUT_DEADMAN_* and the env seeding of switches are removed: teams carry
  their own. Existing development databases must be recreated.

Security and robustness:
- GET /api/users no longer returns other people's email or ntfy topic to
  non-admins.
- The access log records the route pattern, so integration keys and ack
  tokens in the path are not written to the log. Server errors are logged.
- Rate limits take the client address TERDUT_TRUSTED_PROXIES hops from the
  right of X-Forwarded-For instead of trusting the first, forgeable entry.
- /api/bootstrap runs in a transaction under an advisory lock, so two
  concurrent calls cannot both create an administrator.
- API key last_used_at is written at most every five minutes.

Cleanup: remove GET /api/incidents/{id}/alerts, unused exports, SQLite
remnants in comments and config.

Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
2026-10-09 14:56:13 +02:00

61 lines
2.9 KiB
Go

package models
import "time"
// Alert is the machine-owned signal record: what Alertmanager told us, and
// nothing else. It has two states, firing and resolved, and no human ever writes
// to it — acknowledgement, assignment, notes and closure all live on the
// Incident an alert belongs to.
type Alert struct {
ID int64 `json:"id"`
// TeamID is the team whose integration received this alert, and TeamName
// rides along so a combined list can label a row without a second request.
TeamID int64 `json:"team_id"`
TeamName string `json:"team_name,omitempty"`
Fingerprint string `json:"fingerprint"`
Name string `json:"name"`
Status string `json:"status"` // "firing" or "resolved"
Labels map[string]string `json:"labels"`
Annotations map[string]string `json:"annotations"`
StartsAt time.Time `json:"starts_at"`
EndsAt *time.Time `json:"ends_at,omitempty"`
GeneratorURL string `json:"generator_url"`
// ReceivedAt is when the server last accepted a webhook for this
// fingerprint, including the unchanged firing notifications Alertmanager
// re-sends every repeat_interval.
//
// This is a documented part of the public API, not an internal ingest
// detail: StartsAt never changes for an alert instance, so ReceivedAt is
// the only signal a client has that a firing alert is still being
// refreshed. The sweeper stale-dates against it (see expireStale), API
// clients render it, and GET /api/alerts is ordered by it. Anything that
// stops the webhook handler from advancing it on a re-send is a breaking
// change — see "received_at is a liveness heartbeat" in docs/api.md and
// TestWebhook_ResendBumpsReceivedAt.
ReceivedAt time.Time `json:"received_at"`
// IncidentID is the most recent incident this alert belongs to. An alert row
// is reused across occurrences of the same fingerprint, so over its life it
// belongs to a series of incidents; incident_alerts keeps the full history
// and this is only the newest link.
IncidentID *int64 `json:"incident_id,omitempty"`
// ResolutionSource records why a resolved alert left the firing state:
// "alertmanager" for a real resolved webhook, "expiry" when the sweeper
// inferred it after the alert stopped being refreshed. Nil while firing, and
// cleared again by a re-fire under the same fingerprint.
//
// Also public API: it is how a client knows whether EndsAt was observed or
// inferred. Under "expiry" nothing ever reported an end, so EndsAt is only
// an upper bound (see expireStale) and ReceivedAt is the more truthful
// signal. Treat the value set as open — see "resolution_source says how much
// to trust ends_at" in docs/api.md, and TestWebhook_ResolvedSetsSource /
// TestExpiry_StaleFiringAlert.
ResolutionSource *string `json:"resolution_source,omitempty"`
ArchivedAt *time.Time `json:"archived_at,omitempty"`
}