9029d48584
- TERDUT_OPERATOR_KEY creates or re-keys the instance-scoped service account
"terdut-operator" at every start, so terdut-operator needs no bootstrap
handshake. An instance-scoped account now acts as owner of every team's
configuration, but is not a member of any team.
- POST /api/teams takes an external_id (instance service accounts only) and
is idempotent on it, so automation finds its own team again after a crash
instead of adopting by display name. GET /api/teams?name= is removed.
- Integration and dead man's switch names are unique per team (409). The
escalation PUT accepts usernames and resolves them itself.
- The 18 migrations are squashed into 001_schema.sql, with no Default team.
TERDUT_DEADMAN_* and the env seeding of switches are removed: teams carry
their own. Existing development databases must be recreated.
Security and robustness:
- GET /api/users no longer returns other people's email or ntfy topic to
non-admins.
- The access log records the route pattern, so integration keys and ack
tokens in the path are not written to the log. Server errors are logged.
- Rate limits take the client address TERDUT_TRUSTED_PROXIES hops from the
right of X-Forwarded-For instead of trusting the first, forgeable entry.
- /api/bootstrap runs in a transaction under an advisory lock, so two
concurrent calls cannot both create an administrator.
- API key last_used_at is written at most every five minutes.
Cleanup: remove GET /api/incidents/{id}/alerts, unused exports, SQLite
remnants in comments and config.
Claude-Session: https://claude.ai/code/session_016mBLURvJoMuUEr9cB2RpUN
61 lines
2.9 KiB
Go
61 lines
2.9 KiB
Go
package models
|
|
|
|
import "time"
|
|
|
|
// Alert is the machine-owned signal record: what Alertmanager told us, and
|
|
// nothing else. It has two states, firing and resolved, and no human ever writes
|
|
// to it — acknowledgement, assignment, notes and closure all live on the
|
|
// Incident an alert belongs to.
|
|
type Alert struct {
|
|
ID int64 `json:"id"`
|
|
|
|
// TeamID is the team whose integration received this alert, and TeamName
|
|
// rides along so a combined list can label a row without a second request.
|
|
TeamID int64 `json:"team_id"`
|
|
TeamName string `json:"team_name,omitempty"`
|
|
|
|
Fingerprint string `json:"fingerprint"`
|
|
Name string `json:"name"`
|
|
Status string `json:"status"` // "firing" or "resolved"
|
|
Labels map[string]string `json:"labels"`
|
|
Annotations map[string]string `json:"annotations"`
|
|
StartsAt time.Time `json:"starts_at"`
|
|
EndsAt *time.Time `json:"ends_at,omitempty"`
|
|
GeneratorURL string `json:"generator_url"`
|
|
|
|
// ReceivedAt is when the server last accepted a webhook for this
|
|
// fingerprint, including the unchanged firing notifications Alertmanager
|
|
// re-sends every repeat_interval.
|
|
//
|
|
// This is a documented part of the public API, not an internal ingest
|
|
// detail: StartsAt never changes for an alert instance, so ReceivedAt is
|
|
// the only signal a client has that a firing alert is still being
|
|
// refreshed. The sweeper stale-dates against it (see expireStale), API
|
|
// clients render it, and GET /api/alerts is ordered by it. Anything that
|
|
// stops the webhook handler from advancing it on a re-send is a breaking
|
|
// change — see "received_at is a liveness heartbeat" in docs/api.md and
|
|
// TestWebhook_ResendBumpsReceivedAt.
|
|
ReceivedAt time.Time `json:"received_at"`
|
|
|
|
// IncidentID is the most recent incident this alert belongs to. An alert row
|
|
// is reused across occurrences of the same fingerprint, so over its life it
|
|
// belongs to a series of incidents; incident_alerts keeps the full history
|
|
// and this is only the newest link.
|
|
IncidentID *int64 `json:"incident_id,omitempty"`
|
|
|
|
// ResolutionSource records why a resolved alert left the firing state:
|
|
// "alertmanager" for a real resolved webhook, "expiry" when the sweeper
|
|
// inferred it after the alert stopped being refreshed. Nil while firing, and
|
|
// cleared again by a re-fire under the same fingerprint.
|
|
//
|
|
// Also public API: it is how a client knows whether EndsAt was observed or
|
|
// inferred. Under "expiry" nothing ever reported an end, so EndsAt is only
|
|
// an upper bound (see expireStale) and ReceivedAt is the more truthful
|
|
// signal. Treat the value set as open — see "resolution_source says how much
|
|
// to trust ends_at" in docs/api.md, and TestWebhook_ResolvedSetsSource /
|
|
// TestExpiry_StaleFiringAlert.
|
|
ResolutionSource *string `json:"resolution_source,omitempty"`
|
|
|
|
ArchivedAt *time.Time `json:"archived_at,omitempty"`
|
|
}
|