Compare commits

..

13 Commits

Author SHA1 Message Date
Niklas Ye 0f88574a41 Set the chart's placeholder version to 0.39.0
CI / chart (push) Successful in 4s
CI / security (push) Successful in 45s
CI / test (push) Successful in 6m1s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 25s
Release / image (push) Successful in 1m12s
Release / scan-image (push) Successful in 6s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with 584d344 (0.38.0) and df83adf (0.37.2) before it, because a
tree heading for v0.39.0 that still says 0.38.0 tells its reader
something false.
2026-10-08 09:02:07 +02:00
Niklas Ye 3613fd5732 Check rows.Err() in the remaining Next() loops (#27)
Incident list, the three stats breakdowns and the user list could return a
truncated result as if complete when the scan failed partway.
2026-10-08 09:00:54 +02:00
Niklas Ye dc62278788 Check rows.Err() after the timeline scan loop (#27) 2026-10-08 08:58:19 +02:00
Niklas Ye 0aaea8efb5 Record the actor on assign, archive and unarchive (#35)
Assign logged only the assignee (user_id), archive/unarchive logged nothing.
Migration 018 adds actor_user_id/actor_service_account_id to incident_events
for 'assigned'; archive/unarchive now log archived/unarchived events with the
caller via callerActorIDs. Timeline JSON gains actor_* fields; web timeline
renders them. Service accounts are still not assignable.
2026-10-08 08:53:04 +02:00
Niklas Ye 584d3441fc Set the chart's placeholder version to 0.38.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 2m9s
CI / test (push) Successful in 6m7s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 35s
Release / image (push) Successful in 1m10s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with df83adf (0.37.2) and 2b2609e (0.37.1) before it, because a
tree heading for v0.38.0 that still says 0.37.2 tells its reader
something false.
2026-10-07 22:39:00 +02:00
Niklas Ye 926aa2d3ec Add optional API key expiry and a missing way to list them
Part of the same security-hardening pass as the last five commits, and
the last item in its backlog. User API keys had no expiry at all --
unlike service-account keys, visibly distinct only by their "tdsa_"
prefix -- and, it turns out while implementing this, no way to list
them either: only create (returns the raw key once) and delete-by-id
existed, so a key's owner had no way to even discover what keys they
had short of remembering IDs from creation time.

handleCreateAPIKey takes an optional expires_in_days (0, the default,
keeps today's behavior: never expires, so no existing integration is
affected). apiKeyUser's lookup now carries `expires_at IS NULL OR
expires_at > now` as part of the query itself, the same way serveAs's
disabled_at check already works -- an expired key simply fails to
resolve, like a wrong one, rather than resolving and being caught
after the fact. New GET /api/users/{id}/api-keys (requireSelfOrAdmin,
same as create/delete) lists id/name/created_at/last_used_at/expires_at,
never the raw key.

Scoped down from the original plan on request: no web UI change, since
there turned out to be no existing API-keys UI at all to extend --
building one from scratch would have been a real feature addition, not
a hardening tweak.

Mirrored the additive expires_at field in terdut-tui's APIKey struct
(separate commit, separate repo) per this workspace's version-coupling
rule; the TUI does not create or list expiring keys itself yet.

New tests (api_keys_test.go): default never-expires, expires_in_days
sets expires_at, out-of-range values rejected, an expired key fails
auth after a fresh one worked, the listing never includes the raw key.
Also added the new GET route to authz_scope_test.go's self-or-admin
table from the previous commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:31:44 +02:00
Niklas Ye a2ca9c25d0 Add a regression test for the team/admin/self authorization pattern
Part of the same security-hardening pass as the last four commits.
terdut-server's authz is already solid -- centralized predicates
(requireTeamMember, requireTeamOwner, requireSelfOrAdmin, AdminOnly)
rather than ad hoc per-handler checks, confirmed by spot-checking several
handlers while writing this. But it is enforced by convention, not the
type system: a future handler that forgets its guard would compile and
read fine on review, exactly like one that remembers it.

authz_scope_test.go builds two teams and, for every team-scoped route
(members, OIDC groups, invites, escalation, dead man's switches,
integrations, schedule, plus every /api/incidents/{id}/... route, scoped
by the incident's own team_id through incidentIDParam's single
chokepoint), calls it as one team's owner against the other team's
resources and asserts 404 -- requireTeamMember and requireTeamOwner both
answer a non-member that way. Separate tests cover AdminOnly's routes
(403 for a non-admin) and requireSelfOrAdmin's (403 for a non-admin
acting on someone else's account).

Verified the test actually catches a regression, not just that it
passes today: temporarily removed handleListTeamMembers' requireTeamMember
call, confirmed exactly that one subtest failed and nothing else did,
then put it back.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:21:46 +02:00
Niklas Ye 92959cac38 Run the pod as non-root with a read-only filesystem
Part of the same security-hardening pass as the last three commits.
Neither the Dockerfile nor the chart's Deployment set any securityContext
at all, so the container ran as root by default — scratch has no
/etc/passwd for a USER directive to resolve against, so nobody had set one.

Dockerfile now ends with USER 65532:65532 (numeric, since scratch has no
user database; 65532 is the common "nonroot" convention, distroless's own
uid). The chart's Deployment adds a matching pod-level securityContext
(runAsNonRoot, runAsUser/runAsGroup: 65532, seccompProfile: RuntimeDefault)
plus per-container hardening (allowPrivilegeEscalation: false,
capabilities dropped, readOnlyRootFilesystem: true) on both the app
container and the wait-for-postgres init container — neither writes
anything to disk, so the root filesystem can stay read-only.

Verified with helm-lint and a manual `helm template` render of both the
terdut-server and terdut-demo charts. Not yet verified: an actual pod
starting with these in place — readOnlyRootFilesystem is exactly where a
non-obvious write (a temp file, a cache dir) would surface as a crash
rather than a lint error, so that needs a real rollout to confirm.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:16:06 +02:00
Niklas Ye f15db0e20a Add gosec to CI; fix the one real finding it surfaced
Part of the same security-hardening pass as the last two commits. make
lint was go vet only; govulncheck and gitleaks already scanned deps and
secrets on every push, but nothing read this repo's own source for
risky patterns (weak crypto, injection shapes, insecure cookies, ...).

New `make security-code` runs gosec, wired into ci.yaml's security job
alongside the other two. G104 (unchecked error) is excluded at the
Makefile level: every one of its 41 initial hits was this codebase's
existing, deliberate idiom for a best-effort write or an already-
reviewed json.Unmarshal of its own JSONB, predating gosec, and the rule
cannot tell that apart from a mistake -- seventeen individual #nosec
comments would hide a future real G104 regression in the suppression
noise rather than surface it. Reasoning is on the Makefile target.

Of the 12 remaining hits:
  - Genuinely real: oidc.go's callback logged error_description (and,
    two call sites down, identity.Subject) via %s before the request's
    state was even checked against its cookie -- an attacker-reachable
    value going into the log unquoted. Switched to %q, matching
    identity.Username's existing treatment, so a value holding a
    newline can't forge a second log line.
  - False positives, annotated inline rather than globally suppressed:
    4x G124 on cookies that already set Secure via cookieSecure(...)
    (a function call, not the literal `true` the rule wants), 3x G202
    on sqlArgs-built queries that only ever splice in a "$N"
    placeholder, never a value, and the remaining 5x G706 on log lines
    that were already %q-quoted -- gosec's taint analysis doesn't
    model format verbs, so it flags the tainted argument regardless.

Also fixed handleMe's swallowed Scan error (gosec's catch, pre-fix):
a transient DB error left hash/dismissed at their zero values and the
response claimed no password and no onboarding dismissal regardless
of the truth, rather than surfacing a 500.

Checked both workflow files for the injection class letsvisit found
there (a `${{ }}` expression spliced straight into a `run:` block):
every one here already goes through `env:` as a quoted shell variable,
documented in ci.yaml's own header comment. Nothing to fix.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:12:31 +02:00
Niklas Ye b82c10acf4 Back the login/signup/OIDC/device rate limiters with Postgres
Part of the same security-hardening pass as the body-size/header commit.
loginLimiter was an in-memory, per-process sync.Mutex+map -- fine for one
replica, but charts/terdut-server/values.yaml has set replicaCount: 2 in
production since v0.37.0. Each pod counted only its own traffic, so every
limit it guarded (failed logins, sign-ups, OIDC/device-login starts) was
effectively twice as generous as the constants say, not just in theory.

loginLimiter now stores its counters in a new rate_limit_counters table
(migration 016) instead of a map; blocked/fail/clear take a context and
query/upsert/delete a row keyed by the same strings callers already used
(username, client address, "signup:"+address, ...). Semantics are
unchanged -- a fixed window that resets rather than slides -- so no call
site's behavior changes, only where the count lives. Added purgeRateLimits
to the sweeper, alongside purgeSessions/purgeAckTokens, so expired windows
don't accumulate.

New internal (package api) tests in rate_limiter_test.go cover the basic
behavior plus the regression this exists to fix: two loginLimiter values
sharing one database, standing in for two replicas, now see one combined
count instead of each keeping their own.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:03:51 +02:00
Niklas Ye 7cd6fbf571 Cap request body size and add baseline security headers
Part of a security-hardening pass (see wiki for the full backlog).
decodeJSON had no size limit at all, so every JSON endpoint -- including
the two unauthenticated ones (bootstrap, the Alertmanager webhook) --
would buffer an attacker-supplied body of unbounded size before it was
even validated. decodeJSON now wraps the body in http.MaxBytesReader at
a 1 MiB default; the webhook gets its own 8 MiB cap via decodeJSONLimit,
since a real Alertmanager batch can be bigger than an ordinary API body.

Also adds a securityHeaders middleware, applied globally: nosniff on
every response (previously only the static site got it), and HSTS
(180-day max-age, conservative on purpose) whenever cookieSecure's
signal says the browser is on HTTPS. Checked the chart/gateway config
first -- neither sets HSTS anywhere, so this was a real gap.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 21:57:00 +02:00
Niklas Ye df83adfe47 Set the chart's placeholder version to 0.37.2
CI / chart (push) Successful in 1s
CI / security (push) Successful in 20s
CI / test (push) Successful in 5m28s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 49s
Release / image (push) Successful in 1m17s
Release / scan-image (push) Successful in 26s
2026-10-07 21:12:51 +02:00
Niklas Ye 9da913080f Stop 500ing when a service account acts on an incident
Every incident-mutation handler read userFromContext(ctx) and wrote the
result's .ID into acknowledged_by/incident_events.user_id without checking
the ok bool. For a team-scoped service-account caller this returned a
zero-value user id, which violated the users(id) FK and 500'd on
acknowledge, unacknowledge, resolve, snooze, unsnooze and create-note.
handleDeleteNote didn't crash but silently matched zero rows instead
(WHERE user_id = 0), so a service account could never delete its own note.

Add acknowledged_by_service_account_id (incidents) and service_account_id
(incident_events) as nullable FKs to service_accounts(id), parallel to and
mutually exclusive with the existing human columns (migration 015, with a
CHECK enforcing the exclusion). Route every one of the six handlers plus
delete-note through a new callerActorIDs() helper that branches on
Caller.AsHuman()/ServiceAccountID() instead of assuming a human, and thread
a serviceAccountID parameter through logEvent and the new
acknowledgeIncidentAs (acknowledgeIncident itself is untouched: its only
other caller, the push-notification Acknowledge button, is always human).
Render the new actor distinctly from both a human and "the server acted"
in the web UI's incident timeline and facts card.

handleIncidentAssign, handleIncidentArchive and handleIncidentUnarchive are
deliberately not touched here — they track no actor at all today, for
anyone, which is a separate pre-existing gap (follow-up issue to come).

Fixes #25
2026-10-07 21:09:31 +02:00
36 changed files with 1442 additions and 139 deletions
+3
View File
@@ -125,6 +125,9 @@ jobs:
- name: Secret scan (gitleaks)
run: make security-secrets
- name: Code security scan (gosec)
run: make security-code
# Host mode, no `container:`: helm is baked into the runner image, and a container job
# could not install it -- get.helm.sh is unreachable from the dind bridge. Same reason
# release.yaml's chart job runs on the host.
+6
View File
@@ -24,4 +24,10 @@ FROM scratch
COPY --from=builder /etc/ssl/certs/ca-certificates.crt /etc/ssl/certs/ca-certificates.crt
COPY --from=builder /terdut /terdut
EXPOSE 8080
# Numeric, not a name: scratch has no /etc/passwd for one to resolve against,
# and Docker's USER accepts a bare UID:GID without it. 65532 is the common
# "nonroot" convention (distroless's own uid), chosen so the chart's pod
# securityContext (runAsNonRoot, runAsUser: 65532) matches what the image
# already runs as rather than fighting it.
USER 65532:65532
ENTRYPOINT ["/terdut"]
+21
View File
@@ -143,6 +143,7 @@ BUILDX_BUILDER ?= terdut
TRIVY_VERSION := 0.73.0
GOVULNCHECK_VERSION := v1.1.4
GITLEAKS_VERSION := v8.30.0
GOSEC_VERSION := v2.29.0
# --pull, not --no-cache: refresh the base image without discarding the layer cache.
DOCKER_BUILD_FLAGS ?= --pull
@@ -222,6 +223,26 @@ release: push helm-package helm-push ## Publish image + chart (the workflow's on
security-go: ## Scan Go deps for known CVEs (govulncheck)
go run golang.org/x/vuln/cmd/govulncheck@$(GOVULNCHECK_VERSION) ./...
# Code-level, not dependency- or secret-level: gosec reads this repo's own source for
# known-dangerous patterns (weak crypto, SQL/command injection shapes, insecure file
# permissions, …) rather than its module graph or working tree for leaked credentials,
# which is what security-go and security-secrets above already cover.
#
# G104 (unchecked error) is excluded. Every hit it found here on first run was this
# codebase's existing, deliberate idiom for a best-effort write or an already-reviewed
# json.Unmarshal of this server's own JSONB (see the "best-effort" comments in
# middleware.go and the //nolint:errcheck lines in alerts.go/deadman.go) -- a style that
# predates gosec and that G104 cannot distinguish from a mistake. Reaching the same
# green result by adding a dozens of individual #nosec comments would not add
# information; it would just make a future *real* G104 regression one more suppressed
# line instead of a visible one. Same reasoning as the chi-advisories note on
# security-go above: what gosec reports here (nothing, beyond G104) is the useful
# property, not a loophole. -exclude-generated skips web.go's embedded, build-time-only
# assets.
.PHONY: security-code
security-code: ## Scan this repo's own source for risky patterns (gosec)
go run github.com/securego/gosec/v2/cmd/gosec@$(GOSEC_VERSION) -exclude-generated -exclude=G104 ./...
# --no-git scans the working tree rather than the history, so this catches a secret on the
# way in. It is not a history audit and finding nothing here says nothing about what is
# already committed. --redact because the finding is printed into a CI log.
+5 -3
View File
@@ -984,9 +984,11 @@ name: degrade unknown values to "resolved, reason unknown".
| `created_at` | timestamp | |
Types written today: `triggered`, `alert_added`, `alert_resolved`,
`acknowledged`, `unacknowledged`, `assigned`, `snoozed`, `unsnoozed`, `resolved`,
`note`, `notified`, `notify_failed`, `deadman_silent`. On an `assigned` event
`user_id` is the **assignee**, not the actor. New types may be added; render
`acknowledged`, `unacknowledged`, `assigned`, `archived`, `unarchived`, `snoozed`,
`unsnoozed`, `resolved`, `note`, `notified`, `notify_failed`, `deadman_silent`. On an
`assigned` event `user_id` is the **assignee**, not the actor; the actor is in
`actor_user_id`/`actor_username` or `actor_service_account_id`/`actor_service_account_name`
(absent on assignments made before they were recorded). New types may be added; render
unknown ones generically rather than dropping them.
On `notified` and `notify_failed`, `detail` carries the notification kind
+2 -2
View File
@@ -15,5 +15,5 @@ type: application
# appVersion and image.tag in values.yaml no longer agree, and that is not an oversight:
# image.tag stays "latest", which is what a local install actually pulls. appVersion is
# metadata and drives nothing.
version: 0.37.1
appVersion: "v0.37.1"
version: 0.39.0
appVersion: "v0.39.0"
@@ -25,11 +25,30 @@ spec:
{{- include "terdut-server.selectorLabels" . | nindent 8 }}
spec:
enableServiceLinks: false
# Pod-wide default; both containers below run as this UID regardless of
# what their own image would otherwise pick (postgres:17-alpine's
# pg_isready needs no particular user, and 65532 is what the app image
# itself runs as now — see the Dockerfile's USER). seccompProfile here
# rather than per-container: there is no reason it would ever differ
# between them.
securityContext:
runAsNonRoot: true
runAsUser: 65532
runAsGroup: 65532
seccompProfile:
type: RuntimeDefault
{{- if .Values.database.waitForPostgres.enabled }}
initContainers:
- name: wait-for-postgres
image: "{{ .Values.database.waitForPostgres.image.repository }}:{{ .Values.database.waitForPostgres.image.tag }}"
imagePullPolicy: {{ .Values.database.waitForPostgres.image.pullPolicy }}
# No capability this loop needs, and nothing in it writes to disk:
# sh, pg_isready, echo and sleep all run read-only.
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
env:
- name: TERDUT_DB_DSN
value: {{ required "database.dsn is required" .Values.database.dsn | quote }}
@@ -46,6 +65,13 @@ spec:
- name: terdut-server
image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
imagePullPolicy: {{ .Values.image.pullPolicy }}
# scratch, nothing to write: the binary keeps no local state and
# writes nothing to disk, so the root filesystem can stay read-only.
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
ports:
- name: http
containerPort: {{ .Values.service.port }}
+11 -5
View File
@@ -94,10 +94,16 @@ func handleIntegrationWebhook(db *sql.DB, notify NotifyConfig) http.HandlerFunc
}
}
// maxWebhookBodyBytes is larger than maxBodyBytes: a real Alertmanager batch
// can carry many alerts, each with several labels and annotations, and the
// sender is a trusted piece of infrastructure rather than an arbitrary
// caller.
const maxWebhookBodyBytes = 8 << 20
func receiveWebhook(w http.ResponseWriter, r *http.Request, db *sql.DB, notify NotifyConfig, src alertSource) {
teamID := src.teamID
var payload amPayload
if err := decodeJSON(r, &payload); err != nil {
if err := decodeJSONLimit(r, &payload, maxWebhookBodyBytes); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid payload"))
return
}
@@ -169,7 +175,7 @@ func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, src alertSourc
}
touched[id] = true
alertID := a.id
if err := logEvent(ctx, tx, id, evAlertResolved, nil, &alertID, nil); err != nil {
if err := logEvent(ctx, tx, id, evAlertResolved, nil, nil, &alertID, nil); err != nil {
return err
}
}
@@ -417,12 +423,12 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID in
return 0, err
}
if err := logEvent(ctx, q, id, evTriggered, nil, nil, nil); err != nil {
if err := logEvent(ctx, q, id, evTriggered, nil, nil, nil, nil); err != nil {
return 0, err
}
if onCall != nil {
// On an "assigned" event user_id is the assignee, not the actor.
if err := logEvent(ctx, q, id, evAssigned, onCall, nil, nil); err != nil {
if err := logEvent(ctx, q, id, evAssigned, onCall, nil, nil, nil); err != nil {
return 0, err
}
}
@@ -470,5 +476,5 @@ func linkAlert(ctx context.Context, tx *sql.Tx, incidentID, alertID int64) error
if n, _ := res.RowsAffected(); n == 0 {
return nil
}
return logEvent(ctx, tx, incidentID, evAlertAdded, nil, &alertID, nil)
return logEvent(ctx, tx, incidentID, evAlertAdded, nil, nil, &alertID, nil)
}
+135
View File
@@ -0,0 +1,135 @@
package api_test
import (
"net/http"
"testing"
"time"
)
func TestAPIKey_DefaultsToNeverExpiring(t *testing.T) {
s := newTS(t)
var key struct {
Key string `json:"key"`
ExpiresAt *string `json:"expires_at"`
}
decode(t, s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]string{"name": "no-expiry"}), &key)
if key.ExpiresAt != nil {
t.Errorf("expires_at = %v, want nil (unset expires_in_days means never expires)", *key.ExpiresAt)
}
}
func TestAPIKey_ExpiresInDaysSetsExpiresAt(t *testing.T) {
s := newTS(t)
var key struct {
ID int64 `json:"id"`
ExpiresAt *string `json:"expires_at"`
}
decode(t, s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]any{"name": "rotates", "expires_in_days": 30}), &key)
if key.ExpiresAt == nil {
t.Fatal("expires_at = nil, want a timestamp roughly 30 days out")
}
got, err := time.Parse(time.RFC3339, *key.ExpiresAt)
if err != nil {
t.Fatalf("parse expires_at: %v", err)
}
want := time.Now().AddDate(0, 0, 30)
if diff := want.Sub(got).Abs(); diff > time.Hour {
t.Errorf("expires_at = %v, want close to %v (30 days out)", got, want)
}
}
func TestAPIKey_ExpiresInDaysRejectsOutOfRange(t *testing.T) {
s := newTS(t)
for _, days := range []int{-1, 3651} {
resp := s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]any{"name": "bad", "expires_in_days": days})
resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("expires_in_days=%d: status = %d, want %d", days, resp.StatusCode, http.StatusBadRequest)
}
}
}
func TestAPIKey_AnExpiredKeyCannotAuthenticate(t *testing.T) {
s := newTS(t)
var key struct {
ID int64 `json:"id"`
Key string `json:"key"`
}
decode(t, s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]any{"name": "soon-expired", "expires_in_days": 1}), &key)
// A fresh key works...
req, _ := http.NewRequest(http.MethodGet, s.URL+"/api/me", nil)
req.Header.Set("Authorization", "Bearer "+key.Key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET /api/me: %v", err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("fresh key: status = %d, want %d", resp.StatusCode, http.StatusOK)
}
// ...and stops working once its expiry has passed.
s.exec(t, "UPDATE api_keys SET expires_at = $1 WHERE id = $2", time.Now().Add(-time.Hour).Unix(), key.ID)
req2, _ := http.NewRequest(http.MethodGet, s.URL+"/api/me", nil)
req2.Header.Set("Authorization", "Bearer "+key.Key)
resp2, err := http.DefaultClient.Do(req2)
if err != nil {
t.Fatalf("GET /api/me: %v", err)
}
defer resp2.Body.Close()
if resp2.StatusCode != http.StatusUnauthorized {
t.Errorf("expired key: status = %d, want %d", resp2.StatusCode, http.StatusUnauthorized)
}
}
func TestAPIKey_ListNeverReturnsTheRawKey(t *testing.T) {
s := newTS(t)
decode(t, s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]string{"name": "listed"}), new(struct {
Key string `json:"key"`
}))
var keys []struct {
ID int64 `json:"id"`
Name string `json:"name"`
Key string `json:"key"`
}
decode(t, s.req(t, http.MethodGet, "/api/users/1/api-keys", nil), &keys)
found := false
for _, k := range keys {
if k.Name == "listed" {
found = true
}
if k.Key != "" {
t.Errorf("key %d (%s): raw key present in listing", k.ID, k.Name)
}
}
if !found {
t.Error("the key just created does not appear in the listing")
}
}
func TestAPIKey_ListIsSelfOrAdmin(t *testing.T) {
s := newTS(t)
a := newTeam(t, s, "apikeys-a")
resp := a.call(http.MethodGet, "/api/users/1/api-keys", nil)
defer resp.Body.Close()
if resp.StatusCode != http.StatusForbidden {
t.Errorf("status = %d, want %d (not self, not an admin)", resp.StatusCode, http.StatusForbidden)
}
}
+6 -1
View File
@@ -76,6 +76,7 @@ func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Durati
archiveResolvedIncidents(ctx, db, archiveAfter)
purgeAckTokens(ctx, db)
purgeSessions(ctx, db)
purgeRateLimits(ctx, db)
}
// expireStale resolves firing alerts that Alertmanager has stopped refreshing.
@@ -121,6 +122,10 @@ func expireStale(ctx context.Context, db *sql.DB, staleAfter time.Duration, skip
for i, id := range ids {
idList[i] = id
}
// #nosec G202 -- sqlArgs.add/addList only ever splice in the "$N"
// placeholder they hand back, never a value; every value travels through
// args.all() as a bound parameter. See the sqlArgs doc comment in
// helpers.go.
if _, err := db.ExecContext(ctx, `
UPDATE alerts
SET status = 'resolved',
@@ -142,7 +147,7 @@ func expireStale(ctx context.Context, db *sql.DB, staleAfter time.Duration, skip
continue
}
alertID := id
if err := logEvent(ctx, db, incidentID, evAlertResolved, nil, &alertID, nil); err != nil {
if err := logEvent(ctx, db, incidentID, evAlertResolved, nil, nil, &alertID, nil); err != nil {
log.Printf("sweeper: log expiry event: %v", err)
}
}
+83 -42
View File
@@ -45,57 +45,85 @@ var dummyHash = sync.OnceValue(func() []byte {
return h
})
// loginLimiter counts failed logins in a fixed window, per username and per
// client address. The username limit is what stops guessing one account; the
// address limit is looser because every user behind the same gateway or NAT
// shares it.
// loginLimiter counts failed logins (and other unauthenticated attempts:
// sign-up, OIDC/device start) in a fixed window, per key — a username, a
// client address, or both, depending on the caller.
//
// Backed by Postgres rather than an in-memory map: this server runs more
// than one replica in production (v0.37.0), and a counter that only ever
// sees its own pod's traffic would quietly let every limit through
// multiplied by the replica count — two loginLimiter values pointed at the
// same db, standing in for two replicas, now share exactly one count per
// key instead of each keeping their own.
//
// The window resets rather than slides, the same behavior the in-memory
// version it replaces had: once a key's window is older than loginWindow,
// the next fail() starts a fresh one instead of extending the stale one.
type loginLimiter struct {
mu sync.Mutex
failures map[string]*loginWindowCount
db *sql.DB
}
type loginWindowCount struct {
start time.Time
n int
func newLoginLimiter(db *sql.DB) *loginLimiter {
return &loginLimiter{db: db}
}
func newLoginLimiter() *loginLimiter {
return &loginLimiter{failures: map[string]*loginWindowCount{}}
}
func (l *loginLimiter) blocked(key string, max int) bool {
l.mu.Lock()
defer l.mu.Unlock()
c, ok := l.failures[key]
if !ok || time.Since(c.start) > loginWindow {
func (l *loginLimiter) blocked(ctx context.Context, key string, max int) bool {
cutoff := time.Now().Unix() - int64(loginWindow.Seconds())
var count int
err := l.db.QueryRowContext(ctx, `
SELECT count FROM rate_limit_counters
WHERE key = $1 AND window_start > $2`,
key, cutoff,
).Scan(&count)
if err != nil {
// No row (never failed, or its window already expired): not blocked.
// A real query error fails the same way — a rate limiter that locks
// everyone out during a brief database hiccup is worse than one that
// is briefly too generous.
return false
}
return c.n >= max
return count >= max
}
func (l *loginLimiter) fail(keys ...string) {
l.mu.Lock()
defer l.mu.Unlock()
now := time.Now()
for k, c := range l.failures {
if now.Sub(c.start) > loginWindow {
delete(l.failures, k)
}
}
func (l *loginLimiter) fail(ctx context.Context, keys ...string) {
now := time.Now().Unix()
windowSecs := int64(loginWindow.Seconds())
for _, key := range keys {
c, ok := l.failures[key]
if !ok {
c = &loginWindowCount{start: now}
l.failures[key] = c
if _, err := l.db.ExecContext(ctx, `
INSERT INTO rate_limit_counters (key, window_start, count)
VALUES ($1, $2, 1)
ON CONFLICT (key) DO UPDATE SET
window_start = CASE WHEN rate_limit_counters.window_start <= $2 - $3
THEN $2 ELSE rate_limit_counters.window_start END,
count = CASE WHEN rate_limit_counters.window_start <= $2 - $3
THEN 1 ELSE rate_limit_counters.count + 1 END`,
key, now, windowSecs,
); err != nil {
log.Printf("rate limiter: record failure for %q: %v", key, err) // #nosec G706 -- %q
}
c.n++
}
}
func (l *loginLimiter) clear(key string) {
l.mu.Lock()
defer l.mu.Unlock()
delete(l.failures, key)
func (l *loginLimiter) clear(ctx context.Context, key string) {
if _, err := l.db.ExecContext(ctx, "DELETE FROM rate_limit_counters WHERE key = $1", key); err != nil {
log.Printf("rate limiter: clear %q: %v", key, err)
}
}
// purgeRateLimits deletes rate-limit windows that have expired, from the
// sweeper — otherwise every distinct username and address this server has
// ever seen a failed attempt from would stay a row forever.
func purgeRateLimits(ctx context.Context, db *sql.DB) {
cutoff := time.Now().Unix() - int64(loginWindow.Seconds())
res, err := db.ExecContext(ctx,
"DELETE FROM rate_limit_counters WHERE window_start <= $1", cutoff)
if err != nil {
log.Printf("sweeper: purge rate limit counters: %v", err)
return
}
if n, _ := res.RowsAffected(); n > 0 {
log.Printf("sweeper: purged %d expired rate limit counter(s)", n)
}
}
// clientAddr is the address a login is counted against. Behind the gateway
@@ -171,6 +199,9 @@ func startSessionCapped(w http.ResponseWriter, r *http.Request, db *sql.DB, user
return err
}
// #nosec G124 -- HttpOnly/SameSite are literal below; Secure is
// cookieSecure(publicURL, r), not a literal true, which is what trips
// this rule. See cookieSecure's own doc comment above.
http.SetCookie(w, &http.Cookie{
Name: sessionCookie,
Value: raw,
@@ -198,7 +229,7 @@ func handleLogin(db *sql.DB, limiter *loginLimiter, publicURL string) http.Handl
userKey := "user:" + strings.ToLower(username)
addrKey := "addr:" + clientAddr(r)
if limiter.blocked(userKey, loginMaxPerUser) || limiter.blocked(addrKey, loginMaxPerAddr) {
if limiter.blocked(r.Context(), userKey, loginMaxPerUser) || limiter.blocked(r.Context(), addrKey, loginMaxPerAddr) {
w.Header().Set("Retry-After", strconv.Itoa(int(loginWindow.Seconds())))
respond(w, http.StatusTooManyRequests, errResp("too many failed attempts, try again later"))
return
@@ -220,11 +251,11 @@ func handleLogin(db *sql.DB, limiter *loginLimiter, publicURL string) http.Handl
}
match := bcrypt.CompareHashAndPassword(stored, []byte(req.Password)) == nil
if !match || !hash.Valid {
limiter.fail(userKey, addrKey)
limiter.fail(r.Context(), userKey, addrKey)
respond(w, http.StatusUnauthorized, errResp("invalid username or password"))
return
}
limiter.clear(userKey)
limiter.clear(r.Context(), userKey)
if err := startSession(w, r, db, userID, publicURL); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
@@ -252,6 +283,9 @@ func handleLogout(db *sql.DB, publicURL string) http.HandlerFunc {
if c, err := r.Cookie(sessionCookie); err == nil && c.Value != "" {
db.ExecContext(r.Context(), "DELETE FROM sessions WHERE token_hash = $1", hashToken(c.Value))
}
// #nosec G124 -- HttpOnly/SameSite are literal below; Secure is
// cookieSecure(publicURL, r), not a literal true, which is what
// trips this rule. See cookieSecure's own doc comment above.
http.SetCookie(w, &http.Cookie{
Name: sessionCookie,
Value: "",
@@ -291,9 +325,16 @@ func handleMe(db *sql.DB) http.HandlerFunc {
}
var hash sql.NullString
var dismissed *int64
db.QueryRowContext(r.Context(),
if err := db.QueryRowContext(r.Context(),
"SELECT password_hash, onboarding_dismissed_at FROM users WHERE id = $1",
caller.ID).Scan(&hash, &dismissed)
caller.ID).Scan(&hash, &dismissed); err != nil {
// fetchUser above already found this row, so an error here is a
// transient database problem, not a missing user — worth a 500
// rather than silently answering "no password, not dismissed",
// which a client would otherwise take at face value.
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, meResponse{
User: user,
HasPassword: hash.Valid,
+182
View File
@@ -0,0 +1,182 @@
package api_test
import (
"net/http"
"testing"
)
// This file is the regression test for the pattern documented throughout
// middleware.go: every team-scoped handler calls requireTeamMember or
// requireTeamOwner before touching data, every self-or-admin handler calls
// requireSelfOrAdmin, and every admin-only route sits behind AdminOnly. That
// pattern is enforced by convention, not by the type system — a new handler
// that forgets the call would compile and pass review on a quick read just
// as easily as one that remembers it. These tests exercise every route that
// carries one of those guards as a caller who should be refused, so a future
// handler missing its guard fails CI instead of becoming a silent IDOR.
// TestAuthzScope_TeamScopedRoutesRefuseANonMember builds two teams and, for
// every team-scoped route, calls it as team A's owner against team B's
// resources. requireTeamMember and requireTeamOwner both answer a non-member
// with 404 (team.go's own reasoning: whether a team exists is itself
// something only its members should learn), so every one of these must come
// back 404 regardless of which of the two guards its handler uses.
func TestAuthzScope_TeamScopedRoutesRefuseANonMember(t *testing.T) {
s := newTS(t)
a := newTeam(t, s, "authz-a")
b := newTeam(t, s, "authz-b")
// An incident in B, to cover the ID-based routes under /api/incidents —
// scoped by the incident's own team_id rather than a {teamID} path
// segment, but through the same single chokepoint (incidentIDParam).
postToIntegration(t, s, b.key, "fp-authz-scope", "AuthzScopeAlert")
var incidents []struct {
ID int64 `json:"id"`
}
decode(t, b.call(http.MethodGet, "/api/incidents", nil), &incidents)
if len(incidents) == 0 {
t.Fatal("setup: no incident in team B to test against")
}
incidentPath := "/api/incidents/" + id64(incidents[0].ID)
bPath := "/api/teams/" + id64(b.id)
tests := []struct {
method, path string
}{
// Team membership/ownership itself.
{http.MethodPut, bPath},
{http.MethodDelete, bPath},
{http.MethodGet, bPath + "/members"},
{http.MethodPost, bPath + "/members"},
{http.MethodDelete, bPath + "/members/1"},
// OIDC group binding.
{http.MethodGet, bPath + "/oidc-groups"},
{http.MethodPut, bPath + "/oidc-groups"},
// Invites.
{http.MethodGet, bPath + "/invites"},
{http.MethodPost, bPath + "/invites"},
{http.MethodDelete, bPath + "/invites/1"},
// Escalation.
{http.MethodGet, bPath + "/escalation"},
{http.MethodPut, bPath + "/escalation"},
// Dead man's switches.
{http.MethodGet, bPath + "/deadman/switches"},
{http.MethodPost, bPath + "/deadman/switches"},
{http.MethodPut, bPath + "/deadman/switches/1"},
{http.MethodDelete, bPath + "/deadman/switches/1"},
// Integrations.
{http.MethodGet, bPath + "/integrations"},
{http.MethodPost, bPath + "/integrations"},
{http.MethodPatch, bPath + "/integrations/1"},
{http.MethodDelete, bPath + "/integrations/1"},
// Schedule.
{http.MethodGet, bPath + "/schedule"},
{http.MethodPost, bPath + "/schedule"},
{http.MethodDelete, bPath + "/schedule/1"},
// Incidents, scoped by the incident's own team rather than a
// {teamID} segment.
{http.MethodGet, incidentPath},
{http.MethodGet, incidentPath + "/alerts"},
{http.MethodGet, incidentPath + "/timeline"},
{http.MethodGet, incidentPath + "/similar"},
{http.MethodPost, incidentPath + "/acknowledge"},
{http.MethodDelete, incidentPath + "/acknowledge"},
{http.MethodPost, incidentPath + "/resolve"},
{http.MethodPost, incidentPath + "/assign"},
{http.MethodPost, incidentPath + "/snooze"},
{http.MethodDelete, incidentPath + "/snooze"},
{http.MethodPost, incidentPath + "/archive"},
{http.MethodDelete, incidentPath + "/archive"},
{http.MethodPost, incidentPath + "/notes"},
{http.MethodDelete, incidentPath + "/notes/1"},
}
for _, tc := range tests {
t.Run(tc.method+" "+tc.path, func(t *testing.T) {
resp := a.call(tc.method, tc.path, nil)
defer resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("status = %d, want %d (A is not a member of B)", resp.StatusCode, http.StatusNotFound)
}
})
}
}
// TestAuthzScope_AdminOnlyRoutesRefuseANonAdmin exercises AdminOnly's group
// in router.go directly: a signed-in, non-admin caller gets 403 from every
// route in it, before any handler body runs.
func TestAuthzScope_AdminOnlyRoutesRefuseANonAdmin(t *testing.T) {
s := newTS(t)
a := newTeam(t, s, "authz-admin")
tests := []struct {
method, path string
}{
{http.MethodPost, "/api/users"},
{http.MethodDelete, "/api/users/1"},
{http.MethodPut, "/api/users/1/admin"},
{http.MethodPut, "/api/users/1/disabled"},
{http.MethodGet, "/api/admin/teams"},
{http.MethodGet, "/api/admin/teams/" + id64(a.id)},
{http.MethodGet, "/api/admin/settings"},
{http.MethodPut, "/api/admin/settings"},
}
for _, tc := range tests {
t.Run(tc.method+" "+tc.path, func(t *testing.T) {
resp := a.call(tc.method, tc.path, nil)
defer resp.Body.Close()
if resp.StatusCode != http.StatusForbidden {
t.Errorf("status = %d, want %d (not an admin)", resp.StatusCode, http.StatusForbidden)
}
})
}
}
// TestAuthzScope_SelfOrAdminRoutesRefuseAnotherNonAdminUser exercises
// requireSelfOrAdmin's call sites: a non-admin caller acting on a *different*
// user's account must be refused, the same as AdminOnly's routes, even
// though these sit in the general authenticated group rather than behind
// AdminOnly itself.
func TestAuthzScope_SelfOrAdminRoutesRefuseAnotherNonAdminUser(t *testing.T) {
s := newTS(t)
a := newTeam(t, s, "authz-self-a")
b := newTeam(t, s, "authz-self-b")
var members []struct {
UserID int64 `json:"user_id"`
}
decode(t, s.req(t, http.MethodGet, "/api/teams/"+id64(b.id)+"/members", nil), &members)
if len(members) == 0 {
t.Fatal("setup: team B has no members")
}
bUserID := id64(members[0].UserID)
tests := []struct {
method, path string
}{
{http.MethodGet, "/api/users/" + bUserID + "/teams"},
{http.MethodPut, "/api/users/" + bUserID + "/notify"},
{http.MethodPut, "/api/users/" + bUserID + "/password"},
{http.MethodGet, "/api/users/" + bUserID + "/api-keys"},
{http.MethodPost, "/api/users/" + bUserID + "/api-keys"},
{http.MethodDelete, "/api/users/" + bUserID + "/api-keys/1"},
}
for _, tc := range tests {
t.Run(tc.method+" "+tc.path, func(t *testing.T) {
resp := a.call(tc.method, tc.path, nil)
defer resp.Body.Close()
if resp.StatusCode != http.StatusForbidden {
t.Errorf("status = %d, want %d (not self, not an admin)", resp.StatusCode, http.StatusForbidden)
}
})
}
}
+16 -5
View File
@@ -99,13 +99,24 @@ func (c Caller) ServiceAccountID() (int64, bool) {
return c.sa.id, true
}
// ServiceAccountName reports this caller's own service-account name, for a
// handler's synchronous response — the same credential it authenticated
// with, already resolved onto the Caller by serveAsServiceAccount, so no
// extra query is needed.
func (c Caller) ServiceAccountName() (string, bool) {
if c.sa == nil {
return "", false
}
return c.sa.name, true
}
// Identity is a stable, log/audit-facing string distinguishing a human
// caller from a service account — "user:42" or "service-account:7". Not
// wired into any database column today (incidents.go's acknowledged_by/
// assigned_to/user_id are explicitly out of scope for this change — that
// needs its own schema migration, tracked separately), but this is the one
// place in the request path that already knows which kind of caller this
// is, and that follow-up will want exactly this accessor.
// wired into any database column — incidents.go's acknowledged_by/
// incident_events.user_id use AsHuman()/ServiceAccountID() directly against
// the parallel *_service_account_id columns (migration 015) instead, since a
// column needs the id, not this rendered string. assigned_to stays
// human-only and out of scope (terdut-server#25's follow-up).
func (c Caller) Identity() string {
switch {
case c.user != nil:
+6 -2
View File
@@ -283,6 +283,10 @@ func deadmanAlerts(ctx context.Context, db *sql.DB, teamID int64, cfg deadmanSet
nameList[i] = n
}
// #nosec G202 -- sqlArgs.add/addList only ever splice in the "$N"
// placeholder they hand back, never a value; every value travels through
// args.all() as a bound parameter. See the sqlArgs doc comment in
// helpers.go.
rows, err := db.QueryContext(ctx, `
SELECT id, team_id, fingerprint, labels, status, received_at
FROM alerts
@@ -373,7 +377,7 @@ func deadmanDied(ctx context.Context, db *sql.DB, notify NotifyConfig, hb deadma
alertID := hb.id
detail := "last heartbeat " + humanDuration(now.Sub(time.Unix(hb.receivedAt, 0))) + " ago"
if err := logEvent(ctx, tx, incidentID, evDeadmanSilent, nil, &alertID, &detail); err != nil {
if err := logEvent(ctx, tx, incidentID, evDeadmanSilent, nil, nil, &alertID, &detail); err != nil {
return err
}
@@ -415,7 +419,7 @@ func deadmanRecovered(ctx context.Context, db *sql.DB, hb deadmanAlert) error {
time.Now().Unix(), incidentResolutionRecovered, incidentID); err != nil {
return err
}
if err := logEvent(ctx, tx, incidentID, evResolved, nil, nil, nil); err != nil {
if err := logEvent(ctx, tx, incidentID, evResolved, nil, nil, nil, nil); err != nil {
return err
}
// The all-clear goes to whoever was paged, which enqueueResolved works out
+2 -2
View File
@@ -72,12 +72,12 @@ func normalizeUserCode(s string) string {
func handleDeviceStart(db *sql.DB, limiter *loginLimiter, publicURL string) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
addrKey := "device:" + clientAddr(r)
if limiter.blocked(addrKey, deviceStartMaxPerAddr) {
if limiter.blocked(r.Context(), addrKey, deviceStartMaxPerAddr) {
w.Header().Set("Retry-After", strconv.Itoa(int(loginWindow.Seconds())))
respond(w, http.StatusTooManyRequests, errResp("too many sign-in attempts, try again later"))
return
}
limiter.fail(addrKey)
limiter.fail(r.Context(), addrKey)
deviceCode, deviceHash, err := randomToken()
if err != nil {
+2 -2
View File
@@ -229,7 +229,7 @@ func advanceEscalation(ctx context.Context, db *sql.DB, cfg NotifyConfig, policy
// nobody. That is a policy that looks configured and is not.
detail += ": nobody reachable"
}
if err := logEvent(ctx, tx, incidentID, evEscalated, nil, nil, &detail); err != nil {
if err := logEvent(ctx, tx, incidentID, evEscalated, nil, nil, nil, &detail); err != nil {
return err
}
return tx.Commit()
@@ -256,7 +256,7 @@ func escalationExhausted(ctx context.Context, tx *sql.Tx, policy *escalationPoli
incidentID); err != nil {
return err
}
return logEvent(ctx, tx, incidentID, evEscalated, nil, nil, &detail)
return logEvent(ctx, tx, incidentID, evEscalated, nil, nil, nil, &detail)
}
// pageLevel notifies every target of one level and reports who was woken.
+17
View File
@@ -72,8 +72,25 @@ func respond(w http.ResponseWriter, status int, v any) {
json.NewEncoder(w).Encode(v)
}
// maxBodyBytes caps an ordinary JSON request body. 1 MiB is far more than any
// endpoint below needs — it exists so an unauthenticated caller (signup,
// login, bootstrap) can't make the server buffer an arbitrarily large body
// before the request is even validated.
const maxBodyBytes = 1 << 20
func decodeJSON(r *http.Request, v any) error {
return decodeJSONLimit(r, v, maxBodyBytes)
}
// decodeJSONLimit is decodeJSON with an explicit cap, for the one endpoint
// (the Alertmanager webhook, see maxWebhookBodyBytes) whose real payloads can
// legitimately be larger than maxBodyBytes.
func decodeJSONLimit(r *http.Request, v any, limit int64) error {
defer r.Body.Close()
// w is nil: there is no ResponseWriter here to disable keep-alive with,
// which net/http documents as fine — the limit is still enforced, the
// connection just isn't closed early on a request that blows past it.
r.Body = http.MaxBytesReader(nil, r.Body, limit)
return json.NewDecoder(r.Body).Decode(v)
}
+62 -17
View File
@@ -32,6 +32,8 @@ const (
evAcknowledged = "acknowledged"
evUnacknowledged = "unacknowledged"
evAssigned = "assigned"
evArchived = "archived"
evUnarchived = "unarchived"
evSnoozed = "snoozed"
evUnsnoozed = "unsnoozed"
evResolved = "resolved"
@@ -65,11 +67,13 @@ const incidentSelectFrom = `
WHERE el.team_id = i.team_id AND el.position = i.escalation_level),
i.triggered_at,
i.acknowledged_by, i.acknowledged_at, ack.username,
i.acknowledged_by_service_account_id, acksa.name,
i.assigned_to, asg.username, i.snoozed_until,
i.resolved_at, i.resolution_source, i.archived_at
FROM incidents i
JOIN teams t ON t.id = i.team_id
LEFT JOIN users ack ON ack.id = i.acknowledged_by
LEFT JOIN service_accounts acksa ON acksa.id = i.acknowledged_by_service_account_id
LEFT JOIN users asg ON asg.id = i.assigned_to`
func scanIncident(s scanner) (models.Incident, error) {
@@ -83,6 +87,7 @@ func scanIncident(s scanner) (models.Incident, error) {
&i.EscalationLevel, &escalationDue,
&triggeredAt,
&i.AcknowledgedByID, &ackAt, &i.AcknowledgedByUser,
&i.AcknowledgedByServiceAccountID, &i.AcknowledgedByServiceAccountName,
&i.AssignedToID, &i.AssignedToUser, &snoozedUntil,
&resolvedAt, &i.ResolutionSource, &archivedAt,
); err != nil {
@@ -112,13 +117,42 @@ func fetchIncident(ctx context.Context, q querier, id int64) (models.Incident, e
return scanIncident(q.QueryRowContext(ctx, incidentSelectFrom+" WHERE i.id = $1", id))
}
// logEvent appends one entry to an incident's timeline. A nil userID means the
// server acted rather than a person.
func logEvent(ctx context.Context, q querier, incidentID int64, evType string, userID, alertID *int64, detail *string) error {
// callerActorIDs resolves the current request's caller into the pair of
// nilable ids logEvent/acknowledgeIncidentAs expect: exactly one of userID/
// serviceAccountID is set (never both), replacing the unchecked
// userFromContext(ctx) zero-value reads that used to write a human-only id
// of 0 for a service-account caller (terdut-server#25).
func callerActorIDs(ctx context.Context) (userID, serviceAccountID *int64) {
caller, _ := callerFromContext(ctx)
if u, ok := caller.AsHuman(); ok {
return &u.ID, nil
}
if id, ok := caller.ServiceAccountID(); ok {
return nil, &id
}
return nil, nil
}
// logEvent appends one entry to an incident's timeline. userID and
// serviceAccountID are mutually exclusive and both nilable; both nil means
// the server acted rather than any caller (see incident_events_actor_xor_chk,
// migration 015).
func logEvent(ctx context.Context, q querier, incidentID int64, evType string, userID, serviceAccountID, alertID *int64, detail *string) error {
_, err := q.ExecContext(ctx, `
INSERT INTO incident_events (incident_id, type, user_id, alert_id, detail, created_at)
INSERT INTO incident_events (incident_id, type, user_id, service_account_id, alert_id, detail, created_at)
VALUES ($1, $2, $3, $4, $5, $6, $7)`,
incidentID, evType, userID, serviceAccountID, alertID, detail, time.Now().Unix())
return err
}
// logAssignedEvent records an assignment: user_id is the assignee, and the
// caller who performed it goes in the actor_* columns (migration 018), since
// user_id cannot hold both.
func logAssignedEvent(ctx context.Context, q querier, incidentID, assigneeID int64, actorUserID, actorServiceAccountID *int64) error {
_, err := q.ExecContext(ctx, `
INSERT INTO incident_events (incident_id, type, user_id, actor_user_id, actor_service_account_id, created_at)
VALUES ($1, $2, $3, $4, $5, $6)`,
incidentID, evType, userID, alertID, detail, time.Now().Unix())
incidentID, evAssigned, assigneeID, actorUserID, actorServiceAccountID, time.Now().Unix())
return err
}
@@ -251,7 +285,7 @@ func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, e
if err := stopEscalation(ctx, q, incidentID); err != nil {
return false, err
}
if err := logEvent(ctx, q, incidentID, evResolved, nil, nil, nil); err != nil {
if err := logEvent(ctx, q, incidentID, evResolved, nil, nil, nil, nil); err != nil {
return false, err
}
// The all-clear goes only to whoever was paged in the first place, which
@@ -260,19 +294,30 @@ func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, e
return true, enqueueResolved(ctx, q, incidentID)
}
// acknowledgeIncident records that userID has picked an incident up, and reports
// whether it changed anything — an already-resolved or already-acknowledged
// incident is left alone, so a second acknowledge (a retried request, or a
// stale push notification tapped after the web UI already acked it) is a
// no-op rather than a second "acknowledged" timeline entry. Shared by the
// authenticated handler and the Acknowledge button in a push notification,
// so both write the same state and the same timeline entry.
// acknowledgeIncident records that userID — a human — has picked an incident
// up, and reports whether it changed anything — an already-resolved or
// already-acknowledged incident is left alone, so a second acknowledge (a
// retried request, or a stale push notification tapped after the web UI
// already acked it) is a no-op rather than a second "acknowledged" timeline
// entry. Used only by the Acknowledge button in a push notification
// (notify_ack.go), which always resolves a human from
// incident_ack_tokens.user_id — there is no service-account equivalent of
// that flow, so this keeps its human-only signature; the authenticated
// handler goes through acknowledgeIncidentAs below instead.
func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int64) (bool, error) {
return acknowledgeIncidentAs(ctx, q, incidentID, &userID, nil)
}
// acknowledgeIncidentAs is acknowledgeIncident generalized to either actor
// kind. userID and serviceAccountID are mutually exclusive and nilable the
// same way logEvent's are (see incidents_ack_actor_xor_chk, migration 015).
func acknowledgeIncidentAs(ctx context.Context, q querier, incidentID int64, userID, serviceAccountID *int64) (bool, error) {
res, err := q.ExecContext(ctx, `
UPDATE incidents
SET status = 'acknowledged', acknowledged_by = $1, acknowledged_at = $2
WHERE id = $3 AND status = 'triggered'`,
userID, time.Now().Unix(), incidentID)
SET status = 'acknowledged', acknowledged_by = $1, acknowledged_by_service_account_id = $2,
acknowledged_at = $3
WHERE id = $4 AND status = 'triggered'`,
userID, serviceAccountID, time.Now().Unix(), incidentID)
if err != nil {
return false, err
}
@@ -283,7 +328,7 @@ func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int6
if err := stopEscalation(ctx, q, incidentID); err != nil {
return false, err
}
return true, logEvent(ctx, q, incidentID, evAcknowledged, &userID, nil, nil)
return true, logEvent(ctx, q, incidentID, evAcknowledged, userID, serviceAccountID, nil, nil)
}
// openIncidentForAlert returns the open incident an alert currently belongs to,
+61 -25
View File
@@ -100,6 +100,10 @@ func handleListIncidents(db *sql.DB) http.HandlerFunc {
}
incidents = append(incidents, i)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, incidents)
}
}
@@ -157,9 +161,15 @@ func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
rows, err := db.QueryContext(r.Context(), `
SELECT e.id, e.incident_id, e.type, e.user_id, u.username,
e.service_account_id, sa.name,
e.actor_user_id, au.username,
e.actor_service_account_id, asa.name,
e.alert_id, e.detail, e.created_at
FROM incident_events e
LEFT JOIN users u ON u.id = e.user_id
LEFT JOIN service_accounts sa ON sa.id = e.service_account_id
LEFT JOIN users au ON au.id = e.actor_user_id
LEFT JOIN service_accounts asa ON asa.id = e.actor_service_account_id
WHERE e.incident_id = $1
ORDER BY e.created_at ASC, e.id ASC`, id)
if err != nil {
@@ -173,6 +183,9 @@ func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
var e models.IncidentEvent
var ts int64
if err := rows.Scan(&e.ID, &e.IncidentID, &e.Type, &e.UserID, &e.Username,
&e.ServiceAccountID, &e.ServiceAccountName,
&e.ActorUserID, &e.ActorUsername,
&e.ActorServiceAccountID, &e.ActorServiceAccountName,
&e.AlertID, &e.Detail, &ts); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
@@ -180,6 +193,10 @@ func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
e.CreatedAt = time.Unix(ts, 0).UTC()
events = append(events, e)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, events)
}
}
@@ -190,8 +207,8 @@ func handleIncidentAcknowledge(db *sql.DB) http.HandlerFunc {
if !ok {
return
}
user, _ := userFromContext(r.Context())
acked, err := acknowledgeIncident(r.Context(), db, id, user.ID)
userID, saID := callerActorIDs(r.Context())
acked, err := acknowledgeIncidentAs(r.Context(), db, id, userID, saID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
@@ -223,13 +240,14 @@ func handleIncidentUnacknowledge(db *sql.DB) http.HandlerFunc {
if !ok {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
if !updateOpenIncident(w, r, db, id,
`UPDATE incidents SET status = 'triggered', acknowledged_by = NULL, acknowledged_at = NULL
`UPDATE incidents SET status = 'triggered', acknowledged_by = NULL,
acknowledged_by_service_account_id = NULL, acknowledged_at = NULL
WHERE id = $1 AND resolved_at IS NULL`, id) {
return
}
if err := logEvent(r.Context(), db, id, evUnacknowledged, &user.ID, nil, nil); err != nil {
if err := logEvent(r.Context(), db, id, evUnacknowledged, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -247,7 +265,7 @@ func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
if !ok {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
// The body is optional: clients that predate resolution notes send none.
var req struct {
Resolution string `json:"resolution"`
@@ -268,12 +286,12 @@ func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if err := logEvent(r.Context(), db, id, evResolved, &user.ID, nil, nil); err != nil {
if err := logEvent(r.Context(), db, id, evResolved, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if req.Resolution != "" {
if err := logEvent(r.Context(), db, id, evResolutionNote, &user.ID, nil, &req.Resolution); err != nil {
if err := logEvent(r.Context(), db, id, evResolutionNote, userID, saID, nil, &req.Resolution); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -311,8 +329,10 @@ func handleIncidentAssign(db *sql.DB) http.HandlerFunc {
req.UserID, id) {
return
}
// On an "assigned" event user_id is the assignee, not the actor.
if err := logEvent(r.Context(), db, id, evAssigned, &req.UserID, nil, nil); err != nil {
// On an "assigned" event user_id is the assignee; the actor goes in
// the actor_* columns.
actorUserID, actorSAID := callerActorIDs(r.Context())
if err := logAssignedEvent(r.Context(), db, id, req.UserID, actorUserID, actorSAID); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -363,14 +383,14 @@ func handleIncidentSnooze(db *sql.DB) http.HandlerFunc {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
if !updateOpenIncident(w, r, db, id,
"UPDATE incidents SET snoozed_until = $1 WHERE id = $2 AND resolved_at IS NULL",
until.Unix(), id) {
return
}
detail := until.UTC().Format(time.RFC3339)
if err := logEvent(r.Context(), db, id, evSnoozed, &user.ID, nil, &detail); err != nil {
if err := logEvent(r.Context(), db, id, evSnoozed, userID, saID, nil, &detail); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -384,12 +404,12 @@ func handleIncidentUnsnooze(db *sql.DB) http.HandlerFunc {
if !ok {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
if !updateOpenIncident(w, r, db, id,
"UPDATE incidents SET snoozed_until = NULL WHERE id = $1 AND resolved_at IS NULL", id) {
return
}
if err := logEvent(r.Context(), db, id, evUnsnoozed, &user.ID, nil, nil); err != nil {
if err := logEvent(r.Context(), db, id, evUnsnoozed, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -413,6 +433,11 @@ func handleIncidentArchive(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusNotFound, errResp("incident not found"))
return
}
userID, saID := callerActorIDs(r.Context())
if err := logEvent(r.Context(), db, id, evArchived, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respondIncident(w, r, db, id)
}
}
@@ -433,6 +458,11 @@ func handleIncidentUnarchive(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusNotFound, errResp("incident not found"))
return
}
userID, saID := callerActorIDs(r.Context())
if err := logEvent(r.Context(), db, id, evUnarchived, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
@@ -466,27 +496,32 @@ func handleCreateNote(db *sql.DB) http.HandlerFunc {
return
}
user, _ := userFromContext(r.Context())
caller, _ := callerFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
now := time.Now()
var eventID int64
err := db.QueryRowContext(r.Context(), `
INSERT INTO incident_events (incident_id, type, user_id, detail, created_at)
VALUES ($1, $2, $3, $4, $5)
RETURNING id`, id, noteType, user.ID, req.Content, now.Unix()).Scan(&eventID)
INSERT INTO incident_events (incident_id, type, user_id, service_account_id, detail, created_at)
VALUES ($1, $2, $3, $4, $5, $6)
RETURNING id`, id, noteType, userID, saID, req.Content, now.Unix()).Scan(&eventID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusCreated, models.IncidentEvent{
resp := models.IncidentEvent{
ID: eventID,
IncidentID: id,
Type: noteType,
UserID: &user.ID,
Username: &user.Username,
Detail: &req.Content,
CreatedAt: now.UTC().Truncate(time.Second),
})
}
if u, ok := caller.AsHuman(); ok {
resp.UserID, resp.Username = &u.ID, &u.Username
} else if saName, ok := caller.ServiceAccountName(); ok {
resp.ServiceAccountID, resp.ServiceAccountName = saID, &saName
}
respond(w, http.StatusCreated, resp)
}
}
@@ -504,11 +539,12 @@ func handleDeleteNote(db *sql.DB) http.HandlerFunc {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
res, err := db.ExecContext(r.Context(), `
DELETE FROM incident_events
WHERE id = $1 AND incident_id = $2 AND type IN ($3, $4) AND user_id = $5`,
eventID, id, evNote, evResolutionNote, user.ID)
WHERE id = $1 AND incident_id = $2 AND type IN ($3, $4)
AND (user_id = $5 OR service_account_id = $6)`,
eventID, id, evNote, evResolutionNote, userID, saID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
+203
View File
@@ -8,6 +8,7 @@ import (
"time"
"git.ryuvia.com/niklas/terdut-server/internal/api"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// amAlert builds one alert of a webhook payload.
@@ -642,6 +643,106 @@ func TestIncident_ArchiveRoundTrip(t *testing.T) {
}
}
// ---------------------------------------------------------------------------
// Service accounts (terdut-server#25)
// ---------------------------------------------------------------------------
// TestServiceAccount_CanActOnItsTeamsIncidents is #25's regression test.
// Before the fix: acknowledge/resolve/snooze/create-note each 500'd (writing
// acknowledged_by/user_id = 0, violating the users(id) FK for a service
// account), and delete-note silently matched zero rows (WHERE user_id = 0)
// instead of deleting.
func TestServiceAccount_CanActOnItsTeamsIncidents(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
var integration struct {
Key string `json:"key"`
}
decode(t, s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/integrations",
map[string]string{"name": "test"}), &integration)
postToIntegration(t, s, integration.Key, "fp-sa", "SAIncident") // incident 1
// Acknowledge.
resp := s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/acknowledge", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("service account acknowledge: %d", resp.StatusCode)
}
var inc map[string]any
decode(t, resp, &inc)
if inc["acknowledged_by_service_account_id"] == nil {
t.Error("expected acknowledged_by_service_account_id to be set")
}
if inc["acknowledged_by_id"] != nil {
t.Errorf("expected acknowledged_by_id to stay nil for a service-account actor, got %v", inc["acknowledged_by_id"])
}
// Unacknowledge.
resp = s.reqAs(t, keyA, http.MethodDelete, "/api/incidents/1/acknowledge", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Errorf("service account unacknowledge: %d", resp.StatusCode)
}
// Snooze, then unsnooze.
resp = s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/snooze",
map[string]string{"duration": "1h"})
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Errorf("service account snooze: %d", resp.StatusCode)
}
resp = s.reqAs(t, keyA, http.MethodDelete, "/api/incidents/1/snooze", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Errorf("service account unsnooze: %d", resp.StatusCode)
}
// Create, then delete, a note.
var note map[string]any
decode(t, s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/notes",
map[string]string{"content": "looking into it"}), &note)
if note["service_account_id"] == nil {
t.Error("expected service_account_id on the note event")
}
if note["user_id"] != nil {
t.Errorf("expected no user_id on a service-account note, got %v", note["user_id"])
}
noteID := int(note["id"].(float64))
delResp := s.reqAs(t, keyA, http.MethodDelete, fmt.Sprintf("/api/incidents/1/notes/%d", noteID), nil)
delResp.Body.Close()
if delResp.StatusCode != http.StatusNoContent {
t.Errorf("service account deleting its own note: %d", delResp.StatusCode)
}
// Resolve.
resp = s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/resolve", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Errorf("service account resolve: %d", resp.StatusCode)
}
}
// Regression guard: a human actor must still write only the human columns,
// unaffected by the service-account branch added above.
func TestIncident_AcknowledgeStillWritesOnlyHumanColumn(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-human-ack", "Z", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
var inc map[string]any
decode(t, s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil), &inc)
if inc["acknowledged_by_id"] == nil {
t.Error("expected acknowledged_by_id to be set for a human actor")
}
if inc["acknowledged_by_service_account_id"] != nil {
t.Errorf("expected acknowledged_by_service_account_id to stay nil for a human actor, got %v",
inc["acknowledged_by_service_account_id"])
}
}
func TestSweeper_ArchivesResolvedIncidents(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
@@ -769,3 +870,105 @@ func contains(haystack []string, needle string) bool {
}
return false
}
// ---------------------------------------------------------------------------
// Actor on assign / archive / unarchive (terdut-server#35)
// ---------------------------------------------------------------------------
// lastEvent returns the newest timeline event of the given type.
func lastEvent(t *testing.T, events []map[string]any, typ string) map[string]any {
t.Helper()
for i := len(events) - 1; i >= 0; i-- {
if events[i]["type"] == typ {
return events[i]
}
}
t.Fatalf("no %q event in %v", typ, eventTypes(events))
return nil
}
func TestIncident_AssignRecordsHumanActor(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-asg-actor", "Assignable", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.req(t, http.MethodPost, "/api/users",
map[string]string{"username": "alice", "email": "alice@test.com"}).Body.Close()
s.req(t, http.MethodPost, "/api/incidents/1/assign", map[string]any{"user_id": 2}).Body.Close()
ev := lastEvent(t, timeline(t, s, 1), "assigned")
if ev["username"] != "alice" {
t.Errorf("expected the assignee alice in username, got %v", ev["username"])
}
if ev["actor_user_id"] == nil || ev["actor_username"] == nil {
t.Errorf("expected the assigning human in actor_*, got %v", ev)
}
if ev["actor_service_account_id"] != nil {
t.Errorf("expected no service-account actor, got %v", ev["actor_service_account_id"])
}
}
func TestIncident_ArchiveUnarchiveRecordHumanActor(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-arc-actor", "Archivable", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
s.req(t, http.MethodPost, "/api/incidents/1/archive", nil).Body.Close()
s.req(t, http.MethodDelete, "/api/incidents/1/archive", nil).Body.Close()
events := timeline(t, s, 1)
for _, typ := range []string{"archived", "unarchived"} {
ev := lastEvent(t, events, typ)
if ev["user_id"] == nil || ev["service_account_id"] != nil {
t.Errorf("%s: expected only the human actor, got %v", typ, ev)
}
}
}
func TestServiceAccount_AssignArchiveUnarchiveRecordActor(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
var integration struct {
Key string `json:"key"`
}
decode(t, s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/integrations",
map[string]string{"name": "test"}), &integration)
postToIntegration(t, s, integration.Key, "fp-sa-35", "SA35") // incident 1
resp := s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/assign", map[string]any{"user_id": 1})
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("service account assign: %d", resp.StatusCode)
}
resp = s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/resolve", nil)
resp.Body.Close()
resp = s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/archive", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("service account archive: %d", resp.StatusCode)
}
resp = s.reqAs(t, keyA, http.MethodDelete, "/api/incidents/1/archive", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Fatalf("service account unarchive: %d", resp.StatusCode)
}
var events []map[string]any
decode(t, s.reqAs(t, keyA, http.MethodGet, "/api/incidents/1/timeline", nil), &events)
asg := lastEvent(t, events, "assigned")
if asg["actor_service_account_id"] == nil || asg["actor_user_id"] != nil {
t.Errorf("assigned: expected only the service-account actor, got %v", asg)
}
if asg["user_id"] == nil {
t.Errorf("assigned: user_id must stay the assignee, got %v", asg)
}
for _, typ := range []string{"archived", "unarchived"} {
ev := lastEvent(t, events, typ)
if ev["service_account_id"] == nil || ev["user_id"] != nil {
t.Errorf("%s: expected only the service-account actor, got %v", typ, ev)
}
}
}
+29 -1
View File
@@ -79,6 +79,28 @@ func AuthMiddleware(db *sql.DB) func(http.Handler) http.Handler {
}
}
// securityHeaders sets headers that cost nothing to send on every response,
// API or static site alike. nosniff is unconditional; HSTS only fires once
// cookieSecure's signal says the browser is actually looking at this server
// over HTTPS — TLS terminates at the gateway, which (as of this writing) sets
// neither header itself.
//
// max-age is 180 days rather than the usual year-plus: short enough that if
// HTTPS here ever broke for real, the header would age out of a browser's
// cache well within a release cycle instead of locking anyone out of a
// working server. Raise it once this has run clean for a while.
func securityHeaders(publicURL string) func(http.Handler) http.Handler {
return func(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("X-Content-Type-Options", "nosniff")
if cookieSecure(publicURL, r) {
w.Header().Set("Strict-Transport-Security", "max-age=15552000; includeSubDomains")
}
next.ServeHTTP(w, r)
})
}
}
// AdminOnly rejects a caller who is not a system administrator. It runs inside
// AuthMiddleware's group, so by the time it sees a request the caller is known.
//
@@ -113,10 +135,16 @@ func requireSelfOrAdmin(w http.ResponseWriter, r *http.Request, targetID int64)
}
// apiKeyUser resolves an API key to its user and stamps its last use.
// expires_at IS NULL OR > now is part of the lookup itself, the same way
// serveAs's disabled_at check is: an expired key is one that cannot
// authenticate, by construction, rather than one that happens to still
// resolve and has to be caught afterwards.
func apiKeyUser(ctx context.Context, db *sql.DB, token string) (int64, bool) {
var keyID, userID int64
err := db.QueryRowContext(ctx,
"SELECT id, user_id FROM api_keys WHERE key_hash = $1", hashToken(token),
`SELECT id, user_id FROM api_keys
WHERE key_hash = $1 AND (expires_at IS NULL OR expires_at > $2)`,
hashToken(token), time.Now().Unix(),
).Scan(&keyID, &userID)
if err != nil {
return 0, false
+2 -2
View File
@@ -257,7 +257,7 @@ func deliverPending(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
}
// Logged, not returned: the page has already gone out, and treating a
// failed timeline write as a failed delivery would send it again.
if err := logEvent(ctx, db, n.incidentID, eventNotified, n.userID, nil, &n.kind); err != nil {
if err := logEvent(ctx, db, n.incidentID, eventNotified, n.userID, nil, nil, &n.kind); err != nil {
log.Printf("notifier: log delivery of %d: %v", n.id, err)
}
sent++
@@ -312,7 +312,7 @@ func markFailed(ctx context.Context, db *sql.DB, n outboxRow, cause error) {
return
}
detail := fmt.Sprintf("%s: %s", n.kind, cause)
if err := logEvent(ctx, db, n.incidentID, eventNotifyFailed, n.userID, nil, &detail); err != nil {
if err := logEvent(ctx, db, n.incidentID, eventNotifyFailed, n.userID, nil, nil, &detail); err != nil {
log.Printf("notifier: log failure of %d: %v", n.id, err)
}
}
+18 -6
View File
@@ -121,12 +121,12 @@ func ssoRedirect(w http.ResponseWriter, r *http.Request, code ssoError) {
func handleOIDCLogin(db *sql.DB, prov *oidc.Provider, limiter *loginLimiter, publicURL string) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
addrKey := "oidc:" + clientAddr(r)
if limiter.blocked(addrKey, oidcStartMaxPerAddr) {
if limiter.blocked(r.Context(), addrKey, oidcStartMaxPerAddr) {
w.Header().Set("Retry-After", strconv.Itoa(int(loginWindow.Seconds())))
respond(w, http.StatusTooManyRequests, errResp("too many sign-in attempts, try again later"))
return
}
limiter.fail(addrKey)
limiter.fail(r.Context(), addrKey)
state, stateHash, err := randomToken()
if err != nil {
@@ -160,6 +160,9 @@ func handleOIDCLogin(db *sql.DB, prov *oidc.Provider, limiter *loginLimiter, pub
return
}
// #nosec G124 -- HttpOnly/SameSite are literal below; Secure is
// cookieSecure(publicURL, r), not a literal true, which is what
// trips this rule. See cookieSecure's own doc comment in auth.go.
http.SetCookie(w, &http.Cookie{
Name: oidcStateCookie,
Value: state,
@@ -182,6 +185,9 @@ func handleOIDCCallback(db *sql.DB, prov *oidc.Provider, publicURL string) http.
return func(w http.ResponseWriter, r *http.Request) {
// The state cookie has done its job once the callback arrives, whatever
// the outcome.
// #nosec G124 -- HttpOnly/SameSite are literal below; Secure is
// cookieSecure(publicURL, r), not a literal true, which is what
// trips this rule. See cookieSecure's own doc comment in auth.go.
http.SetCookie(w, &http.Cookie{
Name: oidcStateCookie, Value: "", Path: "/api/oidc", MaxAge: -1,
HttpOnly: true, Secure: cookieSecure(publicURL, r), SameSite: http.SameSiteLaxMode,
@@ -189,7 +195,13 @@ func handleOIDCCallback(db *sql.DB, prov *oidc.Provider, publicURL string) http.
q := r.URL.Query()
if e := q.Get("error"); e != "" {
log.Printf("oidc: provider returned error %q: %s", e, q.Get("error_description"))
// %q on both: this runs before state is checked against the
// cookie, so error and error_description are still whatever the
// request's query string says, not yet known to be the real
// provider's. %q keeps a crafted value (say, one holding a
// newline) from forging a second log line rather than just
// being a quoted string within this one.
log.Printf("oidc: provider returned error %q: %q", e, q.Get("error_description")) // #nosec G706 -- both %q
ssoRedirect(w, r, ssoDenied)
return
}
@@ -226,7 +238,7 @@ func handleOIDCCallback(db *sql.DB, prov *oidc.Provider, publicURL string) http.
grants := oidc.ComputeGrants(cfg, identity.Groups)
if !grants.Admitted {
log.Printf("oidc: %q (%s) is in none of the allowed groups", identity.Username, identity.Subject)
log.Printf("oidc: %q (%q) is in none of the allowed groups", identity.Username, identity.Subject) // #nosec G706 -- both %q
ssoRedirect(w, r, ssoNotAllowed)
return
}
@@ -243,11 +255,11 @@ func handleOIDCCallback(db *sql.DB, prov *oidc.Provider, publicURL string) http.
if err != nil {
var se ssoError
if errors.As(err, &se) {
log.Printf("oidc: refused %q (%s): %v", identity.Username, identity.Subject, se)
log.Printf("oidc: refused %q (%q): %v", identity.Username, identity.Subject, se) // #nosec G706 -- both %q
ssoRedirect(w, r, se)
return
}
log.Printf("oidc: sign in %q: %v", identity.Username, err)
log.Printf("oidc: sign in %q: %v", identity.Username, err) // #nosec G706 -- %q
ssoRedirect(w, r, ssoFailed)
return
}
+161
View File
@@ -0,0 +1,161 @@
package api
// This file is internal (package api, not api_test) because loginLimiter and
// its blocked/fail/clear methods are unexported, and TestLoginLimiter_SharedAcrossReplicas
// specifically needs to construct two separate loginLimiter values pointed at
// one database — standing in for two replicas — which only this package can
// do. It duplicates testdb_test.go's newTestDB/withSearchPath rather than
// importing them: those live in the separate api_test package, compiled from
// this directory's external test files, and are not visible here. Same
// reasoning as advisory_lock_test.go, which makes the same trade for the
// same reason.
import (
"context"
"database/sql"
"fmt"
"net/url"
"os"
"strings"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/db"
_ "github.com/jackc/pgx/v5/stdlib"
)
var rateLimiterSchemaSeq int
// rateLimiterTestDB returns a migrated database private to this test.
func rateLimiterTestDB(t *testing.T) *sql.DB {
t.Helper()
dsn := os.Getenv("TERDUT_TEST_DSN")
if dsn == "" {
t.Fatalf("TERDUT_TEST_DSN is not set: these tests need Postgres.\n" +
"Run `make test-db` for a local one, then\n" +
" export TERDUT_TEST_DSN=postgres://terdut:terdut@localhost:5432/terdut_test?sslmode=disable")
}
rateLimiterSchemaSeq++
schema := fmt.Sprintf("test_rl_%d_%d", os.Getpid(), rateLimiterSchemaSeq)
admin, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("connect to TERDUT_TEST_DSN: %v", err)
}
defer admin.Close()
if _, err := admin.Exec("CREATE SCHEMA " + schema); err != nil {
t.Fatalf("create schema %s: %v", schema, err)
}
database, err := db.Open(rateLimiterWithSearchPath(dsn, schema))
if err != nil {
t.Fatalf("open db: %v", err)
}
if err := db.Migrate(database); err != nil {
t.Fatalf("migrate: %v", err)
}
t.Cleanup(func() {
database.Close()
cleanup, err := sql.Open("pgx", dsn)
if err != nil {
return
}
defer cleanup.Close()
if _, err := cleanup.Exec("DROP SCHEMA " + schema + " CASCADE"); err != nil {
t.Logf("drop schema %s: %v", schema, err)
}
})
return database
}
func rateLimiterWithSearchPath(dsn, schema string) string {
opt := "-csearch_path=" + schema
if strings.HasPrefix(dsn, "postgres://") || strings.HasPrefix(dsn, "postgresql://") {
u, err := url.Parse(dsn)
if err == nil {
q := u.Query()
q.Set("options", opt)
u.RawQuery = q.Encode()
return u.String()
}
}
return dsn + " options='" + opt + "'"
}
func TestLoginLimiter_BlocksAtMax(t *testing.T) {
database := rateLimiterTestDB(t)
ctx := context.Background()
l := newLoginLimiter(database)
for range 3 {
if l.blocked(ctx, "k", 3) {
t.Fatal("blocked before reaching max")
}
l.fail(ctx, "k")
}
if !l.blocked(ctx, "k", 3) {
t.Fatal("not blocked after reaching max")
}
}
func TestLoginLimiter_ClearResetsTheCount(t *testing.T) {
database := rateLimiterTestDB(t)
ctx := context.Background()
l := newLoginLimiter(database)
l.fail(ctx, "k")
l.fail(ctx, "k")
l.clear(ctx, "k")
if l.blocked(ctx, "k", 1) {
t.Fatal("still blocked after clear")
}
}
func TestLoginLimiter_KeysAreIndependent(t *testing.T) {
database := rateLimiterTestDB(t)
ctx := context.Background()
l := newLoginLimiter(database)
l.fail(ctx, "a")
if l.blocked(ctx, "b", 1) {
t.Fatal("failing one key blocked an unrelated one")
}
}
// TestLoginLimiter_SharedAcrossReplicas is the regression test for the gap
// this migration closes: an in-memory limiter would let each replica count
// independently, so a caller hitting two different pods could rack up
// max*replicaCount failures before either one blocked. Two loginLimiter
// values sharing one database, standing in for two replicas behind the same
// load balancer, must instead see one combined count.
func TestLoginLimiter_SharedAcrossReplicas(t *testing.T) {
database := rateLimiterTestDB(t)
ctx := context.Background()
replicaA := newLoginLimiter(database)
replicaB := newLoginLimiter(database)
const max = 4
// Alternate which "replica" records the failure, as a real deployment
// would split requests across pods.
for i := range max {
replica := replicaA
if i%2 == 1 {
replica = replicaB
}
if replicaA.blocked(ctx, "k", max) || replicaB.blocked(ctx, "k", max) {
t.Fatalf("blocked after only %d of %d failures", i, max)
}
replica.fail(ctx, "k")
}
if !replicaA.blocked(ctx, "k", max) {
t.Fatal("replica A does not see the combined count as blocked")
}
if !replicaB.blocked(ctx, "k", max) {
t.Fatal("replica B does not see the combined count as blocked")
}
}
+5 -3
View File
@@ -23,13 +23,14 @@ func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config, version strin
// One limiter each, both process-wide for the life of the router: login
// counts failed passwords, sign-up counts account creation, and mixing the
// two would let a burst of sign-ups lock somebody out of logging in.
loginLimit := newLoginLimiter()
signupLimiter := newLoginLimiter()
oidcLimit := newLoginLimiter()
loginLimit := newLoginLimiter(db)
signupLimiter := newLoginLimiter(db)
oidcLimit := newLoginLimiter(db)
r := chi.NewRouter()
r.Use(middleware.Logger)
r.Use(middleware.Recoverer)
r.Use(securityHeaders(notify.PublicURL))
r.Get("/healthz", func(w http.ResponseWriter, r *http.Request) {
respond(w, http.StatusOK, map[string]string{"status": "ok"})
@@ -111,6 +112,7 @@ func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config, version strin
r.Get("/api/users/{id}/teams", handleUserTeams(db))
r.Put("/api/users/{id}/notify", handleSetNotifyTarget(db))
r.Put("/api/users/{id}/password", handleSetPassword(db))
r.Get("/api/users/{id}/api-keys", handleListAPIKeys(db))
r.Post("/api/users/{id}/api-keys", handleCreateAPIKey(db))
r.Delete("/api/users/{id}/api-keys/{keyID}", handleDeleteAPIKey(db))
+3
View File
@@ -235,6 +235,9 @@ func scheduleRange(ctx context.Context, db *sql.DB, teamID int64, from, to strin
clause := strings.Join(where, " AND ")
// #nosec G202 -- clause is built from sqlArgs.add's "$N" placeholders
// only, never a value; every value travels through args.all() as a
// bound parameter. See the sqlArgs doc comment in helpers.go.
rows, err := db.QueryContext(ctx, `
SELECT s.id, s.team_id, t.name, s.user_id, u.username, s.date, s.created_at
FROM schedule_entries s
+115
View File
@@ -0,0 +1,115 @@
package api_test
import (
"bytes"
"net/http"
"strings"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/api"
)
// ---------------------------------------------------------------------------
// Security headers
// ---------------------------------------------------------------------------
func TestSecurityHeaders_NosniffAlwaysSet(t *testing.T) {
s := newTS(t) // no PublicURL: the HTTPS signal is off
resp := s.req(t, http.MethodGet, "/api/me", nil)
defer resp.Body.Close()
if got := resp.Header.Get("X-Content-Type-Options"); got != "nosniff" {
t.Errorf("X-Content-Type-Options = %q, want nosniff", got)
}
if got := resp.Header.Get("Strict-Transport-Security"); got != "" {
t.Errorf("Strict-Transport-Security = %q, want unset without an https PublicURL", got)
}
}
func TestSecurityHeaders_HSTSWhenPublicURLIsHTTPS(t *testing.T) {
s := newTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
resp := s.req(t, http.MethodGet, "/api/me", nil)
defer resp.Body.Close()
got := resp.Header.Get("Strict-Transport-Security")
if !strings.HasPrefix(got, "max-age=") || !strings.Contains(got, "includeSubDomains") {
t.Errorf("Strict-Transport-Security = %q, want a max-age with includeSubDomains", got)
}
}
// ---------------------------------------------------------------------------
// Request body size limits
// ---------------------------------------------------------------------------
// TestBodySizeLimit_OrdinaryEndpointRejectsOversizedBody confirms an
// unauthenticated endpoint can't be made to buffer an arbitrarily large body:
// past maxBodyBytes, decodeJSON fails exactly as it would on any other
// malformed body, rather than the server reading the whole thing first.
func TestBodySizeLimit_OrdinaryEndpointRejectsOversizedBody(t *testing.T) {
s := newTS(t)
huge := bytes.Repeat([]byte("a"), 2<<20) // 2 MiB, past the 1 MiB default
body := []byte(`{"username":"` + string(huge) + `","password":"x"}`)
resp, err := http.Post(s.URL+"/api/login", "application/json", bytes.NewReader(body))
if err != nil {
t.Fatalf("POST /api/login: %v", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("status = %d, want %d (oversized body treated as invalid)", resp.StatusCode, http.StatusBadRequest)
}
}
// TestBodySizeLimit_WebhookAllowsLargerBodyThanDefault confirms the
// Alertmanager webhook's separate, larger cap actually takes effect: a body
// bigger than the ordinary default but within maxWebhookBodyBytes is still
// accepted, not rejected by the smaller limit every other endpoint gets.
func TestBodySizeLimit_WebhookAllowsLargerBodyThanDefault(t *testing.T) {
s := newTS(t)
// Padding kept inside one alert's annotation, comfortably past the 1 MiB
// default and still well under the webhook's 8 MiB cap.
padding := strings.Repeat("a", 3<<20) // 3 MiB
payload := `{"version":"4","status":"firing","groupKey":"big-group",` +
`"groupLabels":{"alertname":"BigAlert"},"alerts":[{"status":"firing",` +
`"labels":{"alertname":"BigAlert"},"annotations":{"note":"` + padding + `"},` +
`"startsAt":"2026-05-20T10:00:00Z","endsAt":"0001-01-01T00:00:00Z",` +
`"fingerprint":"fp-big"}]}`
resp, err := http.Post(s.URL+"/api/integrations/"+s.ingestKey+"/alertmanager",
"application/json", strings.NewReader(payload))
if err != nil {
t.Fatalf("POST webhook: %v", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Errorf("status = %d, want %d (body under the webhook's own cap)", resp.StatusCode, http.StatusOK)
}
}
// TestBodySizeLimit_WebhookRejectsPastItsOwnCap confirms the webhook's larger
// cap is still a cap, not an exemption from one.
func TestBodySizeLimit_WebhookRejectsPastItsOwnCap(t *testing.T) {
s := newTS(t)
huge := strings.Repeat("a", 9<<20) // 9 MiB, past the 8 MiB webhook cap
payload := `{"version":"4","status":"firing","groupKey":"huge-group",` +
`"groupLabels":{"alertname":"HugeAlert"},"alerts":[{"status":"firing",` +
`"labels":{"alertname":"HugeAlert"},"annotations":{"note":"` + huge + `"},` +
`"startsAt":"2026-05-20T10:00:00Z","endsAt":"0001-01-01T00:00:00Z",` +
`"fingerprint":"fp-huge"}]}`
resp, err := http.Post(s.URL+"/api/integrations/"+s.ingestKey+"/alertmanager",
"application/json", strings.NewReader(payload))
if err != nil {
t.Fatalf("POST webhook: %v", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("status = %d, want %d (body past the webhook's own cap)", resp.StatusCode, http.StatusBadRequest)
}
}
+2 -2
View File
@@ -112,7 +112,7 @@ var errInviteUnusable = errors.New("invite is not usable")
func handleSignup(db *sql.DB, limiter *loginLimiter, publicURL string) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
addr := clientAddr(r)
if limiter.blocked("signup:"+addr, maxSignupsPerAddr) {
if limiter.blocked(r.Context(), "signup:"+addr, maxSignupsPerAddr) {
respond(w, http.StatusTooManyRequests, errResp("too many sign-ups from this address"))
return
}
@@ -148,7 +148,7 @@ func handleSignup(db *sql.DB, limiter *loginLimiter, publicURL string) http.Hand
var err error
inv, err = loadInvite(r.Context(), db, req.Invite)
if err != nil {
limiter.fail("signup:" + addr)
limiter.fail(r.Context(), "signup:"+addr)
respond(w, http.StatusForbidden, errResp("this invite link is not usable"))
return
}
+12
View File
@@ -73,6 +73,10 @@ func handleStatsTop(db *sql.DB) http.HandlerFunc {
}
result = append(result, e)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, result)
}
}
@@ -104,6 +108,10 @@ func handleStatsByHour(db *sql.DB) http.HandlerFunc {
}
counts[hr] = cnt
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
type entry struct {
Hour int `json:"hour"`
@@ -146,6 +154,10 @@ func handleStatsByDay(db *sql.DB) http.HandlerFunc {
}
counts[dow] = cnt
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
dayNames := [7]string{"Sunday", "Monday", "Tuesday", "Wednesday", "Thursday", "Friday", "Saturday"}
type entry struct {
+80 -3
View File
@@ -116,6 +116,10 @@ func handleListUsers(db *sql.DB) http.HandlerFunc {
u.DisabledAt = unixPtr(disabled)
users = append(users, u)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, users)
}
}
@@ -234,6 +238,11 @@ func handleDeleteUser(db *sql.DB) http.HandlerFunc {
}
}
// maxAPIKeyExpiryDays bounds expires_in_days: generous enough for any real
// rotation policy, tight enough to reject a typo (a year in hours, say) that
// would otherwise mint a key that outlives the server by decades.
const maxAPIKeyExpiryDays = 3650 // ~10 years
func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
userID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
@@ -247,6 +256,11 @@ func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
var req struct {
Name string `json:"name"`
// ExpiresInDays is optional and, left zero, means the key never
// expires — the only behavior any key had before this field
// existed, so an existing integration that does not send it is
// unaffected.
ExpiresInDays int64 `json:"expires_in_days,omitempty"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
@@ -256,6 +270,10 @@ func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusBadRequest, errResp("name is required"))
return
}
if req.ExpiresInDays < 0 || req.ExpiresInDays > maxAPIKeyExpiryDays {
respond(w, http.StatusBadRequest, errResp("expires_in_days must be 0 (never expires) or up to "+strconv.Itoa(maxAPIKeyExpiryDays)))
return
}
var exists int
if err := db.QueryRowContext(r.Context(), "SELECT 1 FROM users WHERE id = $1", userID).Scan(&exists); err != nil {
@@ -268,18 +286,77 @@ func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
var expiresAt *int64
var expiresAtTime *time.Time
if req.ExpiresInDays > 0 {
t := time.Now().AddDate(0, 0, int(req.ExpiresInDays)).UTC()
u := t.Unix()
expiresAt = &u
expiresAtTime = &t
}
var keyID int64
if err := db.QueryRowContext(r.Context(),
"INSERT INTO api_keys (user_id, key_hash, name) VALUES ($1, $2, $3) RETURNING id",
userID, hash, req.Name).Scan(&keyID); err != nil {
"INSERT INTO api_keys (user_id, key_hash, name, expires_at) VALUES ($1, $2, $3, $4) RETURNING id",
userID, hash, req.Name, expiresAt).Scan(&keyID); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
key := models.APIKey{ID: keyID, UserID: userID, Name: req.Name, Key: raw, CreatedAt: time.Now().UTC()}
key := models.APIKey{
ID: keyID, UserID: userID, Name: req.Name, Key: raw,
CreatedAt: time.Now().UTC(), ExpiresAt: expiresAtTime,
}
respond(w, http.StatusCreated, key)
}
}
// handleListAPIKeys lists a user's own API keys: never the raw key itself
// (only ever returned once, at creation), just enough to tell them apart,
// see which are stale (last_used_at) and which are about to stop working
// (expires_at) — the data handleCreateAPIKey and apiKeyUser's last-use stamp
// already produce, with no endpoint to read it back until now.
func handleListAPIKeys(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
userID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid user id"))
return
}
if !requireSelfOrAdmin(w, r, userID) {
return
}
rows, err := db.QueryContext(r.Context(),
`SELECT id, name, created_at, last_used_at, expires_at
FROM api_keys WHERE user_id = $1 ORDER BY created_at DESC`, userID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer rows.Close()
keys := []models.APIKey{}
for rows.Next() {
var k models.APIKey
var created int64
var lastUsed, expires *int64
if err := rows.Scan(&k.ID, &k.Name, &created, &lastUsed, &expires); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
k.UserID = userID
k.CreatedAt = time.Unix(created, 0).UTC()
k.LastUsedAt = unixPtr(lastUsed)
k.ExpiresAt = unixPtr(expires)
keys = append(keys, k)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, keys)
}
}
func handleDeleteAPIKey(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
userID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
@@ -0,0 +1,39 @@
-- Service-account actors on incident mutations (terdut-server#25). A
-- team-scoped service account acknowledging/resolving/snoozing/noting an
-- incident is not a users row, so it cannot be written into
-- acknowledged_by/incident_events.user_id — doing so either violates the
-- users(id) FK (new rows) or, for incident_events.user_id, silently matches
-- zero rows on delete. These columns are the service-account-shaped parallel
-- to the existing human ones: nullable, mutually exclusive with their human
-- counterpart, ON DELETE SET NULL so a deleted service account doesn't take
-- the incident history with it.
ALTER TABLE incidents
ADD COLUMN acknowledged_by_service_account_id BIGINT
REFERENCES service_accounts(id) ON DELETE SET NULL;
ALTER TABLE incident_events
ADD COLUMN service_account_id BIGINT
REFERENCES service_accounts(id) ON DELETE SET NULL;
-- At most one actor kind per row: both NULL ("the server acted") is valid,
-- exactly one set is valid, both set is a bug this constraint refuses to
-- store rather than silently accepting.
ALTER TABLE incidents
ADD CONSTRAINT incidents_ack_actor_xor_chk CHECK (
acknowledged_by IS NULL OR acknowledged_by_service_account_id IS NULL
);
ALTER TABLE incident_events
ADD CONSTRAINT incident_events_actor_xor_chk CHECK (
user_id IS NULL OR service_account_id IS NULL
);
CREATE INDEX incidents_acknowledged_by_service_account_id_idx
ON incidents(acknowledged_by_service_account_id);
CREATE INDEX incident_events_service_account_id_idx
ON incident_events(service_account_id);
-- assigned_to_service_account_id is deliberately not added here: it would sit
-- unpopulated until handleIncidentAssign itself tracks an actor, which is a
-- separate, pre-existing gap (it records the assignee today, never the
-- actor, for humans either) tracked in its own follow-up issue.
@@ -0,0 +1,16 @@
-- Backs the rate limiters (failed logins, sign-ups, OIDC/device start) with
-- Postgres instead of an in-memory map, now that the server runs more than
-- one replica in production (v0.37.0): a counter that only ever sees its own
-- pod's traffic quietly let every one of these limits through multiplied by
-- the replica count.
--
-- window_start is the start of the current fixed window for key, in the same
-- "unix seconds" shape every other timestamp in this schema uses. The window
-- resets rather than slides, matching the in-memory limiter it replaces:
-- once a key's window is older than the limiter's window length, the next
-- failure starts a fresh one instead of extending the stale one.
CREATE TABLE rate_limit_counters (
key TEXT PRIMARY KEY,
window_start BIGINT NOT NULL,
count INT NOT NULL
);
@@ -0,0 +1,7 @@
-- Optional expiry on a user's own API keys. NULL (the existing default for
-- every row already in this table) means "never expires" -- the same
-- behavior these keys have always had, so no existing integration breaks.
-- Service account keys are deliberately NOT touched: they are a different
-- table, managed by automation, and already distinguished by their own
-- "tdsa_" prefix.
ALTER TABLE api_keys ADD COLUMN expires_at BIGINT;
@@ -0,0 +1,21 @@
-- Who performed an assignment (terdut-server#35). On an 'assigned' event
-- incident_events.user_id is the assignee, so the actor needs columns of its
-- own. Only populated for 'assigned' events; every other event type keeps
-- using user_id/service_account_id for the actor. Older 'assigned' rows stay
-- NULL (the actor was never recorded). Same shape as migration 015: nullable,
-- mutually exclusive, ON DELETE SET NULL.
--
-- assigned_to_service_account_id is still deliberately not added: making
-- service accounts assignable is a separate change (request body, assignee
-- picker, notifier, filters).
ALTER TABLE incident_events
ADD COLUMN actor_user_id BIGINT REFERENCES users(id) ON DELETE SET NULL,
ADD COLUMN actor_service_account_id BIGINT REFERENCES service_accounts(id) ON DELETE SET NULL;
ALTER TABLE incident_events
ADD CONSTRAINT incident_events_assign_actor_xor_chk CHECK (
actor_user_id IS NULL OR actor_service_account_id IS NULL
);
CREATE INDEX incident_events_actor_user_id_idx ON incident_events(actor_user_id);
CREATE INDEX incident_events_actor_service_account_id_idx ON incident_events(actor_service_account_id);
+36 -10
View File
@@ -43,6 +43,13 @@ type Incident struct {
AcknowledgedByUser *string `json:"acknowledged_by,omitempty"`
AcknowledgedAt *time.Time `json:"acknowledged_at,omitempty"`
// AcknowledgedByServiceAccountID/Name are the service-account-shaped
// parallel to AcknowledgedByID/User above: mutually exclusive with it,
// populated when a service account (not a human) acknowledged this
// incident. See migration 015 and terdut-server#25.
AcknowledgedByServiceAccountID *int64 `json:"acknowledged_by_service_account_id,omitempty"`
AcknowledgedByServiceAccountName *string `json:"acknowledged_by_service_account,omitempty"`
AssignedToID *int64 `json:"assigned_to_id,omitempty"`
AssignedToUser *string `json:"assigned_to,omitempty"`
@@ -67,17 +74,36 @@ type Incident struct {
// and is the only history this server keeps — alert rows are mutated in place.
//
// Type is one of: triggered, alert_added, alert_resolved, acknowledged,
// unacknowledged, assigned, snoozed, unsnoozed, resolved, note. A nil UserID
// means the server acted rather than a person.
// unacknowledged, assigned, archived, unarchived, snoozed, unsnoozed,
// resolved, note. UserID and
// ServiceAccountID are mutually exclusive; both nil means the server acted
// rather than any caller.
type IncidentEvent struct {
ID int64 `json:"id"`
IncidentID int64 `json:"incident_id"`
Type string `json:"type"`
UserID *int64 `json:"user_id,omitempty"`
Username *string `json:"username,omitempty"`
AlertID *int64 `json:"alert_id,omitempty"`
Detail *string `json:"detail,omitempty"`
CreatedAt time.Time `json:"created_at"`
ID int64 `json:"id"`
IncidentID int64 `json:"incident_id"`
Type string `json:"type"`
UserID *int64 `json:"user_id,omitempty"`
Username *string `json:"username,omitempty"`
// ServiceAccountID/Name are the service-account-shaped parallel to
// UserID/Username above: mutually exclusive with it, populated when a
// service account (not a human, and not nil-meaning-the-server-acted)
// performed this event. Named Name, not Username — a ServiceAccount has
// a Name field, not a Username. See migration 015 and terdut-server#25.
ServiceAccountID *int64 `json:"service_account_id,omitempty"`
ServiceAccountName *string `json:"service_account_name,omitempty"`
// Actor* name who performed an 'assigned' event, whose UserID is the
// assignee. Mutually exclusive; unset on every other event type and on
// assignments made before migration 018. See terdut-server#35.
ActorUserID *int64 `json:"actor_user_id,omitempty"`
ActorUsername *string `json:"actor_username,omitempty"`
ActorServiceAccountID *int64 `json:"actor_service_account_id,omitempty"`
ActorServiceAccountName *string `json:"actor_service_account_name,omitempty"`
AlertID *int64 `json:"alert_id,omitempty"`
Detail *string `json:"detail,omitempty"`
CreatedAt time.Time `json:"created_at"`
}
// SimilarIncident is an earlier, resolved incident with the same signature as
+9 -1
View File
@@ -37,5 +37,13 @@ type APIKey struct {
Name string `json:"name"`
CreatedAt time.Time `json:"created_at"`
LastUsedAt *time.Time `json:"last_used_at,omitempty"`
Key string `json:"key,omitempty"` // populated only on creation, never stored
// ExpiresAt is nil for a key that never expires, which is every key
// created before this field existed and still the default for a new one
// unless its creator asks otherwise (see handleCreateAPIKey's
// expires_in_days). AuthMiddleware stops accepting a key once this
// passes; nothing deletes the row for it.
ExpiresAt *time.Time `json:"expires_at,omitempty"`
Key string `json:"key,omitempty"` // populated only on creation, never stored
}
+38 -5
View File
@@ -123,6 +123,16 @@ function who(id, name) {
return name || 'someone';
}
// ackActorLabel renders whoever acknowledged inc, human or service account —
// the two are mutually exclusive (migration 015), and a service account is a
// credential, not "you" or "nobody", so it gets its own branch rather than
// going through who()'s id-vs-myID() check.
function ackActorLabel() {
if (inc.acknowledged_by_id != null) return who(inc.acknowledged_by_id, inc.acknowledged_by);
if (inc.acknowledged_by_service_account_id != null) return inc.acknowledged_by_service_account || 'a service account';
return null;
}
function facts() {
const rows = [];
const add = (k, ...v) => rows.push(h('dt', { text: k }), h('dd', {}, ...v));
@@ -132,8 +142,7 @@ function facts() {
// take reading top to bottom to piece together.
const elapsedTo = inc.resolved_at ? Date.parse(inc.resolved_at) : Date.now();
const responsible = inc.assigned_to_id != null ? who(inc.assigned_to_id, inc.assigned_to)
: inc.acknowledged_by_id != null ? who(inc.acknowledged_by_id, inc.acknowledged_by)
: 'Unassigned';
: ackActorLabel() || 'Unassigned';
rows.push(h('dt', { text: 'At a glance' }), h('dd', { class: 'fact-summary' },
h('span', { class: 'fact-chip' }, icon('clock', 'icon fact-icon'), duration(elapsedTo - Date.parse(inc.triggered_at))),
inc.severity && badge(inc.severity, `plain ${severityClass(inc.severity)}`),
@@ -142,7 +151,7 @@ function facts() {
add('Triggered', when(inc.triggered_at), h('span', { class: 'sub', text: ` · ${ago(inc.triggered_at)}` }));
if (inc.acknowledged_at) {
add('Acknowledged', `${who(inc.acknowledged_by_id, inc.acknowledged_by)} · ${when(inc.acknowledged_at)}`);
add('Acknowledged', `${ackActorLabel()} · ${when(inc.acknowledged_at)}`);
}
add('Assigned', inc.assigned_to_id != null ? who(inc.assigned_to_id, inc.assigned_to) : 'Unassigned');
if (inc.status !== 'resolved' && isFuture(inc.snoozed_until)) {
@@ -206,9 +215,28 @@ function alertItem(a) {
// ---------- timeline ----------
// actorLabel renders whoever performed ev, human or service account — the
// two are mutually exclusive (migration 015). null means the server acted:
// ev.user_id == null no longer means that by itself, now that a service
// account's events also leave it null.
function actorLabel(ev, named = false) {
if (ev.user_id != null) return named ? (ev.username || 'someone') : who(ev.user_id, ev.username);
if (ev.service_account_id != null) return ev.service_account_name || 'a service account';
return null;
}
// assignerLabel is actorLabel for an 'assigned' event, whose user_id is the
// assignee: the person who made the assignment is in the actor_* fields
// (null for assignments from before they were recorded).
function assignerLabel(ev, named = false) {
if (ev.actor_user_id != null) return named ? (ev.actor_username || 'someone') : who(ev.actor_user_id, ev.actor_username);
if (ev.actor_service_account_id != null) return ev.actor_service_account_name || 'a service account';
return null;
}
// named spells users out instead of "you", for text that leaves this page.
function eventText(ev, named = false) {
const person = ev.user_id != null ? (named ? ev.username || 'someone' : who(ev.user_id, ev.username)) : null;
const person = actorLabel(ev, named);
const strong = (t) => h('span', { class: 'who', text: t || 'someone' });
const alertName = () => {
const a = (inc.alerts || []).find((x) => x.id === ev.alert_id);
@@ -220,7 +248,12 @@ function eventText(ev, named = false) {
case 'alert_resolved': return [`Alert resolved: ${alertName()}`];
case 'acknowledged': return [strong(person), ' acknowledged'];
case 'unacknowledged': return [strong(person), ' cleared the acknowledgement'];
case 'assigned': return ['Assigned to ', strong(person)];
case 'assigned': {
const by = assignerLabel(ev, named);
return by ? ['Assigned to ', strong(person), ' by ', strong(by)] : ['Assigned to ', strong(person)];
}
case 'archived': return [strong(person), ' archived the incident'];
case 'unarchived': return [strong(person), ' unarchived the incident'];
case 'snoozed': return [strong(person), ` snoozed until ${ev.detail ? when(ev.detail) : '…'}`];
case 'unsnoozed': return [strong(person), ' ended the snooze'];
case 'resolved': return person ? [strong(person), ' resolved the incident'] : ['Resolved: every alert stopped firing'];