Compare commits

..

25 Commits

Author SHA1 Message Date
Niklas Ye 065b557860 Set the chart's placeholder version to 0.41.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 38s
CI / test (push) Successful in 6m4s
Release / test (push) Successful in 7s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 44s
Release / image (push) Successful in 1m19s
Release / scan-image (push) Successful in 33s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with ead5df1 (0.40.0) and 0f88574 (0.39.0) before it, because a
tree heading for v0.41.0 that still says 0.40.0 tells its reader
something false.
2026-10-08 13:24:46 +02:00
Niklas Ye a8dc89e23d Make the web UI easier to read, and add a theme toggle
Secondary text was too dim to read: --faint sat at about 4.0:1 in the dark
theme and 3.3:1 in the light one, and it carries row ages, hints and labels.
It now clears 4.5:1 in both. The dark surfaces and borders are a step
further apart so cards stand out from the page, and the light borders a
touch stronger.

Account has an Appearance section with System, Light and Dark. System is
the old behaviour. The choice is per browser, kept in localStorage, and is
applied by a small js/theme.js loaded from <head> so there is no flash of
the other theme; the CSP allows no inline script, hence a file of its own.

On-call is redesigned: a hero card for who is on call now, with when the
shift ends for the selected team, and the week as seven day cells instead of
grouped rows. Today is marked with a neutral tint and a bar rather than the
accent colour, which now means only things you can act on; your own days are
marked by the "you" badge, not a fill. On a wide screen the hero and week sit
beside your shifts under a page title. The page stays read-only: shifts are
still edited in the TUI.

Smaller fixes: the team switcher no longer appears in the phone's bottom bar
as well as the top bar (a shared rule overrode the one that hides it); the
Team and Admin tab strips fade at the edge and scroll the open section into
view, so "Sources" is no longer cut off; and "All clear" is a green status
badge with a check instead of a grey pill that looked like a button.

Web UI only: no endpoint or JSON shape changed, so nothing to mirror in
terdut-tui.
2026-10-08 13:24:41 +02:00
Niklas Ye ead5df1574 Set the chart's placeholder version to 0.40.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 24s
CI / test (push) Successful in 5m47s
Release / test (push) Successful in 9s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 31s
Release / image (push) Successful in 1m7s
Release / scan-image (push) Successful in 26s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with 0f88574 (0.39.0) and 584d344 (0.38.0) before it, because a
tree heading for v0.40.0 that still says 0.39.0 tells its reader
something false.
2026-10-08 11:46:09 +02:00
Niklas Ye 3ced069134 Open switches, sources and members in a details sheet
The three team lists carried bare buttons on every row (Remove, Rename,
Revoke, Edit), which crowds a row that is meant for scanning and puts a
destructive action one stray click from every entry. A row now opens a
sheet with the facts the row has no room for, and the actions live
there: Edit and Delete for a switch, Rename and Revoke for a source,
Edit and Remove for a member. The guards are unchanged: owners only, the
last owner and SSO-managed members still cannot be removed, and every
destructive action still asks first.

Switches can finally be edited in place. The PUT endpoint has existed
since the operator needed it; the UI simply never called it, so changing
a timeout meant deleting the switch and losing its history. The edit
form is the add form, prefilled.

The first column of all three lists is now headed Name. Web UI only: no
change to any endpoint or JSON shape, so nothing to mirror in the TUI.
2026-10-08 11:46:09 +02:00
Niklas Ye 0f88574a41 Set the chart's placeholder version to 0.39.0
CI / chart (push) Successful in 4s
CI / security (push) Successful in 45s
CI / test (push) Successful in 6m1s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 25s
Release / image (push) Successful in 1m12s
Release / scan-image (push) Successful in 6s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with 584d344 (0.38.0) and df83adf (0.37.2) before it, because a
tree heading for v0.39.0 that still says 0.38.0 tells its reader
something false.
2026-10-08 09:02:07 +02:00
Niklas Ye 3613fd5732 Check rows.Err() in the remaining Next() loops (#27)
Incident list, the three stats breakdowns and the user list could return a
truncated result as if complete when the scan failed partway.
2026-10-08 09:00:54 +02:00
Niklas Ye dc62278788 Check rows.Err() after the timeline scan loop (#27) 2026-10-08 08:58:19 +02:00
Niklas Ye 0aaea8efb5 Record the actor on assign, archive and unarchive (#35)
Assign logged only the assignee (user_id), archive/unarchive logged nothing.
Migration 018 adds actor_user_id/actor_service_account_id to incident_events
for 'assigned'; archive/unarchive now log archived/unarchived events with the
caller via callerActorIDs. Timeline JSON gains actor_* fields; web timeline
renders them. Service accounts are still not assignable.
2026-10-08 08:53:04 +02:00
Niklas Ye 584d3441fc Set the chart's placeholder version to 0.38.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 2m9s
CI / test (push) Successful in 6m7s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 35s
Release / image (push) Successful in 1m10s
Release / scan-image (push) Successful in 2s
Cosmetic: make helm-package passes --version and --app-version from the
tag, so these two fields decide nothing about what gets published. Still
done, as with df83adf (0.37.2) and 2b2609e (0.37.1) before it, because a
tree heading for v0.38.0 that still says 0.37.2 tells its reader
something false.
2026-10-07 22:39:00 +02:00
Niklas Ye 926aa2d3ec Add optional API key expiry and a missing way to list them
Part of the same security-hardening pass as the last five commits, and
the last item in its backlog. User API keys had no expiry at all --
unlike service-account keys, visibly distinct only by their "tdsa_"
prefix -- and, it turns out while implementing this, no way to list
them either: only create (returns the raw key once) and delete-by-id
existed, so a key's owner had no way to even discover what keys they
had short of remembering IDs from creation time.

handleCreateAPIKey takes an optional expires_in_days (0, the default,
keeps today's behavior: never expires, so no existing integration is
affected). apiKeyUser's lookup now carries `expires_at IS NULL OR
expires_at > now` as part of the query itself, the same way serveAs's
disabled_at check already works -- an expired key simply fails to
resolve, like a wrong one, rather than resolving and being caught
after the fact. New GET /api/users/{id}/api-keys (requireSelfOrAdmin,
same as create/delete) lists id/name/created_at/last_used_at/expires_at,
never the raw key.

Scoped down from the original plan on request: no web UI change, since
there turned out to be no existing API-keys UI at all to extend --
building one from scratch would have been a real feature addition, not
a hardening tweak.

Mirrored the additive expires_at field in terdut-tui's APIKey struct
(separate commit, separate repo) per this workspace's version-coupling
rule; the TUI does not create or list expiring keys itself yet.

New tests (api_keys_test.go): default never-expires, expires_in_days
sets expires_at, out-of-range values rejected, an expired key fails
auth after a fresh one worked, the listing never includes the raw key.
Also added the new GET route to authz_scope_test.go's self-or-admin
table from the previous commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:31:44 +02:00
Niklas Ye a2ca9c25d0 Add a regression test for the team/admin/self authorization pattern
Part of the same security-hardening pass as the last four commits.
terdut-server's authz is already solid -- centralized predicates
(requireTeamMember, requireTeamOwner, requireSelfOrAdmin, AdminOnly)
rather than ad hoc per-handler checks, confirmed by spot-checking several
handlers while writing this. But it is enforced by convention, not the
type system: a future handler that forgets its guard would compile and
read fine on review, exactly like one that remembers it.

authz_scope_test.go builds two teams and, for every team-scoped route
(members, OIDC groups, invites, escalation, dead man's switches,
integrations, schedule, plus every /api/incidents/{id}/... route, scoped
by the incident's own team_id through incidentIDParam's single
chokepoint), calls it as one team's owner against the other team's
resources and asserts 404 -- requireTeamMember and requireTeamOwner both
answer a non-member that way. Separate tests cover AdminOnly's routes
(403 for a non-admin) and requireSelfOrAdmin's (403 for a non-admin
acting on someone else's account).

Verified the test actually catches a regression, not just that it
passes today: temporarily removed handleListTeamMembers' requireTeamMember
call, confirmed exactly that one subtest failed and nothing else did,
then put it back.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:21:46 +02:00
Niklas Ye 92959cac38 Run the pod as non-root with a read-only filesystem
Part of the same security-hardening pass as the last three commits.
Neither the Dockerfile nor the chart's Deployment set any securityContext
at all, so the container ran as root by default — scratch has no
/etc/passwd for a USER directive to resolve against, so nobody had set one.

Dockerfile now ends with USER 65532:65532 (numeric, since scratch has no
user database; 65532 is the common "nonroot" convention, distroless's own
uid). The chart's Deployment adds a matching pod-level securityContext
(runAsNonRoot, runAsUser/runAsGroup: 65532, seccompProfile: RuntimeDefault)
plus per-container hardening (allowPrivilegeEscalation: false,
capabilities dropped, readOnlyRootFilesystem: true) on both the app
container and the wait-for-postgres init container — neither writes
anything to disk, so the root filesystem can stay read-only.

Verified with helm-lint and a manual `helm template` render of both the
terdut-server and terdut-demo charts. Not yet verified: an actual pod
starting with these in place — readOnlyRootFilesystem is exactly where a
non-obvious write (a temp file, a cache dir) would surface as a crash
rather than a lint error, so that needs a real rollout to confirm.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:16:06 +02:00
Niklas Ye f15db0e20a Add gosec to CI; fix the one real finding it surfaced
Part of the same security-hardening pass as the last two commits. make
lint was go vet only; govulncheck and gitleaks already scanned deps and
secrets on every push, but nothing read this repo's own source for
risky patterns (weak crypto, injection shapes, insecure cookies, ...).

New `make security-code` runs gosec, wired into ci.yaml's security job
alongside the other two. G104 (unchecked error) is excluded at the
Makefile level: every one of its 41 initial hits was this codebase's
existing, deliberate idiom for a best-effort write or an already-
reviewed json.Unmarshal of its own JSONB, predating gosec, and the rule
cannot tell that apart from a mistake -- seventeen individual #nosec
comments would hide a future real G104 regression in the suppression
noise rather than surface it. Reasoning is on the Makefile target.

Of the 12 remaining hits:
  - Genuinely real: oidc.go's callback logged error_description (and,
    two call sites down, identity.Subject) via %s before the request's
    state was even checked against its cookie -- an attacker-reachable
    value going into the log unquoted. Switched to %q, matching
    identity.Username's existing treatment, so a value holding a
    newline can't forge a second log line.
  - False positives, annotated inline rather than globally suppressed:
    4x G124 on cookies that already set Secure via cookieSecure(...)
    (a function call, not the literal `true` the rule wants), 3x G202
    on sqlArgs-built queries that only ever splice in a "$N"
    placeholder, never a value, and the remaining 5x G706 on log lines
    that were already %q-quoted -- gosec's taint analysis doesn't
    model format verbs, so it flags the tainted argument regardless.

Also fixed handleMe's swallowed Scan error (gosec's catch, pre-fix):
a transient DB error left hash/dismissed at their zero values and the
response claimed no password and no onboarding dismissal regardless
of the truth, rather than surfacing a 500.

Checked both workflow files for the injection class letsvisit found
there (a `${{ }}` expression spliced straight into a `run:` block):
every one here already goes through `env:` as a quoted shell variable,
documented in ci.yaml's own header comment. Nothing to fix.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:12:31 +02:00
Niklas Ye b82c10acf4 Back the login/signup/OIDC/device rate limiters with Postgres
Part of the same security-hardening pass as the body-size/header commit.
loginLimiter was an in-memory, per-process sync.Mutex+map -- fine for one
replica, but charts/terdut-server/values.yaml has set replicaCount: 2 in
production since v0.37.0. Each pod counted only its own traffic, so every
limit it guarded (failed logins, sign-ups, OIDC/device-login starts) was
effectively twice as generous as the constants say, not just in theory.

loginLimiter now stores its counters in a new rate_limit_counters table
(migration 016) instead of a map; blocked/fail/clear take a context and
query/upsert/delete a row keyed by the same strings callers already used
(username, client address, "signup:"+address, ...). Semantics are
unchanged -- a fixed window that resets rather than slides -- so no call
site's behavior changes, only where the count lives. Added purgeRateLimits
to the sweeper, alongside purgeSessions/purgeAckTokens, so expired windows
don't accumulate.

New internal (package api) tests in rate_limiter_test.go cover the basic
behavior plus the regression this exists to fix: two loginLimiter values
sharing one database, standing in for two replicas, now see one combined
count instead of each keeping their own.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 22:03:51 +02:00
Niklas Ye 7cd6fbf571 Cap request body size and add baseline security headers
Part of a security-hardening pass (see wiki for the full backlog).
decodeJSON had no size limit at all, so every JSON endpoint -- including
the two unauthenticated ones (bootstrap, the Alertmanager webhook) --
would buffer an attacker-supplied body of unbounded size before it was
even validated. decodeJSON now wraps the body in http.MaxBytesReader at
a 1 MiB default; the webhook gets its own 8 MiB cap via decodeJSONLimit,
since a real Alertmanager batch can be bigger than an ordinary API body.

Also adds a securityHeaders middleware, applied globally: nosniff on
every response (previously only the static site got it), and HSTS
(180-day max-age, conservative on purpose) whenever cookieSecure's
signal says the browser is on HTTPS. Checked the chart/gateway config
first -- neither sets HSTS anywhere, so this was a real gap.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-10-07 21:57:00 +02:00
Niklas Ye df83adfe47 Set the chart's placeholder version to 0.37.2
CI / chart (push) Successful in 1s
CI / security (push) Successful in 20s
CI / test (push) Successful in 5m28s
Release / test (push) Successful in 8s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 49s
Release / image (push) Successful in 1m17s
Release / scan-image (push) Successful in 26s
2026-10-07 21:12:51 +02:00
Niklas Ye 9da913080f Stop 500ing when a service account acts on an incident
Every incident-mutation handler read userFromContext(ctx) and wrote the
result's .ID into acknowledged_by/incident_events.user_id without checking
the ok bool. For a team-scoped service-account caller this returned a
zero-value user id, which violated the users(id) FK and 500'd on
acknowledge, unacknowledge, resolve, snooze, unsnooze and create-note.
handleDeleteNote didn't crash but silently matched zero rows instead
(WHERE user_id = 0), so a service account could never delete its own note.

Add acknowledged_by_service_account_id (incidents) and service_account_id
(incident_events) as nullable FKs to service_accounts(id), parallel to and
mutually exclusive with the existing human columns (migration 015, with a
CHECK enforcing the exclusion). Route every one of the six handlers plus
delete-note through a new callerActorIDs() helper that branches on
Caller.AsHuman()/ServiceAccountID() instead of assuming a human, and thread
a serviceAccountID parameter through logEvent and the new
acknowledgeIncidentAs (acknowledgeIncident itself is untouched: its only
other caller, the push-notification Acknowledge button, is always human).
Render the new actor distinctly from both a human and "the server acted"
in the web UI's incident timeline and facts card.

handleIncidentAssign, handleIncidentArchive and handleIncidentUnarchive are
deliberately not touched here — they track no actor at all today, for
anyone, which is a separate pre-existing gap (follow-up issue to come).

Fixes #25
2026-10-07 21:09:31 +02:00
Niklas Ye 2b2609e98f Set the chart's placeholder version to 0.37.1
CI / chart (push) Successful in 1s
CI / security (push) Successful in 1m18s
CI / test (push) Successful in 6m25s
Release / test (push) Successful in 7s
Release / chart (push) Successful in 7s
Release / image (push) Successful in 2m22s
Release / scan-image (push) Successful in 26s
Release / binaries (push) Successful in 2m53s
Cosmetic: `make helm-package` passes --version and --app-version from
the tag, so these fields decide nothing about what gets published. But
a tree heading for v0.37.1 that still says 0.37.0 tells a reader
something false. Same as 7efd1bb, which cites 5c4e0bd.
2026-10-04 14:26:31 +02:00
Niklas Ye fa6d82d6e5 Fix desktop layout bugs and show more incident actions directly
On-call, the incident queue and the account page all had latent CSS bugs
that only show up once the browser is wide enough to hit the desktop
breakpoint (900px+):

- On-call: .days reset its own margin to 0, which canceled the page-wide
  auto-centering on just that element, leaving the day list pinned to the
  left edge while every other card on the page centered normally.
- Queue: the list pane stayed capped at 340-420px even with nothing
  selected, leaving the rest of the screen empty. It now fills the width
  until an incident is picked, then goes back to list+detail.
- Account, admin user and admin team: .btn and .back-link are inline-flex,
  and margin:auto only centers a block box, so the Sign-out button and the
  two admin back-links sat left of their sibling cards instead of matching
  their width. Wrapped each in a block div.

The incident detail action bar also folded Assign, Add note, Resolve and
Clear acknowledgement into a "More" sheet sized for a phone's width.
Desktop has the room, so it now shows them as direct buttons and hides
More instead; Copy incident stays out of the bar since the header already
has its own button for it.

Filed as niklas/terdut-server#31, #32, #33, #34, each with a screenshot.
2026-10-04 14:26:16 +02:00
Niklas Ye 7efd1bbba7 Set the chart's placeholder version to 0.37.0
CI / chart (push) Successful in 1s
CI / test (push) Successful in 7s
CI / security (push) Successful in 12s
Release / test (push) Successful in 6s
Release / chart (push) Successful in 1s
Release / binaries (push) Successful in 21s
Release / image (push) Successful in 1m3s
Release / scan-image (push) Successful in 1s
2026-10-03 16:17:20 +02:00
Niklas Ye 3bf94a5d7f Default to 2 replicas and RollingUpdate now that the singleton jobs are locked
CI / chart (push) Successful in 0s
CI / test (push) Successful in 7s
CI / security (push) Successful in 12s
replicas and strategy: Recreate were the chart's only guard against the
archiver, notifier and migration races; v0.36.0 closed all three with
advisory locks and a conflict-resolving incident insert, which made
that guard redundant rather than load-bearing. Expose replicaCount
(new, no values.yaml key existed before) defaulting to 2, and switch
to strategy: RollingUpdate with no explicit maxUnavailable/maxSurge --
the 25%/25% default rounds to 0/1 at replicaCount: 2, which is already
zero-downtime.

The chart does not gate this on image.tag, so pointing it at a
pre-v0.36.0 image with the new default is a foot-gun by omission --
noted in both the values.yaml comment and the deployment.yaml comment,
not guarded in code, same as the chart does for every other
version-coupled assumption today.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 12:30:27 +02:00
Niklas Ye 5c4e0bdd0e Set the chart's placeholder version to 0.36.0
CI / chart (push) Successful in 1s
CI / security (push) Successful in 16s
CI / test (push) Successful in 4m50s
Release / test (push) Successful in 7s
Release / chart (push) Successful in 2s
Release / binaries (push) Successful in 48s
Release / image (push) Successful in 1m12s
Release / scan-image (push) Successful in 2s
2026-10-03 12:03:28 +02:00
Niklas Ye 9e5b085d8b Resolve the new-incident insert conflict instead of dropping the payload
CI / chart (push) Successful in 1s
CI / security (push) Successful in 20s
CI / test (push) Has been cancelled
openIncident's INSERT had no ON CONFLICT clause, relying entirely on
incidentForGroup's earlier SELECT to avoid a duplicate. On more than
one replica, two webhook deliveries for the very first occurrence of
a brand-new groupKey can both pass that SELECT before either INSERTs;
the loser then hit incidents_open_group_key_idx's unique violation,
which rolled back its whole transaction — including that payload's
alert upserts, done earlier in the same transaction. ingest's error is
only logged and receiveWebhook answers 200 regardless, so nothing
retried it: the loser's alerts silently never existed.

Add ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO
NOTHING to the INSERT, matching the partial unique index. Postgres
only resolves that conflict after the winning transaction commits (or
rolls back), so by the time RETURNING comes back empty,
existingOpenIncident's follow-up SELECT is guaranteed to see the
winner's row. The loser attaches to it instead of failing outright,
and the rest of its payload commits normally. Covers both callers,
since the dead man's switch sweeper shares this same function.

New test (package api_test, fires N webhook deliveries for one
groupKey from a synchronized start with distinct fingerprints, so
they aren't accidentally serialized by upsertAlerts' own per-
fingerprint lock) confirmed meaningful: with the ON CONFLICT clause
reverted, it fails 10/10 on a missing alert fingerprint; restored, 0/10.
Note while building it: "exactly one incident" alone cannot
distinguish fixed from broken, since the DB's own unique index already
guarantees that either way — the real signal is the loser's payload
surviving.

Chart comment updated: all three of the chart's original reasons for
Recreate are now addressed in code, though replicas stays at 1 and the
strategy stays Recreate pending a deliberate decision to raise it.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:58:39 +02:00
Niklas Ye 0050738ca0 Advisory-lock DB migrations against concurrent replica startup
Migrate's check-then-apply loop against schema_migrations had no
locking: two replicas booting at once against a fresh or
partially-migrated database could both pass the "not yet applied"
check for the same file and race applying it, crashing whichever lost
the duplicate-key insert (confirmed: reverting the lock fails the new
test 10/10 on a duplicate-key violation, racing as early as the
CREATE TABLE IF NOT EXISTS schema_migrations statement itself).

Hold a Postgres advisory lock for Migrate's whole run, on a dedicated
connection reserved via db.Conn so lock and unlock happen on the same
session. Blocking (pg_advisory_lock), unlike the archiver/notifier's
pg_try_advisory_lock: on boot there's no later tick to defer to, so a
second replica should wait for the first to finish migrating rather
than skip ahead.

Adds internal/db's first test file, exercising two concurrent Migrate
calls against a fresh schema.

Still open: the new-incident-insert race on a webhook for a brand-new
groupKey, noted in the chart's updated comment. Login rate limiting
staying in-process, diluted across replicas, is an accepted tradeoff.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:43:09 +02:00
Niklas Ye 42180948d1 Guard the archiver and notifier passes with a Postgres advisory lock
Both background loops run unconditionally on every instance with no
coordination between them, which the chart's replicas: 1 + strategy:
Recreate exists specifically to paper over: with more than one replica,
every one of them would sweep and deliver notifications independently,
and two overlapping during a rollout would both page for the same
incident.

Add withAdvisoryLock, which takes a Postgres advisory lock on a
dedicated connection and runs a pass only if it gets the lock,
otherwise skipping until the next tick. Wire StartArchiver and
StartNotifier through it with their own lock keys, so Sweep and
NotifySweep themselves are untouched and every existing test calling
them directly keeps working unchanged.

This also closes the notifier's double-delivery race in passing: two
replicas can no longer both be inside deliverPending at once, since
only one can hold notifierLockKey at a time.

Deliberately not addressed here, and still blocking a replica count
above 1: the in-memory login rate limiter, the unlocked migration
runner, and the new-incident-insert race on a webhook for a brand-new
groupKey. Noted in the updated chart comment.

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-03 11:36:48 +02:00
52 changed files with 2453 additions and 356 deletions
+3
View File
@@ -125,6 +125,9 @@ jobs:
- name: Secret scan (gitleaks)
run: make security-secrets
- name: Code security scan (gosec)
run: make security-code
# Host mode, no `container:`: helm is baked into the runner image, and a container job
# could not install it -- get.helm.sh is unreachable from the dind bridge. Same reason
# release.yaml's chart job runs on the host.
+6
View File
@@ -24,4 +24,10 @@ FROM scratch
COPY --from=builder /etc/ssl/certs/ca-certificates.crt /etc/ssl/certs/ca-certificates.crt
COPY --from=builder /terdut /terdut
EXPOSE 8080
# Numeric, not a name: scratch has no /etc/passwd for one to resolve against,
# and Docker's USER accepts a bare UID:GID without it. 65532 is the common
# "nonroot" convention (distroless's own uid), chosen so the chart's pod
# securityContext (runAsNonRoot, runAsUser: 65532) matches what the image
# already runs as rather than fighting it.
USER 65532:65532
ENTRYPOINT ["/terdut"]
+21
View File
@@ -143,6 +143,7 @@ BUILDX_BUILDER ?= terdut
TRIVY_VERSION := 0.73.0
GOVULNCHECK_VERSION := v1.1.4
GITLEAKS_VERSION := v8.30.0
GOSEC_VERSION := v2.29.0
# --pull, not --no-cache: refresh the base image without discarding the layer cache.
DOCKER_BUILD_FLAGS ?= --pull
@@ -222,6 +223,26 @@ release: push helm-package helm-push ## Publish image + chart (the workflow's on
security-go: ## Scan Go deps for known CVEs (govulncheck)
go run golang.org/x/vuln/cmd/govulncheck@$(GOVULNCHECK_VERSION) ./...
# Code-level, not dependency- or secret-level: gosec reads this repo's own source for
# known-dangerous patterns (weak crypto, SQL/command injection shapes, insecure file
# permissions, …) rather than its module graph or working tree for leaked credentials,
# which is what security-go and security-secrets above already cover.
#
# G104 (unchecked error) is excluded. Every hit it found here on first run was this
# codebase's existing, deliberate idiom for a best-effort write or an already-reviewed
# json.Unmarshal of this server's own JSONB (see the "best-effort" comments in
# middleware.go and the //nolint:errcheck lines in alerts.go/deadman.go) -- a style that
# predates gosec and that G104 cannot distinguish from a mistake. Reaching the same
# green result by adding a dozens of individual #nosec comments would not add
# information; it would just make a future *real* G104 regression one more suppressed
# line instead of a visible one. Same reasoning as the chi-advisories note on
# security-go above: what gosec reports here (nothing, beyond G104) is the useful
# property, not a loophole. -exclude-generated skips web.go's embedded, build-time-only
# assets.
.PHONY: security-code
security-code: ## Scan this repo's own source for risky patterns (gosec)
go run github.com/securego/gosec/v2/cmd/gosec@$(GOSEC_VERSION) -exclude-generated -exclude=G104 ./...
# --no-git scans the working tree rather than the history, so this catches a secret on the
# way in. It is not a history audit and finding nothing here says nothing about what is
# already committed. --redact because the finding is printed into a CI log.
+5 -3
View File
@@ -984,9 +984,11 @@ name: degrade unknown values to "resolved, reason unknown".
| `created_at` | timestamp | |
Types written today: `triggered`, `alert_added`, `alert_resolved`,
`acknowledged`, `unacknowledged`, `assigned`, `snoozed`, `unsnoozed`, `resolved`,
`note`, `notified`, `notify_failed`, `deadman_silent`. On an `assigned` event
`user_id` is the **assignee**, not the actor. New types may be added; render
`acknowledged`, `unacknowledged`, `assigned`, `archived`, `unarchived`, `snoozed`,
`unsnoozed`, `resolved`, `note`, `notified`, `notify_failed`, `deadman_silent`. On an
`assigned` event `user_id` is the **assignee**, not the actor; the actor is in
`actor_user_id`/`actor_username` or `actor_service_account_id`/`actor_service_account_name`
(absent on assignments made before they were recorded). New types may be added; render
unknown ones generically rather than dropping them.
On `notified` and `notify_failed`, `detail` carries the notification kind
+2 -2
View File
@@ -15,5 +15,5 @@ type: application
# appVersion and image.tag in values.yaml no longer agree, and that is not an oversight:
# image.tag stays "latest", which is what a local install actually pulls. appVersion is
# metadata and drives nothing.
version: 0.35.1
appVersion: "v0.35.1"
version: 0.41.0
appVersion: "v0.41.0"
+35 -5
View File
@@ -6,26 +6,49 @@ metadata:
labels:
{{- include "terdut-server.labels" . | nindent 4 }}
spec:
replicas: 1
replicas: {{ .Values.replicaCount }}
selector:
matchLabels:
{{- include "terdut-server.selectorLabels" . | nindent 6 }}
# Recreate, not RollingUpdate, even though the PVC that forced it is gone: the
# sweeper and the notifier are unsynchronised singletons, and two replicas
# overlapping during a rollout would both page for the same incident.
# RollingUpdate, not Recreate: the sweeper, notifier and migration runner
# each take a Postgres advisory lock around their own pass, and new-incident
# creation on the first webhook for a brand-new groupKey resolves its own
# insert conflict -- so two replicas overlapping during a rollout no longer
# double-page, race a migration, or drop a webhook payload (v0.36.0). No
# explicit maxUnavailable/maxSurge: the 25%/25% default rounds to 0/1 at
# replicaCount: 2, which is zero-downtime already.
strategy:
type: Recreate
type: RollingUpdate
template:
metadata:
labels:
{{- include "terdut-server.selectorLabels" . | nindent 8 }}
spec:
enableServiceLinks: false
# Pod-wide default; both containers below run as this UID regardless of
# what their own image would otherwise pick (postgres:17-alpine's
# pg_isready needs no particular user, and 65532 is what the app image
# itself runs as now — see the Dockerfile's USER). seccompProfile here
# rather than per-container: there is no reason it would ever differ
# between them.
securityContext:
runAsNonRoot: true
runAsUser: 65532
runAsGroup: 65532
seccompProfile:
type: RuntimeDefault
{{- if .Values.database.waitForPostgres.enabled }}
initContainers:
- name: wait-for-postgres
image: "{{ .Values.database.waitForPostgres.image.repository }}:{{ .Values.database.waitForPostgres.image.tag }}"
imagePullPolicy: {{ .Values.database.waitForPostgres.image.pullPolicy }}
# No capability this loop needs, and nothing in it writes to disk:
# sh, pg_isready, echo and sleep all run read-only.
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
env:
- name: TERDUT_DB_DSN
value: {{ required "database.dsn is required" .Values.database.dsn | quote }}
@@ -42,6 +65,13 @@ spec:
- name: terdut-server
image: "{{ .Values.image.repository }}:{{ .Values.image.tag }}"
imagePullPolicy: {{ .Values.image.pullPolicy }}
# scratch, nothing to write: the binary keeps no local state and
# writes nothing to disk, so the root filesystem can stay read-only.
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
ports:
- name: http
containerPort: {{ .Values.service.port }}
+7
View File
@@ -1,3 +1,10 @@
# Safe above 1 since v0.36.0: the sweeper, notifier and migration runner each
# take a Postgres advisory lock around their own pass, and a webhook that
# loses the race to open a brand-new incident attaches to the winner's row
# instead of dropping its payload. An image older than v0.36.0 does not have
# these guards -- do not raise this against one.
replicaCount: 2
networking:
hostname: "terdut.example.com"
servicePort: 8080
+122
View File
@@ -0,0 +1,122 @@
package api
// This file is internal (package api, not api_test) because withAdvisoryLock is
// unexported and these tests exercise its locking semantics directly rather than
// through the full StartArchiver/StartNotifier loop, which would make the "does
// not run while held" case timing-dependent instead of deterministic. It opens a
// plain connection to TERDUT_TEST_DSN rather than reusing testdb_test.go's
// newTestDB, since that helper lives in the separate, already-compiled
// api_test package and a Postgres advisory lock needs no schema or migration
// to exercise.
import (
"context"
"database/sql"
"os"
"testing"
_ "github.com/jackc/pgx/v5/stdlib"
)
// advisoryTestDB opens a plain, unmigrated connection to the test database. An
// unset DSN fails rather than skips, matching testdb_test.go's rationale: a
// suite that quietly tests nothing is worse than one that does not run.
func advisoryTestDB(t *testing.T) *sql.DB {
t.Helper()
dsn := os.Getenv("TERDUT_TEST_DSN")
if dsn == "" {
t.Fatalf("TERDUT_TEST_DSN is not set: these tests need Postgres.\n" +
"Run `make test-db` for a local one, then\n" +
" export TERDUT_TEST_DSN=postgres://terdut:terdut@localhost:5432/terdut_test?sslmode=disable")
}
db, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("connect to TERDUT_TEST_DSN: %v", err)
}
t.Cleanup(func() { db.Close() })
return db
}
func TestWithAdvisoryLock_RunsWhenFree(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
ran := false
withAdvisoryLock(ctx, db, archiverLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run although the lock was free")
}
}
func TestWithAdvisoryLock_SkipsWhileHeldElsewhere(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
// Hold the lock on a connection of our own, standing in for another
// replica mid-pass.
holder, err := db.Conn(ctx)
if err != nil {
t.Fatalf("acquire holder connection: %v", err)
}
defer holder.Close()
if _, err := holder.ExecContext(ctx, "SELECT pg_advisory_lock($1)", archiverLockKey); err != nil {
t.Fatalf("pre-acquire lock: %v", err)
}
ran := false
withAdvisoryLock(ctx, db, archiverLockKey, "test", func() { ran = true })
if ran {
t.Fatal("fn ran although another connection already held the lock")
}
if _, err := holder.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", archiverLockKey); err != nil {
t.Fatalf("release held lock: %v", err)
}
// Now that the holder released it, the next caller should get it.
ran = false
withAdvisoryLock(ctx, db, archiverLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run after the other connection released the lock")
}
}
func TestWithAdvisoryLock_ReleasesAfterFnReturns(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
withAdvisoryLock(ctx, db, notifierLockKey, "test", func() {})
// If the first call had leaked the lock, this one would see it held and
// skip, leaving ran false.
ran := false
withAdvisoryLock(ctx, db, notifierLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run on a later call: the earlier call leaked its lock")
}
}
func TestWithAdvisoryLock_KeysAreIndependent(t *testing.T) {
db := advisoryTestDB(t)
ctx := context.Background()
holder, err := db.Conn(ctx)
if err != nil {
t.Fatalf("acquire holder connection: %v", err)
}
defer holder.Close()
if _, err := holder.ExecContext(ctx, "SELECT pg_advisory_lock($1)", archiverLockKey); err != nil {
t.Fatalf("pre-acquire archiver lock: %v", err)
}
defer holder.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", archiverLockKey)
// Holding archiverLockKey must not block notifierLockKey.
ran := false
withAdvisoryLock(ctx, db, notifierLockKey, "test", func() { ran = true })
if !ran {
t.Fatal("fn did not run under a different key although only archiverLockKey was held")
}
}
+46 -6
View File
@@ -94,10 +94,16 @@ func handleIntegrationWebhook(db *sql.DB, notify NotifyConfig) http.HandlerFunc
}
}
// maxWebhookBodyBytes is larger than maxBodyBytes: a real Alertmanager batch
// can carry many alerts, each with several labels and annotations, and the
// sender is a trusted piece of infrastructure rather than an arbitrary
// caller.
const maxWebhookBodyBytes = 8 << 20
func receiveWebhook(w http.ResponseWriter, r *http.Request, db *sql.DB, notify NotifyConfig, src alertSource) {
teamID := src.teamID
var payload amPayload
if err := decodeJSON(r, &payload); err != nil {
if err := decodeJSONLimit(r, &payload, maxWebhookBodyBytes); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid payload"))
return
}
@@ -169,7 +175,7 @@ func ingest(ctx context.Context, db *sql.DB, notify NotifyConfig, src alertSourc
}
touched[id] = true
alertID := a.id
if err := logEvent(ctx, tx, id, evAlertResolved, nil, &alertID, nil); err != nil {
if err := logEvent(ctx, tx, id, evAlertResolved, nil, nil, &alertID, nil); err != nil {
return err
}
}
@@ -376,6 +382,17 @@ func incidentForGroup(ctx context.Context, tx *sql.Tx, notify NotifyConfig, team
// its own. Hence the querier rather than a *sql.Tx. A nil severity leaves the
// column for refreshSeverity to fill from the member alerts; the sweeper passes
// one because its incidents have no members to derive it from.
//
// Both callers get here only after their own SELECT found no open incident for
// this group_key — but on more than one replica, two webhook deliveries for the
// very first occurrence of a brand-new group_key can both pass that SELECT
// before either INSERTs. ON CONFLICT DO NOTHING against
// incidents_open_group_key_idx is what makes the loser's INSERT a no-op instead
// of a unique-violation error that would otherwise roll back its entire
// payload; existingOpenIncident then hands it the winner's row. Postgres
// resolves that conflict only once the winner's transaction has committed (or
// rolled back), so by the time this RETURNING comes back empty, the winner's
// row is guaranteed visible to that follow-up SELECT.
func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID int64, groupKey, title string, groupLabels map[string]string, severity *string) (int64, error) {
onCall, err := currentOnCall(ctx, q, teamID)
if err != nil {
@@ -391,19 +408,27 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID in
err = q.QueryRowContext(ctx, `
INSERT INTO incidents (team_id, group_key, title, group_labels, signature, status, severity, triggered_at, assigned_to)
VALUES ($1, $2, $3, $4::jsonb, $5, 'triggered', $6, $7, $8)
ON CONFLICT (team_id, group_key) WHERE resolved_at IS NULL DO NOTHING
RETURNING id`,
teamID, groupKey, title, string(labelsJSON), incidentSignature(groupLabels, title), severity,
time.Now().Unix(), onCall).Scan(&id)
if err != nil {
switch {
case err == sql.ErrNoRows:
// Lost the race: someone else's incident for this group_key exists now.
// Everything below — the trigger event, assignment, page, escalation
// clock — already happened for that row when it was created; attach to
// it rather than fail this call (and the whole payload) outright.
return existingOpenIncident(ctx, q, teamID, groupKey)
case err != nil:
return 0, err
}
if err := logEvent(ctx, q, id, evTriggered, nil, nil, nil); err != nil {
if err := logEvent(ctx, q, id, evTriggered, nil, nil, nil, nil); err != nil {
return 0, err
}
if onCall != nil {
// On an "assigned" event user_id is the assignee, not the actor.
if err := logEvent(ctx, q, id, evAssigned, onCall, nil, nil); err != nil {
if err := logEvent(ctx, q, id, evAssigned, onCall, nil, nil, nil); err != nil {
return 0, err
}
}
@@ -423,6 +448,21 @@ func openIncident(ctx context.Context, q querier, notify NotifyConfig, teamID in
return id, nil
}
// existingOpenIncident looks up the open incident openIncident's own INSERT just
// lost a conflict against — the same lookup incidentForGroup does before ever
// calling openIncident, repeated here for the caller that arrived second.
func existingOpenIncident(ctx context.Context, q querier, teamID int64, groupKey string) (int64, error) {
var id int64
err := q.QueryRowContext(ctx,
"SELECT id FROM incidents WHERE team_id = $1 AND group_key = $2 AND resolved_at IS NULL",
teamID, groupKey,
).Scan(&id)
if err != nil {
return 0, err
}
return id, nil
}
// linkAlert adds an alert to an incident, emitting a timeline entry only the
// first time. Re-sends of an already-linked alert are silent.
func linkAlert(ctx context.Context, tx *sql.Tx, incidentID, alertID int64) error {
@@ -436,5 +476,5 @@ func linkAlert(ctx context.Context, tx *sql.Tx, incidentID, alertID int64) error
if n, _ := res.RowsAffected(); n == 0 {
return nil
}
return logEvent(ctx, tx, incidentID, evAlertAdded, nil, &alertID, nil)
return logEvent(ctx, tx, incidentID, evAlertAdded, nil, nil, &alertID, nil)
}
+112
View File
@@ -0,0 +1,112 @@
package api_test
import (
"bytes"
"encoding/json"
"fmt"
"net/http"
"sync"
"testing"
)
// TestWebhook_ConcurrentFirstOccurrenceOpensOneIncident reproduces two
// replicas racing the very first webhook delivery for a brand-new group_key:
// both see no open incident yet (incidentForGroup's own SELECT finds
// nothing) and race openIncident's INSERT.
//
// The DB's own unique index already guarantees at most one incident either
// way, with or without this fix — so "exactly one incident" alone cannot
// tell a fixed run from a broken one. What ON CONFLICT handling actually
// changes is what happens to the *loser*: before it, the loser's INSERT hit
// incidents_open_group_key_idx's unique violation, which — since
// upsertAlerts ran earlier in that same transaction — rolled back its whole
// payload, alert insert included. ingest's error is only logged and
// receiveWebhook answers 200 regardless, so nothing ever retried it: the
// loser's alert silently never existed. That is the regression signal this
// test checks — every caller's fingerprint must show up in /api/alerts, not
// just the winner's.
func TestWebhook_ConcurrentFirstOccurrenceOpensOneIncident(t *testing.T) {
s := newTS(t)
const callers = 8
const groupKey = "race-group"
// Every caller needs its own fingerprint. A shared one would serialize all
// of them at upsertAlerts' own ON CONFLICT (team_id, fingerprint) row lock,
// long before any of them reached incidentForGroup — which would hide the
// very race this test exists to force.
bodies := make([][]byte, callers)
for i := range callers {
payload := map[string]any{
"version": "4",
"status": "firing",
"groupKey": groupKey,
"groupLabels": map[string]string{"alertname": "RaceAlert"},
"alerts": []map[string]any{amAlert(fmt.Sprintf("fp-race-%d", i), "RaceAlert", "firing",
"2026-05-20T10:00:00Z", "0001-01-01T00:00:00Z", nil)},
}
bodies[i], _ = json.Marshal(payload)
}
// A start line, so every request is fired as close to simultaneously as
// goroutine scheduling allows, rather than trickling out one dial at a
// time — the race window is the gap between incidentForGroup's SELECT and
// openIncident's INSERT, which a staggered start could easily miss.
var ready sync.WaitGroup
start := make(chan struct{})
statuses := make([]int, callers)
var wg sync.WaitGroup
for i := range callers {
ready.Add(1)
wg.Add(1)
go func(i int) {
defer wg.Done()
ready.Done()
<-start
resp, err := http.Post(s.URL+"/api/integrations/"+s.ingestKey+"/alertmanager",
"application/json", bytes.NewReader(bodies[i]))
if err != nil {
t.Errorf("post webhook #%d: %v", i, err)
return
}
defer resp.Body.Close()
statuses[i] = resp.StatusCode
}(i)
}
ready.Wait()
close(start)
wg.Wait()
for i, code := range statuses {
if code != http.StatusOK {
t.Errorf("webhook #%d returned %d, want 200", i, code)
}
}
var matched []any
for _, inc := range listIncidents(t, s, "") {
if inc["group_key"] == groupKey {
matched = append(matched, inc["id"])
}
}
if len(matched) != 1 {
t.Fatalf("expected exactly 1 incident for group_key %q after %d concurrent deliveries, got %d: %v",
groupKey, callers, len(matched), matched)
}
var alerts []map[string]any
decode(t, s.req(t, http.MethodGet, "/api/alerts", nil), &alerts)
seen := map[string]bool{}
for _, a := range alerts {
if fp, ok := a["fingerprint"].(string); ok {
seen[fp] = true
}
}
for i := range callers {
fp := fmt.Sprintf("fp-race-%d", i)
if !seen[fp] {
t.Errorf("alert %q is missing: its delivery's whole payload was silently rolled back "+
"when it lost the race for the incident", fp)
}
}
}
+135
View File
@@ -0,0 +1,135 @@
package api_test
import (
"net/http"
"testing"
"time"
)
func TestAPIKey_DefaultsToNeverExpiring(t *testing.T) {
s := newTS(t)
var key struct {
Key string `json:"key"`
ExpiresAt *string `json:"expires_at"`
}
decode(t, s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]string{"name": "no-expiry"}), &key)
if key.ExpiresAt != nil {
t.Errorf("expires_at = %v, want nil (unset expires_in_days means never expires)", *key.ExpiresAt)
}
}
func TestAPIKey_ExpiresInDaysSetsExpiresAt(t *testing.T) {
s := newTS(t)
var key struct {
ID int64 `json:"id"`
ExpiresAt *string `json:"expires_at"`
}
decode(t, s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]any{"name": "rotates", "expires_in_days": 30}), &key)
if key.ExpiresAt == nil {
t.Fatal("expires_at = nil, want a timestamp roughly 30 days out")
}
got, err := time.Parse(time.RFC3339, *key.ExpiresAt)
if err != nil {
t.Fatalf("parse expires_at: %v", err)
}
want := time.Now().AddDate(0, 0, 30)
if diff := want.Sub(got).Abs(); diff > time.Hour {
t.Errorf("expires_at = %v, want close to %v (30 days out)", got, want)
}
}
func TestAPIKey_ExpiresInDaysRejectsOutOfRange(t *testing.T) {
s := newTS(t)
for _, days := range []int{-1, 3651} {
resp := s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]any{"name": "bad", "expires_in_days": days})
resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("expires_in_days=%d: status = %d, want %d", days, resp.StatusCode, http.StatusBadRequest)
}
}
}
func TestAPIKey_AnExpiredKeyCannotAuthenticate(t *testing.T) {
s := newTS(t)
var key struct {
ID int64 `json:"id"`
Key string `json:"key"`
}
decode(t, s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]any{"name": "soon-expired", "expires_in_days": 1}), &key)
// A fresh key works...
req, _ := http.NewRequest(http.MethodGet, s.URL+"/api/me", nil)
req.Header.Set("Authorization", "Bearer "+key.Key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
t.Fatalf("GET /api/me: %v", err)
}
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("fresh key: status = %d, want %d", resp.StatusCode, http.StatusOK)
}
// ...and stops working once its expiry has passed.
s.exec(t, "UPDATE api_keys SET expires_at = $1 WHERE id = $2", time.Now().Add(-time.Hour).Unix(), key.ID)
req2, _ := http.NewRequest(http.MethodGet, s.URL+"/api/me", nil)
req2.Header.Set("Authorization", "Bearer "+key.Key)
resp2, err := http.DefaultClient.Do(req2)
if err != nil {
t.Fatalf("GET /api/me: %v", err)
}
defer resp2.Body.Close()
if resp2.StatusCode != http.StatusUnauthorized {
t.Errorf("expired key: status = %d, want %d", resp2.StatusCode, http.StatusUnauthorized)
}
}
func TestAPIKey_ListNeverReturnsTheRawKey(t *testing.T) {
s := newTS(t)
decode(t, s.req(t, http.MethodPost, "/api/users/1/api-keys",
map[string]string{"name": "listed"}), new(struct {
Key string `json:"key"`
}))
var keys []struct {
ID int64 `json:"id"`
Name string `json:"name"`
Key string `json:"key"`
}
decode(t, s.req(t, http.MethodGet, "/api/users/1/api-keys", nil), &keys)
found := false
for _, k := range keys {
if k.Name == "listed" {
found = true
}
if k.Key != "" {
t.Errorf("key %d (%s): raw key present in listing", k.ID, k.Name)
}
}
if !found {
t.Error("the key just created does not appear in the listing")
}
}
func TestAPIKey_ListIsSelfOrAdmin(t *testing.T) {
s := newTS(t)
a := newTeam(t, s, "apikeys-a")
resp := a.call(http.MethodGet, "/api/users/1/api-keys", nil)
defer resp.Body.Close()
if resp.StatusCode != http.StatusForbidden {
t.Errorf("status = %d, want %d (not self, not an admin)", resp.StatusCode, http.StatusForbidden)
}
}
+24 -3
View File
@@ -15,6 +15,12 @@ const (
// expiryGrace absorbs clock skew and notification latency before an alert
// whose ends_at watermark has passed is treated as stale.
expiryGrace = 5 * time.Minute
// archiverLockKey is the Postgres advisory lock the sweeper takes for the
// duration of each pass, so that running more than one replica does not run
// the sweep concurrently on all of them. Its value has no meaning beyond
// being distinct from notifierLockKey.
archiverLockKey int64 = 7265_0001
)
// StartArchiver runs the alert sweeper until ctx is cancelled, starting with an
@@ -23,15 +29,25 @@ const (
// the fallback, not the setting: each pass reads the current value from the
// settings table, so an administrator's change takes effect on the next tick
// instead of at the next restart.
//
// Each pass runs under archiverLockKey (see withAdvisoryLock), so that on more
// than one replica only whichever instance's tick takes the lock first actually
// sweeps; the rest skip that tick rather than racing the same pass.
func StartArchiver(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Duration, notify NotifyConfig) {
ticker := time.NewTicker(sweepInterval)
defer ticker.Stop()
Sweep(ctx, db, archiveAfter, staleAfter, notify)
sweep := func() {
withAdvisoryLock(ctx, db, archiverLockKey, "sweeper", func() {
Sweep(ctx, db, archiveAfter, staleAfter, notify)
})
}
sweep()
for {
select {
case <-ticker.C:
Sweep(ctx, db, archiveAfter, staleAfter, notify)
sweep()
case <-ctx.Done():
return
}
@@ -60,6 +76,7 @@ func Sweep(ctx context.Context, db *sql.DB, archiveAfter, staleAfter time.Durati
archiveResolvedIncidents(ctx, db, archiveAfter)
purgeAckTokens(ctx, db)
purgeSessions(ctx, db)
purgeRateLimits(ctx, db)
}
// expireStale resolves firing alerts that Alertmanager has stopped refreshing.
@@ -105,6 +122,10 @@ func expireStale(ctx context.Context, db *sql.DB, staleAfter time.Duration, skip
for i, id := range ids {
idList[i] = id
}
// #nosec G202 -- sqlArgs.add/addList only ever splice in the "$N"
// placeholder they hand back, never a value; every value travels through
// args.all() as a bound parameter. See the sqlArgs doc comment in
// helpers.go.
if _, err := db.ExecContext(ctx, `
UPDATE alerts
SET status = 'resolved',
@@ -126,7 +147,7 @@ func expireStale(ctx context.Context, db *sql.DB, staleAfter time.Duration, skip
continue
}
alertID := id
if err := logEvent(ctx, db, incidentID, evAlertResolved, nil, &alertID, nil); err != nil {
if err := logEvent(ctx, db, incidentID, evAlertResolved, nil, nil, &alertID, nil); err != nil {
log.Printf("sweeper: log expiry event: %v", err)
}
}
+83 -42
View File
@@ -45,57 +45,85 @@ var dummyHash = sync.OnceValue(func() []byte {
return h
})
// loginLimiter counts failed logins in a fixed window, per username and per
// client address. The username limit is what stops guessing one account; the
// address limit is looser because every user behind the same gateway or NAT
// shares it.
// loginLimiter counts failed logins (and other unauthenticated attempts:
// sign-up, OIDC/device start) in a fixed window, per key — a username, a
// client address, or both, depending on the caller.
//
// Backed by Postgres rather than an in-memory map: this server runs more
// than one replica in production (v0.37.0), and a counter that only ever
// sees its own pod's traffic would quietly let every limit through
// multiplied by the replica count — two loginLimiter values pointed at the
// same db, standing in for two replicas, now share exactly one count per
// key instead of each keeping their own.
//
// The window resets rather than slides, the same behavior the in-memory
// version it replaces had: once a key's window is older than loginWindow,
// the next fail() starts a fresh one instead of extending the stale one.
type loginLimiter struct {
mu sync.Mutex
failures map[string]*loginWindowCount
db *sql.DB
}
type loginWindowCount struct {
start time.Time
n int
func newLoginLimiter(db *sql.DB) *loginLimiter {
return &loginLimiter{db: db}
}
func newLoginLimiter() *loginLimiter {
return &loginLimiter{failures: map[string]*loginWindowCount{}}
}
func (l *loginLimiter) blocked(key string, max int) bool {
l.mu.Lock()
defer l.mu.Unlock()
c, ok := l.failures[key]
if !ok || time.Since(c.start) > loginWindow {
func (l *loginLimiter) blocked(ctx context.Context, key string, max int) bool {
cutoff := time.Now().Unix() - int64(loginWindow.Seconds())
var count int
err := l.db.QueryRowContext(ctx, `
SELECT count FROM rate_limit_counters
WHERE key = $1 AND window_start > $2`,
key, cutoff,
).Scan(&count)
if err != nil {
// No row (never failed, or its window already expired): not blocked.
// A real query error fails the same way — a rate limiter that locks
// everyone out during a brief database hiccup is worse than one that
// is briefly too generous.
return false
}
return c.n >= max
return count >= max
}
func (l *loginLimiter) fail(keys ...string) {
l.mu.Lock()
defer l.mu.Unlock()
now := time.Now()
for k, c := range l.failures {
if now.Sub(c.start) > loginWindow {
delete(l.failures, k)
}
}
func (l *loginLimiter) fail(ctx context.Context, keys ...string) {
now := time.Now().Unix()
windowSecs := int64(loginWindow.Seconds())
for _, key := range keys {
c, ok := l.failures[key]
if !ok {
c = &loginWindowCount{start: now}
l.failures[key] = c
if _, err := l.db.ExecContext(ctx, `
INSERT INTO rate_limit_counters (key, window_start, count)
VALUES ($1, $2, 1)
ON CONFLICT (key) DO UPDATE SET
window_start = CASE WHEN rate_limit_counters.window_start <= $2 - $3
THEN $2 ELSE rate_limit_counters.window_start END,
count = CASE WHEN rate_limit_counters.window_start <= $2 - $3
THEN 1 ELSE rate_limit_counters.count + 1 END`,
key, now, windowSecs,
); err != nil {
log.Printf("rate limiter: record failure for %q: %v", key, err) // #nosec G706 -- %q
}
c.n++
}
}
func (l *loginLimiter) clear(key string) {
l.mu.Lock()
defer l.mu.Unlock()
delete(l.failures, key)
func (l *loginLimiter) clear(ctx context.Context, key string) {
if _, err := l.db.ExecContext(ctx, "DELETE FROM rate_limit_counters WHERE key = $1", key); err != nil {
log.Printf("rate limiter: clear %q: %v", key, err)
}
}
// purgeRateLimits deletes rate-limit windows that have expired, from the
// sweeper — otherwise every distinct username and address this server has
// ever seen a failed attempt from would stay a row forever.
func purgeRateLimits(ctx context.Context, db *sql.DB) {
cutoff := time.Now().Unix() - int64(loginWindow.Seconds())
res, err := db.ExecContext(ctx,
"DELETE FROM rate_limit_counters WHERE window_start <= $1", cutoff)
if err != nil {
log.Printf("sweeper: purge rate limit counters: %v", err)
return
}
if n, _ := res.RowsAffected(); n > 0 {
log.Printf("sweeper: purged %d expired rate limit counter(s)", n)
}
}
// clientAddr is the address a login is counted against. Behind the gateway
@@ -171,6 +199,9 @@ func startSessionCapped(w http.ResponseWriter, r *http.Request, db *sql.DB, user
return err
}
// #nosec G124 -- HttpOnly/SameSite are literal below; Secure is
// cookieSecure(publicURL, r), not a literal true, which is what trips
// this rule. See cookieSecure's own doc comment above.
http.SetCookie(w, &http.Cookie{
Name: sessionCookie,
Value: raw,
@@ -198,7 +229,7 @@ func handleLogin(db *sql.DB, limiter *loginLimiter, publicURL string) http.Handl
userKey := "user:" + strings.ToLower(username)
addrKey := "addr:" + clientAddr(r)
if limiter.blocked(userKey, loginMaxPerUser) || limiter.blocked(addrKey, loginMaxPerAddr) {
if limiter.blocked(r.Context(), userKey, loginMaxPerUser) || limiter.blocked(r.Context(), addrKey, loginMaxPerAddr) {
w.Header().Set("Retry-After", strconv.Itoa(int(loginWindow.Seconds())))
respond(w, http.StatusTooManyRequests, errResp("too many failed attempts, try again later"))
return
@@ -220,11 +251,11 @@ func handleLogin(db *sql.DB, limiter *loginLimiter, publicURL string) http.Handl
}
match := bcrypt.CompareHashAndPassword(stored, []byte(req.Password)) == nil
if !match || !hash.Valid {
limiter.fail(userKey, addrKey)
limiter.fail(r.Context(), userKey, addrKey)
respond(w, http.StatusUnauthorized, errResp("invalid username or password"))
return
}
limiter.clear(userKey)
limiter.clear(r.Context(), userKey)
if err := startSession(w, r, db, userID, publicURL); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
@@ -252,6 +283,9 @@ func handleLogout(db *sql.DB, publicURL string) http.HandlerFunc {
if c, err := r.Cookie(sessionCookie); err == nil && c.Value != "" {
db.ExecContext(r.Context(), "DELETE FROM sessions WHERE token_hash = $1", hashToken(c.Value))
}
// #nosec G124 -- HttpOnly/SameSite are literal below; Secure is
// cookieSecure(publicURL, r), not a literal true, which is what
// trips this rule. See cookieSecure's own doc comment above.
http.SetCookie(w, &http.Cookie{
Name: sessionCookie,
Value: "",
@@ -291,9 +325,16 @@ func handleMe(db *sql.DB) http.HandlerFunc {
}
var hash sql.NullString
var dismissed *int64
db.QueryRowContext(r.Context(),
if err := db.QueryRowContext(r.Context(),
"SELECT password_hash, onboarding_dismissed_at FROM users WHERE id = $1",
caller.ID).Scan(&hash, &dismissed)
caller.ID).Scan(&hash, &dismissed); err != nil {
// fetchUser above already found this row, so an error here is a
// transient database problem, not a missing user — worth a 500
// rather than silently answering "no password, not dismissed",
// which a client would otherwise take at face value.
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, meResponse{
User: user,
HasPassword: hash.Valid,
+182
View File
@@ -0,0 +1,182 @@
package api_test
import (
"net/http"
"testing"
)
// This file is the regression test for the pattern documented throughout
// middleware.go: every team-scoped handler calls requireTeamMember or
// requireTeamOwner before touching data, every self-or-admin handler calls
// requireSelfOrAdmin, and every admin-only route sits behind AdminOnly. That
// pattern is enforced by convention, not by the type system — a new handler
// that forgets the call would compile and pass review on a quick read just
// as easily as one that remembers it. These tests exercise every route that
// carries one of those guards as a caller who should be refused, so a future
// handler missing its guard fails CI instead of becoming a silent IDOR.
// TestAuthzScope_TeamScopedRoutesRefuseANonMember builds two teams and, for
// every team-scoped route, calls it as team A's owner against team B's
// resources. requireTeamMember and requireTeamOwner both answer a non-member
// with 404 (team.go's own reasoning: whether a team exists is itself
// something only its members should learn), so every one of these must come
// back 404 regardless of which of the two guards its handler uses.
func TestAuthzScope_TeamScopedRoutesRefuseANonMember(t *testing.T) {
s := newTS(t)
a := newTeam(t, s, "authz-a")
b := newTeam(t, s, "authz-b")
// An incident in B, to cover the ID-based routes under /api/incidents —
// scoped by the incident's own team_id rather than a {teamID} path
// segment, but through the same single chokepoint (incidentIDParam).
postToIntegration(t, s, b.key, "fp-authz-scope", "AuthzScopeAlert")
var incidents []struct {
ID int64 `json:"id"`
}
decode(t, b.call(http.MethodGet, "/api/incidents", nil), &incidents)
if len(incidents) == 0 {
t.Fatal("setup: no incident in team B to test against")
}
incidentPath := "/api/incidents/" + id64(incidents[0].ID)
bPath := "/api/teams/" + id64(b.id)
tests := []struct {
method, path string
}{
// Team membership/ownership itself.
{http.MethodPut, bPath},
{http.MethodDelete, bPath},
{http.MethodGet, bPath + "/members"},
{http.MethodPost, bPath + "/members"},
{http.MethodDelete, bPath + "/members/1"},
// OIDC group binding.
{http.MethodGet, bPath + "/oidc-groups"},
{http.MethodPut, bPath + "/oidc-groups"},
// Invites.
{http.MethodGet, bPath + "/invites"},
{http.MethodPost, bPath + "/invites"},
{http.MethodDelete, bPath + "/invites/1"},
// Escalation.
{http.MethodGet, bPath + "/escalation"},
{http.MethodPut, bPath + "/escalation"},
// Dead man's switches.
{http.MethodGet, bPath + "/deadman/switches"},
{http.MethodPost, bPath + "/deadman/switches"},
{http.MethodPut, bPath + "/deadman/switches/1"},
{http.MethodDelete, bPath + "/deadman/switches/1"},
// Integrations.
{http.MethodGet, bPath + "/integrations"},
{http.MethodPost, bPath + "/integrations"},
{http.MethodPatch, bPath + "/integrations/1"},
{http.MethodDelete, bPath + "/integrations/1"},
// Schedule.
{http.MethodGet, bPath + "/schedule"},
{http.MethodPost, bPath + "/schedule"},
{http.MethodDelete, bPath + "/schedule/1"},
// Incidents, scoped by the incident's own team rather than a
// {teamID} segment.
{http.MethodGet, incidentPath},
{http.MethodGet, incidentPath + "/alerts"},
{http.MethodGet, incidentPath + "/timeline"},
{http.MethodGet, incidentPath + "/similar"},
{http.MethodPost, incidentPath + "/acknowledge"},
{http.MethodDelete, incidentPath + "/acknowledge"},
{http.MethodPost, incidentPath + "/resolve"},
{http.MethodPost, incidentPath + "/assign"},
{http.MethodPost, incidentPath + "/snooze"},
{http.MethodDelete, incidentPath + "/snooze"},
{http.MethodPost, incidentPath + "/archive"},
{http.MethodDelete, incidentPath + "/archive"},
{http.MethodPost, incidentPath + "/notes"},
{http.MethodDelete, incidentPath + "/notes/1"},
}
for _, tc := range tests {
t.Run(tc.method+" "+tc.path, func(t *testing.T) {
resp := a.call(tc.method, tc.path, nil)
defer resp.Body.Close()
if resp.StatusCode != http.StatusNotFound {
t.Errorf("status = %d, want %d (A is not a member of B)", resp.StatusCode, http.StatusNotFound)
}
})
}
}
// TestAuthzScope_AdminOnlyRoutesRefuseANonAdmin exercises AdminOnly's group
// in router.go directly: a signed-in, non-admin caller gets 403 from every
// route in it, before any handler body runs.
func TestAuthzScope_AdminOnlyRoutesRefuseANonAdmin(t *testing.T) {
s := newTS(t)
a := newTeam(t, s, "authz-admin")
tests := []struct {
method, path string
}{
{http.MethodPost, "/api/users"},
{http.MethodDelete, "/api/users/1"},
{http.MethodPut, "/api/users/1/admin"},
{http.MethodPut, "/api/users/1/disabled"},
{http.MethodGet, "/api/admin/teams"},
{http.MethodGet, "/api/admin/teams/" + id64(a.id)},
{http.MethodGet, "/api/admin/settings"},
{http.MethodPut, "/api/admin/settings"},
}
for _, tc := range tests {
t.Run(tc.method+" "+tc.path, func(t *testing.T) {
resp := a.call(tc.method, tc.path, nil)
defer resp.Body.Close()
if resp.StatusCode != http.StatusForbidden {
t.Errorf("status = %d, want %d (not an admin)", resp.StatusCode, http.StatusForbidden)
}
})
}
}
// TestAuthzScope_SelfOrAdminRoutesRefuseAnotherNonAdminUser exercises
// requireSelfOrAdmin's call sites: a non-admin caller acting on a *different*
// user's account must be refused, the same as AdminOnly's routes, even
// though these sit in the general authenticated group rather than behind
// AdminOnly itself.
func TestAuthzScope_SelfOrAdminRoutesRefuseAnotherNonAdminUser(t *testing.T) {
s := newTS(t)
a := newTeam(t, s, "authz-self-a")
b := newTeam(t, s, "authz-self-b")
var members []struct {
UserID int64 `json:"user_id"`
}
decode(t, s.req(t, http.MethodGet, "/api/teams/"+id64(b.id)+"/members", nil), &members)
if len(members) == 0 {
t.Fatal("setup: team B has no members")
}
bUserID := id64(members[0].UserID)
tests := []struct {
method, path string
}{
{http.MethodGet, "/api/users/" + bUserID + "/teams"},
{http.MethodPut, "/api/users/" + bUserID + "/notify"},
{http.MethodPut, "/api/users/" + bUserID + "/password"},
{http.MethodGet, "/api/users/" + bUserID + "/api-keys"},
{http.MethodPost, "/api/users/" + bUserID + "/api-keys"},
{http.MethodDelete, "/api/users/" + bUserID + "/api-keys/1"},
}
for _, tc := range tests {
t.Run(tc.method+" "+tc.path, func(t *testing.T) {
resp := a.call(tc.method, tc.path, nil)
defer resp.Body.Close()
if resp.StatusCode != http.StatusForbidden {
t.Errorf("status = %d, want %d (not self, not an admin)", resp.StatusCode, http.StatusForbidden)
}
})
}
}
+16 -5
View File
@@ -99,13 +99,24 @@ func (c Caller) ServiceAccountID() (int64, bool) {
return c.sa.id, true
}
// ServiceAccountName reports this caller's own service-account name, for a
// handler's synchronous response — the same credential it authenticated
// with, already resolved onto the Caller by serveAsServiceAccount, so no
// extra query is needed.
func (c Caller) ServiceAccountName() (string, bool) {
if c.sa == nil {
return "", false
}
return c.sa.name, true
}
// Identity is a stable, log/audit-facing string distinguishing a human
// caller from a service account — "user:42" or "service-account:7". Not
// wired into any database column today (incidents.go's acknowledged_by/
// assigned_to/user_id are explicitly out of scope for this change — that
// needs its own schema migration, tracked separately), but this is the one
// place in the request path that already knows which kind of caller this
// is, and that follow-up will want exactly this accessor.
// wired into any database column — incidents.go's acknowledged_by/
// incident_events.user_id use AsHuman()/ServiceAccountID() directly against
// the parallel *_service_account_id columns (migration 015) instead, since a
// column needs the id, not this rendered string. assigned_to stays
// human-only and out of scope (terdut-server#25's follow-up).
func (c Caller) Identity() string {
switch {
case c.user != nil:
+6 -2
View File
@@ -283,6 +283,10 @@ func deadmanAlerts(ctx context.Context, db *sql.DB, teamID int64, cfg deadmanSet
nameList[i] = n
}
// #nosec G202 -- sqlArgs.add/addList only ever splice in the "$N"
// placeholder they hand back, never a value; every value travels through
// args.all() as a bound parameter. See the sqlArgs doc comment in
// helpers.go.
rows, err := db.QueryContext(ctx, `
SELECT id, team_id, fingerprint, labels, status, received_at
FROM alerts
@@ -373,7 +377,7 @@ func deadmanDied(ctx context.Context, db *sql.DB, notify NotifyConfig, hb deadma
alertID := hb.id
detail := "last heartbeat " + humanDuration(now.Sub(time.Unix(hb.receivedAt, 0))) + " ago"
if err := logEvent(ctx, tx, incidentID, evDeadmanSilent, nil, &alertID, &detail); err != nil {
if err := logEvent(ctx, tx, incidentID, evDeadmanSilent, nil, nil, &alertID, &detail); err != nil {
return err
}
@@ -415,7 +419,7 @@ func deadmanRecovered(ctx context.Context, db *sql.DB, hb deadmanAlert) error {
time.Now().Unix(), incidentResolutionRecovered, incidentID); err != nil {
return err
}
if err := logEvent(ctx, tx, incidentID, evResolved, nil, nil, nil); err != nil {
if err := logEvent(ctx, tx, incidentID, evResolved, nil, nil, nil, nil); err != nil {
return err
}
// The all-clear goes to whoever was paged, which enqueueResolved works out
+2 -2
View File
@@ -72,12 +72,12 @@ func normalizeUserCode(s string) string {
func handleDeviceStart(db *sql.DB, limiter *loginLimiter, publicURL string) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
addrKey := "device:" + clientAddr(r)
if limiter.blocked(addrKey, deviceStartMaxPerAddr) {
if limiter.blocked(r.Context(), addrKey, deviceStartMaxPerAddr) {
w.Header().Set("Retry-After", strconv.Itoa(int(loginWindow.Seconds())))
respond(w, http.StatusTooManyRequests, errResp("too many sign-in attempts, try again later"))
return
}
limiter.fail(addrKey)
limiter.fail(r.Context(), addrKey)
deviceCode, deviceHash, err := randomToken()
if err != nil {
+2 -2
View File
@@ -229,7 +229,7 @@ func advanceEscalation(ctx context.Context, db *sql.DB, cfg NotifyConfig, policy
// nobody. That is a policy that looks configured and is not.
detail += ": nobody reachable"
}
if err := logEvent(ctx, tx, incidentID, evEscalated, nil, nil, &detail); err != nil {
if err := logEvent(ctx, tx, incidentID, evEscalated, nil, nil, nil, &detail); err != nil {
return err
}
return tx.Commit()
@@ -256,7 +256,7 @@ func escalationExhausted(ctx context.Context, tx *sql.Tx, policy *escalationPoli
incidentID); err != nil {
return err
}
return logEvent(ctx, tx, incidentID, evEscalated, nil, nil, &detail)
return logEvent(ctx, tx, incidentID, evEscalated, nil, nil, nil, &detail)
}
// pageLevel notifies every target of one level and reports who was woken.
+56
View File
@@ -1,8 +1,11 @@
package api
import (
"context"
"database/sql"
"encoding/json"
"errors"
"log"
"net/http"
"strconv"
"strings"
@@ -69,11 +72,64 @@ func respond(w http.ResponseWriter, status int, v any) {
json.NewEncoder(w).Encode(v)
}
// maxBodyBytes caps an ordinary JSON request body. 1 MiB is far more than any
// endpoint below needs — it exists so an unauthenticated caller (signup,
// login, bootstrap) can't make the server buffer an arbitrarily large body
// before the request is even validated.
const maxBodyBytes = 1 << 20
func decodeJSON(r *http.Request, v any) error {
return decodeJSONLimit(r, v, maxBodyBytes)
}
// decodeJSONLimit is decodeJSON with an explicit cap, for the one endpoint
// (the Alertmanager webhook, see maxWebhookBodyBytes) whose real payloads can
// legitimately be larger than maxBodyBytes.
func decodeJSONLimit(r *http.Request, v any, limit int64) error {
defer r.Body.Close()
// w is nil: there is no ResponseWriter here to disable keep-alive with,
// which net/http documents as fine — the limit is still enforced, the
// connection just isn't closed early on a request that blows past it.
r.Body = http.MaxBytesReader(nil, r.Body, limit)
return json.NewDecoder(r.Body).Decode(v)
}
func errResp(msg string) map[string]string {
return map[string]string{"error": msg}
}
// withAdvisoryLock runs fn only if it can take the named Postgres advisory lock on a
// dedicated connection, and skips fn otherwise. This is what keeps the archiver and
// notifier safe to run on more than one replica: whichever instance's tick gets there
// first does the work; the rest see the lock held and simply wait for their next tick
// instead of running the same pass concurrently.
//
// pg_try_advisory_lock is session-scoped, so taking and releasing it must happen on the
// same connection, reserved via db.Conn rather than borrowed from the pool's shared
// connections fn itself may use — and released (unlocked, then closed) before returning,
// since a session lock otherwise outlives this call and leaks onto whatever reuses the
// pooled connection next.
func withAdvisoryLock(ctx context.Context, db *sql.DB, key int64, name string, fn func()) {
conn, err := db.Conn(ctx)
if err != nil {
log.Printf("%s: advisory lock: acquire connection: %v", name, err)
return
}
defer conn.Close()
var locked bool
if err := conn.QueryRowContext(ctx, "SELECT pg_try_advisory_lock($1)", key).Scan(&locked); err != nil {
log.Printf("%s: advisory lock: %v", name, err)
return
}
if !locked {
return // another replica is already running this pass
}
defer func() {
if _, err := conn.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", key); err != nil {
log.Printf("%s: advisory unlock: %v", name, err)
}
}()
fn()
}
+62 -17
View File
@@ -32,6 +32,8 @@ const (
evAcknowledged = "acknowledged"
evUnacknowledged = "unacknowledged"
evAssigned = "assigned"
evArchived = "archived"
evUnarchived = "unarchived"
evSnoozed = "snoozed"
evUnsnoozed = "unsnoozed"
evResolved = "resolved"
@@ -65,11 +67,13 @@ const incidentSelectFrom = `
WHERE el.team_id = i.team_id AND el.position = i.escalation_level),
i.triggered_at,
i.acknowledged_by, i.acknowledged_at, ack.username,
i.acknowledged_by_service_account_id, acksa.name,
i.assigned_to, asg.username, i.snoozed_until,
i.resolved_at, i.resolution_source, i.archived_at
FROM incidents i
JOIN teams t ON t.id = i.team_id
LEFT JOIN users ack ON ack.id = i.acknowledged_by
LEFT JOIN service_accounts acksa ON acksa.id = i.acknowledged_by_service_account_id
LEFT JOIN users asg ON asg.id = i.assigned_to`
func scanIncident(s scanner) (models.Incident, error) {
@@ -83,6 +87,7 @@ func scanIncident(s scanner) (models.Incident, error) {
&i.EscalationLevel, &escalationDue,
&triggeredAt,
&i.AcknowledgedByID, &ackAt, &i.AcknowledgedByUser,
&i.AcknowledgedByServiceAccountID, &i.AcknowledgedByServiceAccountName,
&i.AssignedToID, &i.AssignedToUser, &snoozedUntil,
&resolvedAt, &i.ResolutionSource, &archivedAt,
); err != nil {
@@ -112,13 +117,42 @@ func fetchIncident(ctx context.Context, q querier, id int64) (models.Incident, e
return scanIncident(q.QueryRowContext(ctx, incidentSelectFrom+" WHERE i.id = $1", id))
}
// logEvent appends one entry to an incident's timeline. A nil userID means the
// server acted rather than a person.
func logEvent(ctx context.Context, q querier, incidentID int64, evType string, userID, alertID *int64, detail *string) error {
// callerActorIDs resolves the current request's caller into the pair of
// nilable ids logEvent/acknowledgeIncidentAs expect: exactly one of userID/
// serviceAccountID is set (never both), replacing the unchecked
// userFromContext(ctx) zero-value reads that used to write a human-only id
// of 0 for a service-account caller (terdut-server#25).
func callerActorIDs(ctx context.Context) (userID, serviceAccountID *int64) {
caller, _ := callerFromContext(ctx)
if u, ok := caller.AsHuman(); ok {
return &u.ID, nil
}
if id, ok := caller.ServiceAccountID(); ok {
return nil, &id
}
return nil, nil
}
// logEvent appends one entry to an incident's timeline. userID and
// serviceAccountID are mutually exclusive and both nilable; both nil means
// the server acted rather than any caller (see incident_events_actor_xor_chk,
// migration 015).
func logEvent(ctx context.Context, q querier, incidentID int64, evType string, userID, serviceAccountID, alertID *int64, detail *string) error {
_, err := q.ExecContext(ctx, `
INSERT INTO incident_events (incident_id, type, user_id, alert_id, detail, created_at)
INSERT INTO incident_events (incident_id, type, user_id, service_account_id, alert_id, detail, created_at)
VALUES ($1, $2, $3, $4, $5, $6, $7)`,
incidentID, evType, userID, serviceAccountID, alertID, detail, time.Now().Unix())
return err
}
// logAssignedEvent records an assignment: user_id is the assignee, and the
// caller who performed it goes in the actor_* columns (migration 018), since
// user_id cannot hold both.
func logAssignedEvent(ctx context.Context, q querier, incidentID, assigneeID int64, actorUserID, actorServiceAccountID *int64) error {
_, err := q.ExecContext(ctx, `
INSERT INTO incident_events (incident_id, type, user_id, actor_user_id, actor_service_account_id, created_at)
VALUES ($1, $2, $3, $4, $5, $6)`,
incidentID, evType, userID, alertID, detail, time.Now().Unix())
incidentID, evAssigned, assigneeID, actorUserID, actorServiceAccountID, time.Now().Unix())
return err
}
@@ -251,7 +285,7 @@ func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, e
if err := stopEscalation(ctx, q, incidentID); err != nil {
return false, err
}
if err := logEvent(ctx, q, incidentID, evResolved, nil, nil, nil); err != nil {
if err := logEvent(ctx, q, incidentID, evResolved, nil, nil, nil, nil); err != nil {
return false, err
}
// The all-clear goes only to whoever was paged in the first place, which
@@ -260,19 +294,30 @@ func resolveIfSettled(ctx context.Context, q querier, incidentID int64) (bool, e
return true, enqueueResolved(ctx, q, incidentID)
}
// acknowledgeIncident records that userID has picked an incident up, and reports
// whether it changed anything — an already-resolved or already-acknowledged
// incident is left alone, so a second acknowledge (a retried request, or a
// stale push notification tapped after the web UI already acked it) is a
// no-op rather than a second "acknowledged" timeline entry. Shared by the
// authenticated handler and the Acknowledge button in a push notification,
// so both write the same state and the same timeline entry.
// acknowledgeIncident records that userID — a human — has picked an incident
// up, and reports whether it changed anything — an already-resolved or
// already-acknowledged incident is left alone, so a second acknowledge (a
// retried request, or a stale push notification tapped after the web UI
// already acked it) is a no-op rather than a second "acknowledged" timeline
// entry. Used only by the Acknowledge button in a push notification
// (notify_ack.go), which always resolves a human from
// incident_ack_tokens.user_id — there is no service-account equivalent of
// that flow, so this keeps its human-only signature; the authenticated
// handler goes through acknowledgeIncidentAs below instead.
func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int64) (bool, error) {
return acknowledgeIncidentAs(ctx, q, incidentID, &userID, nil)
}
// acknowledgeIncidentAs is acknowledgeIncident generalized to either actor
// kind. userID and serviceAccountID are mutually exclusive and nilable the
// same way logEvent's are (see incidents_ack_actor_xor_chk, migration 015).
func acknowledgeIncidentAs(ctx context.Context, q querier, incidentID int64, userID, serviceAccountID *int64) (bool, error) {
res, err := q.ExecContext(ctx, `
UPDATE incidents
SET status = 'acknowledged', acknowledged_by = $1, acknowledged_at = $2
WHERE id = $3 AND status = 'triggered'`,
userID, time.Now().Unix(), incidentID)
SET status = 'acknowledged', acknowledged_by = $1, acknowledged_by_service_account_id = $2,
acknowledged_at = $3
WHERE id = $4 AND status = 'triggered'`,
userID, serviceAccountID, time.Now().Unix(), incidentID)
if err != nil {
return false, err
}
@@ -283,7 +328,7 @@ func acknowledgeIncident(ctx context.Context, q querier, incidentID, userID int6
if err := stopEscalation(ctx, q, incidentID); err != nil {
return false, err
}
return true, logEvent(ctx, q, incidentID, evAcknowledged, &userID, nil, nil)
return true, logEvent(ctx, q, incidentID, evAcknowledged, userID, serviceAccountID, nil, nil)
}
// openIncidentForAlert returns the open incident an alert currently belongs to,
+61 -25
View File
@@ -100,6 +100,10 @@ func handleListIncidents(db *sql.DB) http.HandlerFunc {
}
incidents = append(incidents, i)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, incidents)
}
}
@@ -157,9 +161,15 @@ func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
rows, err := db.QueryContext(r.Context(), `
SELECT e.id, e.incident_id, e.type, e.user_id, u.username,
e.service_account_id, sa.name,
e.actor_user_id, au.username,
e.actor_service_account_id, asa.name,
e.alert_id, e.detail, e.created_at
FROM incident_events e
LEFT JOIN users u ON u.id = e.user_id
LEFT JOIN service_accounts sa ON sa.id = e.service_account_id
LEFT JOIN users au ON au.id = e.actor_user_id
LEFT JOIN service_accounts asa ON asa.id = e.actor_service_account_id
WHERE e.incident_id = $1
ORDER BY e.created_at ASC, e.id ASC`, id)
if err != nil {
@@ -173,6 +183,9 @@ func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
var e models.IncidentEvent
var ts int64
if err := rows.Scan(&e.ID, &e.IncidentID, &e.Type, &e.UserID, &e.Username,
&e.ServiceAccountID, &e.ServiceAccountName,
&e.ActorUserID, &e.ActorUsername,
&e.ActorServiceAccountID, &e.ActorServiceAccountName,
&e.AlertID, &e.Detail, &ts); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
@@ -180,6 +193,10 @@ func handleIncidentTimeline(db *sql.DB) http.HandlerFunc {
e.CreatedAt = time.Unix(ts, 0).UTC()
events = append(events, e)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, events)
}
}
@@ -190,8 +207,8 @@ func handleIncidentAcknowledge(db *sql.DB) http.HandlerFunc {
if !ok {
return
}
user, _ := userFromContext(r.Context())
acked, err := acknowledgeIncident(r.Context(), db, id, user.ID)
userID, saID := callerActorIDs(r.Context())
acked, err := acknowledgeIncidentAs(r.Context(), db, id, userID, saID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
@@ -223,13 +240,14 @@ func handleIncidentUnacknowledge(db *sql.DB) http.HandlerFunc {
if !ok {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
if !updateOpenIncident(w, r, db, id,
`UPDATE incidents SET status = 'triggered', acknowledged_by = NULL, acknowledged_at = NULL
`UPDATE incidents SET status = 'triggered', acknowledged_by = NULL,
acknowledged_by_service_account_id = NULL, acknowledged_at = NULL
WHERE id = $1 AND resolved_at IS NULL`, id) {
return
}
if err := logEvent(r.Context(), db, id, evUnacknowledged, &user.ID, nil, nil); err != nil {
if err := logEvent(r.Context(), db, id, evUnacknowledged, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -247,7 +265,7 @@ func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
if !ok {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
// The body is optional: clients that predate resolution notes send none.
var req struct {
Resolution string `json:"resolution"`
@@ -268,12 +286,12 @@ func handleIncidentResolve(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if err := logEvent(r.Context(), db, id, evResolved, &user.ID, nil, nil); err != nil {
if err := logEvent(r.Context(), db, id, evResolved, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
if req.Resolution != "" {
if err := logEvent(r.Context(), db, id, evResolutionNote, &user.ID, nil, &req.Resolution); err != nil {
if err := logEvent(r.Context(), db, id, evResolutionNote, userID, saID, nil, &req.Resolution); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -311,8 +329,10 @@ func handleIncidentAssign(db *sql.DB) http.HandlerFunc {
req.UserID, id) {
return
}
// On an "assigned" event user_id is the assignee, not the actor.
if err := logEvent(r.Context(), db, id, evAssigned, &req.UserID, nil, nil); err != nil {
// On an "assigned" event user_id is the assignee; the actor goes in
// the actor_* columns.
actorUserID, actorSAID := callerActorIDs(r.Context())
if err := logAssignedEvent(r.Context(), db, id, req.UserID, actorUserID, actorSAID); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -363,14 +383,14 @@ func handleIncidentSnooze(db *sql.DB) http.HandlerFunc {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
if !updateOpenIncident(w, r, db, id,
"UPDATE incidents SET snoozed_until = $1 WHERE id = $2 AND resolved_at IS NULL",
until.Unix(), id) {
return
}
detail := until.UTC().Format(time.RFC3339)
if err := logEvent(r.Context(), db, id, evSnoozed, &user.ID, nil, &detail); err != nil {
if err := logEvent(r.Context(), db, id, evSnoozed, userID, saID, nil, &detail); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -384,12 +404,12 @@ func handleIncidentUnsnooze(db *sql.DB) http.HandlerFunc {
if !ok {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
if !updateOpenIncident(w, r, db, id,
"UPDATE incidents SET snoozed_until = NULL WHERE id = $1 AND resolved_at IS NULL", id) {
return
}
if err := logEvent(r.Context(), db, id, evUnsnoozed, &user.ID, nil, nil); err != nil {
if err := logEvent(r.Context(), db, id, evUnsnoozed, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
@@ -413,6 +433,11 @@ func handleIncidentArchive(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusNotFound, errResp("incident not found"))
return
}
userID, saID := callerActorIDs(r.Context())
if err := logEvent(r.Context(), db, id, evArchived, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respondIncident(w, r, db, id)
}
}
@@ -433,6 +458,11 @@ func handleIncidentUnarchive(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusNotFound, errResp("incident not found"))
return
}
userID, saID := callerActorIDs(r.Context())
if err := logEvent(r.Context(), db, id, evUnarchived, userID, saID, nil, nil); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
w.WriteHeader(http.StatusNoContent)
}
}
@@ -466,27 +496,32 @@ func handleCreateNote(db *sql.DB) http.HandlerFunc {
return
}
user, _ := userFromContext(r.Context())
caller, _ := callerFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
now := time.Now()
var eventID int64
err := db.QueryRowContext(r.Context(), `
INSERT INTO incident_events (incident_id, type, user_id, detail, created_at)
VALUES ($1, $2, $3, $4, $5)
RETURNING id`, id, noteType, user.ID, req.Content, now.Unix()).Scan(&eventID)
INSERT INTO incident_events (incident_id, type, user_id, service_account_id, detail, created_at)
VALUES ($1, $2, $3, $4, $5, $6)
RETURNING id`, id, noteType, userID, saID, req.Content, now.Unix()).Scan(&eventID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusCreated, models.IncidentEvent{
resp := models.IncidentEvent{
ID: eventID,
IncidentID: id,
Type: noteType,
UserID: &user.ID,
Username: &user.Username,
Detail: &req.Content,
CreatedAt: now.UTC().Truncate(time.Second),
})
}
if u, ok := caller.AsHuman(); ok {
resp.UserID, resp.Username = &u.ID, &u.Username
} else if saName, ok := caller.ServiceAccountName(); ok {
resp.ServiceAccountID, resp.ServiceAccountName = saID, &saName
}
respond(w, http.StatusCreated, resp)
}
}
@@ -504,11 +539,12 @@ func handleDeleteNote(db *sql.DB) http.HandlerFunc {
return
}
user, _ := userFromContext(r.Context())
userID, saID := callerActorIDs(r.Context())
res, err := db.ExecContext(r.Context(), `
DELETE FROM incident_events
WHERE id = $1 AND incident_id = $2 AND type IN ($3, $4) AND user_id = $5`,
eventID, id, evNote, evResolutionNote, user.ID)
WHERE id = $1 AND incident_id = $2 AND type IN ($3, $4)
AND (user_id = $5 OR service_account_id = $6)`,
eventID, id, evNote, evResolutionNote, userID, saID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
+203
View File
@@ -8,6 +8,7 @@ import (
"time"
"git.ryuvia.com/niklas/terdut-server/internal/api"
"git.ryuvia.com/niklas/terdut-server/internal/models"
)
// amAlert builds one alert of a webhook payload.
@@ -642,6 +643,106 @@ func TestIncident_ArchiveRoundTrip(t *testing.T) {
}
}
// ---------------------------------------------------------------------------
// Service accounts (terdut-server#25)
// ---------------------------------------------------------------------------
// TestServiceAccount_CanActOnItsTeamsIncidents is #25's regression test.
// Before the fix: acknowledge/resolve/snooze/create-note each 500'd (writing
// acknowledged_by/user_id = 0, violating the users(id) FK for a service
// account), and delete-note silently matched zero rows (WHERE user_id = 0)
// instead of deleting.
func TestServiceAccount_CanActOnItsTeamsIncidents(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
var integration struct {
Key string `json:"key"`
}
decode(t, s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/integrations",
map[string]string{"name": "test"}), &integration)
postToIntegration(t, s, integration.Key, "fp-sa", "SAIncident") // incident 1
// Acknowledge.
resp := s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/acknowledge", nil)
if resp.StatusCode != http.StatusOK {
t.Fatalf("service account acknowledge: %d", resp.StatusCode)
}
var inc map[string]any
decode(t, resp, &inc)
if inc["acknowledged_by_service_account_id"] == nil {
t.Error("expected acknowledged_by_service_account_id to be set")
}
if inc["acknowledged_by_id"] != nil {
t.Errorf("expected acknowledged_by_id to stay nil for a service-account actor, got %v", inc["acknowledged_by_id"])
}
// Unacknowledge.
resp = s.reqAs(t, keyA, http.MethodDelete, "/api/incidents/1/acknowledge", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Errorf("service account unacknowledge: %d", resp.StatusCode)
}
// Snooze, then unsnooze.
resp = s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/snooze",
map[string]string{"duration": "1h"})
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Errorf("service account snooze: %d", resp.StatusCode)
}
resp = s.reqAs(t, keyA, http.MethodDelete, "/api/incidents/1/snooze", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Errorf("service account unsnooze: %d", resp.StatusCode)
}
// Create, then delete, a note.
var note map[string]any
decode(t, s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/notes",
map[string]string{"content": "looking into it"}), &note)
if note["service_account_id"] == nil {
t.Error("expected service_account_id on the note event")
}
if note["user_id"] != nil {
t.Errorf("expected no user_id on a service-account note, got %v", note["user_id"])
}
noteID := int(note["id"].(float64))
delResp := s.reqAs(t, keyA, http.MethodDelete, fmt.Sprintf("/api/incidents/1/notes/%d", noteID), nil)
delResp.Body.Close()
if delResp.StatusCode != http.StatusNoContent {
t.Errorf("service account deleting its own note: %d", delResp.StatusCode)
}
// Resolve.
resp = s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/resolve", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Errorf("service account resolve: %d", resp.StatusCode)
}
}
// Regression guard: a human actor must still write only the human columns,
// unaffected by the service-account branch added above.
func TestIncident_AcknowledgeStillWritesOnlyHumanColumn(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-human-ack", "Z", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
var inc map[string]any
decode(t, s.req(t, http.MethodPost, "/api/incidents/1/acknowledge", nil), &inc)
if inc["acknowledged_by_id"] == nil {
t.Error("expected acknowledged_by_id to be set for a human actor")
}
if inc["acknowledged_by_service_account_id"] != nil {
t.Errorf("expected acknowledged_by_service_account_id to stay nil for a human actor, got %v",
inc["acknowledged_by_service_account_id"])
}
}
func TestSweeper_ArchivesResolvedIncidents(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
@@ -769,3 +870,105 @@ func contains(haystack []string, needle string) bool {
}
return false
}
// ---------------------------------------------------------------------------
// Actor on assign / archive / unarchive (terdut-server#35)
// ---------------------------------------------------------------------------
// lastEvent returns the newest timeline event of the given type.
func lastEvent(t *testing.T, events []map[string]any, typ string) map[string]any {
t.Helper()
for i := len(events) - 1; i >= 0; i-- {
if events[i]["type"] == typ {
return events[i]
}
}
t.Fatalf("no %q event in %v", typ, eventTypes(events))
return nil
}
func TestIncident_AssignRecordsHumanActor(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-asg-actor", "Assignable", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.req(t, http.MethodPost, "/api/users",
map[string]string{"username": "alice", "email": "alice@test.com"}).Body.Close()
s.req(t, http.MethodPost, "/api/incidents/1/assign", map[string]any{"user_id": 2}).Body.Close()
ev := lastEvent(t, timeline(t, s, 1), "assigned")
if ev["username"] != "alice" {
t.Errorf("expected the assignee alice in username, got %v", ev["username"])
}
if ev["actor_user_id"] == nil || ev["actor_username"] == nil {
t.Errorf("expected the assigning human in actor_*, got %v", ev)
}
if ev["actor_service_account_id"] != nil {
t.Errorf("expected no service-account actor, got %v", ev["actor_service_account_id"])
}
}
func TestIncident_ArchiveUnarchiveRecordHumanActor(t *testing.T) {
s := newTS(t)
postWebhook(t, s, []map[string]any{
amAlert("fp-arc-actor", "Archivable", "firing", "2026-05-20T10:00:00Z", zeroTime, nil),
})
s.req(t, http.MethodPost, "/api/incidents/1/resolve", nil).Body.Close()
s.req(t, http.MethodPost, "/api/incidents/1/archive", nil).Body.Close()
s.req(t, http.MethodDelete, "/api/incidents/1/archive", nil).Body.Close()
events := timeline(t, s, 1)
for _, typ := range []string{"archived", "unarchived"} {
ev := lastEvent(t, events, typ)
if ev["user_id"] == nil || ev["service_account_id"] != nil {
t.Errorf("%s: expected only the human actor, got %v", typ, ev)
}
}
}
func TestServiceAccount_AssignArchiveUnarchiveRecordActor(t *testing.T) {
s := newTS(t)
instanceKey := createServiceAccount(t, s, s.key, "operator", models.ServiceAccountScopeInstance, 0)
teamA := createTeamAs(t, s, instanceKey, "team-a")
keyA := createServiceAccount(t, s, instanceKey, "team-a-sa", models.ServiceAccountScopeTeam, teamA)
var integration struct {
Key string `json:"key"`
}
decode(t, s.reqAs(t, keyA, http.MethodPost, "/api/teams/"+id64(teamA)+"/integrations",
map[string]string{"name": "test"}), &integration)
postToIntegration(t, s, integration.Key, "fp-sa-35", "SA35") // incident 1
resp := s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/assign", map[string]any{"user_id": 1})
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("service account assign: %d", resp.StatusCode)
}
resp = s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/resolve", nil)
resp.Body.Close()
resp = s.reqAs(t, keyA, http.MethodPost, "/api/incidents/1/archive", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("service account archive: %d", resp.StatusCode)
}
resp = s.reqAs(t, keyA, http.MethodDelete, "/api/incidents/1/archive", nil)
resp.Body.Close()
if resp.StatusCode != http.StatusNoContent {
t.Fatalf("service account unarchive: %d", resp.StatusCode)
}
var events []map[string]any
decode(t, s.reqAs(t, keyA, http.MethodGet, "/api/incidents/1/timeline", nil), &events)
asg := lastEvent(t, events, "assigned")
if asg["actor_service_account_id"] == nil || asg["actor_user_id"] != nil {
t.Errorf("assigned: expected only the service-account actor, got %v", asg)
}
if asg["user_id"] == nil {
t.Errorf("assigned: user_id must stay the assignee, got %v", asg)
}
for _, typ := range []string{"archived", "unarchived"} {
ev := lastEvent(t, events, typ)
if ev["service_account_id"] == nil || ev["user_id"] != nil {
t.Errorf("%s: expected only the service-account actor, got %v", typ, ev)
}
}
}
+29 -1
View File
@@ -79,6 +79,28 @@ func AuthMiddleware(db *sql.DB) func(http.Handler) http.Handler {
}
}
// securityHeaders sets headers that cost nothing to send on every response,
// API or static site alike. nosniff is unconditional; HSTS only fires once
// cookieSecure's signal says the browser is actually looking at this server
// over HTTPS — TLS terminates at the gateway, which (as of this writing) sets
// neither header itself.
//
// max-age is 180 days rather than the usual year-plus: short enough that if
// HTTPS here ever broke for real, the header would age out of a browser's
// cache well within a release cycle instead of locking anyone out of a
// working server. Raise it once this has run clean for a while.
func securityHeaders(publicURL string) func(http.Handler) http.Handler {
return func(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("X-Content-Type-Options", "nosniff")
if cookieSecure(publicURL, r) {
w.Header().Set("Strict-Transport-Security", "max-age=15552000; includeSubDomains")
}
next.ServeHTTP(w, r)
})
}
}
// AdminOnly rejects a caller who is not a system administrator. It runs inside
// AuthMiddleware's group, so by the time it sees a request the caller is known.
//
@@ -113,10 +135,16 @@ func requireSelfOrAdmin(w http.ResponseWriter, r *http.Request, targetID int64)
}
// apiKeyUser resolves an API key to its user and stamps its last use.
// expires_at IS NULL OR > now is part of the lookup itself, the same way
// serveAs's disabled_at check is: an expired key is one that cannot
// authenticate, by construction, rather than one that happens to still
// resolve and has to be caught afterwards.
func apiKeyUser(ctx context.Context, db *sql.DB, token string) (int64, bool) {
var keyID, userID int64
err := db.QueryRowContext(ctx,
"SELECT id, user_id FROM api_keys WHERE key_hash = $1", hashToken(token),
`SELECT id, user_id FROM api_keys
WHERE key_hash = $1 AND (expires_at IS NULL OR expires_at > $2)`,
hashToken(token), time.Now().Unix(),
).Scan(&keyID, &userID)
if err != nil {
return 0, false
+21 -4
View File
@@ -38,6 +38,13 @@ const (
ackTokenTTL = 24 * time.Hour
)
// notifierLockKey is the Postgres advisory lock the notifier takes for the
// duration of each pass, so that running more than one replica does not
// deliver (or double-deliver) the same notification from more than one of
// them at once. Its value has no meaning beyond being distinct from
// archiverLockKey.
const notifierLockKey int64 = 7265_0002
// Notification kinds, recording why a push was sent.
const (
notifyTriggered = "triggered"
@@ -97,6 +104,10 @@ var notifyClient = &http.Client{Timeout: 10 * time.Second}
// StartNotifier delivers queued notifications until ctx is cancelled, starting
// with an immediate pass so a restart flushes whatever the last one left behind.
//
// Each pass runs under notifierLockKey (see withAdvisoryLock), so that on more
// than one replica only whichever instance's tick takes the lock first actually
// delivers; the rest skip that tick rather than racing the same pass.
func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
if !cfg.enabled() {
log.Print("notifier: disabled (no ntfy URL configured)")
@@ -107,11 +118,17 @@ func StartNotifier(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
ticker := time.NewTicker(notifyInterval)
defer ticker.Stop()
NotifySweep(ctx, db, cfg)
sweep := func() {
withAdvisoryLock(ctx, db, notifierLockKey, "notifier", func() {
NotifySweep(ctx, db, cfg)
})
}
sweep()
for {
select {
case <-ticker.C:
NotifySweep(ctx, db, cfg)
sweep()
case <-ctx.Done():
return
}
@@ -240,7 +257,7 @@ func deliverPending(ctx context.Context, db *sql.DB, cfg NotifyConfig) {
}
// Logged, not returned: the page has already gone out, and treating a
// failed timeline write as a failed delivery would send it again.
if err := logEvent(ctx, db, n.incidentID, eventNotified, n.userID, nil, &n.kind); err != nil {
if err := logEvent(ctx, db, n.incidentID, eventNotified, n.userID, nil, nil, &n.kind); err != nil {
log.Printf("notifier: log delivery of %d: %v", n.id, err)
}
sent++
@@ -295,7 +312,7 @@ func markFailed(ctx context.Context, db *sql.DB, n outboxRow, cause error) {
return
}
detail := fmt.Sprintf("%s: %s", n.kind, cause)
if err := logEvent(ctx, db, n.incidentID, eventNotifyFailed, n.userID, nil, &detail); err != nil {
if err := logEvent(ctx, db, n.incidentID, eventNotifyFailed, n.userID, nil, nil, &detail); err != nil {
log.Printf("notifier: log failure of %d: %v", n.id, err)
}
}
+18 -6
View File
@@ -121,12 +121,12 @@ func ssoRedirect(w http.ResponseWriter, r *http.Request, code ssoError) {
func handleOIDCLogin(db *sql.DB, prov *oidc.Provider, limiter *loginLimiter, publicURL string) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
addrKey := "oidc:" + clientAddr(r)
if limiter.blocked(addrKey, oidcStartMaxPerAddr) {
if limiter.blocked(r.Context(), addrKey, oidcStartMaxPerAddr) {
w.Header().Set("Retry-After", strconv.Itoa(int(loginWindow.Seconds())))
respond(w, http.StatusTooManyRequests, errResp("too many sign-in attempts, try again later"))
return
}
limiter.fail(addrKey)
limiter.fail(r.Context(), addrKey)
state, stateHash, err := randomToken()
if err != nil {
@@ -160,6 +160,9 @@ func handleOIDCLogin(db *sql.DB, prov *oidc.Provider, limiter *loginLimiter, pub
return
}
// #nosec G124 -- HttpOnly/SameSite are literal below; Secure is
// cookieSecure(publicURL, r), not a literal true, which is what
// trips this rule. See cookieSecure's own doc comment in auth.go.
http.SetCookie(w, &http.Cookie{
Name: oidcStateCookie,
Value: state,
@@ -182,6 +185,9 @@ func handleOIDCCallback(db *sql.DB, prov *oidc.Provider, publicURL string) http.
return func(w http.ResponseWriter, r *http.Request) {
// The state cookie has done its job once the callback arrives, whatever
// the outcome.
// #nosec G124 -- HttpOnly/SameSite are literal below; Secure is
// cookieSecure(publicURL, r), not a literal true, which is what
// trips this rule. See cookieSecure's own doc comment in auth.go.
http.SetCookie(w, &http.Cookie{
Name: oidcStateCookie, Value: "", Path: "/api/oidc", MaxAge: -1,
HttpOnly: true, Secure: cookieSecure(publicURL, r), SameSite: http.SameSiteLaxMode,
@@ -189,7 +195,13 @@ func handleOIDCCallback(db *sql.DB, prov *oidc.Provider, publicURL string) http.
q := r.URL.Query()
if e := q.Get("error"); e != "" {
log.Printf("oidc: provider returned error %q: %s", e, q.Get("error_description"))
// %q on both: this runs before state is checked against the
// cookie, so error and error_description are still whatever the
// request's query string says, not yet known to be the real
// provider's. %q keeps a crafted value (say, one holding a
// newline) from forging a second log line rather than just
// being a quoted string within this one.
log.Printf("oidc: provider returned error %q: %q", e, q.Get("error_description")) // #nosec G706 -- both %q
ssoRedirect(w, r, ssoDenied)
return
}
@@ -226,7 +238,7 @@ func handleOIDCCallback(db *sql.DB, prov *oidc.Provider, publicURL string) http.
grants := oidc.ComputeGrants(cfg, identity.Groups)
if !grants.Admitted {
log.Printf("oidc: %q (%s) is in none of the allowed groups", identity.Username, identity.Subject)
log.Printf("oidc: %q (%q) is in none of the allowed groups", identity.Username, identity.Subject) // #nosec G706 -- both %q
ssoRedirect(w, r, ssoNotAllowed)
return
}
@@ -243,11 +255,11 @@ func handleOIDCCallback(db *sql.DB, prov *oidc.Provider, publicURL string) http.
if err != nil {
var se ssoError
if errors.As(err, &se) {
log.Printf("oidc: refused %q (%s): %v", identity.Username, identity.Subject, se)
log.Printf("oidc: refused %q (%q): %v", identity.Username, identity.Subject, se) // #nosec G706 -- both %q
ssoRedirect(w, r, se)
return
}
log.Printf("oidc: sign in %q: %v", identity.Username, err)
log.Printf("oidc: sign in %q: %v", identity.Username, err) // #nosec G706 -- %q
ssoRedirect(w, r, ssoFailed)
return
}
+161
View File
@@ -0,0 +1,161 @@
package api
// This file is internal (package api, not api_test) because loginLimiter and
// its blocked/fail/clear methods are unexported, and TestLoginLimiter_SharedAcrossReplicas
// specifically needs to construct two separate loginLimiter values pointed at
// one database — standing in for two replicas — which only this package can
// do. It duplicates testdb_test.go's newTestDB/withSearchPath rather than
// importing them: those live in the separate api_test package, compiled from
// this directory's external test files, and are not visible here. Same
// reasoning as advisory_lock_test.go, which makes the same trade for the
// same reason.
import (
"context"
"database/sql"
"fmt"
"net/url"
"os"
"strings"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/db"
_ "github.com/jackc/pgx/v5/stdlib"
)
var rateLimiterSchemaSeq int
// rateLimiterTestDB returns a migrated database private to this test.
func rateLimiterTestDB(t *testing.T) *sql.DB {
t.Helper()
dsn := os.Getenv("TERDUT_TEST_DSN")
if dsn == "" {
t.Fatalf("TERDUT_TEST_DSN is not set: these tests need Postgres.\n" +
"Run `make test-db` for a local one, then\n" +
" export TERDUT_TEST_DSN=postgres://terdut:terdut@localhost:5432/terdut_test?sslmode=disable")
}
rateLimiterSchemaSeq++
schema := fmt.Sprintf("test_rl_%d_%d", os.Getpid(), rateLimiterSchemaSeq)
admin, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("connect to TERDUT_TEST_DSN: %v", err)
}
defer admin.Close()
if _, err := admin.Exec("CREATE SCHEMA " + schema); err != nil {
t.Fatalf("create schema %s: %v", schema, err)
}
database, err := db.Open(rateLimiterWithSearchPath(dsn, schema))
if err != nil {
t.Fatalf("open db: %v", err)
}
if err := db.Migrate(database); err != nil {
t.Fatalf("migrate: %v", err)
}
t.Cleanup(func() {
database.Close()
cleanup, err := sql.Open("pgx", dsn)
if err != nil {
return
}
defer cleanup.Close()
if _, err := cleanup.Exec("DROP SCHEMA " + schema + " CASCADE"); err != nil {
t.Logf("drop schema %s: %v", schema, err)
}
})
return database
}
func rateLimiterWithSearchPath(dsn, schema string) string {
opt := "-csearch_path=" + schema
if strings.HasPrefix(dsn, "postgres://") || strings.HasPrefix(dsn, "postgresql://") {
u, err := url.Parse(dsn)
if err == nil {
q := u.Query()
q.Set("options", opt)
u.RawQuery = q.Encode()
return u.String()
}
}
return dsn + " options='" + opt + "'"
}
func TestLoginLimiter_BlocksAtMax(t *testing.T) {
database := rateLimiterTestDB(t)
ctx := context.Background()
l := newLoginLimiter(database)
for range 3 {
if l.blocked(ctx, "k", 3) {
t.Fatal("blocked before reaching max")
}
l.fail(ctx, "k")
}
if !l.blocked(ctx, "k", 3) {
t.Fatal("not blocked after reaching max")
}
}
func TestLoginLimiter_ClearResetsTheCount(t *testing.T) {
database := rateLimiterTestDB(t)
ctx := context.Background()
l := newLoginLimiter(database)
l.fail(ctx, "k")
l.fail(ctx, "k")
l.clear(ctx, "k")
if l.blocked(ctx, "k", 1) {
t.Fatal("still blocked after clear")
}
}
func TestLoginLimiter_KeysAreIndependent(t *testing.T) {
database := rateLimiterTestDB(t)
ctx := context.Background()
l := newLoginLimiter(database)
l.fail(ctx, "a")
if l.blocked(ctx, "b", 1) {
t.Fatal("failing one key blocked an unrelated one")
}
}
// TestLoginLimiter_SharedAcrossReplicas is the regression test for the gap
// this migration closes: an in-memory limiter would let each replica count
// independently, so a caller hitting two different pods could rack up
// max*replicaCount failures before either one blocked. Two loginLimiter
// values sharing one database, standing in for two replicas behind the same
// load balancer, must instead see one combined count.
func TestLoginLimiter_SharedAcrossReplicas(t *testing.T) {
database := rateLimiterTestDB(t)
ctx := context.Background()
replicaA := newLoginLimiter(database)
replicaB := newLoginLimiter(database)
const max = 4
// Alternate which "replica" records the failure, as a real deployment
// would split requests across pods.
for i := range max {
replica := replicaA
if i%2 == 1 {
replica = replicaB
}
if replicaA.blocked(ctx, "k", max) || replicaB.blocked(ctx, "k", max) {
t.Fatalf("blocked after only %d of %d failures", i, max)
}
replica.fail(ctx, "k")
}
if !replicaA.blocked(ctx, "k", max) {
t.Fatal("replica A does not see the combined count as blocked")
}
if !replicaB.blocked(ctx, "k", max) {
t.Fatal("replica B does not see the combined count as blocked")
}
}
+5 -3
View File
@@ -23,13 +23,14 @@ func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config, version strin
// One limiter each, both process-wide for the life of the router: login
// counts failed passwords, sign-up counts account creation, and mixing the
// two would let a burst of sign-ups lock somebody out of logging in.
loginLimit := newLoginLimiter()
signupLimiter := newLoginLimiter()
oidcLimit := newLoginLimiter()
loginLimit := newLoginLimiter(db)
signupLimiter := newLoginLimiter(db)
oidcLimit := newLoginLimiter(db)
r := chi.NewRouter()
r.Use(middleware.Logger)
r.Use(middleware.Recoverer)
r.Use(securityHeaders(notify.PublicURL))
r.Get("/healthz", func(w http.ResponseWriter, r *http.Request) {
respond(w, http.StatusOK, map[string]string{"status": "ok"})
@@ -111,6 +112,7 @@ func NewRouter(db *sql.DB, notify NotifyConfig, cfg config.Config, version strin
r.Get("/api/users/{id}/teams", handleUserTeams(db))
r.Put("/api/users/{id}/notify", handleSetNotifyTarget(db))
r.Put("/api/users/{id}/password", handleSetPassword(db))
r.Get("/api/users/{id}/api-keys", handleListAPIKeys(db))
r.Post("/api/users/{id}/api-keys", handleCreateAPIKey(db))
r.Delete("/api/users/{id}/api-keys/{keyID}", handleDeleteAPIKey(db))
+3
View File
@@ -235,6 +235,9 @@ func scheduleRange(ctx context.Context, db *sql.DB, teamID int64, from, to strin
clause := strings.Join(where, " AND ")
// #nosec G202 -- clause is built from sqlArgs.add's "$N" placeholders
// only, never a value; every value travels through args.all() as a
// bound parameter. See the sqlArgs doc comment in helpers.go.
rows, err := db.QueryContext(ctx, `
SELECT s.id, s.team_id, t.name, s.user_id, u.username, s.date, s.created_at
FROM schedule_entries s
+115
View File
@@ -0,0 +1,115 @@
package api_test
import (
"bytes"
"net/http"
"strings"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/api"
)
// ---------------------------------------------------------------------------
// Security headers
// ---------------------------------------------------------------------------
func TestSecurityHeaders_NosniffAlwaysSet(t *testing.T) {
s := newTS(t) // no PublicURL: the HTTPS signal is off
resp := s.req(t, http.MethodGet, "/api/me", nil)
defer resp.Body.Close()
if got := resp.Header.Get("X-Content-Type-Options"); got != "nosniff" {
t.Errorf("X-Content-Type-Options = %q, want nosniff", got)
}
if got := resp.Header.Get("Strict-Transport-Security"); got != "" {
t.Errorf("Strict-Transport-Security = %q, want unset without an https PublicURL", got)
}
}
func TestSecurityHeaders_HSTSWhenPublicURLIsHTTPS(t *testing.T) {
s := newTS(t, api.NotifyConfig{PublicURL: "https://terdut.example.com"})
resp := s.req(t, http.MethodGet, "/api/me", nil)
defer resp.Body.Close()
got := resp.Header.Get("Strict-Transport-Security")
if !strings.HasPrefix(got, "max-age=") || !strings.Contains(got, "includeSubDomains") {
t.Errorf("Strict-Transport-Security = %q, want a max-age with includeSubDomains", got)
}
}
// ---------------------------------------------------------------------------
// Request body size limits
// ---------------------------------------------------------------------------
// TestBodySizeLimit_OrdinaryEndpointRejectsOversizedBody confirms an
// unauthenticated endpoint can't be made to buffer an arbitrarily large body:
// past maxBodyBytes, decodeJSON fails exactly as it would on any other
// malformed body, rather than the server reading the whole thing first.
func TestBodySizeLimit_OrdinaryEndpointRejectsOversizedBody(t *testing.T) {
s := newTS(t)
huge := bytes.Repeat([]byte("a"), 2<<20) // 2 MiB, past the 1 MiB default
body := []byte(`{"username":"` + string(huge) + `","password":"x"}`)
resp, err := http.Post(s.URL+"/api/login", "application/json", bytes.NewReader(body))
if err != nil {
t.Fatalf("POST /api/login: %v", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("status = %d, want %d (oversized body treated as invalid)", resp.StatusCode, http.StatusBadRequest)
}
}
// TestBodySizeLimit_WebhookAllowsLargerBodyThanDefault confirms the
// Alertmanager webhook's separate, larger cap actually takes effect: a body
// bigger than the ordinary default but within maxWebhookBodyBytes is still
// accepted, not rejected by the smaller limit every other endpoint gets.
func TestBodySizeLimit_WebhookAllowsLargerBodyThanDefault(t *testing.T) {
s := newTS(t)
// Padding kept inside one alert's annotation, comfortably past the 1 MiB
// default and still well under the webhook's 8 MiB cap.
padding := strings.Repeat("a", 3<<20) // 3 MiB
payload := `{"version":"4","status":"firing","groupKey":"big-group",` +
`"groupLabels":{"alertname":"BigAlert"},"alerts":[{"status":"firing",` +
`"labels":{"alertname":"BigAlert"},"annotations":{"note":"` + padding + `"},` +
`"startsAt":"2026-05-20T10:00:00Z","endsAt":"0001-01-01T00:00:00Z",` +
`"fingerprint":"fp-big"}]}`
resp, err := http.Post(s.URL+"/api/integrations/"+s.ingestKey+"/alertmanager",
"application/json", strings.NewReader(payload))
if err != nil {
t.Fatalf("POST webhook: %v", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Errorf("status = %d, want %d (body under the webhook's own cap)", resp.StatusCode, http.StatusOK)
}
}
// TestBodySizeLimit_WebhookRejectsPastItsOwnCap confirms the webhook's larger
// cap is still a cap, not an exemption from one.
func TestBodySizeLimit_WebhookRejectsPastItsOwnCap(t *testing.T) {
s := newTS(t)
huge := strings.Repeat("a", 9<<20) // 9 MiB, past the 8 MiB webhook cap
payload := `{"version":"4","status":"firing","groupKey":"huge-group",` +
`"groupLabels":{"alertname":"HugeAlert"},"alerts":[{"status":"firing",` +
`"labels":{"alertname":"HugeAlert"},"annotations":{"note":"` + huge + `"},` +
`"startsAt":"2026-05-20T10:00:00Z","endsAt":"0001-01-01T00:00:00Z",` +
`"fingerprint":"fp-huge"}]}`
resp, err := http.Post(s.URL+"/api/integrations/"+s.ingestKey+"/alertmanager",
"application/json", strings.NewReader(payload))
if err != nil {
t.Fatalf("POST webhook: %v", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusBadRequest {
t.Errorf("status = %d, want %d (body past the webhook's own cap)", resp.StatusCode, http.StatusBadRequest)
}
}
+2 -2
View File
@@ -112,7 +112,7 @@ var errInviteUnusable = errors.New("invite is not usable")
func handleSignup(db *sql.DB, limiter *loginLimiter, publicURL string) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
addr := clientAddr(r)
if limiter.blocked("signup:"+addr, maxSignupsPerAddr) {
if limiter.blocked(r.Context(), "signup:"+addr, maxSignupsPerAddr) {
respond(w, http.StatusTooManyRequests, errResp("too many sign-ups from this address"))
return
}
@@ -148,7 +148,7 @@ func handleSignup(db *sql.DB, limiter *loginLimiter, publicURL string) http.Hand
var err error
inv, err = loadInvite(r.Context(), db, req.Invite)
if err != nil {
limiter.fail("signup:" + addr)
limiter.fail(r.Context(), "signup:"+addr)
respond(w, http.StatusForbidden, errResp("this invite link is not usable"))
return
}
+12
View File
@@ -73,6 +73,10 @@ func handleStatsTop(db *sql.DB) http.HandlerFunc {
}
result = append(result, e)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, result)
}
}
@@ -104,6 +108,10 @@ func handleStatsByHour(db *sql.DB) http.HandlerFunc {
}
counts[hr] = cnt
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
type entry struct {
Hour int `json:"hour"`
@@ -146,6 +154,10 @@ func handleStatsByDay(db *sql.DB) http.HandlerFunc {
}
counts[dow] = cnt
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
dayNames := [7]string{"Sunday", "Monday", "Tuesday", "Wednesday", "Thursday", "Friday", "Saturday"}
type entry struct {
+80 -3
View File
@@ -116,6 +116,10 @@ func handleListUsers(db *sql.DB) http.HandlerFunc {
u.DisabledAt = unixPtr(disabled)
users = append(users, u)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, users)
}
}
@@ -234,6 +238,11 @@ func handleDeleteUser(db *sql.DB) http.HandlerFunc {
}
}
// maxAPIKeyExpiryDays bounds expires_in_days: generous enough for any real
// rotation policy, tight enough to reject a typo (a year in hours, say) that
// would otherwise mint a key that outlives the server by decades.
const maxAPIKeyExpiryDays = 3650 // ~10 years
func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
userID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
@@ -247,6 +256,11 @@ func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
var req struct {
Name string `json:"name"`
// ExpiresInDays is optional and, left zero, means the key never
// expires — the only behavior any key had before this field
// existed, so an existing integration that does not send it is
// unaffected.
ExpiresInDays int64 `json:"expires_in_days,omitempty"`
}
if err := decodeJSON(r, &req); err != nil {
respond(w, http.StatusBadRequest, errResp("invalid request body"))
@@ -256,6 +270,10 @@ func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusBadRequest, errResp("name is required"))
return
}
if req.ExpiresInDays < 0 || req.ExpiresInDays > maxAPIKeyExpiryDays {
respond(w, http.StatusBadRequest, errResp("expires_in_days must be 0 (never expires) or up to "+strconv.Itoa(maxAPIKeyExpiryDays)))
return
}
var exists int
if err := db.QueryRowContext(r.Context(), "SELECT 1 FROM users WHERE id = $1", userID).Scan(&exists); err != nil {
@@ -268,18 +286,77 @@ func handleCreateAPIKey(db *sql.DB) http.HandlerFunc {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
var expiresAt *int64
var expiresAtTime *time.Time
if req.ExpiresInDays > 0 {
t := time.Now().AddDate(0, 0, int(req.ExpiresInDays)).UTC()
u := t.Unix()
expiresAt = &u
expiresAtTime = &t
}
var keyID int64
if err := db.QueryRowContext(r.Context(),
"INSERT INTO api_keys (user_id, key_hash, name) VALUES ($1, $2, $3) RETURNING id",
userID, hash, req.Name).Scan(&keyID); err != nil {
"INSERT INTO api_keys (user_id, key_hash, name, expires_at) VALUES ($1, $2, $3, $4) RETURNING id",
userID, hash, req.Name, expiresAt).Scan(&keyID); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
key := models.APIKey{ID: keyID, UserID: userID, Name: req.Name, Key: raw, CreatedAt: time.Now().UTC()}
key := models.APIKey{
ID: keyID, UserID: userID, Name: req.Name, Key: raw,
CreatedAt: time.Now().UTC(), ExpiresAt: expiresAtTime,
}
respond(w, http.StatusCreated, key)
}
}
// handleListAPIKeys lists a user's own API keys: never the raw key itself
// (only ever returned once, at creation), just enough to tell them apart,
// see which are stale (last_used_at) and which are about to stop working
// (expires_at) — the data handleCreateAPIKey and apiKeyUser's last-use stamp
// already produce, with no endpoint to read it back until now.
func handleListAPIKeys(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
userID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
if err != nil {
respond(w, http.StatusBadRequest, errResp("invalid user id"))
return
}
if !requireSelfOrAdmin(w, r, userID) {
return
}
rows, err := db.QueryContext(r.Context(),
`SELECT id, name, created_at, last_used_at, expires_at
FROM api_keys WHERE user_id = $1 ORDER BY created_at DESC`, userID)
if err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
defer rows.Close()
keys := []models.APIKey{}
for rows.Next() {
var k models.APIKey
var created int64
var lastUsed, expires *int64
if err := rows.Scan(&k.ID, &k.Name, &created, &lastUsed, &expires); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
k.UserID = userID
k.CreatedAt = time.Unix(created, 0).UTC()
k.LastUsedAt = unixPtr(lastUsed)
k.ExpiresAt = unixPtr(expires)
keys = append(keys, k)
}
if err := rows.Err(); err != nil {
respond(w, http.StatusInternalServerError, errResp("internal error"))
return
}
respond(w, http.StatusOK, keys)
}
}
func handleDeleteAPIKey(db *sql.DB) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
userID, err := strconv.ParseInt(chi.URLParam(r, "id"), 10, 64)
+27
View File
@@ -1,6 +1,7 @@
package db
import (
"context"
"database/sql"
"embed"
"fmt"
@@ -62,6 +63,16 @@ func Open(dsn string) (*sql.DB, error) {
}
}
// migrationLockKey is the Postgres advisory lock Migrate holds for its whole
// run. Two replicas starting at once would otherwise race the check-then-apply
// loop below against schema_migrations: the loser could crash on a
// duplicate-key insert, or contend with the winner's uncommitted DDL. Blocking
// (pg_advisory_lock, not pg_try_advisory_lock as the archiver and notifier
// use): on boot there is no later tick to defer to, so the right behaviour is
// to wait for the other replica to finish migrating, not to skip ahead and
// start serving against an unmigrated schema.
const migrationLockKey int64 = 7265_0003
// Migrate applies every embedded migration that has not been applied yet, in
// filename order, recording each in schema_migrations.
//
@@ -69,6 +80,22 @@ func Open(dsn string) (*sql.DB, error) {
// migration that failed half way used to leave the schema in whatever state it
// had reached. Postgres has transactional DDL, so the rollback is real.
func Migrate(db *sql.DB) error {
ctx := context.Background()
conn, err := db.Conn(ctx)
if err != nil {
return fmt.Errorf("migrate: acquire connection: %w", err)
}
defer conn.Close()
if _, err := conn.ExecContext(ctx, "SELECT pg_advisory_lock($1)", migrationLockKey); err != nil {
return fmt.Errorf("migrate: acquire advisory lock: %w", err)
}
defer func() {
if _, err := conn.ExecContext(ctx, "SELECT pg_advisory_unlock($1)", migrationLockKey); err != nil {
log.Printf("migrate: release advisory lock: %v", err)
}
}()
if _, err := db.Exec(`CREATE TABLE IF NOT EXISTS schema_migrations (
version TEXT PRIMARY KEY,
applied_at BIGINT NOT NULL DEFAULT FLOOR(EXTRACT(EPOCH FROM now()))::bigint
+125
View File
@@ -0,0 +1,125 @@
package db_test
import (
"database/sql"
"fmt"
"net/url"
"os"
"strings"
"sync"
"testing"
"git.ryuvia.com/niklas/terdut-server/internal/db"
_ "github.com/jackc/pgx/v5/stdlib"
)
// TERDUT_TEST_DSN must point at a database the test role may create schemas
// in; see internal/api/testdb_test.go for the fuller rationale this mirrors.
// An unset DSN fails rather than skips, deliberately.
const testDSNEnv = "TERDUT_TEST_DSN"
// TestMigrate_ConcurrentCallersDoNotRace reproduces two replicas starting at
// once against a brand-new, unmigrated schema: both call db.Migrate at the
// same time. Before migrationLockKey, the loser could crash on a
// duplicate-key insert into schema_migrations, or contend with the winner's
// uncommitted DDL; with the advisory lock, one blocks until the other
// finishes and both return cleanly.
func TestMigrate_ConcurrentCallersDoNotRace(t *testing.T) {
dsn := os.Getenv(testDSNEnv)
if dsn == "" {
t.Fatalf("%s is not set: these tests need Postgres.\n"+
"Run `make test-db` for a local one, then\n"+
" export %s=postgres://terdut:terdut@localhost:5432/terdut_test?sslmode=disable",
testDSNEnv, testDSNEnv)
}
schema := fmt.Sprintf("migrate_race_%d", os.Getpid())
admin, err := sql.Open("pgx", dsn)
if err != nil {
t.Fatalf("connect to %s: %v", testDSNEnv, err)
}
defer admin.Close()
if _, err := admin.Exec("CREATE SCHEMA " + schema); err != nil {
t.Fatalf("create schema %s: %v", schema, err)
}
t.Cleanup(func() {
cleanup, err := sql.Open("pgx", dsn)
if err != nil {
return
}
defer cleanup.Close()
if _, err := cleanup.Exec("DROP SCHEMA " + schema + " CASCADE"); err != nil {
t.Logf("drop schema %s: %v", schema, err)
}
})
scoped := withSearchPath(dsn, schema)
const callers = 2
errs := make([]error, callers)
var wg sync.WaitGroup
for i := range callers {
wg.Add(1)
go func(i int) {
defer wg.Done()
database, err := db.Open(scoped)
if err != nil {
errs[i] = fmt.Errorf("open: %w", err)
return
}
defer database.Close()
errs[i] = db.Migrate(database)
}(i)
}
wg.Wait()
for i, err := range errs {
if err != nil {
t.Fatalf("Migrate #%d: %v", i, err)
}
}
entries, err := os.ReadDir("migrations")
if err != nil {
t.Fatalf("read migrations dir: %v", err)
}
var want int
for _, e := range entries {
if !e.IsDir() && strings.HasSuffix(e.Name(), ".sql") {
want++
}
}
check, err := sql.Open("pgx", scoped)
if err != nil {
t.Fatalf("connect for verification: %v", err)
}
defer check.Close()
var got int
if err := check.QueryRow("SELECT COUNT(*) FROM schema_migrations").Scan(&got); err != nil {
t.Fatalf("count schema_migrations: %v", err)
}
if got != want {
t.Fatalf("schema_migrations has %d row(s) after two concurrent Migrate calls, want %d (one per migration file, no duplicates)", got, want)
}
}
// withSearchPath pins a DSN to one schema. Copied from
// internal/api/testdb_test.go rather than shared: that helper lives in the
// api_test package, a separate compiled package this one cannot import.
func withSearchPath(dsn, schema string) string {
opt := "-csearch_path=" + schema
if strings.HasPrefix(dsn, "postgres://") || strings.HasPrefix(dsn, "postgresql://") {
u, err := url.Parse(dsn)
if err == nil {
q := u.Query()
q.Set("options", opt)
u.RawQuery = q.Encode()
return u.String()
}
}
return dsn + " options='" + opt + "'"
}
@@ -0,0 +1,39 @@
-- Service-account actors on incident mutations (terdut-server#25). A
-- team-scoped service account acknowledging/resolving/snoozing/noting an
-- incident is not a users row, so it cannot be written into
-- acknowledged_by/incident_events.user_id — doing so either violates the
-- users(id) FK (new rows) or, for incident_events.user_id, silently matches
-- zero rows on delete. These columns are the service-account-shaped parallel
-- to the existing human ones: nullable, mutually exclusive with their human
-- counterpart, ON DELETE SET NULL so a deleted service account doesn't take
-- the incident history with it.
ALTER TABLE incidents
ADD COLUMN acknowledged_by_service_account_id BIGINT
REFERENCES service_accounts(id) ON DELETE SET NULL;
ALTER TABLE incident_events
ADD COLUMN service_account_id BIGINT
REFERENCES service_accounts(id) ON DELETE SET NULL;
-- At most one actor kind per row: both NULL ("the server acted") is valid,
-- exactly one set is valid, both set is a bug this constraint refuses to
-- store rather than silently accepting.
ALTER TABLE incidents
ADD CONSTRAINT incidents_ack_actor_xor_chk CHECK (
acknowledged_by IS NULL OR acknowledged_by_service_account_id IS NULL
);
ALTER TABLE incident_events
ADD CONSTRAINT incident_events_actor_xor_chk CHECK (
user_id IS NULL OR service_account_id IS NULL
);
CREATE INDEX incidents_acknowledged_by_service_account_id_idx
ON incidents(acknowledged_by_service_account_id);
CREATE INDEX incident_events_service_account_id_idx
ON incident_events(service_account_id);
-- assigned_to_service_account_id is deliberately not added here: it would sit
-- unpopulated until handleIncidentAssign itself tracks an actor, which is a
-- separate, pre-existing gap (it records the assignee today, never the
-- actor, for humans either) tracked in its own follow-up issue.
@@ -0,0 +1,16 @@
-- Backs the rate limiters (failed logins, sign-ups, OIDC/device start) with
-- Postgres instead of an in-memory map, now that the server runs more than
-- one replica in production (v0.37.0): a counter that only ever sees its own
-- pod's traffic quietly let every one of these limits through multiplied by
-- the replica count.
--
-- window_start is the start of the current fixed window for key, in the same
-- "unix seconds" shape every other timestamp in this schema uses. The window
-- resets rather than slides, matching the in-memory limiter it replaces:
-- once a key's window is older than the limiter's window length, the next
-- failure starts a fresh one instead of extending the stale one.
CREATE TABLE rate_limit_counters (
key TEXT PRIMARY KEY,
window_start BIGINT NOT NULL,
count INT NOT NULL
);
@@ -0,0 +1,7 @@
-- Optional expiry on a user's own API keys. NULL (the existing default for
-- every row already in this table) means "never expires" -- the same
-- behavior these keys have always had, so no existing integration breaks.
-- Service account keys are deliberately NOT touched: they are a different
-- table, managed by automation, and already distinguished by their own
-- "tdsa_" prefix.
ALTER TABLE api_keys ADD COLUMN expires_at BIGINT;
@@ -0,0 +1,21 @@
-- Who performed an assignment (terdut-server#35). On an 'assigned' event
-- incident_events.user_id is the assignee, so the actor needs columns of its
-- own. Only populated for 'assigned' events; every other event type keeps
-- using user_id/service_account_id for the actor. Older 'assigned' rows stay
-- NULL (the actor was never recorded). Same shape as migration 015: nullable,
-- mutually exclusive, ON DELETE SET NULL.
--
-- assigned_to_service_account_id is still deliberately not added: making
-- service accounts assignable is a separate change (request body, assignee
-- picker, notifier, filters).
ALTER TABLE incident_events
ADD COLUMN actor_user_id BIGINT REFERENCES users(id) ON DELETE SET NULL,
ADD COLUMN actor_service_account_id BIGINT REFERENCES service_accounts(id) ON DELETE SET NULL;
ALTER TABLE incident_events
ADD CONSTRAINT incident_events_assign_actor_xor_chk CHECK (
actor_user_id IS NULL OR actor_service_account_id IS NULL
);
CREATE INDEX incident_events_actor_user_id_idx ON incident_events(actor_user_id);
CREATE INDEX incident_events_actor_service_account_id_idx ON incident_events(actor_service_account_id);
+36 -10
View File
@@ -43,6 +43,13 @@ type Incident struct {
AcknowledgedByUser *string `json:"acknowledged_by,omitempty"`
AcknowledgedAt *time.Time `json:"acknowledged_at,omitempty"`
// AcknowledgedByServiceAccountID/Name are the service-account-shaped
// parallel to AcknowledgedByID/User above: mutually exclusive with it,
// populated when a service account (not a human) acknowledged this
// incident. See migration 015 and terdut-server#25.
AcknowledgedByServiceAccountID *int64 `json:"acknowledged_by_service_account_id,omitempty"`
AcknowledgedByServiceAccountName *string `json:"acknowledged_by_service_account,omitempty"`
AssignedToID *int64 `json:"assigned_to_id,omitempty"`
AssignedToUser *string `json:"assigned_to,omitempty"`
@@ -67,17 +74,36 @@ type Incident struct {
// and is the only history this server keeps — alert rows are mutated in place.
//
// Type is one of: triggered, alert_added, alert_resolved, acknowledged,
// unacknowledged, assigned, snoozed, unsnoozed, resolved, note. A nil UserID
// means the server acted rather than a person.
// unacknowledged, assigned, archived, unarchived, snoozed, unsnoozed,
// resolved, note. UserID and
// ServiceAccountID are mutually exclusive; both nil means the server acted
// rather than any caller.
type IncidentEvent struct {
ID int64 `json:"id"`
IncidentID int64 `json:"incident_id"`
Type string `json:"type"`
UserID *int64 `json:"user_id,omitempty"`
Username *string `json:"username,omitempty"`
AlertID *int64 `json:"alert_id,omitempty"`
Detail *string `json:"detail,omitempty"`
CreatedAt time.Time `json:"created_at"`
ID int64 `json:"id"`
IncidentID int64 `json:"incident_id"`
Type string `json:"type"`
UserID *int64 `json:"user_id,omitempty"`
Username *string `json:"username,omitempty"`
// ServiceAccountID/Name are the service-account-shaped parallel to
// UserID/Username above: mutually exclusive with it, populated when a
// service account (not a human, and not nil-meaning-the-server-acted)
// performed this event. Named Name, not Username — a ServiceAccount has
// a Name field, not a Username. See migration 015 and terdut-server#25.
ServiceAccountID *int64 `json:"service_account_id,omitempty"`
ServiceAccountName *string `json:"service_account_name,omitempty"`
// Actor* name who performed an 'assigned' event, whose UserID is the
// assignee. Mutually exclusive; unset on every other event type and on
// assignments made before migration 018. See terdut-server#35.
ActorUserID *int64 `json:"actor_user_id,omitempty"`
ActorUsername *string `json:"actor_username,omitempty"`
ActorServiceAccountID *int64 `json:"actor_service_account_id,omitempty"`
ActorServiceAccountName *string `json:"actor_service_account_name,omitempty"`
AlertID *int64 `json:"alert_id,omitempty"`
Detail *string `json:"detail,omitempty"`
CreatedAt time.Time `json:"created_at"`
}
// SimilarIncident is an earlier, resolved incident with the same signature as
+9 -1
View File
@@ -37,5 +37,13 @@ type APIKey struct {
Name string `json:"name"`
CreatedAt time.Time `json:"created_at"`
LastUsedAt *time.Time `json:"last_used_at,omitempty"`
Key string `json:"key,omitempty"` // populated only on creation, never stored
// ExpiresAt is nil for a key that never expires, which is every key
// created before this field existed and still the default for a new one
// unless its creator asks otherwise (see handleCreateAPIKey's
// expires_in_days). AuthMiddleware stops accepting a key once this
// passes; nothing deletes the row for it.
ExpiresAt *time.Time `json:"expires_at,omitempty"`
Key string `json:"key,omitempty"` // populated only on creation, never stored
}
+135 -36
View File
@@ -10,11 +10,11 @@
--surface: #ffffff;
--surface-2: #eff1f4;
--surface-hover: #f7f8fa;
--border: #e2e5ea;
--border-strong: #cfd3da;
--border: #d9dde4;
--border-strong: #c5cad3;
--text: #16181d;
--muted: #5b626e;
--faint: #8a909b;
--faint: #676d79;
--accent: #2f5bd3;
--accent-text: #ffffff;
@@ -72,17 +72,21 @@
--safe-bottom: env(safe-area-inset-bottom, 0px);
}
/* Dark palette. Applied by the OS preference unless the user chose Light in
Account (data-theme="light" on <html>, set by js/theme.js), and always when
they chose Dark. The two blocks must stay identical: CSS has no way to share
a declaration list between a media query and an attribute selector. */
@media (prefers-color-scheme: dark) {
:root {
:root:not([data-theme="light"]) {
--bg: #0f1115;
--surface: #171a20;
--surface-2: #1f232b;
--surface-hover: #1c2027;
--border: #2a2f38;
--border-strong: #394050;
--surface: #1a1e26;
--surface-2: #232834;
--surface-hover: #20252e;
--border: #343b49;
--border-strong: #444d5f;
--text: #e7e9ed;
--muted: #a0a7b3;
--faint: #737a87;
--faint: #8a92a0;
--accent: #6d8ff0;
--accent-text: #0b0d12;
@@ -108,6 +112,42 @@
--shadow-lg: 0 16px 40px rgb(0 0 0 / 55%);
}
}
:root[data-theme="dark"] {
--bg: #0f1115;
--surface: #1a1e26;
--surface-2: #232834;
--surface-hover: #20252e;
--border: #343b49;
--border-strong: #444d5f;
--text: #e7e9ed;
--muted: #a0a7b3;
--faint: #8a92a0;
--accent: #6d8ff0;
--accent-text: #0b0d12;
--accent-soft: #1d2640;
--crit: #ff6b61;
--crit-soft: #3a1c1b;
--warn: #f0b140;
--warn-soft: #362a14;
--info: #74a3ff;
--info-soft: #1a2640;
--ok: #4cc488;
--ok-soft: #15301f;
--snooze: #a89bff;
--snooze-soft: #262245;
--teal: #4fc2d4;
--teal-soft: #0f2e33;
--pink: #f07fb8;
--pink-soft: #3a1c2d;
--shadow: 0 1px 2px rgb(0 0 0 / 40%);
--shadow-lg: 0 16px 40px rgb(0 0 0 / 55%);
}
:root[data-theme="light"] { color-scheme: light; }
:root[data-theme="dark"] { color-scheme: dark; }
*, *::before, *::after { box-sizing: border-box; }
[hidden] { display: none !important; }
@@ -237,6 +277,10 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.open-pill.has-triggered { background: var(--crit-soft); color: var(--crit); }
.open-pill.has-triggered::before { background: var(--crit); }
.open-pill.all-acked::before { background: var(--warn); }
/* "All clear" is a status, not an action: green with a check, no dot. */
.open-pill.all-clear { background: var(--ok-soft); color: var(--ok); }
.open-pill.all-clear::before { content: none; }
.open-pill .icon { width: 14px; height: 14px; }
/* Bottom tab bar on phones (Queue, On-call, Alerts, Team, More); becomes the
left sidebar from 900px, where the desktop block below redeclares display
@@ -290,13 +334,13 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
.nav-team-selector,
.team-selector-mobile {
display: inline-flex; align-items: center; gap: 8px;
align-items: center; gap: 8px;
border: 1px solid var(--border-strong); border-radius: 999px;
background: var(--surface); color: var(--text);
font-size: 13px; font-weight: 600; cursor: pointer;
padding: 4px 12px; max-width: 100%;
}
.team-selector-mobile { padding: 4px 10px; font-size: 12px; max-width: 120px; }
.team-selector-mobile { display: inline-flex; padding: 4px 10px; font-size: 12px; max-width: 120px; }
.team-selector-label { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
.team-selector-chevron { width: 14px; height: 14px; flex: none; color: var(--faint); margin-left: -2px; }
@@ -340,7 +384,8 @@ input:focus, textarea:focus { outline: none; border-color: var(--accent); box-sh
background: var(--surface); color: var(--muted);
font-size: 13px; font-weight: 600; cursor: pointer;
}
.chip[aria-selected="true"] { background: var(--text); border-color: var(--text); color: var(--bg); }
.chip[aria-selected="true"], .chip[aria-checked="true"] { background: var(--text); border-color: var(--text); color: var(--bg); }
.theme-picker { padding: 0 0 8px; }
.chip .count { margin-left: 4px; opacity: 0.7; }
/* An overlay, not a flex item: absolute against .chips' own (non-scrolling)
box stays flush with its real right edge regardless of scroll position,
@@ -583,6 +628,10 @@ details[open] > summary { margin-bottom: 8px; }
}
.actionbar .btn { min-height: 48px; }
.actionbar .btn-primary { flex: 1; font-size: 16px; }
/* The extra actions that only phones hide behind "More" — see actionBar() and
extraActions() in incident.js. Hidden by default; the desktop block below
shows them and hides the now-redundant More button instead. */
.action-extra { display: none; }
.detail-placeholder {
display: grid; place-items: center; height: 100%;
@@ -646,7 +695,6 @@ details[open] > summary { margin-bottom: 8px; }
.page-head { display: flex; align-items: center; justify-content: space-between; gap: 8px; margin: 16px auto 12px; }
.page-head h2 { font-size: 13px; font-weight: 700; text-transform: uppercase; letter-spacing: 0.06em; color: var(--muted); }
.now-card { display: flex; align-items: center; gap: 14px; padding: 16px; margin-top: 16px; }
.avatar {
flex: none; display: grid; place-items: center;
width: 44px; height: 44px; border-radius: 50%;
@@ -654,8 +702,6 @@ details[open] > summary { margin-bottom: 8px; }
font-weight: 750; font-size: 17px; text-transform: uppercase;
}
.avatar.none { background: var(--surface-2); color: var(--faint); }
.now-label { color: var(--muted); font-size: 13px; font-weight: 600; }
.now-name { font-size: 20px; font-weight: 750; }
.you { color: var(--accent); font-weight: 650; font-size: 13px; margin-left: 6px; }
/* Oncall's own "you" indicator only (see you() in oncall.js) — a pill badge
is easier to spot there than this plain accent-coloured text. */
@@ -669,25 +715,49 @@ details[open] > summary { margin-bottom: 8px; }
caught by the unrelated .label > span styling meant for label chips. */
.week-nav .week-label { font-size: 14px; font-weight: 650; min-width: 11.5em; text-align: center; white-space: nowrap; }
.week-label small { color: var(--faint); font-weight: 600; font-size: 11px; margin-left: 2px; }
.days { list-style: none; margin: 0; padding: 0; }
.day { display: grid; grid-template-columns: 3.2em 4.2em 1fr; align-items: center; gap: 8px; min-height: 50px; padding: 0 14px; }
.day + .day { border-top: 1px solid var(--border); }
.day-name { font-weight: 650; }
.day-date { color: var(--faint); font-size: 13px; }
.day-who { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
.day-who.nobody { color: var(--faint); font-style: italic; }
/* A run of several days held by the same person (or left empty), replacing
what used to be one identical row per day — see weekRuns() in oncall.js. */
.day.range { grid-template-columns: 1fr auto; }
.day-range { font-weight: 650; }
.day.today { background: var(--accent-soft); }
.day.today:first-child { border-radius: var(--radius) var(--radius) 0 0; }
.day.today:last-child { border-radius: 0 0 var(--radius) var(--radius); }
.day.today .day-name, .day.today .day-range { color: var(--accent); }
.day.past { opacity: 0.6; }
/* Highlights whichever row is yours, same soft tint as .today — they already
read fine layered (today's own row is almost always one of yours too). */
.day.mine { background: var(--accent-soft); }
.oncall-title h1 { font-size: 24px; font-weight: 600; letter-spacing: -0.01em; margin: 8px 0 0; }
.oncall-title p { margin: 2px 0 0; font-size: 14px; }
.oncall-grid { display: grid; gap: 0; }
.oncall-main, .oncall-side { min-width: 0; }
.hero-list { display: grid; gap: 12px; }
.hero { display: grid; gap: 14px; padding: 18px; margin-top: 16px; }
.hero-top { display: flex; align-items: center; gap: 16px; }
.hero-label { font-size: 12px; font-weight: 700; text-transform: uppercase; letter-spacing: 0.07em; color: var(--muted); }
.hero-avatar { width: 64px; height: 64px; font-size: 24px; }
.hero-name { font-size: 24px; font-weight: 650; letter-spacing: -0.01em; display: flex; align-items: center; flex-wrap: wrap; gap: 4px; }
.hero-until { color: var(--muted); font-size: 15px; margin-top: 2px; }
.week-card { margin-top: 16px; padding: 16px 12px 12px; }
.week-head { display: flex; align-items: center; justify-content: space-between; gap: 8px; margin-bottom: 12px; padding: 0 4px; }
.week-head h2 { margin: 0; font-size: 15px; font-weight: 650; }
/* Bordered, so the arrows read as buttons. 40px: a touch target without
crowding the date between them. */
.week-arrow { width: 40px; min-height: 40px; border: 1px solid var(--border-strong); }
.strip { list-style: none; margin: 0; padding: 0; display: grid; grid-template-columns: repeat(7, minmax(0, 1fr)); gap: 2px; }
.strip-day {
position: relative; min-width: 0;
display: flex; flex-direction: column; align-items: center; gap: 4px;
padding: 10px 0 8px; border: 1px solid transparent; border-radius: var(--radius);
}
.strip-name { font-size: 12px; font-weight: 650; text-transform: uppercase; letter-spacing: 0.04em; color: var(--muted); }
.strip-date { font-size: 15px; font-weight: 650; }
.strip-avatar { width: 32px; height: 32px; font-size: 13px; }
.strip-who { max-width: 100%; padding: 0 2px; font-size: 12px; color: var(--muted); overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
.strip-who.nobody { font-style: italic; color: var(--faint); }
.strip-day.past { opacity: 0.6; }
/* "Current", not "selected": a neutral tint, a bar and aria-current, so it
reads as a marker and not as a pressed button. The accent stays for things
you can act on. */
.strip-day.today { background: var(--surface-2); border-color: var(--border-strong); }
.strip-day.today::before {
content: ""; position: absolute; top: -1px; left: 12px; right: 12px; height: 3px;
border-radius: 0 0 3px 3px; background: var(--text);
}
.strip-day.today .strip-name { color: var(--text); }
.strip-legend { display: flex; flex-wrap: wrap; gap: 4px 14px; margin-top: 10px; padding: 10px 4px 0; border-top: 1px solid var(--border); font-size: 13px; color: var(--muted); }
.strip-legend span { display: inline-flex; align-items: center; gap: 6px; }
.strip-legend .strip-dot { width: 8px; height: 8px; border-radius: 50%; background: currentColor; }
.shift-list { list-style: none; margin: 0; padding: 0; }
.shift-list li { display: flex; justify-content: space-between; padding: 12px 14px; }
@@ -773,7 +843,15 @@ kbd {
.pane-list .chips { flex-wrap: wrap; overflow-x: visible; }
.pane-list .chip-sep { display: none; }
.chips-fade { display: none; }
.view-queue:not(.has-detail) .pane-detail { display: block; }
/* With nothing selected there is no detail to show next to, so the list
takes the whole row instead of leaving the second column as dead space
around the placeholder text. Selecting an incident (.has-detail) drops
back to the base minmax(340,420) 1fr rule above. */
.view-queue:not(.has-detail) { grid-template-columns: 1fr; }
.view-queue:not(.has-detail) .pane-list { border-right: 0; }
/* Full width reads better capped than edge-to-edge on a very wide monitor,
matching .detail's own cap below. */
.view-queue:not(.has-detail) .list { max-width: 900px; margin: 0 auto; }
/* On desktop the list stays visible next to the detail. */
.app.detail-open .nav { display: flex; }
@@ -787,8 +865,12 @@ kbd {
.actionbar {
position: sticky; bottom: 0;
padding: 12px 32px;
flex-wrap: wrap;
}
.actionbar .btn-primary { flex: 0 1 240px; }
/* Room enough to show every action, so More is fully redundant here. */
.action-extra { display: inline-flex; }
.more-btn { display: none; }
.sheet {
width: min(440px, calc(100% - 32px));
@@ -829,6 +911,11 @@ kbd {
.status-table th, .status-table td { white-space: nowrap; }
.status-table td.wrap { white-space: normal; min-width: 12em; }
.status-table .source-row td { border-bottom-style: dashed; }
.status-table tr.clickable { cursor: pointer; }
.status-table tr.clickable:hover td, .status-table tr.clickable:focus-visible td { background: rgba(127, 127, 127, .1); }
.switch-facts { display: grid; grid-template-columns: max-content 1fr; gap: 6px 16px; margin: 12px 0; }
.switch-facts dt { opacity: .7; }
.switch-facts dd { margin: 0; }
.status-table .source-row td:first-child { padding-left: 16px; }
.source-labels { display: flex; flex-wrap: wrap; gap: 4px; align-items: center; }
@@ -1017,6 +1104,11 @@ button.rota-week:hover { background: var(--surface-2); color: var(--text); }
overflow-x: auto; scrollbar-width: none;
}
.subnav::-webkit-scrollbar { display: none; }
/* A fade on the right edge says "there is more" where the strip overflows;
the desktop width fits every entry, so it is phone-only. */
@media (max-width: 899px) {
.subnav { -webkit-mask-image: linear-gradient(90deg, #000 calc(100% - 28px), transparent); mask-image: linear-gradient(90deg, #000 calc(100% - 28px), transparent); }
}
.subnav-link {
flex: none;
padding: 8px 12px; margin-bottom: -1px;
@@ -1076,3 +1168,10 @@ button.rota-week:hover { background: var(--surface-2); color: var(--text); }
@media (min-width: 900px) {
.stat-tiles { grid-template-columns: repeat(3, 1fr); }
}
/* On-call, wide: the hero and the week on the left, your shifts on the right.
.view-page > * caps its children at 760px, so the grid lifts that. */
@media (min-width: 900px) {
#view-oncall > .oncall-grid { max-width: 1080px; grid-template-columns: minmax(0, 1.6fr) minmax(0, 1fr); gap: 0 24px; align-items: start; }
#view-oncall > .oncall-title { max-width: 1080px; }
}
+1
View File
@@ -14,6 +14,7 @@
<link rel="icon" href="/icon.svg" type="image/svg+xml">
<link rel="apple-touch-icon" href="/apple-touch-icon.png">
<link rel="stylesheet" href="/app.css">
<script src="/js/theme.js"></script>
<script type="module" src="/js/app.js"></script>
</head>
<body>
+23 -1
View File
@@ -28,15 +28,37 @@ function render() {
...passwordSection(user, hasPassword),
h('div', { class: 'page-head' }, h('h2', { text: 'Appearance' })),
themePicker(),
h('div', { class: 'only-desktop' },
h('div', { class: 'page-head' }, h('h2', { text: 'Keyboard' })),
h('div', { class: 'card' }, shortcuts())),
h('div', { class: 'page-head' }),
h('button', { class: 'btn btn-block', type: 'button', onclick: signOut }, icon('logout'), 'Sign out'),
// Wrapped in a div: .btn is inline-flex, and only a block-level element
// picks up .view-page > *'s margin:auto centering (see app.css:316).
h('div', {}, h('button', { class: 'btn btn-block', type: 'button', onclick: signOut }, icon('logout'), 'Sign out')),
);
}
// System / Light / Dark. Per browser, not per account: it lives in
// localStorage (see js/theme.js), the same place the selected team does.
const THEMES = [['system', 'System'], ['light', 'Light'], ['dark', 'Dark']];
function themePicker() {
const theme = window.terdutTheme;
const buttons = THEMES.map(([value, text]) => h('button', {
class: 'chip', type: 'button', role: 'radio', text,
onclick: () => { theme.set(value); sync(); },
}));
const sync = () => buttons.forEach((b, i) => b.setAttribute('aria-checked', String(THEMES[i][0] === theme.get())));
sync();
return h('div', { class: 'card card-pad' },
h('div', { class: 'chips theme-picker', role: 'radiogroup', 'aria-label': 'Theme' }, ...buttons),
h('p', { class: 'muted small', text: 'System follows your device. This applies to this browser only.' }));
}
// Where this user's pages go. The onboarding checklist's first step sends
// people here for it, and until now there was nothing here to send them to:
// the topic could only be set with curl or by an administrator.
+6 -1
View File
@@ -102,13 +102,18 @@ function section() {
// buttons, because these are four URLs: app.js intercepts the click, the
// browser's Back walks them, and a reload lands where you were.
function subnav() {
return h('nav', { class: 'subnav', 'aria-label': 'Administration' },
const nav = h('nav', { class: 'subnav', 'aria-label': 'Administration' },
TABS.map((t) => h('a', {
class: 'subnav-link',
href: t.path,
text: t.label,
'aria-current': t.tab === tab ? 'page' : null,
})));
// On a phone the strip overflows; bring the open section into view so a
// tab past the edge (Sources, Switches) is never the one that is hidden.
requestAnimationFrame(() => nav.querySelector('[aria-current]')
?.scrollIntoView({ inline: 'center', block: 'nearest' }));
return nav;
}
// --- overview --------------------------------------------------------------
+3 -1
View File
@@ -96,7 +96,9 @@ function render() {
}
function backLink() {
return h('a', { class: 'back-link', href: '/admin/teams' }, icon('chevronLeft'), h('span', { text: 'Teams' }));
// Wrapped in a div: .back-link is inline-flex, and only a block-level
// element picks up .view-page > *'s margin:auto centering (app.css:316).
return h('div', {}, h('a', { class: 'back-link', href: '/admin/teams' }, icon('chevronLeft'), h('span', { text: 'Teams' })));
}
// --- identity --------------------------------------------------------------
+3 -1
View File
@@ -83,7 +83,9 @@ function render() {
}
function backLink() {
return h('a', { class: 'back-link', href: '/admin/users' }, icon('chevronLeft'), h('span', { text: 'Users' }));
// Wrapped in a div: .back-link is inline-flex, and only a block-level
// element picks up .view-page > *'s margin:auto centering (app.css:316).
return h('div', {}, h('a', { class: 'back-link', href: '/admin/users' }, icon('chevronLeft'), h('span', { text: 'Users' })));
}
// --- identity --------------------------------------------------------------
+2
View File
@@ -150,6 +150,8 @@ export const deleteIntegration = (id, integrationID) =>
export const deadmanSwitches = (id) => call('GET', `/teams/${id}/deadman/switches`);
export const createDeadmanSwitch = (id, body) =>
call('POST', `/teams/${id}/deadman/switches`, { body });
export const updateDeadmanSwitch = (id, switchID, body) =>
call('PUT', `/teams/${id}/deadman/switches/${switchID}`, { body });
export const deleteDeadmanSwitch = (id, switchID) =>
call('DELETE', `/teams/${id}/deadman/switches/${switchID}`);
+2 -1
View File
@@ -202,7 +202,8 @@ function updateBadges() {
const pill = $('open-pill');
pill.hidden = false;
pill.textContent = open ? `${open} open` : 'All clear';
pill.replaceChildren(...(open ? [`${open} open`] : [ui.icon('checkCircle'), 'All clear']));
pill.classList.toggle('all-clear', open === 0);
pill.classList.toggle('has-triggered', triggered > 0);
pill.classList.toggle('all-acked', open > 0 && triggered === 0);
+56 -7
View File
@@ -123,6 +123,16 @@ function who(id, name) {
return name || 'someone';
}
// ackActorLabel renders whoever acknowledged inc, human or service account —
// the two are mutually exclusive (migration 015), and a service account is a
// credential, not "you" or "nobody", so it gets its own branch rather than
// going through who()'s id-vs-myID() check.
function ackActorLabel() {
if (inc.acknowledged_by_id != null) return who(inc.acknowledged_by_id, inc.acknowledged_by);
if (inc.acknowledged_by_service_account_id != null) return inc.acknowledged_by_service_account || 'a service account';
return null;
}
function facts() {
const rows = [];
const add = (k, ...v) => rows.push(h('dt', { text: k }), h('dd', {}, ...v));
@@ -132,8 +142,7 @@ function facts() {
// take reading top to bottom to piece together.
const elapsedTo = inc.resolved_at ? Date.parse(inc.resolved_at) : Date.now();
const responsible = inc.assigned_to_id != null ? who(inc.assigned_to_id, inc.assigned_to)
: inc.acknowledged_by_id != null ? who(inc.acknowledged_by_id, inc.acknowledged_by)
: 'Unassigned';
: ackActorLabel() || 'Unassigned';
rows.push(h('dt', { text: 'At a glance' }), h('dd', { class: 'fact-summary' },
h('span', { class: 'fact-chip' }, icon('clock', 'icon fact-icon'), duration(elapsedTo - Date.parse(inc.triggered_at))),
inc.severity && badge(inc.severity, `plain ${severityClass(inc.severity)}`),
@@ -142,7 +151,7 @@ function facts() {
add('Triggered', when(inc.triggered_at), h('span', { class: 'sub', text: ` · ${ago(inc.triggered_at)}` }));
if (inc.acknowledged_at) {
add('Acknowledged', `${who(inc.acknowledged_by_id, inc.acknowledged_by)} · ${when(inc.acknowledged_at)}`);
add('Acknowledged', `${ackActorLabel()} · ${when(inc.acknowledged_at)}`);
}
add('Assigned', inc.assigned_to_id != null ? who(inc.assigned_to_id, inc.assigned_to) : 'Unassigned');
if (inc.status !== 'resolved' && isFuture(inc.snoozed_until)) {
@@ -206,9 +215,28 @@ function alertItem(a) {
// ---------- timeline ----------
// actorLabel renders whoever performed ev, human or service account — the
// two are mutually exclusive (migration 015). null means the server acted:
// ev.user_id == null no longer means that by itself, now that a service
// account's events also leave it null.
function actorLabel(ev, named = false) {
if (ev.user_id != null) return named ? (ev.username || 'someone') : who(ev.user_id, ev.username);
if (ev.service_account_id != null) return ev.service_account_name || 'a service account';
return null;
}
// assignerLabel is actorLabel for an 'assigned' event, whose user_id is the
// assignee: the person who made the assignment is in the actor_* fields
// (null for assignments from before they were recorded).
function assignerLabel(ev, named = false) {
if (ev.actor_user_id != null) return named ? (ev.actor_username || 'someone') : who(ev.actor_user_id, ev.actor_username);
if (ev.actor_service_account_id != null) return ev.actor_service_account_name || 'a service account';
return null;
}
// named spells users out instead of "you", for text that leaves this page.
function eventText(ev, named = false) {
const person = ev.user_id != null ? (named ? ev.username || 'someone' : who(ev.user_id, ev.username)) : null;
const person = actorLabel(ev, named);
const strong = (t) => h('span', { class: 'who', text: t || 'someone' });
const alertName = () => {
const a = (inc.alerts || []).find((x) => x.id === ev.alert_id);
@@ -220,7 +248,12 @@ function eventText(ev, named = false) {
case 'alert_resolved': return [`Alert resolved: ${alertName()}`];
case 'acknowledged': return [strong(person), ' acknowledged'];
case 'unacknowledged': return [strong(person), ' cleared the acknowledgement'];
case 'assigned': return ['Assigned to ', strong(person)];
case 'assigned': {
const by = assignerLabel(ev, named);
return by ? ['Assigned to ', strong(person), ' by ', strong(by)] : ['Assigned to ', strong(person)];
}
case 'archived': return [strong(person), ' archived the incident'];
case 'unarchived': return [strong(person), ' unarchived the incident'];
case 'snoozed': return [strong(person), ` snoozed until ${ev.detail ? when(ev.detail) : '…'}`];
case 'unsnoozed': return [strong(person), ' ended the snooze'];
case 'resolved': return person ? [strong(person), ' resolved the incident'] : ['Resolved: every alert stopped firing'];
@@ -453,9 +486,25 @@ function quickActions() {
return div;
}
// Desktop has room to show what a phone folds into the More sheet below — see
// the .action-extra/.more-btn rules in app.css. Resolved/archived incidents
// already say everything via primaryAction()/secondaryAction(), so there is
// nothing extra to surface for them.
function extraActions() {
if (!isOpen()) return [];
const out = [
h('button', { class: 'btn btn-sm action-extra', type: 'button', onclick: assign }, icon('user'), 'Assign…'),
h('button', { class: 'btn btn-sm action-extra', type: 'button', onclick: addNote }, icon('note'), 'Add note…'),
];
out.push(inc.status === 'acknowledged'
? h('button', { class: 'btn btn-sm action-extra', type: 'button', onclick: unacknowledge }, icon('undo'), 'Clear ack')
: h('button', { class: 'btn btn-sm action-extra', type: 'button', onclick: resolve }, icon('checkCircle'), 'Resolve…'));
return out;
}
function actionBar() {
const more = h('button', { class: 'btn', type: 'button', 'aria-label': 'More actions', onclick: moreMenu }, icon('more'), 'More');
const bar = h('div', { class: 'actionbar' }, primaryAction(), secondaryAction(), more);
const more = h('button', { class: 'btn more-btn', type: 'button', 'aria-label': 'More actions', onclick: moreMenu }, icon('more'), 'More');
const bar = h('div', { class: 'actionbar' }, primaryAction(), secondaryAction(), ...extraActions(), more);
if (busy) for (const b of bar.querySelectorAll('button')) b.disabled = true;
return bar;
}
+73 -67
View File
@@ -1,5 +1,6 @@
// On-call: who is on duty now, the week around it, and your own next shifts.
// Read-only for now; the TUI edits the schedule.
// Read-only for now; the TUI edits the schedule, so there is no add or swap
// button here.
//
// One team's rota at a time — the viewer's first team, since a viewer in one
// team has nothing to choose between. "On call now" is the exception and shows
@@ -59,11 +60,17 @@ function render() {
clear(view(), error ? h('div', { class: 'load-error', text: error }) : spinner());
return;
}
const team = currentTeam();
clear(view(),
// The top bar carries the title on a phone; the desktop has none.
h('div', { class: 'only-desktop oncall-title' },
h('h1', { text: 'On-call' }),
team && h('p', { class: 'muted', text: `${team.name} · who is on call, the week ahead and your shifts` })),
error && h('div', { class: 'load-error', text: `Showing older data: ${error}` }),
nowCard(),
weekCard(),
myShifts(),
h('div', { class: 'oncall-grid' },
h('div', { class: 'oncall-main' }, nowCard(), weekCard()),
h('div', { class: 'oncall-side' }, myShifts()),
),
);
}
@@ -78,75 +85,73 @@ function nowCard() {
const entries = data.now || [];
const showTeam = entries.length > 1;
if (entries.length === 0) {
return h('div', { class: 'card now-card' },
h('div', { class: 'avatar none', text: '–' }),
h('div', {},
h('div', { class: 'now-label', text: 'On call now' }),
h('div', { class: 'now-name', text: 'Nobody' }),
),
);
return h('div', { class: 'card hero' },
h('div', { class: 'hero-label', text: 'On call now' }),
h('div', { class: 'hero-top' },
h('div', { class: 'avatar hero-avatar none', text: '–' }),
h('div', { class: 'hero-name', text: 'Nobody' })));
}
return h('div', {}, ...entries.map((n) =>
h('div', { class: 'card now-card' },
h('div', { class: 'avatar', text: initial(n.username) }),
h('div', {},
h('div', {
class: 'now-label',
text: showTeam ? `On call now · ${n.team_name}` : 'On call now',
}),
h('div', { class: 'now-name' }, n.username, you(n.user_id)),
),
)));
return h('div', { class: 'hero-list' }, ...entries.map((n) => {
const until = shiftEnd(n);
return h('div', { class: 'card hero' },
h('div', { class: 'hero-label', text: showTeam ? `On call now · ${n.team_name}` : 'On call now' }),
h('div', { class: 'hero-top' },
h('div', { class: 'avatar hero-avatar', text: initial(n.username) }),
h('div', {},
h('div', { class: 'hero-name' }, n.username, you(n.user_id)),
until && h('div', { class: 'hero-until', text: until }))));
}));
}
// Groups the week's 7 days into runs held by the same person (or the same
// empty slot) — the week's own version of the consecutive-day grouping
// myShifts does for a single person's own dates, below. Seven identical rows
// for one person all week collapses to the one bar this way.
function weekRuns(byDate) {
const runs = [];
for (let i = 0; i < 7; i++) {
const date = isoDate(addDays(weekStart, i));
const e = byDate.get(date) || null;
const uid = e ? e.user_id : null;
const last = runs[runs.length - 1];
if (last && last.uid === uid) last.to = date;
else runs.push({ uid, entry: e, from: date, to: date });
}
return runs;
// "until Mon 12 Oct · ends in 3d 20h", for the current team only: the shift's
// end is read off the rota (the run of consecutive days from today held by the
// same person), and that is only loaded for one team. The same midnight
// boundary the "Current shift" row below uses.
function shiftEnd(n) {
const team = currentTeam();
if (!team || n.team_id !== team.id) return null;
const mine = new Map(data.upcoming.map((e) => [e.date, e.user_id]));
let day = new Date();
if (mine.get(isoDate(day)) !== n.user_id) return null;
while (mine.get(isoDate(addDays(day, 1))) === n.user_id) day = addDays(day, 1);
// Still holding the last day loaded: the shift may run on past it.
if (isoDate(day) >= data.upcoming.reduce((m, e) => (e.date > m ? e.date : m), '')) return null;
const end = addDays(parse(isoDate(day)), 1);
return `until ${dayName.format(end)} ${dayDate.format(end)} · ends in ${duration(end - Date.now())}`;
}
// A stable colour per person from the six-colour rcN palette. The week page
// has no member list to take an index from (team.js does), so the id decides.
const personClass = (userID) => `rc${(userID % 6) + 1}`;
// The week as seven cells, Monday to Sunday, each showing who holds that day.
// A hand-over in the middle of the week is visible without reading anything.
function weekCard() {
const byDate = new Map(data.week.map((e) => [e.date, e]));
const today = isoDate(new Date());
const mine = myID();
const days = weekRuns(byDate).map((r) => {
const single = r.from === r.to;
const cls = [
'day',
single ? '' : 'range',
r.from <= today && today <= r.to ? 'today' : '',
r.to < today ? 'past' : '',
r.uid === mine ? 'mine' : '',
].filter(Boolean).join(' ');
const label = single
? [h('span', { class: 'day-name', text: dayName.format(parse(r.from)) }),
h('span', { class: 'day-date', text: dayDate.format(parse(r.from)) })]
: [h('span', {
class: 'day-range',
text: `${dayName.format(parse(r.from))} ${dayDate.format(parse(r.from))} – ${dayName.format(parse(r.to))} ${dayDate.format(parse(r.to))}`,
})];
return h('li', { class: cls },
...label,
h('span', { class: `day-who ${r.entry ? '' : 'nobody'}` }, r.entry ? r.entry.username : 'nobody', r.entry && you(r.entry.user_id)),
);
});
const seen = new Map();
const cells = [];
for (let i = 0; i < 7; i++) {
const d = addDays(weekStart, i);
const date = isoDate(d);
const e = byDate.get(date) || null;
if (e) seen.set(e.user_id, e.username);
cells.push(h('li', {
class: ['strip-day', date === today && 'today', date < today && 'past'].filter(Boolean).join(' '),
'aria-current': date === today ? 'date' : null,
},
h('span', { class: 'strip-name', text: dayName.format(d) }),
h('span', { class: 'strip-date', text: String(d.getDate()) }),
h('span', { class: `avatar strip-avatar ${e ? personClass(e.user_id) : 'none'}`, text: e ? initial(e.username) : '–' }),
h('span', { class: `strip-who${e ? '' : ' nobody'}`, text: e ? (e.user_id === mine ? 'You' : e.username) : 'nobody' })));
}
const thisWeek = isoDate(weekStart) === isoDate(mondayOf(new Date()));
return [
h('div', { class: 'page-head' },
return h('div', { class: 'card week-card' },
h('div', { class: 'week-head' },
h('h2', { text: thisWeek ? 'This week' : 'Week' }),
h('div', { class: 'week-nav' },
h('button', { class: 'btn btn-ghost btn-icon', type: 'button', 'aria-label': 'Previous week', onclick: () => shiftWeek(-1) },
h('button', { class: 'btn btn-icon week-arrow', type: 'button', 'aria-label': 'Previous week', onclick: () => shiftWeek(-1) },
icon('chevronLeft')),
h('button', {
class: 'btn btn-ghost week-label',
@@ -155,12 +160,13 @@ function weekCard() {
onclick: () => { weekStart = mondayOf(new Date()); refresh(); },
text: `${dayDate.format(weekStart)} – ${dayDate.format(addDays(weekStart, 6))}`,
}, h('small', { text: ` Week ${isoWeek(weekStart)}` })),
h('button', { class: 'btn btn-ghost btn-icon', type: 'button', 'aria-label': 'Next week', onclick: () => shiftWeek(1) },
icon('chevronRight')),
),
),
h('ul', { class: 'card days' }, days),
];
h('button', { class: 'btn btn-icon week-arrow', type: 'button', 'aria-label': 'Next week', onclick: () => shiftWeek(1) },
icon('chevronRight')))),
h('ul', { class: 'strip' }, cells),
seen.size > 0 && h('div', { class: 'strip-legend' },
[...seen].map(([id, name]) => h('span', {},
h('i', { class: `strip-dot ${personClass(id)}` }),
id === mine ? `${name} (you)` : name))));
}
// myShifts groups your upcoming dates into runs of consecutive days, then
+187 -97
View File
@@ -168,13 +168,18 @@ function section() {
// buttons, because these are six URLs: app.js intercepts the click, the
// browser's Back walks them, and a reload lands where you were.
function subnav() {
return h('nav', { class: 'subnav', 'aria-label': 'Team' },
const nav = h('nav', { class: 'subnav', 'aria-label': 'Team' },
TABS.map((t) => h('a', {
class: 'subnav-link',
href: t.path,
text: t.label,
'aria-current': t.tab === tab ? 'page' : null,
})));
// On a phone the strip overflows; bring the open section into view so a
// tab past the edge (Sources, Switches) is never the one that is hidden.
requestAnimationFrame(() => nav.querySelector('[aria-current]')
?.scrollIntoView({ inline: 'center', block: 'nearest' }));
return nav;
}
// Names which team's settings the six sections below belong to. It used to be
@@ -735,36 +740,19 @@ function minutesInput(seconds, onChange) {
// --- integrations ----------------------------------------------------------
function integrationsCard() {
const rows = (data.integrations || []).map((i) =>
h('tr', {},
h('td', {}, sourceBadge(i.status)),
h('td', { class: 'wrap' },
h('strong', { text: i.name }),
h('div', { class: 'muted small', text: i.kind })),
// When the key last posted, and when an alert last arrived on it. They
// differ: a payload with nothing usable in it stamps only the first.
h('td', { class: 'muted small' }, timeCell(i.last_used_at)),
h('td', { class: 'muted small' }, timeCell(i.last_alert_at)),
h('td', { class: 'muted small num', title: 'Distinct alerts refreshed in the last 24 hours',
text: String(i.alerts_24h ?? 0) }),
h('td', { class: 'muted small' }, h('span', { title: when(i.created_at), text: ago(i.created_at) })),
h('td', {}, isOwner() && h('div', { class: 'row-actions' },
h('button', {
class: 'btn-sm', type: 'button', text: 'Rename', onclick: () => openRenameSource(i),
}),
h('button', {
class: 'btn-sm danger', type: 'button', text: 'Revoke',
onclick: async () => {
if (!(await confirm({
title: `Revoke ${i.name}?`,
text: 'Anything posting with this key stops delivering immediately. Alerts it already delivered stay.',
confirmLabel: 'Revoke',
danger: true,
}))) return;
act(() => api.deleteIntegration(teamID, i.id));
},
}))),
));
const rows = (data.integrations || []).map((i) => clickableRow(h('tr', {},
h('td', {}, sourceBadge(i.status)),
h('td', { class: 'wrap' },
h('strong', { text: i.name }),
h('div', { class: 'muted small', text: i.kind })),
// When the key last posted, and when an alert last arrived on it. They
// differ: a payload with nothing usable in it stamps only the first.
h('td', { class: 'muted small' }, timeCell(i.last_used_at)),
h('td', { class: 'muted small' }, timeCell(i.last_alert_at)),
h('td', { class: 'muted small num', title: 'Distinct alerts refreshed in the last 24 hours',
text: String(i.alerts_24h ?? 0) }),
h('td', { class: 'muted small' }, h('span', { title: when(i.created_at), text: ago(i.created_at) })),
), () => openSourceDetail(i)));
return h('div', { class: 'card' },
h('div', { class: 'card-head' },
@@ -780,9 +768,9 @@ function integrationsCard() {
? h('div', { class: 'table-scroll' },
h('table', { class: 'admin-table status-table' },
h('thead', {}, h('tr', {},
h('th', { text: 'Status' }), h('th', { text: 'Source' }),
h('th', { text: 'Status' }), h('th', { text: 'Name' }),
h('th', { text: 'Last webhook' }), h('th', { text: 'Last alert' }),
h('th', { class: 'num', text: 'Alerts 24h' }), h('th', { text: 'Created' }), h('th'))),
h('th', { class: 'num', text: 'Alerts 24h' }), h('th', { text: 'Created' }))),
h('tbody', {}, rows)))
: h('p', { class: 'muted', text: 'No alert source yet, so nothing can reach this team.' }),
);
@@ -844,6 +832,34 @@ function openNameSheet({ title, submit, value, run }) {
openSheet(() => [h('h2', { class: 'sheet-title', text: title }), form]);
}
function openSourceDetail(i) {
openDetailSheet(i.name, [
sheetFact('Status', sourceBadge(i.status)),
sheetFact('Kind', i.kind),
sheetFact('Last webhook', timeCell(i.last_used_at)),
sheetFact('Last alert', timeCell(i.last_alert_at)),
sheetFact('Alerts, 24h', String(i.alerts_24h ?? 0)),
sheetFact('Created', h('span', { title: when(i.created_at), text: ago(i.created_at) })),
],
isOwner() && h('button', {
class: 'btn', type: 'button', text: 'Rename',
onclick: () => { closeSheet(); openRenameSource(i); },
}),
isOwner() && h('button', {
class: 'btn btn-danger', type: 'button', text: 'Revoke',
onclick: async () => {
closeSheet();
if (!(await confirm({
title: `Revoke ${i.name}?`,
text: 'Anything posting with this key stops delivering immediately. Alerts it already delivered stay.',
confirmLabel: 'Revoke',
danger: true,
}))) return;
act(() => api.deleteIntegration(teamID, i.id));
},
}));
}
function openNewSource() {
openNameSheet({
title: 'New source', submit: 'Add source', value: '',
@@ -899,6 +915,34 @@ const triggeredCell = (iso, incidentID) => {
: h('span', { title: when(iso), text: ago(iso) });
};
// A list row that opens its details: the list is for finding the thing, the
// sheet it opens is where the buttons are. Links inside the row keep working.
function clickableRow(tr, open) {
tr.classList.add('clickable');
tr.tabIndex = 0;
tr.addEventListener('click', (e) => {
if (!e.target.closest('a')) open();
});
tr.addEventListener('keydown', (e) => {
if (e.key === 'Enter' && e.target === tr) open();
});
return tr;
}
// Label/value rows for a details sheet, and the sheet itself: facts above, then
// Close and whatever the viewer may do. Falsy actions (a non-owner's) drop out.
const sheetFact = (label, value) => [h('dt', { text: label }), h('dd', {}, value)];
function openDetailSheet(title, facts, ...actions) {
openSheet(() => [
h('h2', { class: 'sheet-title', text: title }),
h('dl', { class: 'switch-facts' }, facts),
h('div', { class: 'sheet-actions' },
h('button', { class: 'btn', type: 'button', text: 'Close', onclick: () => closeSheet() }),
...actions),
]);
}
function switchRows(sw) {
const main = h('tr', {},
h('td', {}, switchBadge(sw.status)),
@@ -908,19 +952,8 @@ function switchRows(sw) {
h('td', { class: 'muted small' }, timeCell(sw.last_heartbeat_at)),
h('td', { class: 'muted small' }, triggeredCell(sw.last_triggered_at, sw.open_incident_id)),
h('td', { class: 'muted small', text: duration(sw.timeout_seconds * 1000) }),
h('td', {}, isOwner() && h('button', {
class: 'btn-sm danger', type: 'button', text: 'Remove',
onclick: async () => {
if (!(await confirm({
title: `Remove ${sw.name}?`,
text: 'It stops being watched. An incident it already opened stays open until it is resolved.',
confirmLabel: 'Remove',
danger: true,
}))) return;
act(() => api.deleteDeadmanSwitch(teamID, sw.id));
},
})),
);
clickableRow(main, () => openSwitchDetail(sw));
// One heartbeat is the switch's own times; several are worth telling apart,
// since a live cluster must not hide a dead one.
@@ -935,7 +968,7 @@ function switchRows(sw) {
&& h('code', { class: 'small', text: src.fingerprint })),
h('td', { class: 'muted small' }, timeCell(src.last_heartbeat_at)),
h('td', { class: 'muted small' }, triggeredCell(src.last_triggered_at, src.incident_id)),
h('td'), h('td')))
h('td')))
: [];
return [main, ...sources];
}
@@ -947,7 +980,7 @@ function deadmanCard() {
h('div', { class: 'card-head' },
h('h2', { text: 'Dead man’s switches' }),
isOwner() && h('button', {
class: 'btn', type: 'button', text: 'New switch', onclick: openNewSwitch,
class: 'btn', type: 'button', text: 'New switch', onclick: () => openSwitchForm(),
})),
h('p', { class: 'muted small' },
'Alerts whose ABSENCE is the signal. Receiving one opens nothing; going ',
@@ -956,9 +989,9 @@ function deadmanCard() {
? h('div', { class: 'table-scroll' },
h('table', { class: 'admin-table status-table' },
h('thead', {}, h('tr', {},
h('th', { text: 'Status' }), h('th', { text: 'Switch' }),
h('th', { text: 'Status' }), h('th', { text: 'Name' }),
h('th', { text: 'Last heartbeat' }), h('th', { text: 'Last triggered' }),
h('th', { text: 'Silent after' }), h('th'))),
h('th', { text: 'Silent after' }))),
h('tbody', {}, switches.flatMap(switchRows))))
: h('p', { class: 'muted', text: 'Nothing watched.' }),
);
@@ -966,16 +999,22 @@ function deadmanCard() {
// The form lives in the sheet, not on the page: most visits are to look at the
// list, and a form that is always open is the page this replaced.
function openNewSwitch() {
function openSwitchForm(existing) {
const name = h('input', { type: 'text', placeholder: 'Prod Watchdog', autofocus: true });
const matcher = h('input', {
type: 'text', placeholder: 'alertname=Watchdog,cluster=prod', class: 'wide', required: true,
});
const timeout = h('input', {
type: 'number', min: '1', value: '15', class: 'setting-value', required: true,
type: 'number', min: '1', step: 'any', value: '15', class: 'setting-value', required: true,
});
const severity = h('select', {},
...['critical', 'error', 'warning', 'info'].map((s) => h('option', { value: s, text: s })));
if (existing) {
name.value = existing.name;
matcher.value = existing.matcher;
timeout.value = String(existing.timeout_seconds / 60);
severity.value = existing.severity;
}
const problem = h('p', { class: 'load-error', hidden: true });
const form = h('form', { class: 'stacked-form' },
@@ -990,17 +1029,20 @@ function openNewSwitch() {
problem,
h('div', { class: 'sheet-actions' },
h('button', { class: 'btn', type: 'button', text: 'Cancel', onclick: () => closeSheet(false) }),
h('button', { class: 'btn btn-primary', type: 'submit', text: 'Add switch' })));
h('button', { class: 'btn btn-primary', type: 'submit', text: existing ? 'Save' : 'Add switch' })));
form.addEventListener('submit', async (e) => {
e.preventDefault();
try {
await api.createDeadmanSwitch(teamID, {
const body = {
name: name.value.trim(),
matcher: matcher.value.trim(),
timeout_seconds: Math.round(Number(timeout.value) * 60),
severity: severity.value,
});
};
await (existing
? api.updateDeadmanSwitch(teamID, existing.id, body)
: api.createDeadmanSwitch(teamID, body));
} catch (err) {
problem.textContent = err.message;
problem.hidden = false;
@@ -1010,7 +1052,47 @@ function openNewSwitch() {
refresh();
});
openSheet(() => [h('h2', { class: 'sheet-title', text: 'New switch' }), form]);
openSheet(() => [
h('h2', { class: 'sheet-title', text: existing ? 'Edit switch' : 'New switch' }), form]);
}
// What the list row has no room for, and where Edit and Delete live.
function openSwitchDetail(sw) {
const sources = sw.sources.length > 1
? [sheetFact('Sources', h('div', {}, ...sw.sources.map((src) => h('div', { class: 'source-labels' },
switchBadge(src.status),
...Object.entries(src.labels || {})
.filter(([k]) => k !== 'alertname')
.map(([k, v]) => labelChip(k, v)),
h('span', { class: 'muted small' }, ' ', timeCell(src.last_heartbeat_at))))))]
: [];
openDetailSheet(sw.name, [
sheetFact('Status', switchBadge(sw.status)),
sheetFact('Matcher', h('code', { text: sw.matcher })),
sheetFact('Silent after', duration(sw.timeout_seconds * 1000)),
sheetFact('Severity', sw.severity),
sheetFact('Last heartbeat', timeCell(sw.last_heartbeat_at)),
sheetFact('Last triggered', triggeredCell(sw.last_triggered_at, sw.open_incident_id)),
sources,
],
isOwner() && h('button', {
class: 'btn', type: 'button', text: 'Edit',
onclick: () => { closeSheet(); openSwitchForm(sw); },
}),
isOwner() && h('button', {
class: 'btn btn-danger', type: 'button', text: 'Delete',
onclick: async () => {
closeSheet();
if (!(await confirm({
title: `Remove ${sw.name}?`,
text: 'It stops being watched. An incident it already opened stays open until it is resolved.',
confirmLabel: 'Remove',
danger: true,
}))) return;
act(() => api.deleteDeadmanSwitch(teamID, sw.id));
},
}));
}
// --- members ---------------------------------------------------------------
@@ -1114,45 +1196,18 @@ function membersCard() {
const members = data.members || [];
const owners = members.filter((m) => m.role === 'owner').length;
const rows = members.map((m) => {
const lastOwner = m.role === 'owner' && owners === 1;
return h('tr', {},
h('td', {}, memberBadge(m)),
h('td', { class: 'wrap' },
h('strong', { text: m.username }),
m.user_id === myID() && h('span', { class: 'muted small', text: ' (you)' }),
m.problem && h('div', { class: 'target-problem', text: m.problem })),
h('td', { class: 'muted small' }, m.role, m.source === 'oidc' && ssoBadge()),
h('td', { class: 'muted small' }, shiftCell(m)),
h('td', { class: 'muted small' }, timeCell(m.last_active_at)),
h('td', { class: 'muted small' },
h('span', { title: when(m.joined_at), text: ago(m.joined_at) })),
h('td', {}, isOwner() && h('div', { class: 'row-actions' },
h('button', {
class: 'btn-sm', type: 'button', text: 'Edit',
// The server refuses to edit a membership the groups grant.
disabled: m.source === 'oidc',
title: m.source === 'oidc' ? SSO_MANAGED : null,
onclick: () => openEditMember(m),
}),
h('button', {
class: 'btn-sm danger', type: 'button', text: 'Remove',
disabled: lastOwner || m.source === 'oidc',
title: m.source === 'oidc' ? SSO_MANAGED
: lastOwner ? 'A team needs an owner. Make somebody else one first.' : null,
onclick: async () => {
if (!(await confirm({
title: `Remove ${m.username}?`,
text: 'They lose access to this team. Rota days already assigned to them are not '
+ 'changed, so reassign those from the Rota tab.',
confirmLabel: 'Remove',
danger: true,
}))) return;
act(() => api.removeTeamMember(teamID, m.user_id));
},
}))),
);
});
const rows = members.map((m) => clickableRow(h('tr', {},
h('td', {}, memberBadge(m)),
h('td', { class: 'wrap' },
h('strong', { text: m.username }),
m.user_id === myID() && h('span', { class: 'muted small', text: ' (you)' }),
m.problem && h('div', { class: 'target-problem', text: m.problem })),
h('td', { class: 'muted small' }, m.role, m.source === 'oidc' && ssoBadge()),
h('td', { class: 'muted small' }, shiftCell(m)),
h('td', { class: 'muted small' }, timeCell(m.last_active_at)),
h('td', { class: 'muted small' },
h('span', { title: when(m.joined_at), text: ago(m.joined_at) })),
), () => openMemberDetail(m, owners)));
return [oidcGroupsCard(), h('div', { class: 'card' },
h('div', { class: 'card-head' },
@@ -1167,14 +1222,49 @@ function membersCard() {
? h('div', { class: 'table-scroll' },
h('table', { class: 'admin-table status-table' },
h('thead', {}, h('tr', {},
h('th', { text: 'Status' }), h('th', { text: 'Member' }), h('th', { text: 'Role' }),
h('th', { text: 'Rota' }), h('th', { text: 'Last active' }), h('th', { text: 'Joined' }),
h('th'))),
h('th', { text: 'Status' }), h('th', { text: 'Name' }), h('th', { text: 'Role' }),
h('th', { text: 'Rota' }), h('th', { text: 'Last active' }), h('th', { text: 'Joined' }))),
h('tbody', {}, rows)))
: h('p', { class: 'muted', text: 'Nobody is in this team.' }),
)];
}
function openMemberDetail(m, owners) {
const lastOwner = m.role === 'owner' && owners === 1;
openDetailSheet(m.username, [
sheetFact('Status', memberBadge(m)),
sheetFact('Role', h('span', {}, m.role, m.source === 'oidc' && ssoBadge())),
sheetFact('Rota', shiftCell(m)),
sheetFact('Last active', timeCell(m.last_active_at)),
sheetFact('Joined', h('span', { title: when(m.joined_at), text: ago(m.joined_at) })),
m.problem && sheetFact('Problem', h('span', { class: 'target-problem', text: m.problem })),
],
isOwner() && h('button', {
class: 'btn', type: 'button', text: 'Edit',
// The server refuses to edit a membership the groups grant.
disabled: m.source === 'oidc',
title: m.source === 'oidc' ? SSO_MANAGED : null,
onclick: () => { closeSheet(); openEditMember(m); },
}),
isOwner() && h('button', {
class: 'btn btn-danger', type: 'button', text: 'Remove',
disabled: lastOwner || m.source === 'oidc',
title: m.source === 'oidc' ? SSO_MANAGED
: lastOwner ? 'A team needs an owner. Make somebody else one first.' : null,
onclick: async () => {
closeSheet();
if (!(await confirm({
title: `Remove ${m.username}?`,
text: 'They lose access to this team. Rota days already assigned to them are not '
+ 'changed, so reassign those from the Rota tab.',
confirmLabel: 'Remove',
danger: true,
}))) return;
act(() => api.removeTeamMember(teamID, m.user_id));
},
}));
}
// One sheet for both jobs a member's row has: who, and as what. Adding is
// choosing a person and a role; editing is the same with the person fixed. The
// API is one call either way — POST upserts the role.
+45
View File
@@ -0,0 +1,45 @@
// Theme override: System (no attribute, the OS decides), Light or Dark.
//
// A classic script loaded from <head>, not a module, so the saved choice is on
// <html> before the first paint and a Dark user on a light OS never sees a
// flash. The CSP allows no inline script, hence a file of its own. Account
// (account.js) reads and writes the choice through window.terdutTheme.
(function () {
var KEY = 'terdut.theme';
var COLORS = { light: '#f5f6f8', dark: '#0f1115' };
function get() {
try {
var v = localStorage.getItem(KEY);
return v === 'light' || v === 'dark' ? v : 'system';
} catch (e) {
return 'system'; // storage unavailable
}
}
// The browser chrome colour follows the choice too; the two <meta>s are
// media-keyed to the OS, so a forced theme sets both to the same colour.
function apply(theme) {
var root = document.documentElement;
if (theme === 'system') root.removeAttribute('data-theme');
else root.setAttribute('data-theme', theme);
var metas = document.querySelectorAll('meta[name="theme-color"]');
for (var i = 0; i < metas.length; i++) {
var scheme = /dark/.test(metas[i].media) ? 'dark' : 'light';
metas[i].content = COLORS[theme === 'system' ? scheme : theme];
}
}
function set(theme) {
try {
if (theme === 'system') localStorage.removeItem(KEY);
else localStorage.setItem(KEY, theme);
} catch (e) {
/* storage unavailable: applies for this page view only */
}
apply(theme);
}
window.terdutTheme = { get: get, set: set };
apply(get());
})();